Model training methods, audio recognition methods and related equipment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2026-08-14
AI Technical Summary
对于研究而言,标注数据(或称有标签数据)匮乏,要想获取大量的标注数据,就意味着要在数据分类和标注环节付诸大量的人工成本和时间,耗力又耗时
[0022]考虑到了实际场景中,改版音频与源音频间存在差异,故选择对无标签音频的note序列做增强处理,以期丰富训练样本及提高模型落地时的鲁棒性;其中,通过大量的无标签note特征对初始模型进行初步训练,有助于提高模型的自监督学习效果,并降低对有标签样本数据的成本投入和用时。此外,使用少量的有标签note特征对模型再训练,能进一步增强模型在音频识别任务中的实际预测性能,提升用户体验。
Smart Images

Figure CN117727327B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio technology, and in particular to model training methods, audio recognition methods, and related devices. Background Technology
[0002] In daily life, users have an increasingly wide demand for audio recognition functions, such as song recognition and humming recognition.
[0003] With the massive increase in audio data, more researchers are focusing on using deep learning networks to solve audio recognition problems. However, training these networks relies heavily on labeled data. For research, labeled data is scarce. Obtaining large amounts of labeled data means investing significant manpower and time in data classification and labeling, which is both labor-intensive and time-consuming. Therefore, it is necessary to provide an effective solution. Summary of the Invention
[0004] This application provides a model training method, an audio recognition method, and related equipment to improve the accuracy of the audio recognition method with a small amount of labeled data.
[0005] The first aspect of this application provides a model training method, including:
[0006] The main melody feature note sequence of each unlabeled audio in the library is enhanced to obtain an enhanced note sequence. The enhanced note sequence differs from the unenhanced note sequence in terms of signal-to-noise ratio, sequence length, or speed of sound.
[0007] The enhanced note sequence is used to train the initial model to obtain the intermediate model;
[0008] The modified audio note sequence and the source audio note sequence are input into the intermediate model, and the feature similarity between the modified audio embedding features and the source audio embedding features output by the intermediate model is calculated; wherein, both the modified audio and the source audio are labeled audio, and the number of labeled audio is less than the number of unlabeled audio;
[0009] The intermediate model is trained based on the feature similarity to obtain the target model.
[0010] In practice, the method described in the first aspect of this application may be implemented using the content described in the second aspect of this application.
[0011] A second aspect of this application provides an audio recognition method, including:
[0012] The main melody feature note sequence of the audio to be identified is input into the target model to obtain the embedding features corresponding to the note sequence. The target model is trained according to the model training method described in the first aspect or any specific implementation of the first aspect.
[0013] Calculate the feature similarity between the embedding features of the note sequence and the embedding features of each source audio;
[0014] The source audio corresponding to the one with the largest similarity value among the aforementioned features is determined as the target audio to which the audio to be identified belongs.
[0015] A third aspect of this application provides an electronic device, including:
[0016] Central processing unit, memory, and input / output interfaces;
[0017] The memory is either a short-term storage memory or a persistent storage memory;
[0018] The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the method described in the first aspect of the embodiments of this application or any specific implementation thereof.
[0019] A fourth aspect of this application provides a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the method described in the first aspect of this application or any specific implementation thereof.
[0020] The fifth aspect of this application provides a computer program product comprising instructions or a computer program, which, when run on a computer, causes the computer to perform the method described in the first aspect of this application or any specific implementation thereof.
[0021] As can be seen from the above technical solutions, the embodiments of this application have at least the following advantages:
[0022] Considering the differences between modified and original audio in real-world scenarios, we chose to enhance the note sequences of unlabeled audio to enrich the training samples and improve the robustness of the model during deployment. Initial training of the model with a large number of unlabeled note features helps improve the model's self-supervised learning performance and reduces the cost and time required for labeled sample data. Furthermore, retraining the model using a small number of labeled note features further enhances its actual predictive performance in audio recognition tasks, improving the user experience. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings.
[0024] It should be noted that although the steps in the flowcharts (if any) involved in the embodiments are drawn sequentially according to the arrows, unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts involved in the embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.
[0025] Figure 1 This is a schematic diagram of a system architecture according to an embodiment of this application;
[0026] Figure 2 This is a schematic flowchart of the model training method in an embodiment of this application;
[0027] Figure 3 This is another schematic diagram of the model training method in the embodiments of this application;
[0028] Figure 4 This is a schematic diagram of note sequence segmentation in the model training method of this application embodiment;
[0029] Figure 5 This is a flowchart illustrating the audio recognition method according to an embodiment of this application;
[0030] Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0032] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0033] In the following description, expressions such as "one specific implementation" or "one specific example" describe a subset of all possible embodiments. However, it is understood that "one specific implementation" or "one specific example" can be the same or different subset of all possible embodiments and can be combined with each other without conflict. In the following description, the term "multiple" means at least two. When a certain value mentioned in this application reaches a threshold (if it exists), in some specific examples, it may include the former being greater than the latter. When "any" or "at least one" or similar expressions are mentioned, it specifically refers to any one of the listed examples or any combination of these examples.
[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0035] For ease of understanding and explanation, before providing a further detailed description of this application, the nouns and terms involved in the embodiments or drawings of this application will be explained, and the nouns and terms involved in the embodiments or drawings of this application shall be interpreted as follows.
[0036] Contrastive learning is a self-supervised learning method used to learn general features of a dataset by teaching the model which data points are similar or different in the presence of unlabeled data. In other words, this machine learning technique learns to distinguish between similar and dissimilar data points.
[0037] MIDI, or Musical Instrument Digital Interface, is a real-time communication protocol used between electronic musical instruments, synthesizers, and other performance devices for the transmission of real-time performance data between hardware. The initial goal of MIDI was to achieve interoperability between devices from different instrument manufacturers through a unified communication protocol, such as connecting Roland keyboards to Yamaha synthesizers. The MIDI protocol's encoding has been extended to also serve as a file format for recording musical information, known as the "standard MIDI file format."
[0038] The multilayer perceptron (MLP), also known as an artificial neural network (ANN), is primarily used for non-linear classification.
[0039] Please see as follows Figure 1 The diagram shows a system architecture of an embodiment of this application. The audio recognition method provided in this application can be applied to, for example... Figure 1 In the illustrated application environment, terminal 102 communicates with server 101 via a network, and data storage system 100 stores the data that server 101 needs to process. Data storage system 100 can be integrated onto server 101 or placed in the cloud or on other network servers. Terminal 102 can acquire user-recorded audio to be identified, such as a modified audio clip or adapted section of the song "Ten Years" (or the original audio clip), and transmit this clip to server 101. Server 101 can output the embeddings corresponding to the note sequence of the humming clip using a trained target model, and identify the source audio of "Ten Years" based on the similarity of the embeddings. Then, server 101 can return the audio information of the source audio (such as the original song playback resource, the original song name, and the original singer, etc.) to terminal 102, thereby fulfilling the need for song recognition from modified audio clips such as humming clips and enhancing the user experience. Of course, the audio mentioned in this application can also be sound signals other than songs, such as crosstalk or recitation; the source audio mentioned in this application can be the original version released first or the cover version released later, but the original version is generally preferred to be identified.
[0040] The aforementioned terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc. The server 101 can be implemented using a standalone server or a server cluster composed of multiple servers. It should be noted that the method provided in this application embodiment can be implemented jointly by the terminal device and the server as described above, or it can be implemented entirely on the server side, or it can be implemented entirely on the terminal device side. The specific implementation can be determined according to the actual application scenario, and no limitation is made here.
[0041] The following will use humming a song as an example to further explain the method of this application in detail.
[0042] Please see Figure 2 The first aspect of this application provides some specific embodiments of a model training method, which include the following operational steps:
[0043] Step 21: Enhance the main melody feature note sequence of each unlabeled audio in the library to obtain the enhanced note sequence.
[0044] Untagged audio includes copied audio. A note, which can be translated as a musical note, has a specific pitch and duration; in other words, a note can be used to represent pitch and its duration. Typically, a song is composed of multiple notes to represent the main melody changes. In this application's embodiments, the tag can be used to mark which source audio a segment of audio (such as a humming audio segment) belongs to.
[0045] The purpose of step 21 is to ensure that the note sequences before and after enhancement differ at least in signal-to-noise ratio, sequence length, or speed of sound. This is because, in actual audio recognition scenarios, the humming audio to be recognized is usually not exactly the same as the original audio (i.e., the source audio). There are often random differences such as noise, varying durations of singing the same lyrics, and varying speeds of singing the same lyrics. Therefore, it is necessary to construct these enhanced audio samples based on the actual situation, so that the training data for model learning is richer and closer to reality, and the robustness is better.
[0046] Step 22: Train the initial model using the enhanced note sequence to obtain the intermediate model.
[0047] In step 22, the training process can optionally add the original note sequence before enhancement. This original note sequence is the enhanced note sequence (which can be called the original sequence). The sequence obtained by enhancing the original sequence can be called the new sequence or the enhanced note sequence. Generally, a single original sequence can generate multiple new sequences depending on the enhancement method. Enhancement methods include any of the following: adding noise, data shifting, data pruning, and data scaling (scaling, i.e., variable speed). In other words, one or more enhancement methods can be selected (or randomly) to produce the enhanced note sequence required in this application, either individually or in combination; there are no restrictions here. Using the enhanced (pre-)note sequence to train the initial model helps improve the model's self-supervised learning (or unsupervised learning) performance, enhances prediction accuracy, and significantly reduces the cost of sample labeling.
[0048] Step 23: Input the note sequence of the modified audio and the note sequence of the source audio of the modified audio into the intermediate model, and calculate the feature similarity between the embedded features of the modified audio and the embedded features of the source audio output by the intermediate model.
[0049] In this application, both the modified audio and the source audio are labeled audio, with the number of labeled audio being less than the number of unlabeled audio. Embedding can be understood as the model's feature prediction and representation for each input note sequence. This embedding can specifically be a multi-dimensional vector or a representation other than a vector, such as a scalar, used to reflect the melodic characteristics of the audio in virtual space. The feature similarity in this application can specifically be one or more similarity metrics such as cosine similarity, L2 norm distance, and Euclidean distance.
[0050] Step 24: Train the intermediate model based on feature similarity to obtain the target model.
[0051] Compared to steps 21 and 22, the training processes in steps 23 and 24 can be understood as supervised learning, or fine-tuning and optimizing the intermediate model obtained in step 22. In other words, the initial model in step 22 is initially trained into an intermediate model using a large amount of unsupervised data (i.e., note features of unlabeled audio). Then, to better apply the model to specific downstream tasks, such as humming-midi recognition, a small number of labeled humming audio samples and their corresponding source song note sequences can be prepared as labeled samples. These labeled samples are then used to further train the intermediate model to obtain the optimized target model. In short, to significantly reduce the cost of labeling samples and achieve high accuracy in model prediction during real-world tasks, the main idea of this application's embodiments is to perform initial training and retraining of the initial model.
[0052] By using the feature similarity between labeled audios, the model is iterated in reverse during the intermediate stage, and the model parameters are adjusted until the target model with optimized performance is obtained. In the downstream task, the target model can predict the audio note sequence in the library that is most similar to the humming note sequence, thereby achieving the goal of accurate song recognition.
[0053] In summary, this application's embodiments take into account the differences between modified and original audio in real-world scenarios. Therefore, it chooses to enhance the note sequences of unlabeled audio to enrich training samples and improve the robustness of the model during deployment. Specifically, initial training of the model using a large number of unlabeled note features helps improve the model's self-supervised learning performance and reduces the cost and time required for labeled sample data. Furthermore, retraining the model using a small number of labeled note features can further enhance the model's actual prediction performance in downstream audio recognition tasks, improving user experience.
[0054] Based on the examples above, some specific possible implementation examples will be provided below. In practical applications, the implementation content of these examples can be combined or implemented separately as needed according to the corresponding functional principles and application logic, depending on the actual scenario.
[0055] Please see Figures 2 to 4 This application provides some other specific embodiments of the model training method, which include the following operation steps:
[0056] As one possible implementation, before step 22, the model training method of this application may further include segmenting the note sequence: obtaining the full note sequence for each unlabeled audio; segmenting each full note sequence into multiple note subsequences with a preset step size; wherein, the first note in two adjacent note subsequences (originating from the same full note sequence) is spaced by a displacement length, and the preset step size and displacement length are determined based on the length characteristics of the text information in the unlabeled audio. The segmented note subsequences can be used as note sequences for training. In practical applications, the execution order between "segmenting the note sequence" and step 21 is not limited, that is, the full note sequence can be enhanced first and then segmented, which can be determined by the needs. The step size and displacement length of different full note sequences can be the same or different. Figure 3 The square-shaped broken-line spectrum and the wavy spectrum shown can be regarded as the spectrum of the note sequence before and after enhancement. The main difference is that the spectrum obtained by enhancement is more irregular than the spectrum that was enhanced.
[0057] like Figure 4As shown, based on the results of length features such as lyric length (i.e., text length) and / or lyric timestamp (time length) of a song's text information, the step size for segmentation is determined to be 3 lines of lyric length and the displacement is 1 line of lyric length, so that the entire note sequence (i.e., the original sequence) is segmented into i note subsequences N. i (i.e., the new sequence), where i is greater than or equal to 2; the shift of 1 here means that for every lyric, the notes of 3 lyrics will be extracted as a subsequence, that is, 1-3 lyrics as one segment, 2-4 lyrics as one segment, and so on. If the shift is 3 lyrics, it means that for every 3 lyrics, the note sequence of 3 lyrics will be extracted as a subsequence. The purpose of such segmentation and interval shifting is that in actual humming recognition scenarios, the data hummed by users is generally only two or three lyrics. Therefore, the training process should also use audio features of about three lyrics as sample data, and the shifting is to construct more training data to participate in model training, so as to ensure that the model fits reality when it is deployed.
[0058] It should be noted that the step size and displacement used for segmenting the entire sequence of each note can be the same or different, depending on the user's requirements.
[0059] Step 21: Enhance the main melody feature note sequence of each unlabeled audio in the library to obtain the enhanced note sequence.
[0060] Generally, the note features of songs in the music library can be pre-calculated offline, and a note can represent the main melody variation of a song. In some specific examples, the implementation of step 21 may include performing at least one of the following enhancement operations: inserting noise signals into each note sequence (i.e., adding noise); adding, deleting, or modifying some note features in each note sequence (i.e., data pruning); adjusting the occurrence time of note features in each note sequence (i.e., data scaling) to obtain a variable-speed note sequence.
[0061] For example, you can selectively add noise features such as wind sounds, car sounds, or male and female bass voices to the original note sequence; or you can add, delete, or modify certain note features in the original note sequence; or you can advance or delay the appearance time of certain note features in the original note sequence to enhance the modified audio, such as humming segments that might exist in actual situations, thereby enriching the training data. Of course, you can choose to use one of the above enhancement operations or combine them randomly, without any specific restrictions.
[0062] As one possible implementation, if the operation of "segmenting the note sequence" is performed first, then the object of the enhancement operation in step 21 is the segmented note subsequence N. iThe augmented note subsequence can be denoted as N_aug i Of course, you can also choose to first randomly enhance the entire sequence, and then split it.
[0063] Step 22: Train the initial model using the enhanced note sequence to obtain the intermediate model.
[0064] In some specific examples, step 22 may include: inputting the note sequences before and after enhancement into the initial model f(·) to output the embedding features corresponding to the note sequences before and after enhancement respectively; taking each enhanced (i.e., before enhancement) original note sequence and the new note sequence enhanced from the original note sequence (i.e. after enhancement) as a positive sample pair, and taking the original note sequence and the new note sequence enhanced from the non-original note sequence as a negative sample pair; if the feature dimension corresponding to each embedding feature does not meet the preset dimension, then perform dimension correction on each embedding feature; when the note sequences before and after enhancement both meet the preset dimension, calculate the embedding feature similarity within each positive sample pair and the embedding feature similarity within each negative sample pair; iteratively train the initial model based on the embedding feature similarity of each pair until the embedding feature similarity within each positive sample pair approaches the similarity threshold, and the embedding feature similarity within each negative sample pair is inversely related to the similarity threshold, then stop training.
[0065] Specifically, the original and enhanced sequences are input into an initial model, which predicts corresponding embeddings for each input note sequence, serving as a representation of the note features. The model mentioned in this application can employ some commonly used backbone networks, such as CNN, LSTM, or ResNet.
[0066] During training, for each batch of samples, the original sequence and its augmented version can be paired as positive samples (i.e., forming a positive sample pair). This is because the sequences before and after augmentation can have the same label. Conversely, within the same batch, two sequences with different labels, "the original sequence and a sequence other than its augmented version," can be paired as negative samples (i.e., forming a negative sample pair). Here, "same label" can refer to two note sequences originating from the same source audio. Positive and negative samples describe the relationship between samples, not simply a single sample. For example, for a humming audio 1 of the song "Ten Years," the note sequence of audio 1 and the note sequence augmented from audio 1 can be paired as a positive sample pair, while the note sequence of audio 1 and the note sequence augmented from the humming song "If" can be paired as a negative sample pair.
[0067] As one possible implementation, if the feature dimensions of the initial model output embeddings do not conform to the preset dimensions, then these embeddings are dimensionally corrected. This is because this application first uses a self-supervised pre-training method, and the training data at this time is augmented, differing from the user's humming data in actual applications, which can be understood as insufficient realism. Therefore, it is necessary to perform dimensional mapping on the model output embeddings to achieve the purpose of dimensional transformation. Optionally, an MLP layer can be added after the model output layer. This MLP layer can map the embeddings to the latent feature space for learning. Specifically, the MLP layer can map (e.g., reduce the dimensionality) of 512-dimensional embeddings to 128-dimensional embeddings. Afterwards, similarity calculation can be performed on the dimensionality-reduced embeddings. Of course, the dimensionality can also be increased, whichever is required.
[0068] Ideally, the similarity, or convergence condition, should be such that the feature similarity between positive samples is as high as possible, ideally approaching a similarity threshold of 1, while the feature similarity between negative samples is as low as possible, ideally moving away from the similarity threshold of 1 and approaching a threshold of 0. The intermediate model trained in this way has a certain optimization effect, but to further enhance the accuracy of the model's predictions, iterative optimization of the intermediate model can be performed, such as by executing step 23.
[0069] As one possible implementation, audio samples with evaluation metrics exceeding a preset threshold (specifically, there may be a distinction between labeled and unlabeled audio) can be classified as labeled or unlabeled audio. The audio evaluation metrics may include playback rate and / or search rate. Here, selecting audio samples with high evaluation metrics for training reflects the preference or demand of a wide range of users for audio of interest, enhancing the model's prediction efficiency and user experience in practical scenarios. Of course, to enrich the training samples, the preset threshold does not need to be set very high, or other factors besides search rate and / or playback rate can be used to select the audio for training; the specific choice is up to the user and is not limited here. In some examples, if an audio segment or at least one fragment of its source audio has been searched or played, it can be counted as a search volume +1 or a playback volume +1.
[0070] Step 23: Input the note sequence of the modified audio and the note sequence of the source audio of the modified audio into the intermediate model, and calculate the feature similarity between the embedded features of the modified audio and the embedded features of the source audio output by the intermediate model.
[0071] Specifically, after pre-training the intermediate model using a large amount of unsupervised data in step 22 above, the model can then be used to further train for downstream tasks. For example, for the humming recognition task, what it really needs to match is the pitch sequence of the user's humming and the note sequence in the music library. Therefore, a batch of labeled humming data and the corresponding source song note sequence can be prepared in advance as labeled samples.
[0072] Here, pitch refers to the level of a note, used to describe the highness or lowness of a tone, and the unit is Hz. In this embodiment of the application, extracting the pitch of the user's humming reveals the pitch changes of the vocal melody. The pitch value and the note value can be converted using a formula; therefore, it can be understood that note and pitch have something in common, and note can be considered to contain pitch.
[0073] Step 24: Train the intermediate model based on feature similarity to obtain the target model.
[0074] In some specific examples, the implementation of step 24 may include: adjusting the model parameters of the intermediate model based on feature similarity until the feature similarity between the latest output modified audio embedding features and the source audio embedding features is maximized, and then stopping training.
[0075] During downstream training, an intermediate model is used to predict embeddings for the humming pitch sequence and its source song note sequence. The feature distance between these embeddings, i.e., the loss function, is calculated. The intermediate model is then iterated backward until the loss is minimized (i.e., the feature similarity is maximized), at which point the target model is obtained. Subsequently, in the validation or application phase, the target model can accurately retrieve the source song note sequence that is closest to the humming pitch to be identified (i.e., the one with the highest similarity), thus meeting the user's need for humming-based song recognition. The relationship between feature similarity and feature distance is: Feature Similarity = 1 - Feature Distance.
[0076] In summary, addressing the issues of scarce labeled data and high labeling costs in humming recognition, this application adopts a contrastive learning approach. It utilizes a large number of unlabeled humming samples for model pre-training (i.e., self-supervised or unsupervised learning). Subsequently, for downstream tasks related to humming recognition, it uses the pitch features of a small number of labeled humming audio samples and the note features of the songs in the database to retrain and fine-tune the model (i.e., supervised learning) to improve the accuracy of the humming recognition algorithm.
[0077] Compared to Figure 2The example shown illustrates that the additional operation "segment the note sequence" may not necessarily be executed in practice. If more than two operations are added, these operations can be implemented in combination or individually, depending on the actual scenario.
[0078] Please see Figure 5 The second aspect of this application provides some specific embodiments of an audio recognition method, which include the following operation steps:
[0079] Step 51: Input the main melody feature note sequence of the audio to be identified into the target model to obtain the embedded features corresponding to the note sequence.
[0080] The target model is trained according to the training method described in the first aspect above. The input note sequence here can be the pitch sequence of the user humming a segment. The pitch can be calculated separately and there are no restrictions on the specifics.
[0081] Step 52: Calculate the feature similarity between the embedding features of the note sequence and the embedding features of each source audio.
[0082] The source audio here is usually the original audio from the music library, but it can also be a modified version.
[0083] Step 53: The source audio corresponding to the one with the largest similarity value among all features is determined as the target audio to which the audio to be identified belongs.
[0084] As one possible implementation, when the feature similarity includes multiple similarity measures, step 53 may include: for each source audio, calculating the feature similarity of different types with weights to obtain a similarity fusion value for the source audio; and determining the source audio corresponding to the largest similarity fusion value as the target audio. For example, the aforementioned similarity measures may include values such as cosine similarity or L2 norm, and the similarity fusion value = cosine distance * weight coefficient_1 + L2 norm distance * weight coefficient_2, where weight coefficient_1 and weight coefficient_2 can be set based on practical experience.
[0085] In summary, this application proposes a self-supervised humming recognition method that can solve the problem of insufficient labeled data in deep learning tasks for humming recognition. At the same time, it can significantly improve the recognition performance of the model even with a small amount of labeled data.
[0086] In the second aspect of this application, the operations performed by the audio recognition method are similar to those described in the first aspect or any specific embodiment of the first aspect, and will not be repeated here. Of course, the specific implementation process of each operation in the first aspect of this application can also be found in the relevant description of the second aspect.
[0087] Please see Figure 6 The electronic device 600 of this application embodiment may include one or more central processing units (CPUs) 601 and a memory 605, wherein the memory 605 stores one or more applications or data.
[0088] The memory 605 can be volatile or persistent storage. The program stored in the memory 605 can include one or more modules, each module including a series of instruction operations on the electronic device. Furthermore, the central processing unit 601 can be configured to communicate with the memory 605 and execute the series of instruction operations stored in the memory 605 on the electronic device 600.
[0089] The electronic device 600 may also include one or more power supplies 602, one or more wired or wireless network interfaces 603, one or more input / output interfaces 604, and / or one or more operating systems, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc. The central processing unit 601 can perform the operations described in the first or second aspect above, which will not be elaborated further.
[0090] This application provides a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the methods described in the first or second aspect above.
[0091] This application provides a computer program product containing instructions or computer programs, which, when run on a computer, causes the computer to perform the methods described in the first or second aspect above.
[0092] It is understood that, in the various embodiments of this application, the sequence number of each step does not imply the order of execution. The execution order of each step should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0093] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system (if it exists) and device described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0094] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system or apparatus, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0095] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0096] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0097] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product (computer program product) is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a business server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A model training method, characterized in that, include: The main melody feature note sequence of each unlabeled audio in the library is enhanced to obtain an enhanced note sequence. The enhanced note sequence differs from the unenhanced note sequence in terms of signal-to-noise ratio, sequence length, or speed of sound. The enhanced note sequence is used to train the initial model to obtain the intermediate model; The modified audio note sequence and the source audio note sequence are input into the intermediate model, and the feature similarity between the modified audio embedding features and the source audio embedding features output by the intermediate model is calculated; wherein, both the modified audio and the source audio are labeled audio, and the number of labeled audio is less than the number of unlabeled audio; The intermediate model is trained based on the aforementioned feature similarity to obtain the target model; The step of training the initial model using the enhanced note sequence includes: The note sequences before and after enhancement are input into the initial model to output the embedding features corresponding to the note sequences before and after enhancement, respectively. Each enhanced original note sequence and the new note sequence enhanced from the original note sequence are taken as positive sample pairs, and the original note sequence and the new note sequence enhanced from a non-original note sequence are taken as negative sample pairs. If the feature dimension corresponding to each of the embedded features does not conform to the preset dimension, then the dimension of each of the embedded features is corrected. When both the note sequences before and after enhancement conform to the preset dimension, calculate the embedding feature similarity within each pair of positive samples and the embedding feature similarity within each pair of negative samples. The initial model is iteratively trained based on the similarity of each set of embedded features until the similarity of embedded features in each positive sample pair approaches the similarity threshold, and the similarity of embedded features in each negative sample pair moves in the opposite direction to the similarity threshold, at which point training stops.
2. The model training method according to claim 1, characterized in that, The process of enhancing the main melody feature note sequence of each unlabeled audio in the library includes performing at least one of the following operations: Noise signals are inserted into each of the aforementioned note sequences; Add, delete, or modify some note features in each of the aforementioned note sequences; Adjust the occurrence time of note features in each note sequence to obtain a variable-speed note sequence.
3. The model training method according to claim 1 or 2, characterized in that, Before training the initial model using the enhanced note sequence, the method further includes: Obtain the complete note sequence for each of the aforementioned unlabeled audio tracks; Each full sequence of notes is divided into multiple subsequences of notes with a preset step size; wherein the first note in two adjacent subsequences of notes is spaced apart by a displacement length, and the preset step size and the displacement length are determined based on the length characteristics of the text information in the unlabeled audio.
4. The model training method according to claim 1, characterized in that, The method further includes: Audio files whose evaluation criteria in the library are higher than a preset threshold are classified as either labeled audio files or unlabeled audio files.
5. The model training method according to claim 1, characterized in that, Training the intermediate model based on the feature similarity includes: The model parameters of the intermediate model are adjusted based on the feature similarity until the feature similarity between the latest output modified audio embedding features and the original audio embedding features is maximized, at which point training stops.
6. An audio recognition method, characterized in that, include: The main melody feature note sequence of the audio to be identified is input into the target model to obtain the embedding features corresponding to the note sequence. The target model is trained according to the model training method of any one of claims 1 to 5. Calculate the feature similarity between the embedding features of the note sequence and the embedding features of each source audio; The source audio corresponding to the one with the largest similarity value among the aforementioned features is determined as the target audio to which the audio to be identified belongs.
7. The audio recognition method according to claim 6, characterized in that, If the feature similarity includes multiple similarity measures, then determining the source audio corresponding to the largest value among the feature similarities as the target audio to which the audio to be identified belongs includes: For each source audio, the feature similarity is weighted and calculated for different types of feature similarities to obtain the similarity fusion value of the source audio. The source audio corresponding to the largest value among the various similarity fusion values is determined as the target audio.
8. An electronic device, characterized in that, include: Central processing unit, memory, and input / output interfaces; The memory is either a short-term storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, Includes instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Singing sound recognition model training method, singing sound recognition method and related device
CN115273826A
Music instruction system
US20110003638A1