Audio data set construction method, and audio recognition model training method and device

Through the editing process of accompaniment separation and dry sound audio, a more accurate cover audio data set is generated, which solves the problem that cover audio data in the prior art cannot accurately reflect the real scene, and improves the training accuracy of the audio recognition model.

CN120048252APending Publication Date: 2025-05-27HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510205951.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing cover audio data construction method cannot accurately reflect the real cover scene, which affects the training accuracy of the audio recognition model.

Method used

The dry sound audio and the initial accompaniment audio are obtained by accompaniment separation, the dry sound audio is edited to enhance its content diversity, and then a cover audio dataset is generated based on the enhanced dry sound audio and reference accompaniment audio.

Benefits of technology

Improve the accuracy of the cover audio dataset and enhance the training accuracy of the audio recognition model, so that the model can more accurately identify real cover audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048252A_ABST
    Figure CN120048252A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and relates to an audio data set construction method and device, an audio recognition model training method and device, a computer program product and electronic equipment. The audio data set construction method comprises the following steps: carrying out accompaniment separation on an original audio to obtain a dry audio and an initial accompaniment audio; editing a dry sound segment in the dry sound audio to obtain a dry sound enhanced audio; generating enhanced audio data of the original audio according to the dry sound enhanced audio and the reference accompaniment audio, so as to construct a turning-singing audio data set according to the multiple pieces of enhanced audio data; wherein the source audio of the reference accompaniment audio does not have a turning and singing relationship with the original audio. According to the method, the accuracy of constructing the turning singing audio data set can be improved, so that the training precision of the audio recognition model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and more particularly, to a method for constructing an audio data set, a method for training an audio recognition model, an apparatus for constructing an audio data set, an apparatus for training an audio recognition model, a computer program product, and an electronic device. Background Art

[0002] With the increasing popularity of live performances and song adaptations, users' demands for song recognition in such scenarios are becoming stronger and stronger. By recognizing cover audio, users can audition, view audio details, or collect audio. An audio recognition model can be trained with cover audio data, and then the cover audio can be recognized using the audio recognition model.

[0003] Currently, cover audio data is obtained by screening and manual annotation. For example, songs with the same lyrics but different singers are selected as cover songs. However, in actual cover scenarios, song recognition is for audio sounds in a real recording environment. During the process from the sound source to the recording device, the audio is affected by environmental noise on the one hand, and reverberation occurs after multiple reflections on the other hand. Therefore, the current method for constructing cover audio data cannot accurately reflect the real cover scenario, which to a certain extent also affects the training accuracy of the audio recognition model trained based on the cover audio data.

[0004] It should be noted that the information disclosed in the above background art is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] An object of the present disclosure is to provide a method and apparatus for constructing an audio data set, a method and apparatus for training an audio recognition model, a computer program product, and an electronic device, so as to at least improve the accuracy of constructing a cover audio data set to a certain extent, thereby improving the training accuracy of the audio recognition model.

[0006] Other features and advantages of the present disclosure will become apparent through the following detailed description, or be learned in part through the practice of the present disclosure.

[0007] According to one aspect of the present disclosure, there is provided a method for constructing an audio data set, including: separating the accompaniment from the original audio to obtain a dry audio and an initial accompaniment audio; editing and processing the dry audio segments in the dry audio to obtain an enhanced dry audio; generating enhanced audio data of the original audio based on the enhanced dry audio and a reference accompaniment audio, so as to construct a cover audio data set according to a plurality of enhanced audio data; wherein, there is no cover relationship between the source audio of the reference accompaniment audio and the original audio.

[0008] In an exemplary embodiment of the present disclosure, the dry sound segments in the dry sound audio are edited to obtain enhanced dry sound audio, including: randomly sampling the dry sound audio along the time axis to determine a plurality of dry sound segments, where there is no overlapping relationship between the dry sound segments; editing each dry sound segment in the dry sound audio to obtain enhanced dry sound audio.

[0009] In an exemplary embodiment of the present disclosure, randomly sampling the dry sound audio along the time axis to determine a plurality of dry sound segments includes: randomly sampling along the time axis of the dry sound audio to determine the time start point of the dry sound segment; randomly sampling within a preset duration range to determine the time length of the dry sound segment; determining a plurality of dry sound segments according to the time start point and time length of each dry sound segment.

[0010] In an exemplary embodiment of the present disclosure, editing each dry sound segment in the dry sound audio to obtain enhanced dry sound audio includes: determining the target enhancement type corresponding to each of the plurality of dry sound segments from a preset dry sound enhancement type; respectively editing each dry sound segment in the dry sound audio based on the target enhancement type corresponding to each of the plurality of dry sound segments to obtain enhanced dry sound audio.

[0011] In an exemplary embodiment of the present disclosure, the preset dry sound enhancement types at least include volume adjustment, pitch adjustment, and dry sound speed adjustment.

[0012] In an exemplary embodiment of the present disclosure, editing the dry sound segments in the dry sound audio to obtain enhanced dry sound audio further includes: obtaining the first duration of the dry sound audio and the second duration of the enhanced dry sound audio; adjusting the audio speed of the enhanced dry sound audio according to the first duration and the second duration so that the duration of the enhanced dry sound audio is the same as that of the dry sound audio.

[0013] In an exemplary embodiment of the present disclosure, generating enhanced audio data of the original audio according to the enhanced dry sound audio and the reference accompaniment audio includes: randomly sampling within a preset weight range to determine the fusion weight; fusing the enhanced dry sound audio and the reference accompaniment audio based on the fusion weight to obtain the enhanced audio data.

[0014] According to one aspect of the present disclosure, there is provided a method for training an audio recognition model, including: obtaining audio sample pairs, the audio sample pairs including positive sample pairs and negative sample pairs, where the positive sample pairs are obtained from a cover audio dataset, the cover audio dataset is constructed according to the method of any one of the above exemplary embodiments, and there is no cover relationship between the negative sample pairs; using the audio sample pairs to train a model to be trained to obtain an audio recognition model.

[0015] According to one aspect of the present disclosure, there is provided an audio dataset construction device, including: a vocal separation module for separating vocals from the original audio to obtain a dry audio and an initial accompaniment audio; an audio editing module for editing and processing the dry audio segments in the dry audio to obtain an enhanced dry audio; an audio generation module for generating enhanced audio data of the original audio based on the enhanced dry audio and a reference accompaniment audio, so as to construct a cover audio dataset according to multiple pieces of enhanced audio data; wherein, there is no cover relationship between the source audio of the reference accompaniment audio and the original audio.

[0016] According to one aspect of the present disclosure, there is provided a training device for an audio recognition model, including: a sample acquisition module for acquiring audio sample pairs, the audio sample pairs including positive sample pairs and negative sample pairs, wherein the positive sample pairs are acquired from a cover audio dataset, the cover audio dataset is constructed according to the method of any one of the above exemplary embodiments, and there is no cover relationship between the negative sample pairs; a model training module for training a model to be trained using the audio sample pairs to obtain an audio recognition model.

[0017] According to one aspect of the present disclosure, there is provided a computer program product including a computer program, which when executed by a processor implements the method of any one of the above.

[0018] According to one aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the method of any one of the above by executing the executable instructions.

[0019] The technical solution in the exemplary embodiment of the present disclosure, when constructing an audio data set, separates the original audio from the accompaniment track to obtain the dry audio and the initial accompaniment audio, edits and processes the dry segments in the dry audio to obtain the enhanced dry audio, generates enhanced audio data of the original audio based on the enhanced dry audio and the reference accompaniment audio, so as to construct a cover audio data set based on multiple enhanced audio data; wherein, there is no cover relationship between the source audio of the reference accompaniment audio and the original audio. On the one hand, through the accompaniment separation operation, the dry track of the original audio is separated from the accompaniment track, and then the dry segments in the dry audio are edited and processed to obtain the enhanced dry audio, thereby increasing the diversity of the dry audio content to simulate more possible changes in interpretation style, and providing diversified dry audio support for generating cover audio. On the other hand, enhanced audio data of the original audio is generated based on the dry voice enhanced audio and the reference accompaniment audio, and there is no cover relationship between the source audio of the reference accompaniment audio and the original audio, simulating the situation in which the accompaniment is changed, replaced or even missing in the real cover scene, avoiding interference to the model by using audio with similar accompaniment as cover audio, which is conducive to the model learning to distinguish whether it is a cover based on the dry voice singing part during training, rather than just based on the accompaniment. In the model application stage, the audio recognition model trained based on the cover audio data set obtained by fully simulating the real cover scene can accurately identify more diverse real cover audio and improve the user experience.

[0020] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the following detailed description with reference to the accompanying drawings, in which several embodiments of the present disclosure are shown in an exemplary and non-limiting manner.

[0022] Figure 1 An application environment diagram related to the technical solution of the exemplary embodiment of the present disclosure is shown.

[0023] Figure 2 A flowchart of an audio recognition method based on cover audio according to an exemplary embodiment of the present disclosure is shown.

[0024] Figure 3 A flowchart of a method for constructing an audio dataset according to an exemplary embodiment of the present disclosure is shown.

[0025] Figure 4 A flowchart of an implementation method for acquiring dry sound enhanced audio according to an exemplary embodiment of the present disclosure is shown.

[0026] Figure 5 Shows a schematic diagram of randomly sampling dry audio according to an exemplary embodiment of the present disclosure.

[0027] Figure 6 Shows a flowchart of another implementation manner of obtaining enhanced dry audio according to an exemplary embodiment of the present disclosure.

[0028] Figure 7 Shows a complete flowchart of a method for constructing an audio data set according to an exemplary embodiment of the present disclosure.

[0029] Figure 8 Shows a flowchart of a method for training an audio recognition model according to an exemplary embodiment of the present disclosure.

[0030] Figure 9 Shows a schematic diagram of the composition of an audio data set construction device according to an exemplary embodiment of the present disclosure.

[0031] Figure 10 Shows a schematic diagram of the composition of an audio recognition model training device according to an exemplary embodiment of the present disclosure.

[0032] Figure 11 Shows a block diagram of an electronic device according to an exemplary embodiment of the present disclosure.

[0033] In the drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Embodiments

[0034] Now, exemplary embodiments will be described more fully with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The same reference numerals in the figures denote the same or similar structures, and thus their detailed descriptions will be omitted.

[0035] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be employed. In other cases, well-known structures, methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present disclosure.

[0036] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more software-hardened modules, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0037] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0038] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0039] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and the theory of algorithm complexity. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.

[0040] The technical solution provided by the exemplary embodiments of the present disclosure relates to the machine learning technology of artificial intelligence. By constructing a cover audio dataset and using the audio samples containing the cover audio dataset to train the model to be trained, an audio recognition model is obtained, improving the training accuracy of the model and achieving accurate recognition of cover audio.

[0041] It should be noted that the exemplary embodiments of the present disclosure can be applied to the field of audio recognition. For example, the song recognition function provided by a video and audio APP (Application), the song recognition function provided by a social software, etc., and there is no limitation thereto.

[0042] The technical solutions provided by the exemplary embodiments of the present disclosure can be applied to an application environment as Figure 1 shown. Among them, the terminal 101 communicates with the server 102 through a network. The data storage system can store the data that the server 102 needs to process. The data storage system can be integrated on the server 102, or can be placed in the cloud or on other network servers.

[0043] In an exemplary embodiment, the technical solutions provided by the exemplary embodiments of the present disclosure can be executed by the server 102, and the corresponding device is set in the server 102. Correspondingly, in this manner executed by the server 102, the server 102 can start executing the steps in the technical solutions of the exemplary embodiments of the present disclosure in response to a trigger execution order, where the trigger execution order can be sent by the terminal used by the user, or can be locally triggered by the server in response to some automated events.

[0044] The server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server 102 can execute background tasks.

[0045] Furthermore, in another exemplary embodiment, the terminal 101 can also have a similar function to the server 102, so as to execute the technical solutions provided by the exemplary embodiments of the present disclosure.

[0046] Among them, the terminal 101 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, an Internet of Things device, and a portable wearable device. The Internet of Things device can be a smart speaker, a smart TV, a smart air conditioner, and a smart vehicle-mounted device, etc. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The terminal 101 can also be referred to as a mobile terminal, a terminal device, a mobile device, etc. The exemplary embodiments of the present disclosure do not limit the type of the terminal 101.

[0047] In addition, the technical solutions of the exemplary embodiments of the present disclosure can also be executed collaboratively by the terminal 101 and the server 102. In this manner of collaborative execution by the terminal 101 and the server 102, some steps in the technical solutions provided by the exemplary embodiments of the present disclosure are executed by the terminal 101, while other steps are executed by the server 102. For example, the server 102 executes the model training phase, and the terminal 101 executes the model application phase. It should be noted that in this manner of collaborative execution by the terminal 101 and the server 102, the steps respectively executed by the terminal 101 and the server 102 can be dynamically adjusted according to actual conditions, and there is no special restriction on this.

[0048] The terminal 101 and the server 102 may be connected directly or indirectly via wireless communication, and the exemplary embodiments of the present disclosure do not impose any special limitation on this.

[0049] First, concepts or nouns involved in exemplary embodiments of the present disclosure are explained.

[0050] Cover: The main melody and lyrics of the audio itself have not changed significantly, but there are one or more changes in the singer, performance method (such as tone, rhythm, transition technique, etc.), accompaniment style, etc. For example, a song originally sung by singer Zhang was covered by singer Teng. The main melody and lyrics of the cover song are basically the same as the original song, but the singing style and accompaniment have changed to a large extent. After listening to the song, the general public can identify that the two songs actually correspond to the same source song. For another example, in the street, Livehouse (music performance space) or bar environment, the singer sings an existing song in the form of self-playing and self-singing.

[0051] Identify songs by listening to them: Use a recording device to record the audio sounds in the scene, and match and retrieve the songs corresponding to the recorded audio sounds in the pre-built music library. For example, in the audio function provided by the music APP, the user triggers the identification function key on the terminal device, and the terminal device turns on the recording function to record and try to identify the song. When the song corresponding to the recorded audio sound is identified, the song details are displayed on the result page for users to view, listen to or collect. If no song is retrieved until the recording time threshold (for example, 15 to 20 seconds) is reached, the process ends.

[0052] Currently, it is possible to identify songs by audio fingerprinting based on signal features. Audio fingerprinting based on signal features is a typical song identification algorithm technology. Its basic principle is to extract the basic signal features that can represent each song in the song library from the audio of each song through signal processing technology, convert them into so-called fingerprint information (usually stored as numerical values, numerical vectors, or sequences of numerical vectors), and assemble the fingerprint information of these songs into an efficient index library. When in use, for the recorded query request to be retrieved, fingerprint information is extracted from it and retrieved and matched in the index library (i.e., fingerprint comparison) to find the song corresponding to the recorded audio to be retrieved. Audio fingerprinting based on signal features has low computational complexity and high noise resistance. However, this technology has a typical drawback: due to its precise comparison based on signal features, it cannot work properly in scenarios where the signal features change. For example, in live singing and playing (such as in bars, concerts, street performances, etc.), cover scenarios such as adaptation and reproduction, compared with the original songs in the song library, the songs recorded by users have changed to a certain extent. Even if the change is very slight, it will cause the audio fingerprint to fail and cannot be recognized.

[0053] For identifying songs from cover audio of the original audio (i.e., audio cover song identification), in live singing and playing, adaptation and reset scenarios, identifying the audio that the user likes through cover audio song identification can solve the problems existing in the audio fingerprint technology.

[0054] In the process of training an audio recognition model through contrastive learning, similar samples can be pulled closer in the feature space, while dissimilar samples can be pushed farther away, so as to learn effective representations of the similarities and differences between the comparison samples. Among them, for a pair of similar samples (positive sample pairs), the loss should be minimized, which is achieved by calculating their distances in the feature space. For a pair of dissimilar samples (negative sample pairs), the loss should be increased to pull them farther apart in the feature space.

[0055] Exemplarily, Figure 2 shows a flowchart of an audio recognition method based on cover audio, as Figure 2 , for the input audio segment (with a duration of t seconds), perform short-time Fourier transform (STFT) spectrum analysis on it to extract acoustic features (such as Mel spectrum, linear spectrum, CQT (Constant-Q Transform) features, etc.). Then input the acoustic features into the cover feature network to map them into embedding vectors with a fixed dimension. Among them, the cover feature network is, for example, a siamese network, which shares network parameters for multiple inputs.

[0056] In the training stage, cover sample pairs are used as positive sample pairs, and non-cover sample pairs are used as negative sample pairs. Contrastive learning loss is adopted for training to obtain the trained cover feature network.

[0057] In the testing stage, the songs in the music library are sliced offline to obtain multiple t-second segments, which are input into the trained cover feature network to obtain the embedding vectors corresponding to each segment and stored in the music library song vector library. For the given audio to be recognized, similarly, it is mapped to an embedding vector, and then a nearest neighbor search is performed in the music library song vector library to obtain possible candidate songs. Among them, the nearest neighbor search is based on the cosine distance or Euclidean distance between vectors. Finally, the candidate songs are sorted and threshold judged to return the segment cover matching result.

[0058] It should be noted that the above cover feature network structure and loss function are both examples and can be flexibly adjusted according to training requirements.

[0059] From Figure 2 As can be seen from the shown process, in the idea of contrastive learning, positive sample pairs and negative sample pairs are required during the training process. In the audio cover scenario, the positive sample pairs are cover sample pairs, which can be collected by screening + manual annotation. For example, select songs in the music library with the same song name and lyrics but different singers. These songs can be considered as cover songs, and then based on auxiliary information such as lyric timestamps, time alignment is performed on the song segments. Finally, through manual annotation, verification and fine-tuning are carried out to obtain cover sample pairs as the positive sample pairs for model training. The negative sample pairs can be generated by simply extracting different songs.

[0060] In the above method of obtaining positive sample pairs, segments are directly intercepted from the audio, but it is not advisable to directly use such audio segments for training because in the real scenario, when users identify songs by listening, it is the audio sound in the real recording environment. During the process from the sound source to the audio acquisition device, on the one hand, the sound will be interfered by environmental noise, and on the other hand, after multiple reflections, reverberation will also occur. Therefore, it is necessary to consider data augmentation for the training data before use.

[0061] Exemplarily, the source audio can be optionally speed-changed, pitch-changed, etc. to generate audio with acoustic features close to real recordings for audio fingerprint song recognition. However, although this method generates audio with reverberation and noise interference, it ignores the particularity of the cover song scenario. In many live acoustic performances and adaptation scenarios, the performing style and accompanying instruments will change significantly. For example, when a cover singer performs, they may slightly adjust the rhythm and pitch of certain words in the lyrics to form their own unique performing style, or they may play and sing with a guitar or other instrument, completely abandoning the original accompaniment. When preparing a training set by batch exporting cover songs from a song library, a large number of cover songs use the original accompaniment and only change the singer. If such positive sample pairs with identical accompaniments are directly used to train an audio recognition model based on cover audio, not only will a suitable model not be trained, but it will also mislead the learning direction of the model.

[0062] Specifically, deep learning itself is data-driven and gradually automatically analyzes the hidden features of the training data through the training process. A correct audio recognition model should originally gradually learn the following rule: when the melodies and lyrics of the vocal parts of two audio are very similar, the two mapped embedding vectors should be close. However, when using positive sample pairs with identical accompaniments but different singers to train the model, the model discovers during learning that the accompaniment parts of the positive sample pairs are very similar, while the accompaniment parts of the negative sample pairs are not. Therefore, the model will tend to make decisions directly based on the accompaniment. As a result, the model will gradually learn that for two segments with more similar accompaniments, the mapped embedding vectors should be closer, and for two segments with dissimilar accompaniments, the mapped embedding vectors should be farther apart. Based on this, the model can obtain a relatively low loss on the training set. However, obviously, the learning direction of the model is deviated, resulting in the model learning something that is not the real desired learning result. If such a model is used for audio recognition based on cover songs, when the accompaniment of the cover song changes (e.g., the cover singer plays and sings by themselves), the corresponding song cannot be recognized.

[0063] It can be seen from this that the current method for constructing cover audio data cannot accurately reflect the real cover song scenario, and to a certain extent, it also affects the training accuracy of the audio recognition model trained based on cover audio data. At the same time, the applicant discovers that the key to determining whether two audio are cover versions of each other is to look at the lyrics and the singing melody of the vocal part, that is, the criterion by which the public judges whether it is a cover song through the human ear.

[0064] Based on this, the exemplary embodiments of the present disclosure first provide a method for constructing an audio data set based on data augmentation, as Figure 3 shown in the flowchart of the audio data set construction method of the exemplary embodiments of the present disclosure. The audio data set construction method includes steps S310 to S330, which are specifically as follows:

[0065] In step S310, the original audio is separated into vocals to obtain dry audio and initial accompaniment audio.

[0066] In an exemplary embodiment of the present disclosure, the original audio is separated into vocals to obtain dry audio and initial accompaniment audio.

[0067] Among them, the original audio refers to the source audio segment used to train the audio recognition model, and the original audio can be a song or pure music (without lyrics) directly exported from a music library. Vocal separation means separating the original audio into tracks to obtain a vocal track (also called a dry track) and an accompaniment track, that is, obtaining dry audio and initial accompaniment audio.

[0068] Among them, a vocal separation and splitting algorithm or existing vocal separation editing software can be used to separate the original audio, and there is no limitation on this. For example, let A represent an original audio, R A represent the dry audio of the original audio, and B A represent the initial accompaniment audio of the original audio. Each track corresponds to an audio segment. Among them, the dry audio is used for subsequent data augmentation, and the initial accompaniment audio can be used as a material to provide accompaniment material for the processing of other original audios.

[0069] In step S320, the dry segments in the dry audio are edited to obtain enhanced dry audio.

[0070] In an exemplary embodiment of the present disclosure, data augmentation refers to generating new training samples by performing various transformations on the training data. The purpose is to increase the diversity of the training set, make the enhanced training set closer to the real scenario, reduce overfitting, improve the generalization ability of the model, and improve the learning effect of the model. In an exemplary embodiment of the present disclosure, the enhanced processing of the dry audio is realized by editing several dry segments of the dry audio to simulate the diverse cover dry sounds in the real scenario.

[0071] In step S330, enhanced audio data of the original audio is generated based on the enhanced dry audio and the reference accompaniment audio, so as to construct a cover audio dataset according to multiple enhanced audio data.

[0072] In an exemplary embodiment of the present disclosure, there is no cover relationship between the source audio referring to the accompaniment audio and the original audio. That is to say, the reference accompaniment audio is completely different from the initial accompaniment audio, so that the obtained enhanced audio data is completely different from the accompaniment part of the original audio, while the lyrics content and the main melody sung by the vocal part still remain similar. During the subsequent model training process based on the obtained cover audio data, it is possible to learn to judge whether it is a cover according to the vocal singing part, rather than just judging according to the accompaniment. For example, the reference accompaniment audio can be the accompaniment audio obtained by separating the backing vocals from other audio. Correspondingly, the initial accompaniment audio of the original audio in step S310 can also be used as the reference accompaniment audio when generating the enhanced audio data of other audio.

[0073] Exemplarily, several vocal enhanced audio is R ′ A , a reference accompaniment audio B can be randomly selected from an audio library that has no cover relationship with the original audio A X , so that the vocal enhanced audio is R ′ A is fused with the reference accompaniment audio B X Since the reference accompaniment audio B X has been prepared in advance in step S310, no additional operation is required in this step.

[0074] Optionally, a pure audio (without lyrics) can also be directly selected as the reference accompaniment audio B X , and there is no limitation on this.

[0075] Furthermore, the vocal enhanced audio can be fused with the reference accompaniment audio to obtain enhanced audio data, so as to construct a cover audio dataset according to multiple enhanced audio data.

[0076] In the technical solution of the exemplary embodiments of the present disclosure, on the one hand, through the accompaniment separation operation, the dry vocal track and the accompaniment track of the original audio are separated, and then the dry vocal segments in the dry vocal audio are edited to obtain enhanced dry vocal audio, increasing the diversity of the dry vocal audio content to simulate more possible changes in the performance style and providing support for diverse dry vocal audio for generating cover audio. On the other hand, enhanced audio data of the original audio is generated based on the enhanced dry vocal audio and the reference accompaniment audio, and there is no cover relationship between the source audio of the reference accompaniment audio and the original audio, simulating the situation where the accompaniment changes, is replaced or even missing in a real cover scenario, avoiding interference with the model by using audio with similar accompaniments as cover audio, and facilitating the model to learn to distinguish whether it is a cover based on the dry vocal singing part during the training process, rather than just based on the accompaniment. In the model application stage, an audio recognition model trained based on a cover audio dataset obtained by fully simulating a real cover scenario can accurately identify more diverse real cover audios, improving the user experience.

[0077] In an exemplary embodiment, an implementation method for obtaining enhanced dry vocal audio is provided. As Figure 4 shown, editing the dry vocal segments in the dry vocal audio to obtain enhanced dry vocal audio may include:

[0078] Step S410: Randomly sample the dry vocal audio along the time axis to determine multiple dry vocal segments, and there is no overlapping relationship between the dry vocal segments.

[0079] As Figure 5 shown is a schematic diagram of randomly sampling the dry vocal audio. The duration of the dry vocal audio is T. Randomly sample along the time axis of the dry vocal audio to obtain multiple non-overlapping dry vocal segments. Among them, the time distance between adjacent dry vocal segments may be the same or different, and the time lengths of different dry vocal segments may also be the same or different.

[0080] Specifically, randomly sampling the dry vocal audio along the time axis to determine multiple dry vocal segments may further include:

[0081] First, randomly sample along the time axis of the dry vocal audio to determine the time start point of the dry vocal segment, then randomly sample within a preset duration range to determine the time length of the dry vocal segment, and finally determine multiple dry vocal segments according to the time start point and time length of each dry vocal segment.

[0082] Among them, the time start point of the dry vocal segment is determined by random sampling. Correspondingly, the time length of the dry vocal segment is also determined by random sampling, thereby increasing the diversity of processing the dry vocal audio.

[0083] Continue to refer to Figure 5As shown, for the dry audio with a duration of T, random sampling is performed along the time axis to obtain k non-overlapping time windows. The starting point "start_t" of each time window is randomly sampled from the dry audio with a duration of T, and the time length dur of the time window is randomly sampled within the preset duration range [min_duration, max_duration]. Here, min_duration and max_duration are preset parameter values, and [min_duration, max_duration] is less than the duration T. When it is necessary to increase the time length of the time window, the numerical values of the preset parameters can be increased. When it is necessary to reduce the time length of the time window, the numerical values of the preset parameters can be decreased. The specific numerical values of the preset parameters are not limited in the exemplary embodiments of the present disclosure. Each time window serves as the area to be edited (i.e., the dry sound segment), and the areas outside the time window will not be edited.

[0084] Since the starting point and time length of each dry sound segment can be obtained through random sampling, the diversity of the dry sound segments in the dry audio can be increased, thereby increasing the diversity of editing the dry audio.

[0085] Step S420: Edit each dry sound segment in the dry audio to obtain enhanced dry audio.

[0086] For different dry sound segments of the dry audio, the editing methods used can be the same or different. By editing the dry sound segments, reverberation, noise, etc. of the cover audio in the real scenario can be simulated, improving the authenticity and diversity of the enhanced dry audio.

[0087] Specifically, as Figure 6 shown, editing each dry sound segment in the dry audio to obtain enhanced dry audio may include:

[0088] Step S610: Determine the target enhancement type corresponding to each of the multiple dry sound segments from the preset dry sound enhancement types.

[0089] The preset dry sound enhancement types at least include volume adjustment, pitch adjustment, and dry sound speed adjustment. Of course, other types may also be included.

[0090] Among them, volume adjustment is to adjust the dry sound volume in the window. The volume gain value gain is randomly sampled and determined within the range of [min_gain, max_gain], where min_gain and max_gain are preset parameter values. For example, min_gain = 0.1 and max_gain = 2.0, and there is no restriction on this. Pitch adjustment is to adjust the dry sound pitch of the dry sound segment. The pitch adjustment value is randomly sampled and determined within the range of [min_semitone, max_semitone]. Among them, semitone is a half tone in the music field and is usually used as a unit to describe pitch. min_semitone and max_semitone are preset parameter values. For example, min_semitone = -2 and max_semitone = 3, indicating a random selection between lowering 2 half tones and raising 3 half tones. Pitch adjustment can use common audio processing tools such as FFMpeg (Fast Forward Mpeg), SoX (SoundeXchange), etc., and there is no restriction on this. Dry sound speed adjustment is to adjust the dry sound speed of the dry sound segment. The speed adjustment value is randomly sampled and determined within the range of [min_speed, max_speed], where min_speed and max_speed are preset parameter values. For example, min_speed = 0.8 and max_speed = 1.2. Dry sound speed adjustment can also use common audio processing tools, and there is no restriction on this.

[0091] For each dry sound segment, one or more target enhancement types can be randomly selected or specified from the preset dry sound enhancement types.

[0092] Step S620: Based on the target enhancement types corresponding to the multiple dry sound segments respectively, edit and process each dry sound segment in the dry sound audio to obtain the enhanced dry sound audio.

[0093] Then, use the target enhancement types corresponding to each dry sound segment to process each dry sound segment respectively to obtain the enhanced dry sound audio.

[0094] It should be understood that in the above editing process, volume adjustment and pitch adjustment will not affect the duration of the dry sound segment, while speed adjustment will. For example, after processing a dry sound segment at 1.1 times the speed, the duration of the dry sound segment will change from dur to dur / 1.1.

[0095] Based on this, if the editing operation of a certain or certain dry sound segments includes dry sound speed adjustment, in order to ensure that the time length of the edited dry sound audio is still T, it is also necessary to eliminate the influence of speed adjustment on the duration of the dry sound audio. Editing and processing the dry sound segments in the dry sound audio to obtain the enhanced dry sound audio further includes:

[0096] First, obtain the first duration of the dry audio and the second duration of the enhanced dry audio. Then, according to the first duration and the second duration, adjust the audio speed of the enhanced dry audio so that the durations of the enhanced dry audio and the dry audio are the same.

[0097] Specifically, the first duration of the dry audio is its original duration T. After editing, the second duration of the enhanced dry audio is T'. Then, adjust the audio speed of the enhanced dry audio according to the first duration T and the second duration T'. The speed value of the speed change is T' / T. Based on this speed value, process the second duration T' of the enhanced dry audio so that the duration of the edited enhanced dry audio changes back to the value T.

[0098] After editing the dry audio segments of the dry audio R A the enhanced dry audio R ′ A is obtained, with a time length of T.

[0099] By editing different dry audio segments of the dry audio according to the target enhancement type, the characteristics of cover singing in a real scenario can be automatically simulated to a great extent, which promotes the generalization of the audio recognition model.

[0100] In an exemplary embodiment, generating the enhanced audio data of the original audio based on the enhanced dry audio and the reference accompaniment audio may include:

[0101] First, randomly sample within a preset weight range to determine the fusion weight. Then, based on the fusion weight, fuse the enhanced dry audio and the reference accompaniment audio to obtain the enhanced audio data.

[0102] Among them, the fusion weight is used to control the ratio of the enhanced dry audio and the reference accompaniment audio.

[0103] Specifically, the fusion of the enhanced dry audio R ′ A and the reference accompaniment audio B X can be achieved through the following formula 1:

[0104] A ′ = R ′ A + αB X Formula 1

[0105] Among them, A' is the enhanced audio data of the original audio A, and α is the fusion weight, which can be randomly sampled and determined within the preset weight range [min_α, max_α]. Among them, min_α and max_α are preset parameter values. For example, min_α = 0 and max_α = 1.5. The preset parameter values can be flexibly adjusted according to the actual scenario, and no limitation is imposed on this.

[0106] By adjusting the fusion weight, it is convenient to simulate the situation where the accompaniment in the real scene changes and is replaced, improving the authenticity and diversity of the enhanced audio data (cover audio).

[0107] Such as Figure 7 shows a complete flowchart of an audio dataset construction method. The following will describe the audio dataset construction method of the exemplary embodiments of the present disclosure in conjunction with Figure 7 explain the audio dataset construction method of the exemplary embodiments of the present disclosure.

[0108] First, input the original audio.

[0109] Secondly, separate the vocals from the original audio to obtain the dry audio and the initial accompaniment audio.

[0110] Among them, the dry audio is used for subsequent data enhancement processing, and the initial accompaniment audio can be used for fusion with the dry audio of other audio.

[0111] Then, edit the dry voice segments in the dry audio to obtain the enhanced dry audio.

[0112] Among them, randomly sample along the time axis of the dry audio to determine the starting time point of the dry voice segment, randomly sample within the preset duration range to determine the time length of the dry voice segment, determine multiple dry voice segments according to the starting time point and time length of each dry voice segment, determine the target enhancement type corresponding to each of the multiple dry voice segments from the preset dry voice enhancement types, and respectively edit each dry voice segment in the dry audio based on the target enhancement type corresponding to each of the multiple dry voice segments to obtain the enhanced dry audio.

[0113] Finally, generate the enhanced audio data of the original audio according to the enhanced dry audio and the reference accompaniment audio, so as to construct a cover audio dataset according to multiple enhanced audio data.

[0114] Among them, there is no cover relationship between the source audio of the reference accompaniment audio and the original audio. Randomly sample within the preset weight range to determine the fusion weight, and based on the fusion weight, fuse the enhanced dry audio with the reference accompaniment audio to obtain the enhanced audio data.

[0115] Among them, the specific content of each execution step has been recorded in the above exemplary embodiments and will not be elaborated here.

[0116] Through the above processing process, enhanced audio data A' is obtained. Compared with the original audio A, the vocal part has changed to a certain extent, simulating possible changes in the performance style in a real cover singing scenario. The accompaniment part is replaced with an accompaniment that has no relation to the original audio A, so that the two audio tracks of such cover singing audio (used as positive sample pairs for model training) do not have similarities in terms of the accompaniment, simulating the phenomenon of accompaniment adaptation, replacement or even deletion in a real cover singing scenario, and providing a more accurate cover singing audio dataset for model training.

[0117] Furthermore, an exemplary embodiment of the present disclosure also provides a training method for an audio recognition model. As Figure 8 shown, the training method of the recognition model may include:

[0118] Step S810: Obtain audio sample pairs, where the audio sample pairs include positive sample pairs and negative sample pairs. Among them, positive sample pairs are obtained from the cover singing audio dataset, and the cover singing audio dataset is constructed according to the method of any one of the above exemplary embodiments. There is no cover singing relationship between negative sample pairs.

[0119] Step S820: Use the audio sample pairs to train the model to be trained to obtain an audio recognition model.

[0120] Specifically, referring to Figure 2 shown, during the training process of the audio recognition model based on contrastive learning, it is necessary to obtain positive sample pairs and negative sample pairs. Among them, positive sample pairs are cover singing sample pairs, which can be obtained from the cover singing audio dataset. Negative sample pairs can be randomly selected from audio that has no cover singing relationship, and there is no limitation on this.

[0121] Furthermore, the positive sample pairs and negative sample pairs can be used to train the model to be trained to obtain an audio recognition model. The model to be trained can be a network structure as Figure 2 shown, and there is no limitation on this. Other model training steps are no different from conventional model training, and will not be elaborated here.

[0122] Using the cover singing audio dataset constructed by the exemplary embodiment of the present disclosure to obtain cover singing audio pairs as positive sample pairs can solve the problem of similar accompaniment in positive sample pairs in the training data of cover singing audio, enabling the model to learn the main features for cover singing discrimination based on the vocal singing part during the contrastive learning training process, that is, learning to make decisions based on the similarity of the vocal singing part. At the same time, diversity changes (dry audio editing) are also introduced in the vocal part, which can significantly improve the generalization of the audio recognition model based on cover singing segments, enabling the model to achieve better recognition effects in various cover singing scenarios.

[0123] The technical solution in the exemplary embodiment of the present disclosure, on the one hand, separates the dry voice track and the accompaniment track of the original audio through the accompaniment separation operation, and then edits and processes the dry voice segment in the dry voice audio to obtain the dry voice enhanced audio, which increases the diversity of the dry voice audio content, so as to simulate more possible changes in the interpretation style, and provide diversified dry voice audio support for the generation of cover audio. On the other hand, the enhanced audio data of the original audio is generated according to the dry voice enhanced audio and the reference accompaniment audio, and there is no cover relationship between the source audio of the reference accompaniment audio and the original audio, simulating the situation in which the accompaniment is changed, replaced or even missing in the real cover scene, avoiding the interference of the model by using the audio with the same accompaniment as the cover audio, which is conducive to the model learning to distinguish whether it is a cover according to the dry voice singing part during the training process, rather than just according to the accompaniment. In the model application stage, the audio recognition model trained based on the cover audio data set obtained by fully simulating the real cover scene can accurately identify more diverse real cover audio and improve the user experience.

[0124] In an exemplary embodiment of the present disclosure, a device for constructing an audio data set is also provided. Figure 9 As shown, the audio data set construction device 900 may include a chorus separation module 910, an audio editing module 920 and an audio generation module 930. Specifically:

[0125] The accompaniment separation module 910 is used to separate the accompaniment from the original audio to obtain the dry audio and the initial accompaniment audio; the audio editing module 920 is used to edit the dry audio segment in the dry audio to obtain the dry enhanced audio; the audio generation module 930 is used to generate enhanced audio data of the original audio based on the dry enhanced audio and the reference accompaniment audio, so as to construct a cover audio data set based on multiple enhanced audio data; wherein, there is no cover relationship between the source audio of the reference accompaniment audio and the original audio.

[0126] In an exemplary embodiment of the present disclosure, the audio editing module 920 is configured to perform: randomly sampling the dry sound audio along the time axis to determine multiple dry sound segments, and there is no overlapping relationship between the dry sound segments; editing each dry sound segment in the dry sound audio to obtain dry sound enhanced audio.

[0127] In an exemplary embodiment of the present disclosure, the audio editing module 920 is configured to perform: random sampling along the time axis of the dry sound audio to determine the time starting point of the dry sound segment; random sampling within a preset time range to determine the time length of the dry sound segment; and determining multiple dry sound segments based on the time starting point and time length of each dry sound segment.

[0128] In an exemplary embodiment of the present disclosure, the audio editing module 920 is configured to perform: determining the target enhancement type corresponding to each of the multiple dry sound segments from a preset dry sound enhancement type; and respectively editing each dry sound segment in the dry sound audio based on the target enhancement type corresponding to each of the multiple dry sound segments to obtain a dry sound enhanced audio.

[0129] In an exemplary embodiment of the present disclosure, the preset dry sound enhancement types at least include volume adjustment, pitch adjustment, and dry sound speed adjustment.

[0130] In an exemplary embodiment of the present disclosure, the audio editing module 920 is configured to perform: obtaining a first duration of the dry sound audio and a second duration of the dry sound enhanced audio; and adjusting the audio speed of the dry sound enhanced audio according to the first duration and the second duration so that the durations of the dry sound enhanced audio and the dry sound audio are the same.

[0131] In an exemplary embodiment of the present disclosure, the audio generation module 930 is configured to perform: randomly sampling within a preset weight range to determine a fusion weight; and based on the fusion weight, fusing the dry sound enhanced audio with a reference accompaniment audio to obtain enhanced audio data.

[0132] Since the detailed content of each functional module of the audio dataset construction device according to the exemplary embodiment of the present disclosure has been described in the exemplary embodiment of the above audio dataset construction method, it will not be elaborated herein.

[0133] Furthermore, an exemplary embodiment of the present disclosure further provides a training device 1000 for an audio recognition model, as Figure 10 shown. The training device 1000 for the audio recognition model may include:

[0134] A sample acquisition module 1010, configured to acquire audio sample pairs, where the audio sample pairs include positive sample pairs and negative sample pairs. Among them, the positive sample pairs are acquired from a cover audio dataset, and the cover audio dataset is constructed according to the method of any one of the above exemplary embodiments, and there is no cover relationship between the negative sample pairs; a model training module 1020, configured to train a model to be trained using the audio sample pairs to obtain an audio recognition model.

[0135] Since the detailed content of each functional module of the training device for the audio recognition model according to the exemplary embodiment of the present disclosure has been described in the exemplary embodiment of the above audio recognition model training method, it will not be elaborated herein.

[0136] It should be noted that although several modules or units of the audio dataset construction device or the training device of the audio recognition model are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0137] Exemplary embodiments of the present disclosure also provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the above-described method is implemented.

[0138] In one embodiment, the computer program product may be a tangible product containing the computer program, such as a computer-readable storage medium storing the computer program. The readable storage medium may be a storage medium based on signals such as electricity, magnetism, light, electromagnetic, infrared, etc., including but not limited to: random access memory (RAM), read-only memory (ROM), magnetic tape, floppy disk, flash memory (Flash), mechanical hard disk (HDD), solid-state drive (SSD), and so on. Exemplarily, the computer program product may be implemented as a non-volatile storage medium storing the computer program, such as read-only memory, Nand Flash, etc.

[0139] In one embodiment, the computer program product may be an intangible product containing the computer program. Exemplarily, the computer program product may be implemented as a virtual digital product, such as an executable file storing the computer program, a digital file such as an installation package.

[0140] The code of the computer program can be written in one or more programming languages. Programming languages such as C language, Java, C++, etc. The program code can be executed entirely on the user's computing device, or partially on the user's computing device, or executed as an independent software package, or partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, such as a local area network (LAN), a wide area network (WAN), etc., or can be connected to an external computing device (for example, through an Internet connection provided by an operator).

[0141] The computer program may be carried or transmitted via electrical, magnetic, optical, electromagnetic, infrared, or other signals. The electronic device may convert the signal carrying the computer program into a digital signal, thereby running the computer program. When the computer program runs on an electronic device, its code is used to enable the electronic device to execute (more specifically, enable the processor of the electronic device to execute) the method steps of various exemplary embodiments of the present disclosure, such as the method described above, for example, including the following steps:

[0142] The original audio is separated from the accompaniment to obtain the dry audio and the initial accompaniment audio; the dry segment in the dry audio is edited to obtain the dry enhanced audio; the enhanced audio data of the original audio is generated according to the dry enhanced audio and the reference accompaniment audio, so as to construct a cover audio data set according to the multiple enhanced audio data; wherein there is no cover relationship between the source audio of the reference accompaniment audio and the original audio.

[0143] The above method steps are performed by a computer program. On the one hand, the dry voice track of the original audio is separated from the accompaniment track through the accompaniment separation operation, and then the dry voice segment in the dry voice audio is edited and processed to obtain the dry voice enhanced audio, which increases the diversity of the dry voice audio content, so as to simulate more possible changes in the interpretation style, and provide diversified dry voice audio support for the generation of cover audio. On the other hand, the enhanced audio data of the original audio is generated according to the dry voice enhanced audio and the reference accompaniment audio, and there is no cover relationship between the source audio of the reference accompaniment audio and the original audio, simulating the situation where the accompaniment is changed, replaced or even missing in the real cover scene, avoiding the interference of the model by using the audio with the same accompaniment as the cover audio, which is conducive to the model learning to distinguish whether it is a cover according to the dry voice singing part during the training process, rather than just according to the accompaniment. In the model application stage, the audio recognition model obtained by training the cover audio data set obtained by fully simulating the real cover scene can accurately identify more diverse real cover audio and improve the user experience.

[0144] In an exemplary embodiment of the present disclosure, dry sound segments in dry sound audio are edited to obtain dry sound enhanced audio, including: randomly sampling the dry sound audio along a time axis to determine multiple dry sound segments, and there is no overlapping relationship between the dry sound segments; editing each dry sound segment in the dry sound audio to obtain dry sound enhanced audio.

[0145] In an exemplary embodiment of the present disclosure, dry sound audio is randomly sampled along the time axis to determine a plurality of dry sound segments, including: performing random sampling along the time axis of the dry sound audio to determine the time starting point of the dry sound segment; performing random sampling within a preset time range to determine the time length of the dry sound segment; and determining a plurality of dry sound segments based on the time starting point and time length of each dry sound segment.

[0146] In an exemplary embodiment of the present disclosure, each dry sound segment in the dry sound audio is edited to obtain enhanced dry sound audio, including: determining the target enhancement type corresponding to each of the multiple dry sound segments from a preset dry sound enhancement type; and respectively editing each dry sound segment in the dry sound audio based on the target enhancement type corresponding to each of the multiple dry sound segments to obtain enhanced dry sound audio.

[0147] In an exemplary embodiment of the present disclosure, the preset dry sound enhancement type at least includes volume adjustment, pitch adjustment, and dry sound speed adjustment.

[0148] In an exemplary embodiment of the present disclosure, when editing the dry sound segments in the dry sound audio to obtain enhanced dry sound audio, it further includes: obtaining the first duration of the dry sound audio and the second duration of the enhanced dry sound audio; and adjusting the audio speed of the enhanced dry sound audio according to the first duration and the second duration so that the durations of the enhanced dry sound audio and the dry sound audio are the same.

[0149] In an exemplary embodiment of the present disclosure, generating enhanced audio data of the original audio according to the enhanced dry sound audio and the reference accompaniment audio includes: randomly sampling within a preset weight range to determine a fusion weight; and fusing the enhanced dry sound audio and the reference accompaniment audio based on the fusion weight to obtain enhanced audio data.

[0150] In addition, in an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is further provided. Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, which can be collectively referred to as "circuitry", "module", or "system" here.

[0151] Next, refer to Figure 11 to describe the electronic device 1100 according to this embodiment of the present disclosure. Figure 11 The shown electronic device 1100 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0152] As Figure 11 shown, the electronic device 1100 is presented in the form of a general-purpose computing device. The components of the electronic device 1100 may include, but are not limited to: the at least one processing unit 1110 described above, the at least one storage unit 1120 described above, a bus 1130 connecting different system components (including the storage unit 1120 and the processing unit 1110), and a display unit 1140.

[0153] Among them, the storage unit stores program code, which can be executed by the processing unit 1110, so that the processing unit 1110 executes the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section above of this specification. For example, the processing unit 1110 can execute the following steps:

[0154] Separate the background vocals from the original audio to obtain the dry audio and the initial accompaniment audio; edit the dry audio segments in the dry audio to obtain enhanced dry audio; generate enhanced audio data for the original audio based on the enhanced dry audio and a reference accompaniment audio, so as to construct a cover audio dataset according to multiple pieces of enhanced audio data; wherein, there is no cover relationship between the source audio of the reference accompaniment audio and the original audio.

[0155] In an exemplary embodiment of the present disclosure, editing the dry audio segments in the dry audio to obtain enhanced dry audio includes: randomly sampling the dry audio along the time axis to determine multiple dry audio segments, and there is no overlapping relationship between the dry audio segments; editing each dry audio segment in the dry audio to obtain enhanced dry audio.

[0156] In an exemplary embodiment of the present disclosure, randomly sampling the dry audio along the time axis to determine multiple dry audio segments includes: randomly sampling along the time axis of the dry audio to determine the time starting point of the dry audio segment; randomly sampling within a preset duration range to determine the time length of the dry audio segment; determining multiple dry audio segments according to the time starting point and time length of each dry audio segment.

[0157] In an exemplary embodiment of the present disclosure, editing each dry audio segment in the dry audio to obtain enhanced dry audio includes: determining the target enhancement type corresponding to each of the multiple dry audio segments from a preset dry audio enhancement type; respectively editing each dry audio segment in the dry audio based on the target enhancement type corresponding to each dry audio segment to obtain enhanced dry audio.

[0158] In an exemplary embodiment of the present disclosure, the preset dry audio enhancement types at least include volume adjustment, pitch adjustment, and dry audio speed adjustment.

[0159] In an exemplary embodiment of the present disclosure, editing the dry audio segments in the dry audio to obtain enhanced dry audio further includes: obtaining the first duration of the dry audio and the second duration of the enhanced dry audio; adjusting the audio speed of the enhanced dry audio according to the first duration and the second duration so that the duration of the enhanced dry audio is the same as that of the dry audio.

[0160] In an exemplary embodiment of the present disclosure, generating enhanced audio data of the original audio based on the dry voice enhanced audio and the reference accompaniment audio includes: randomly sampling within a preset weight range to determine a fusion weight; and based on the fusion weight, fusing the dry voice enhanced audio with the reference accompaniment audio to obtain the enhanced audio data.

[0161] The storage unit 1120 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 1121 and / or a cache storage unit 1122, and may further include a read-only storage unit (ROM) 1123.

[0162] The storage unit 1120 may further include a program / utilities 1124 having a set (at least one) of program modules 1125. Such program modules 1125 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.

[0163] The bus 1130 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.

[0164] The electronic device 1100 may also communicate with one or more external devices 1200 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 1100, and / or may communicate with any device that enables the electronic device 1100 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through an input / output (I / O) interface 1150. And, the electronic device 1100 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 1160. As shown in the figure, the network adapter 1160 communicates with other modules of the electronic device 1100 through the bus 1130. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1100, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0165] From the description of the above embodiments, it is easy for those skilled in the art to understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0166] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously, for example, in multiple modules.

[0167] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.

Claims

1. A method for constructing an audio data set, characterized in that: include: Separate the original audio from the accompaniment to obtain the dry audio and the initial accompaniment audio; Editing the dry sound segment in the dry sound audio to obtain dry sound enhanced audio; Enhanced audio data of the original audio is generated according to the dry sound enhanced audio and the reference accompaniment audio, so as to construct a cover audio data set according to a plurality of the enhanced audio data; wherein, there is no cover relationship between the source audio of the reference accompaniment audio and the original audio.

2. The method according to claim 1, characterized in that The step of editing the dry sound segment in the dry sound audio to obtain the dry sound enhanced audio includes: Randomly sampling the dry sound audio along the time axis to determine a plurality of dry sound segments, wherein there is no overlap between the dry sound segments; Each of the dry sound segments in the dry sound audio is edited to obtain the dry sound enhanced audio.

3. The method according to claim 2, characterized in that The step of randomly sampling the dry sound audio along the time axis to determine a plurality of dry sound segments includes: Random sampling is performed along the time axis of the dry audio to determine the time starting point of the dry audio segment; Perform random sampling within a preset duration range to determine the duration of the dry sound segment; A plurality of dry sound segments are determined according to the time starting point and time length of each dry sound segment.

4. The method according to claim 2, characterized in that: The editing and processing of each of the dry sound segments in the dry sound audio to obtain the dry sound enhanced audio includes: Determining target enhancement types corresponding to each of the plurality of dry sound segments from preset dry sound enhancement types; Based on the target enhancement types corresponding to the plurality of dry sound segments, the dry sound segments in the dry sound audio are edited and processed respectively to obtain the dry sound enhanced audio.

5. The method according to claim 4, characterized in that The preset dry sound enhancement types include at least volume adjustment, pitch adjustment and dry sound speed adjustment.

6. The method according to claim 1, characterized in that The editing and processing of the dry sound segment in the dry sound audio to obtain the dry sound enhanced audio further includes: Acquire a first duration of the dry sound audio and a second duration of the dry sound enhanced audio; The audio speed of the dry sound enhanced audio is adjusted according to the first duration and the second duration so that the durations of the dry sound enhanced audio and the dry sound audio are the same.

7. The method according to claim 1, characterized in that The step of generating the enhanced audio data of the original audio according to the dry sound enhanced audio and the reference accompaniment audio comprises: Random sampling is performed within a preset weight range to determine the fusion weight; Based on the fusion weight, the dry sound enhanced audio is fused with the reference accompaniment audio to obtain the enhanced audio data.

8. A method for training an audio recognition model, characterized in that: include: Acquire audio sample pairs, the audio sample pairs comprising positive sample pairs and negative sample pairs, wherein the positive sample pairs are acquired from a cover audio dataset, the cover audio dataset being constructed according to the method of any one of claims 1 to 7, and there is no cover relationship between the negative sample pairs; The audio sample pairs are used to train the model to be trained to obtain an audio recognition model.

9. An audio data set construction device, characterized in that: include: The accompaniment separation module is used to separate the accompaniment from the original audio to obtain the dry audio and the initial accompaniment audio; An audio editing module, used for editing the dry sound segment in the dry sound audio to obtain dry sound enhanced audio; An audio generation module is used to generate enhanced audio data of the original audio based on the dry sound enhanced audio and the reference accompaniment audio, so as to construct a cover audio data set based on multiple enhanced audio data; wherein there is no cover relationship between the source audio of the reference accompaniment audio and the original audio.

10. A training device for an audio recognition model, characterized in that: include: A sample acquisition module, used to acquire audio sample pairs, wherein the audio sample pairs include positive sample pairs and negative sample pairs, wherein the positive sample pairs are acquired from a cover audio data set, wherein the cover audio data set is constructed according to the method of any one of claims 1 to 7, and there is no cover relationship between the negative sample pairs; The model training module is used to train the model to be trained using the audio sample pairs to obtain an audio recognition model.