Live broadcast audio data cleaning method and related device

By using the preset audio cleaning frame and text similarity to clean again in live audio data cleaning, the problem of cleaning time and low efficiency in the prior art is solved, efficient and automated audio data cleaning is achieved, and the tone cloning effect is improved.

CN119993159APending Publication Date: 2025-05-13GUANGZHOU QUYAN NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510210416.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, the cleaning of live audio data takes a long time and is low in cleaning efficiency, resulting in poor tone cloning effect.

Method used

The original live audio data is initially cleaned using the preset first audio cleaning framework and the second audio cleaning framework. The two frameworks are used to clean the host's vocal data and the inconsistency of the background vocal data, and the cleaning process is automatically reduced to reduce the proportion of dirty data.

Benefits of technology

Through automated cleaning, the proportion of dirty data in the audio data is effectively reduced, and high-quality available vocal data is obtained, which shortens cleaning time, improves cleaning efficiency, and reduces the degree of manual participation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993159A_ABST
    Figure CN119993159A_ABST
Patent Text Reader

Abstract

The invention provides a live broadcast audio data cleaning method and a related device, and the method comprises the steps: carrying out the preliminary cleaning of original live broadcast audio data through employing a first audio cleaning frame and a second audio cleaning frame, and carrying out the cleaning of the original live broadcast audio data through employing the processing consistency of the two frames for anchor voice data, and the processing inconsistency of the two frames for background voice data. And cleaning again from the perspective of text similarity. Therefore, the proportion of dirty data in the audio data can be effectively reduced through an automatic cleaning mode, and high-quality available human voice data can be obtained, so that the manual participation degree can be effectively reduced, even manual cleaning is not needed, the cleaning time can be shortened, and the cleaning efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data cleaning, and in particular to a method and related device for cleaning live audio data. Background Art

[0002] With the rapid development of AI (Artificial Intelligence) generation technology, speech synthesis, as one of the important branches of AI generation, has become one of the important areas of exploration and advancement in the industry. At present, the existing technology has proposed a variety of speech synthesis models, each with its own strengths, and provides a basis for the realization of more voice business forms.

[0003] In the live voice broadcasting business, the existing technology proposes a solution to clone the host's timbre using a large voice synthesis model to provide support and assistance to the host. With the host's authorization, the host's historical live broadcast audio is collected, and the host's timbre is cloned based on the historical live broadcast audio, thereby obtaining a synthesized voice with consistent timbre.

[0004] However, in most cases, historical live audio includes not only the host's vocal data, but also background vocal data caused by background music, assistants or emergencies, which is noisy. Considering that the data quality of the training vocal data determines the upper limit of voice cloning to a certain extent, the less noise the training vocal data contains, the better the voice cloning effect. Therefore, before performing voice cloning, it is necessary to clean the historical live audio data to obtain the host's pure vocal data.

[0005] However, existing cleaning tools, such as Emilia-Pipeline and ClearerVoiceStudio, can perform noise reduction to a certain extent, but the cleaned audio data still contains a lot of background human voice data, and the data quality is low, which has a great negative impact on voice cloning. In order to ensure the voice cloning effect, the existing technology requires manual cleaning of audio data, which takes a long time and has low cleaning efficiency. Summary of the invention

[0006] The purpose of the present application is to solve at least one of the above-mentioned technical defects, especially the technical defects of long cleaning time and low cleaning efficiency in the prior art.

[0007] In a first aspect, an embodiment of the present application provides a method for cleaning live audio data, comprising:

[0008] Acquire original live audio data, where the original live audio data includes background human voice data;

[0009] The original live audio data is cleaned using a preset first audio cleaning framework to obtain first intermediate audio data, and a first intermediate text corresponding to the first intermediate audio data is generated;

[0010] The original live audio data is cleaned by using a preset second audio cleaning framework to obtain second intermediate audio data, and a second intermediate text corresponding to the second intermediate audio data is generated; wherein the first audio cleaning framework and the second audio cleaning framework have different degrees of suppression for background vocal data;

[0011] According to the text similarity between the first intermediate text and the second intermediate text, the background vocal data in the first intermediate audio data and / or the second intermediate audio data is cleaned to obtain the target audio data.

[0012] In some embodiments, the cleaning of the background vocal data in the first intermediate audio data and / or the second intermediate audio data according to the text similarity between the first intermediate text and the second intermediate text, and obtaining the target audio data, includes:

[0013] Determine a reference text and a comparison text; wherein the reference text is one of the first intermediate text and the second intermediate text, and includes N first sentence texts; and the comparison text is the other of the first intermediate text and the second intermediate text;

[0014] Calculating text similarity between each of the first sentence texts and the comparison text, and obtaining sentence similarity corresponding to each of the first sentence texts;

[0015] According to the similarities of each of the sentences, the background vocal data in the benchmark audio data is cleaned to obtain the target audio data; wherein the benchmark audio data is the intermediate audio data corresponding to the benchmark text.

[0016] In some embodiments, the comparison text includes M second sentence texts;

[0017] The calculating text similarity between each of the first sentence text and the comparison text respectively, and obtaining the sentence similarity corresponding to each of the first sentence text, includes:

[0018] For each of the first sentence texts, according to the sentence position of the first sentence text in the benchmark text, K sentences of sentence texts to be compared are determined in the M sentences of the second sentence texts, and the text similarity of the first sentence text is calculated with the K sentences of the sentence texts to be compared respectively, to obtain the sentence similarity corresponding to the first sentence text; wherein K<M.

[0019] In some embodiments, for each of the first sentence texts, determining K sentences to be compared in the M sentences of the second sentence texts according to the sentence position of the first sentence text in the reference text includes:

[0020] For each of the first sentence texts, if the sentence position i of the first sentence text in the reference text is less than or equal to the preset sentence matching number j, the second sentence texts of the 1st to i+jth sentences of the comparison text are used as the sentence texts to be compared;

[0021] For each of the first sentence texts, if the sentence position i of the first sentence text in the reference text is greater than the sentence matching number j, the second sentence texts of the ij~i+jth sentences of the comparison text are used as the sentence texts to be compared.

[0022] In some embodiments, the background vocal data in the reference audio data is cleaned according to the similarities of the respective sentences, and the target audio data is obtained, including:

[0023] According to the similarities of each of the sentences, taking the first sentence text whose sentence similarity is greater than or equal to a preset similarity threshold as a pure sentence text;

[0024] In the reference audio data, respectively determining the clean audio segment corresponding to each of the clean sentence texts;

[0025] The target audio data is obtained based on each of the clean audio segments, so that the target audio data includes each of the clean audio segments.

[0026] In some embodiments, the first audio cleaning framework is an Emilia-Pipeline framework; the second audio cleaning framework is an audio cleaning framework built based on Emilia-Pipeline, using the vocal enhancement module of ClearerVoiceStudio as a vocal separation module.

[0027] In some embodiments, the target audio data includes first target audio data obtained by cleaning the first intermediate audio data, and second target audio data obtained by cleaning the second intermediate audio data;

[0028] After obtaining the target audio data, the method further includes:

[0029] Using the first target audio data and the first target text corresponding to the first target audio data, the initial timbre cloning model is trained to obtain a trained first timbre cloning model;

[0030] The initial timbre cloning model is trained using the second target audio data and the second target text corresponding to the second target audio data, and a trained second timbre cloning model is obtained.

[0031] In a second aspect, an embodiment of the present application provides a device for cleaning live audio data, comprising:

[0032] An original data acquisition module, used to acquire original live audio data, wherein the original live audio data includes background human voice data;

[0033] A first cleaning module, used to clean the original live audio data using a preset first audio cleaning framework to obtain first intermediate audio data, and generate a first intermediate text corresponding to the first intermediate audio data;

[0034] A second cleaning module is used to clean the original live audio data using a preset second audio cleaning framework to obtain second intermediate audio data, and generate a second intermediate text corresponding to the second intermediate audio data; wherein the first audio cleaning framework and the second audio cleaning framework have different degrees of suppression for background vocal data;

[0035] The third cleaning module is used to clean the background vocal data in the first intermediate audio data and / or the second intermediate audio data according to the text similarity between the first intermediate text and the second intermediate text, and obtain the target audio data.

[0036] In a third aspect, an embodiment of the present application provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the live audio data cleaning method described in any of the above embodiments.

[0037] In a fourth aspect, an embodiment of the present application provides a computer device, the computer device comprising: one or more processors, and a memory;

[0038] The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the live audio data cleaning method described in any of the above embodiments are performed.

[0039] In the cleaning method and related device of live audio data provided in some embodiments of the present application, the original live audio data is preliminarily cleaned by using the first audio cleaning framework and the second audio cleaning framework, and the consistency of the two frameworks in processing the host voice data and the inconsistency of the two frameworks in processing the background voice data are used to perform another cleaning from the perspective of text similarity. In this way, the proportion of dirty data in the audio data can be effectively reduced through automated cleaning, and high-quality usable voice data can be obtained, thereby effectively reducing the degree of manual participation, and even eliminating the need for manual participation in cleaning, thereby shortening the cleaning time and improving the cleaning efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0041] Figure 1 A schematic diagram of a flow chart of a method for cleaning live audio data in one embodiment;

[0042] Figure 2 A schematic diagram of a flow chart of a step of obtaining target audio data in one embodiment;

[0043] Figure 3 A schematic diagram of a framework of a method for cleaning live audio data in one embodiment;

[0044] Figure 4 A schematic diagram of the structure of a cleaning device for live audio data in one embodiment;

[0045] Figure 5 FIG. 1 is a diagram of the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0046] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0047] In some embodiments, the embodiments of the present application provide a method for cleaning live audio data. The following embodiments are described by taking the method applied to a computer device as an example. The computer device may be any device with data receiving and data processing functions, and may be, but is not limited to, a server, a desktop computer, a notebook computer, a laptop, a tablet computer, a smart terminal, etc.

[0048] It should be noted that the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in this application are in compliance with relevant laws and regulations and do not violate public order and good morals.

[0049] like Figure 1 As shown, the method for cleaning live audio data provided by the present application may include the following steps:

[0050] S102: Obtain original live audio data.

[0051] The original live audio data may include background vocal data and target speaker vocal data. In the description of this application, background vocal data may be audio data of non-target speakers, including but not limited to audio data of background music, non-target speaker vocal data, environmental noise data, interference audio data, etc.

[0052] It is understandable that, depending on the application scenario, the original live audio data may be different types of audio data, and may be obtained through various public, legal and compliant methods. For example, the original live audio data may be the anchor's historical live audio data. With the anchor's authorization, this application may obtain the original live audio data by real-time recording and / or pulling the live data stream.

[0053] S104: Use a preset first audio cleaning framework to clean the original live audio data to obtain first intermediate audio data, and generate a first intermediate text corresponding to the first intermediate audio data.

[0054] The first audio cleaning framework may be a processing framework for cleaning background human voice data in the original live audio data, which may be implemented in any manner. The first intermediate audio data refers to the audio data obtained after the original live audio data is cleaned using the first audio cleaning framework. The first intermediate text records the voice content in the first intermediate audio data in text form.

[0055] In this step, when the original live audio data is obtained, the computer device can use the first audio cleaning framework to perform audio cleaning on the original live audio data to preliminarily filter out the background human voice data in the original live audio data and obtain the first intermediate audio data.

[0056] The computer device may perform speech-to-text processing on the first intermediate audio data and obtain a first intermediate text corresponding to the first intermediate audio data. It is understood that the present application may use any speech-to-text algorithm to generate the first intermediate text, for example, an ASR (Automatic Speech Recognition) algorithm may be used to generate the first intermediate text.

[0057] S106: Use a preset second audio cleaning framework to clean the original live audio data to obtain second intermediate audio data, and generate a second intermediate text corresponding to the second intermediate audio data.

[0058] Among them, the second audio cleaning framework can be a processing framework for cleaning the background vocal data in the original live audio data, which can be implemented in any way. It should be noted that the first audio cleaning framework and the second audio cleaning framework have different degrees of suppression of background vocal data. For example, the first audio cleaning framework may have a higher degree of suppression of background vocal data than the second audio cleaning framework. For another example, the second audio cleaning framework may have a higher degree of suppression of background vocal data than the first audio cleaning framework.

[0059] The second intermediate audio data refers to the audio data obtained by cleaning the original live audio data using the second audio cleaning framework. The second intermediate text records the voice content in the second intermediate audio data in text form.

[0060] In this step, when the original live audio data is obtained, the present application may use a second audio cleaning framework to perform audio cleaning on the original live audio data to preliminarily filter out the background human voice data in the original live audio data and obtain the second intermediate audio data. It should be noted that there is no specific execution order relationship between step S104 and step S106. For example, the computer device may execute S104 first and then execute S106; for another example, the computer device may execute S106 first and then execute S106; for another example, the computer device may execute S104 and S106 in parallel.

[0061] The computer device may perform speech-to-text processing on the second intermediate audio data and obtain a second intermediate text corresponding to the second intermediate audio data. It is understood that the present application may use any speech-to-text algorithm to generate the second intermediate text, and the present application may use the same or different speech-to-text algorithms to generate the first intermediate text and the second intermediate text respectively. The specific implementation method may be determined according to the actual situation, and the present application does not impose specific restrictions on this. In some examples, the present application may use an ASR algorithm to process the second intermediate audio data to generate the second intermediate text.

[0062] S108: According to the text similarity between the first intermediate text and the second intermediate text, the background vocal data in the first intermediate audio data and / or the second intermediate audio data is cleaned to obtain the target audio data.

[0063] Specifically, since the first audio cleaning framework and the second audio cleaning framework have consistent processing for the vocal data of the target speaker, after the first audio cleaning framework and the second audio cleaning framework are used for audio cleaning respectively, the vocal data of the target speaker are retained to the same degree / similar degree.

[0064] As for the background vocal data, the first audio cleaning framework and the second audio cleaning framework have different suppression levels, so the same background vocal data will have different results after being processed by the two audio cleaning frameworks. For example, when the suppression level of the first audio cleaning framework is lower than that of the second audio cleaning framework, if the background vocal data in the original live audio data is "The great river flows east, the waves wash away all the heroes and heroines throughout the ages", then after the first audio cleaning framework is used for audio cleaning, the background vocal data may be completely retained, so that the first intermediate text includes the complete sentence "The great river flows east, the waves wash away all the heroes and heroines throughout the ages". After the second audio cleaning framework is used for audio cleaning, the background vocal data may be missing or directly lost, so that the second intermediate text includes part of the sentence (for example, only "The great river flows east") or the second intermediate text does not include the sentence at all.

[0065] Based on this, the computer device can combine the differences between the two audio cleaning frameworks and screen the first intermediate audio data and / or the second intermediate audio data in a similarity matching manner according to the text similarity between the first intermediate text and the second intermediate text, thereby effectively reducing the proportion of dirty data in the audio data. For example, the present application can perform audio cleaning on the first intermediate audio data or the second intermediate audio according to the text similarity between the first intermediate text and the second intermediate text. Alternatively, the present application can perform synchronous cleaning on the first intermediate audio data and the second intermediate audio data according to the text similarity between the first intermediate text and the second intermediate text.

[0066] After testing, compared with the prior art, the method provided by this application can reduce the proportion of dirty data in the intermediate audio data from more than 50% to less than 10%, greatly improving the quality of the target audio data. In this way, the original live audio data can be processed in an automated manner, and high-quality pure human voice data can be obtained, which can effectively reduce the degree of manual participation, and even eliminate the need for manual participation in cleaning, thereby greatly improving the frequency data processing efficiency and the quality of human voice data.

[0067] In this application, the original live audio data is initially cleaned using the first audio cleaning framework and the second audio cleaning framework, and the consistency of the two frameworks in processing the host's voice data and the inconsistency of the two frameworks in processing the background voice data are used to perform another cleaning from the perspective of text similarity. In this way, the proportion of dirty data in the audio data can be effectively reduced through automated cleaning, and high-quality usable voice data can be obtained, thereby effectively reducing the degree of manual participation, or even eliminating the need for manual participation in cleaning, thereby shortening the cleaning time and improving the cleaning efficiency.

[0068] In some embodiments, Figure 2 As shown, according to the text similarity between the first intermediate text and the second intermediate text, the background human voice data in the first intermediate audio data and / or the second intermediate audio data is cleaned, and the target audio data is obtained, including:

[0069] S202: Determine a reference text and a comparison text.

[0070] Among them, the reference text can be the text to be screened, which can be one of the first intermediate text and the second intermediate text. The comparison text is the other of the first intermediate text and the second intermediate text. It can be understood that the reference text and the comparison text can be carried out in any way, for example, they can be determined according to preset rules or specified by technicians, and this application does not impose specific restrictions on this. In some examples, the reference text can be the first intermediate text, and the comparison text can be the second intermediate text. In other examples, the reference text can be the second intermediate text, and the comparison text can be the first intermediate text.

[0071] The reference text may include N first sentence texts, where N is a positive integer. It should be noted that the present application may perform text segmentation according to actual needs and obtain N first sentence texts. For example, the present application may perform text segmentation according to preset punctuation marks (such as commas, periods and / or semicolons, etc.). For another example, the present application may perform text segmentation according to the line division of the reference text, for example, one line of text is regarded as one sentence.

[0072] S204: Calculate text similarity between each first sentence text and the comparison text respectively, and obtain sentence similarity corresponding to each first sentence text.

[0073] In this step, the computer device can perform text similarity matching on a sentence basis, and obtain the sentence similarity corresponding to each first sentence text in the first intermediate text. It can be understood that the sentence similarity of the first sentence text is used to reflect the text similarity between the first sentence text and a part of the sentence text / the whole sentence text of the comparison text.

[0074] For example, when N=3, the computer device may calculate the text similarity between the first sentence text A and the comparison text, and obtain the sentence similarity corresponding to the first sentence text A; the computer device may calculate the text similarity between the first sentence text B and the comparison text, and obtain the sentence similarity corresponding to the first sentence text B; the computer device may calculate the text similarity between the first sentence text C and the comparison text, and obtain the sentence similarity corresponding to the first sentence text C.

[0075] It should be noted that this step can be implemented using any text similarity algorithm, including but not limited to string-based similarity algorithms, word vector-based similarity algorithms, deep learning-based similarity algorithms, etc.

[0076] S206: Clean the background human voice data in the reference audio data according to the similarity of each sentence, and obtain the target audio data.

[0077] The reference audio data is the intermediate audio data corresponding to the reference text. For example, when the reference text is the first intermediate text, the reference audio data is the first intermediate audio data. For another example, when the reference text is the second intermediate text, the reference audio data is the second intermediate audio data.

[0078] In this step, since the sentence similarity of the first sentence text can reflect the text similarity between the first sentence text and the partial sentence text / the whole sentence text of the comparison text, and thus can reflect whether the first sentence text is background voice, the computer device can perform another cleaning of the intermediate audio data from the text perspective by differential means. Specifically, the present application can determine whether the N first sentence texts belong to background voices one by one according to the N sentence similarities, and filter the audio data belonging to the background voices, thereby obtaining the target audio data.

[0079] In this embodiment, by calculating text similarity using sentences as matching units and performing audio cleaning based on the sentence similarity of each sentence, background vocal data can be more accurately identified and purer target audio data can be obtained, thereby further improving the quality of the target audio data.

[0080] In some embodiments, the comparative text includes M second sentence texts, where M is a positive integer, and there is no necessary connection between N and M. For example, N may be greater than M, or N may be less than M, or N may be equal to M. It should be noted that the present application may perform text sentence segmentation according to actual needs, and obtain M second sentence texts. For the description of the text sentence segmentation of the comparative text, please refer to the text sentence segmentation description of the above-mentioned benchmark text, and the present application will not repeat it here.

[0081] The text similarity of each first sentence text and the comparison text is calculated respectively, and the sentence similarity corresponding to each first sentence text is obtained, including:

[0082] For each first sentence text, according to the sentence position of the first sentence text in the reference text, K sentences of sentence texts to be compared are determined in the M sentences of the second sentence texts, and the text similarity of the first sentence text and the K sentences of sentence texts to be compared are calculated respectively to obtain the sentence similarity corresponding to the first sentence text; wherein K is a positive integer, and K<M.

[0083] In this embodiment, for the i-th first sentence text of the first intermediate text, if the second intermediate text has a second sentence text associated with the first sentence text, then the associated second sentence text is mostly located at a specific sentence position of the second intermediate text, and the specific sentence position is associated with i. Based on this, in order to improve the cleaning efficiency, the computer device can use part of the sentences of the second intermediate text as the sentences to be compared based on the sentence position of the first sentence text in the first intermediate text, and obtain the sentence similarity corresponding to the first sentence text accordingly.

[0084] Specifically, for each first sentence text tgt_text of the first intermediate text, the computer device may perform the following steps to calculate the sentence similarity corresponding to the first sentence text:

[0085] Step A1: determine the sentence position of tgt_text in the benchmark text;

[0086] Step A2: according to the sentence position of tgt_text in the reference text, determining a portion of the second sentence text of the comparison text as the sentence text to be compared cmp_text;

[0087] Step A3: Calculate the text similarity between tgt_text and K sentences of cmp_text respectively, and obtain the sentence similarity corresponding to tgt_text accordingly.

[0088] For example, when K=3, the computer device may respectively calculate the text similarity between tgt_text and the first sentence cmp_text, the text similarity between tgt_text and the second sentence cmp_text, and the text similarity between tgt_text and the third sentence cmp_text.

[0089] In some examples, in the above step A3, after calculating the text similarities between tgt_text and K sentences of cmp_text, the computer device may obtain K text similarities, and may use the maximum value of the K text similarities as the sentence similarity corresponding to tgt_text.

[0090] In some embodiments, for each first sentence text, determining K sentences of sentence texts to be compared from the M sentences of second sentence texts according to the sentence position of the first sentence text in the reference text includes:

[0091] For each first sentence text, if the first sentence text at the sentence position i of the reference text is less than or equal to the preset sentence matching number j, the second sentence texts of the 1st to i+jth sentences of the comparison text are used as the sentence texts to be compared;

[0092] For each first sentence text, if the sentence position i of the first sentence text in the reference text is greater than the sentence matching number j, the second sentence text of the ij~i+jth sentence of the comparison text is used as the sentence text to be compared.

[0093] In this embodiment, for the i-th first sentence text of the first intermediate text, the computer device may use the j second sentence texts before and after i in the second intermediate text as the texts to be compared, so as to avoid missing possible matching text data as much as possible, thereby improving the recognition accuracy of background voice data and the cleaning efficiency.

[0094] Specifically, for the i-th first sentence text of the first intermediate text, if i≤j, the computer device may use the 1st to i+j-th second sentence texts of the second intermediate text as the text to be compared with the i-th first sentence text. In this case, K=i+j.

[0095] If i>j, the computer device may use the second sentence texts of the ij~i+jth sentences of the second intermediate text as the texts to be compared with the first sentence text of the ith sentence. In this case, K=2j.

[0096] If i+j>M, the computer device may use the second sentence texts of the ij-Mth sentences of the second intermediate text as the texts to be compared with the first sentence text of the ith sentence. In this case, K=M-i+j+1.

[0097] It can be understood that the specific value of the statement matching number j can be preset according to actual conditions, and this application does not impose any specific restrictions on this.

[0098] In some embodiments, the background vocal data in the reference audio data is cleaned according to the similarity of each sentence, and the target audio data is obtained, including:

[0099] According to the similarities of each sentence, the first sentence text whose sentence similarity is greater than or equal to a preset similarity threshold is taken as a pure sentence text;

[0100] In the reference audio data, respectively determine the clean audio segment corresponding to each clean sentence text;

[0101] The target audio data is obtained based on each clean audio segment, so that the target audio data includes each clean audio segment.

[0102] In this embodiment, after obtaining the text similarity corresponding to each first sentence text, the computer device can determine whether each first sentence text belongs to the background voice one by one according to the N text similarities. If the sentence similarity corresponding to the first sentence text is greater than or equal to the preset similarity threshold, it indicates that there is a second sentence text that is highly similar to the first sentence text in the comparison text, and further indicates that the first audio cleaning framework and the second audio cleaning framework have processing consistency for the audio segment corresponding to the first sentence text. Therefore, it can be determined that the first sentence text reflects the speech content of the target speaker, that is, the pure sentence text reflects the speech content of the target speaker.

[0103] On the contrary, if the sentence similarity corresponding to the first sentence text is less than the preset similarity threshold, it indicates that there is no second sentence text highly similar to the first sentence text in the comparison text, which further indicates that the first audio cleaning framework and the second audio cleaning framework have different suppression degrees for the audio segment corresponding to the first sentence text. Therefore, it can be determined that the first sentence text is an audio segment of a non-target speaker, and the audio segment corresponding to the first sentence text belongs to the background voice and should be filtered out.

[0104] It is understood that the specific value of the preset similarity threshold can be set according to the actual situation, and this application does not impose specific restrictions on this. Exemplarily, the preset similarity threshold can be 90% to comprehensively consider the recall rate and accuracy.

[0105] When the clean sentence text is determined, the computer device determines the audio segments corresponding to each clean sentence text in the intermediate audio data corresponding to the reference text (i.e., the reference audio data described above), and obtains each clean audio segment. It can be understood that the clean audio segment is the human voice audio of the target speaker.

[0106] The computer device obtains the target audio data based on each clean audio segment, so that the target audio data includes each clean audio segment, and the audio cleaning is completed. In some examples, the computer device can extract each clean audio segment from the reference audio data, and recombines each clean audio segment to obtain the target audio data.

[0107] In this embodiment, the background voice data and the voice data of the target speaker are distinguished by presetting a similarity threshold, thereby improving the screening efficiency and further improving the cleaning efficiency.

[0108] In one embodiment, the first audio cleaning framework is an Emilia-Pipeline framework. The second audio cleaning framework is an audio cleaning framework built on Emilia-Pipeline and using the human voice enhancement module of ClearerVoiceStudio as a human voice separation module.

[0109] In this embodiment, the open source framework Emilia-Pipeline can be used as the first audio cleaning framework. Considering that Emilia-Pipeline is relatively conservative in the task of separating the target speaker's vocal data from the background sound, for the original live audio data with obvious background music, the audio data cleaned by Emilia-Pipeline still has a lot of background vocal data. Therefore, ClearerVoiceStudio with a stronger noise suppression effect can be used to replace the original vocal separation module of Emilia-Pipeline, to optimize the vocal separation module and improve the suppression effect on background vocal data.

[0110] In some embodiments, the target audio data includes first target audio data obtained by cleaning the first intermediate audio data, and second target audio data obtained by cleaning the second intermediate audio data. That is, the computer device may perform another cleaning on the first intermediate audio data and the second intermediate audio data, respectively, to obtain the first target audio data and the second target audio data.

[0111] After obtaining the target audio data, it also includes:

[0112] Using the first target audio data and the first target text corresponding to the first target audio data, the initial timbre cloning model is trained to obtain a trained first timbre cloning model;

[0113] The initial timbre cloning model is trained using the second target audio data and the second target text corresponding to the second target audio data, and a trained second timbre cloning model is obtained.

[0114] It can be understood that in most cases, the higher the degree of suppression of background vocal data by the audio cleaning framework, the greater the damage to the vocal data of the target speaker. Conversely, the lower the degree of suppression of background vocal data by the audio cleaning framework, the less damage to the vocal data of the target speaker.

[0115] Considering the diversity of user needs, the present application can use target audio data with greater damage and target audio data with less damage for model training, respectively, to obtain a first timbre cloning model and a second timbre cloning model. In this way, the timbre cloning service can be provided through the first timbre cloning model and the second timbre cloning model, thereby meeting user needs in multiple scenarios and improving practicality.

[0116] To facilitate understanding of the solution of the present application, a specific example is provided below. In this example, the first audio cleaning framework is the Emilia-Pipeline framework. The second audio cleaning framework is an audio cleaning framework built on Emilia-Pipeline, using the voice enhancement module of ClearerVoiceStudio as the voice separation module. That is, the second audio cleaning framework is a hybrid data service of Emilia-Pipeline and ClearerVoiceStudio.

[0117] See also Figure 3 The computer device can input the original live audio data into the first audio cleaning framework and the second audio cleaning framework respectively, and obtain the first intermediate audio data output by the first audio cleaning framework and the second intermediate audio data output by the second audio cleaning framework, and respectively generate the first intermediate text corresponding to the first intermediate audio data and the second intermediate text corresponding to the second intermediate audio data.

[0118] When the first intermediate text and the second intermediate text are obtained, the computer device can clean the first intermediate audio data again based on the text similarity matching strategy, and obtain the first target audio data, and generate the first target text corresponding to the first target audio data. Similarly, the computer device can clean the second intermediate audio data again based on the text similarity matching strategy, and obtain the second target audio data, and generate the second target text corresponding to the second target audio data. Among them, the text similarity matching strategy refers to the relevant description of step S108 above, and its specific description can be found in the above embodiment, and this application will not repeat it here.

[0119] The first target audio data tends to have less damage, and the second target audio data tends to have less background vocal data. The computer device can provide the two target audio data to the downstream voice cloning service for model training.

[0120] The cleaning method provided in this example greatly improves the quality of the data used for voice cloning, including audio quality and text quality. At the same time, it also effectively improves the efficiency of audio cleaning.

[0121] The following is a description of a cleaning device for live audio data provided in an embodiment of the present application. The cleaning device for live audio data described below and the cleaning method for live audio data described above can be referenced to each other.

[0122] In some embodiments, Figure 4 As shown, the present application provides a cleaning device 400 for live audio data, comprising:

[0123] The original data acquisition module 402 is used to acquire original live audio data, wherein the original live audio data includes background human voice data;

[0124] A first cleaning module 404 is used to clean the original live audio data using a preset first audio cleaning framework to obtain first intermediate audio data, and generate a first intermediate text corresponding to the first intermediate audio data;

[0125] A second cleaning module 406 is used to clean the original live audio data using a preset second audio cleaning framework to obtain second intermediate audio data, and generate a second intermediate text corresponding to the second intermediate audio data; wherein the first audio cleaning framework and the second audio cleaning framework have different degrees of suppression for background vocal data;

[0126] The third cleaning module 408 is used to clean the background vocal data in the first intermediate audio data and / or the second intermediate audio data according to the text similarity between the first intermediate text and the second intermediate text, and obtain the target audio data.

[0127] In some embodiments, the third cleaning module 408 of the present application may include:

[0128] A reference determination unit, configured to determine a reference text and a comparison text; wherein the reference text is one of the first intermediate text and the second intermediate text, and includes N first sentence texts; and the comparison text is the other of the first intermediate text and the second intermediate text;

[0129] A text similarity calculation unit, used to calculate the text similarity between each of the first sentence texts and the comparison text, and obtain the sentence similarity corresponding to each of the first sentence texts;

[0130] The sentence cleaning unit is used to clean the background vocal data in the reference audio data according to the similarity of each sentence and obtain the target audio data; wherein the reference audio data is the intermediate audio data corresponding to the reference text.

[0131] In some embodiments, the comparison text includes M second sentence texts. The text similarity calculation unit of the present application may include:

[0132] The text comparison unit is used to determine K sentences to be compared in the M sentences of the second sentence text for each sentence of the first sentence text according to the sentence position of the first sentence text in the reference text, and calculate the text similarity between the first sentence text and the K sentences of the sentence text to be compared, so as to obtain the sentence similarity corresponding to the first sentence text; wherein K<M.

[0133] In some embodiments, the text comparison unit of the present application includes:

[0134] A first sentence text to be compared determining unit is used for, for each of the first sentence texts, if the sentence position i of the first sentence text in the reference text is less than or equal to a preset sentence matching number j, to use the second sentence texts of the 1st to i+jth sentences of the comparison text as the sentence text to be compared;

[0135] The second sentence text to be compared determining unit is used for, for each sentence of the first sentence text, if the sentence position i of the first sentence text in the reference text is greater than the sentence matching number j, then taking the second sentence text of the ij~i+jth sentence of the comparison text as the sentence text to be compared.

[0136] In some embodiments, the statement cleaning unit of the present application includes:

[0137] A clean sentence text screening unit, used for selecting, according to the similarities of each of the sentences, a first sentence text whose sentence similarity is greater than or equal to a preset similarity threshold as a clean sentence text;

[0138] A clean audio segment determining unit, configured to determine, in the reference audio data, a clean audio segment corresponding to each of the clean sentence texts;

[0139] The target audio data acquisition unit is used to obtain the target audio data based on each of the clean audio segments, so that the target audio data includes each of the clean audio segments.

[0140] In some embodiments, the first audio cleaning framework is an Emilia-Pipeline framework; the second audio cleaning framework is an audio cleaning framework built based on Emilia-Pipeline, using the vocal enhancement module of ClearerVoiceStudio as a vocal separation module.

[0141] In some embodiments, the target audio data includes first target audio data obtained by cleaning the first intermediate audio data, and second target audio data obtained by cleaning the second intermediate audio data.

[0142] The live audio data cleaning device 400 of the present application also includes:

[0143] A first training module, configured to perform model training on an initial timbre cloning model using the first target audio data and a first target text corresponding to the first target audio data, and obtain a trained first timbre cloning model;

[0144] The second training module is used to use the second target audio data and the second target text corresponding to the second target audio data to perform model training on the initial timbre cloning model and obtain a trained second timbre cloning model.

[0145] In one embodiment, the present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the live audio data cleaning method in any embodiment.

[0146] In one embodiment, the present application also provides a computer device having computer-readable instructions stored therein. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the live audio data cleaning method in any embodiment.

[0147] Indicatively, Figure 5 The internal structure diagram of a computer device provided in an embodiment of the present application is shown in FIG. 1 . In one example, the computer device may be a server. Figure 5 , the computer device 900 includes a processing component 902, which further includes one or more processors, and a memory resource represented by a memory 901, for storing instructions that can be executed by the processing component 902, such as an application. The application stored in the memory 901 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 902 is configured to execute instructions to perform the steps of the method described in any of the above embodiments.

[0148] The computer device 900 may further include a power supply component 903 configured to perform power management of the computer device 900, a wired or wireless network interface 904 configured to connect the computer device 900 to a network, and an input / output (I / O) interface 905. The computer device 900 may operate based on an operating system stored in the memory 901, such as Windows Server TM, Mac OS X TM, Unix TM, Linux TM, Free BSD TM, or the like.

[0149] Those skilled in the art will understand that the internal structure of the computer device shown in the present application is merely a block diagram of a partial structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0150] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the term "include", "comprise" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements not only includes those elements, but also includes other elements not clearly listed, or also includes elements inherent to such process, method, article or equipment. In the absence of more restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the process, method, article or equipment including the elements. Herein, "one", "one", "said", "the" and "it" may also include plural forms, unless the context clearly indicates another way. A plurality refers to at least two cases, such as 2, 3, 5 or 8, etc. "And / or" includes any and all combinations of the relevant listed items.

[0151] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can refer to each other.

[0152] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for cleaning live audio data, characterized in that: include: Acquire original live audio data, where the original live audio data includes background human voice data; The original live audio data is cleaned using a preset first audio cleaning framework to obtain first intermediate audio data, and a first intermediate text corresponding to the first intermediate audio data is generated; The original live audio data is cleaned by using a preset second audio cleaning framework to obtain second intermediate audio data, and a second intermediate text corresponding to the second intermediate audio data is generated; wherein the first audio cleaning framework and the second audio cleaning framework have different degrees of suppression for background vocal data; According to the text similarity between the first intermediate text and the second intermediate text, the background vocal data in the first intermediate audio data and / or the second intermediate audio data is cleaned to obtain the target audio data.

2. The method according to claim 1, characterized in that The step of cleaning the background human voice data in the first intermediate audio data and / or the second intermediate audio data according to the text similarity between the first intermediate text and the second intermediate text, and obtaining the target audio data, includes: Determine a reference text and a comparison text; wherein the reference text is one of the first intermediate text and the second intermediate text, and includes N first sentence texts; and the comparison text is the other of the first intermediate text and the second intermediate text; Calculating text similarity between each of the first sentence texts and the comparison text, and obtaining sentence similarity corresponding to each of the first sentence texts; According to the similarities of each of the sentences, the background vocal data in the benchmark audio data is cleaned to obtain the target audio data; wherein the benchmark audio data is the intermediate audio data corresponding to the benchmark text.

3. The method according to claim 2, characterized in that The comparison text includes M second sentence texts; The calculating text similarity between each of the first sentence text and the comparison text respectively, and obtaining the sentence similarity corresponding to each of the first sentence text, includes: For each of the first sentence texts, according to the sentence position of the first sentence text in the benchmark text, K sentences of sentence texts to be compared are determined in the M sentences of the second sentence texts, and the text similarity of the first sentence text is calculated with the K sentences of the sentence texts to be compared respectively, to obtain the sentence similarity corresponding to the first sentence text; wherein K<M.

4. The method according to claim 3, characterized in that For each of the first sentence texts, determining K sentences to be compared in the M sentences of the second sentence texts according to the sentence position of the first sentence text in the reference text, including: For each of the first sentence texts, if the sentence position i of the first sentence text in the reference text is less than or equal to the preset sentence matching number j, the second sentence texts of the 1st to i+jth sentences of the comparison text are used as the sentence texts to be compared; For each of the first sentence texts, if the sentence position i of the first sentence text in the reference text is greater than the sentence matching number j, the second sentence texts of the ij~i+jth sentences of the comparison text are used as the sentence texts to be compared.

5. The method according to claim 2, characterized in that: The step of cleaning the background human voice data in the reference audio data according to the similarities of the respective sentences and obtaining the target audio data comprises: According to the similarities of each of the sentences, taking the first sentence text whose sentence similarity is greater than or equal to a preset similarity threshold as a pure sentence text; In the reference audio data, respectively determining the clean audio segment corresponding to each of the clean sentence texts; The target audio data is obtained based on each of the clean audio segments, so that the target audio data includes each of the clean audio segments.

6. The method according to claim 1, characterized in that The first audio cleaning framework is the Emilia-Pipeline framework; the second audio cleaning framework is built based on Emilia-Pipeline, and uses the vocal enhancement module of ClearerVoiceStudio as the audio cleaning framework of the vocal separation module.

7. The method according to any one of claims 1 to 6, characterized in that: The target audio data includes first target audio data obtained by cleaning the first intermediate audio data, and second target audio data obtained by cleaning the second intermediate audio data; After obtaining the target audio data, the method further includes: Using the first target audio data and the first target text corresponding to the first target audio data, the initial timbre cloning model is trained to obtain a trained first timbre cloning model; The initial timbre cloning model is trained using the second target audio data and the second target text corresponding to the second target audio data, and a trained second timbre cloning model is obtained.

8. A cleaning device for live audio data, characterized in that: include: An original data acquisition module, used to acquire original live audio data, wherein the original live audio data includes background human voice data; A first cleaning module, used to clean the original live audio data using a preset first audio cleaning framework to obtain first intermediate audio data, and generate a first intermediate text corresponding to the first intermediate audio data; A second cleaning module is used to clean the original live audio data using a preset second audio cleaning framework to obtain second intermediate audio data, and generate a second intermediate text corresponding to the second intermediate audio data; wherein the first audio cleaning framework and the second audio cleaning framework have different degrees of suppression for background vocal data; The third cleaning module is used to clean the background vocal data in the first intermediate audio data and / or the second intermediate audio data according to the text similarity between the first intermediate text and the second intermediate text, and obtain the target audio data.

9. A storage medium, characterized in that: The storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the live audio data cleaning method as described in any one of claims 1 to 7.

10. A computer device, characterized in that: include: one or more processors, and memory; The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the method for cleaning live audio data as described in any one of claims 1 to 7 are performed.