Audio processing method and device

By identifying the song information sung by the singer and inputting it along with reference audio data into the noise reduction model, the problem of poor audio data quality in complex acoustic scenarios is solved, achieving a higher purity audio processing effect.

CN121260171APending Publication Date: 2026-01-02VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511346421.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

In complex acoustic scenarios such as concerts, existing technologies struggle to accurately identify and effectively separate the main melody, vocals, or key instrument tracks. Traditional recording equipment, lacking the capability to separate specific acoustic elements, results in the loss of effective signals or noise residue in the audio data, leading to poor output audio data quality.

Method used

By collecting mixed audio data of the singer's performance, identifying song information, and inputting this data along with reference audio data into a noise reduction model for processing, the model analyzes the mixed audio data to perform noise reduction and outputs the first audio data. Through data analysis and comparison of the characteristic differences between the mixed audio data and the reference audio data, the noise reduction model can more accurately locate and suppress noise components while preserving the singer's clear voice.

Benefits of technology

It effectively improves audio quality, with the output first audio data having higher purity, significantly reducing noise interference in mixed audio and improving the overall audio quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121260171A_ABST
    Figure CN121260171A_ABST
Patent Text Reader

Abstract

The invention discloses an audio processing method and device, and belongs to the technical field of data processing. The method comprises the following steps: collecting mixed audio data sung by a singer; identifying first song information sung by the singer according to the mixed audio data; and inputting the mixed audio data and reference audio data corresponding to the first song information into a noise reduction model, performing noise reduction processing on the mixed audio data, and outputting first audio data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of data processing, and particularly relates to an audio processing method and device. BACKGROUND

[0002] With the continuous evolution of intelligent wearable devices, intelligent glasses have shown significant advantages in the field of shooting creation due to their highly integrated optical display and camera acquisition system, especially providing new possibilities for users to record high-quality content in complex acoustic scenes such as concerts. Although users can conveniently capture live performances with the help of intelligent glasses, they generally face the challenge of excessive environmental noise interference.

[0003] There are a large number of acoustic interferences in the concert environment, including high-frequency shouting of the audience, background diffuse noise, and reverberation superposition of non-target musical instruments. Traditional recording devices are difficult to accurately identify the song being performed and effectively separate the main melody, vocals or key instrument track due to the lack of analysis and enhancement capabilities for specific sound sources. Existing noise reduction algorithms cannot adapt to the spectral structure, instrument composition and dynamic range changes of different songs.

[0004] Therefore, there is a problem of loss of effective signals or residual noise in the output audio data, resulting in poor quality of the output audio data. SUMMARY

[0005] The purpose of the embodiments of the application is to provide an audio processing method and device, which can solve the problem of poor quality of audio data.

[0006] In a first aspect, the embodiments of the application provide an audio processing method, which comprises:

[0007] acquiring mixed audio data sung by a singer;

[0008] identifying first song information sung by the singer according to the mixed audio data;

[0009] inputting the mixed audio data and reference audio data corresponding to the first song information into a noise reduction model, performing noise reduction processing on the mixed audio data, and outputting first audio data.

[0010] In a second aspect, the embodiments of the application provide an audio processing device, which comprises:

[0011] a collection module configured to acquire mixed audio data sung by a singer;

[0012] an identification module configured to identify first song information sung by the singer according to the mixed audio data;

[0013] The input module is configured to input the mixed audio data and reference audio data corresponding to the first song information into a noise reduction model, perform noise reduction processing on the mixed audio data, and output first audio data.

[0014] In a third aspect, an electronic device is provided. The electronic device includes a processor and a memory. The memory stores programs or instructions executable on the processor. The programs or instructions, when executed by the processor, implement the steps of the method of the first aspect.

[0015] In a fourth aspect, a readable storage medium is provided. The readable storage medium stores programs or instructions. The programs or instructions, when executed by a processor, implement the steps of the method of the first aspect.

[0016] In a fifth aspect, a chip is provided. The chip includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is configured to execute programs or instructions to implement the method of the first aspect.

[0017] In a sixth aspect, a computer program product is provided. The program product is stored in a storage medium. The program product is executed by at least one processor to implement the method of the first aspect.

[0018] In the embodiments of the present application, by collecting mixed audio data sung by a singer, the first song information sung by the singer is identified according to the mixed audio data. After determining the first song information by analyzing the features in the mixed audio data, the reference audio data corresponding to the first song information is determined, which provides a clear reference standard for the noise reduction process. The mixed audio data and the reference audio data corresponding to the first song information are input into a noise reduction model. The reference audio data can help the noise reduction model to distinguish the effective signal in the mixed audio data from the noise. The effective signal includes the singer's voice and the corresponding accompaniment, and the noise includes environmental noise. The noise reduction model can more accurately locate and suppress the noise component by comparing the feature differences between the mixed audio data and the reference audio data, while retaining the clear singer's voice. The final output of the first audio data has higher purity, which can effectively improve the audio quality. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 is a flowchart of an audio processing method provided by the embodiments of the present application;

[0020] Figure 2 is a schematic diagram of another audio processing method provided by the embodiments of the present application;

[0021] Figure 3 is a flowchart of another audio processing method provided by the embodiments of the present application;

[0022] Figure 4 is a structural diagram of an audio processing device provided by an embodiment of the present application;

[0023] Figure 5 is one of hardware structure schematic diagrams of an electronic device of an embodiment of the present application;

[0024] Figure 6 is another hardware structure schematic diagram of an electronic device of an embodiment of the present application. DETAILED DESCRIPTION

[0025] The technical solutions of the embodiments of the present application will be described clearly below in conjunction with the drawings of the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0026] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", and the like are generally of a kind and are not limited to the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims indicates at least one of the connected objects, and the character " / ", generally indicates that the objects before and after are in an "or" relationship.

[0027] To solve the problems in the related art, the embodiments of the present application provide an audio processing method and device, which can solve the problem of poor quality of audio data in the related art.

[0028] The audio processing method provided by the embodiments of the present application will be described in detail below in conjunction with the drawings and specific embodiments and application scenarios.

[0029] Figure 1 is a flowchart of an audio processing method provided by an embodiment of the present application.

[0030] As shown in Figure 1 , the audio processing method can include steps 110-130, and the method is applied to an audio processing device, as shown below:

[0031] Step 110, collecting mixed audio data sung by a singer;

[0032] Mixed audio data refers to the audio information collected containing various sound components such as the singer's singing voice, environmental noise, and background music;

[0033] The singing process of the singer is captured by a microphone or other sound pickup device, and the sound wave vibration in the air is converted into an electrical signal, and then an analog-to-digital conversion is performed to form digital mixed audio data. The core of this step is to record the sound information during singing, whether it contains noise or not, to provide original materials for subsequent processing.

[0034] Step 120, according to the mixed audio data, identifying the first song information sung by the singer;

[0035] The first song information refers to the specific identification of the song sung by the singer identified from the mixed audio, including the song name, original singer information, song duration, etc.

[0036] The mixed audio data is analyzed to extract information related to song characteristics such as melody, rhythm, and lyrics fragments. These characteristics are compared and matched with the information in the preset song database to determine the first song information sung by the singer, providing a key reference for subsequent noise reduction processing.

[0037] Step 130, inputting the mixed audio data and the reference audio data corresponding to the first song information into the noise reduction model to perform noise reduction processing on the mixed audio data, and outputting the first audio data.

[0038] The reference audio data is the standard audio corresponding to the first song information, usually the original version of the song accompaniment or the original singing audio without noise; the noise reduction model is an algorithm model trained through machine learning, which can distinguish the singer's singing voice and noise in the audio; the first audio data is the audio result after noise reduction processing, which retains the clear singer's voice.

[0039] The mixed audio data and the corresponding reference audio data are input into the noise reduction model, which compares the differences between the two. The reference audio data serves as a template to help the model identify the noise components in the mixed audio that do not belong to the target song, and through algorithms to filter and suppress these noises, finally output the first audio data. By utilizing the relevance between the reference audio and the mixed audio, the model can more accurately distinguish between noise and target sound, thereby effectively improving the noise reduction effect.

[0040] Thus, by identifying the first song information sung by the singer according to the mixed audio data, the noise reduction process can be targeted to use the reference audio, by inputting the mixed audio data and the reference audio data corresponding to the first song information into the noise reduction model, the mixed audio data is processed to output the first audio data, which significantly reduces the noise interference in the mixed audio, makes the singer's voice clearer, and improves the overall quality of the audio.

[0041] In a possible embodiment, the step 120 can specifically include the following steps:

[0042] Step 210, identifying human voice audio data and non-human voice audio data in the mixed audio data;

[0043] Step 220, identifying the human voice audio data to obtain the first lyric text;

[0044] Step 230, determining the first song information from the index library according to the first lyric text, the index library including: a plurality of corresponding lyric texts and song information.

[0045] The human voice audio data is the sound signal produced by the human singing in the mixed audio, which contains the singing information such as lyrics and melody; the non-human voice audio data refers to other sounds in the mixed audio except the human voice, such as background noise and musical instrument sound; the first lyric text is the text content identified from the human voice audio; the index library is a database storing a plurality of lyric texts and corresponding song information, which is used for fast matching query.

[0046] By using the feature difference of the audio signal, the human voice and the non-human voice are separated by signal processing technology. The human voice usually has a specific frequency range and spectral characteristics, while the non-human voice such as noise and musical instrument sound has different spectral distribution. By analyzing these feature differences, the mixed audio data is separated into human voice audio data and non-human voice audio data, which can exclude the interference of non-human voice and focus on the human voice part containing the key information of the song, providing a clearer audio basis for subsequent lyric recognition.

[0047] Based on the speech recognition technology, the separated human voice audio data is processed to convert the audio signal into a corresponding text sequence. By using the trained speech recognition model, the lyrics contained in the human voice are identified to form the first lyric text, which converts the singing information in the form of audio into a text form that can be directly processed, providing conditions for subsequent matching with the index library.

[0048] The obtained first song lyrics text is compared with the song lyrics texts stored in the index library by using a text matching technology, the most matched song lyrics text group is found, and the corresponding first song information is determined. The existence of the index library makes the matching process efficient and accurate, accurately locates the song sung by the singer through text matching, realizes the conversion from song lyrics to specific song information, and provides accurate reference basis for subsequent noise reduction processing.

[0049] Exemplarily, in a concert scene, after the concert starts, the collection module of the intelligent glasses records the live sound in real time to obtain mixed audio data containing singer voice, main instrument sound, audience cheers, non-main direction drum sound and other components. At the same time, the system relies on a pre-constructed song library quick matching system to identify the first song information. The construction process of the song library first selects a candidate song set from the song library according to the concert information, which is equivalent to narrowing the search range of the index library and reducing the complexity of subsequent matching. Then, the lyrics sentences of each song in the candidate song set are simplified and an n-gram set is constructed, which is similar to establishing structured features for the song lyrics texts in the index library, facilitating subsequent fast similarity calculation with the first song lyrics text extracted from the mixed audio.

[0050] The human voice audio data and non-human voice audio data, such as audience cheers and non-main direction drum sound, are separated from the mixed audio data. Then, the first song lyrics text is obtained from the human voice audio data through voice recognition, and then the first song lyrics text is compared with the processed lyrics features of the candidate songs in the song library, i.e., the n-gram set, to calculate the similarity score, and finally determine the first song information being sung.

[0051] After the user selects to start recording, the signal is recorded using the microphone sensor of the intelligent glasses, and the signal is recorded as x(t). The voice signal is separated and then multiple feature matching is performed. The singer sound is separated using a discriminative network, which is split into human voice audio data x singer (t) and non-human voice audio data x other (t):

[0052] x(t) = x singer (t) + x other (t)

[0053] In one possible embodiment, step 230 can specifically include the following steps:

[0054] Similarity calculation is performed on the first song lyrics text and the song lyrics text of each song in the index library to obtain multiple similarity scores;

[0055] According to the multiple similarity scores, the first song information is determined from the index library.

[0056] The similarity calculation refers to a process of quantitatively evaluating the similarity between two texts by a specific algorithm, and the result is usually presented in the form of a score, with a higher score indicating that the text content is closer; the similarity score is the quantitative result of similarity calculation, which is used to intuitively reflect the matching degree of the first lyric text and the lyric texts of each song in the index library.

[0057] Based on the feature extraction and comparison of the text, the first lyric text and the lyric texts of each song in the index library are decomposed into smaller language units, and the similarity between the two texts is calculated by counting the common occurrence frequency and arrangement order of each language unit in the two texts, combined with the preset algorithm, to obtain multiple corresponding similarity scores.

[0058] According to the multiple similarity scores, the first song information is determined from the index library, that is, the song information corresponding to the lyric text with the highest score is selected as the result, because the highest similarity score means that the content overlap of the two lyric texts is the highest, and the corresponding song is the most likely song sung by the singer.

[0059] Therefore, through similarity calculation, accurate comparison between the first lyric text and the lyrics in the index library is realized, avoiding direct matching failure caused by possible lyric omission and pronunciation deviation during singing. The way of determining the song information based on the highest similarity score improves the accuracy of recognition and ensures that the first song information most matching the song sung by the singer can be quickly found from the index library, providing a reliable foundation for subsequent audio processing.

[0060] Among the above steps of determining the first song information from the index library according to the multiple similarity scores, the steps can specifically include the following steps:

[0061] At least one song information with a similarity score exceeding a preset threshold is determined as candidate song information;

[0062] The first song information is determined from the candidate song information according to the non-human voice audio data.

[0063] The preset threshold refers to a reference value preset in the similarity judgment, which is used to filter out song information similar enough to the first lyric text; the candidate song information refers to multiple song information with a similarity score meeting the standard after being filtered by the preset threshold; and the non-human voice audio data refers to the sound components other than human voice in the mixed audio, such as song accompaniment and instrument sound.

[0064] The songs with low similarity are filtered out by a preset threshold, only those songs with lyrics text similar enough to the first lyrics text are reserved as objects for further screening, the preset threshold can be 0.9 or 0.8, and the setting of the preset threshold needs to balance the screening efficiency and accuracy, avoiding retaining too many low-similarity songs to increase the subsequent processing burden, and preventing missing possible correct songs.

[0065] The audio features such as accompaniment and instruments that may be contained in the non-human audio data are utilized, the audio features in the non-human audio data have consistency with the original accompaniment features of the corresponding song, the original accompaniment audio corresponding to the candidate song information is compared with the non-human audio data, the candidate song with the most matched features is found, and thus the first song information is determined.

[0066] The singer's human voice audio data x singer (t) applying speech recognition technology to transcribe lyrics text, fuzzy matching with n-gram constructed by the song library lyrics and calculating similarity, at least one song information with a similarity score exceeding a preset threshold is set as candidate song information.

[0067] Commonly used text rough matching algorithms include Dice coefficient algorithm, Jaccard similarity algorithm and Levenshtein distance algorithm. The calculation process of Levenshtein distance needs to use dynamic programming and matrix. Compared with Levenshtein distance, fuzzy matching algorithm applied to embedded end uses Dice coefficient and Jaccard similarity, which consumes less computing resources. Dice coefficient is more suitable for short text fuzzy matching than Jaccard similarity in concert scenarios, and is more matched with lyrics search, and has better recall rate in fault-tolerant recognition tasks. The calculation formula of Dice coefficient is as follows:

[0068]

[0069] Wherein, A and B are character sets, for example, when divided by 2-gram, "hello"→{"he", "el", "ll", "lo"}. Meanwhile, the non-human audio data x other (t) is separated again a (t) extracts audio fingerprint F x , and matches with the alternative audio segment, finally determines whether it matches and the current performance progress.

[0070] Thus, by determining at least one song information with a plurality of similarity scores exceeding a preset threshold as candidate song information, the filtering based on the preset threshold effectively narrows down the candidate range, improving the efficiency of subsequent processing; in combination with secondary screening of non-human voice data, the uniqueness of non-human voice features such as accompaniment is utilized, further improving the accuracy of the first song information recognition, ensuring that the finally determined song information is completely matched with the song sung by the singer, and providing accurate reference basis for subsequent noise reduction processing.

[0071] In a possible embodiment, before step 130, the following steps can also be included:

[0072] Obtaining a plurality of sets of training data, the training data including: label audio data corresponding to sample song information and sample mixed audio data, the sample mixed audio data being audio data obtained by adding noise audio data to the label audio data;

[0073] Inputting the label audio data and the sample mixed audio data into an initial noise reduction model to output sample audio data after noise reduction;

[0074] Determining a loss value according to the label audio data and the sample audio data after noise reduction;

[0075] Training the initial noise reduction model according to the loss value until a training stop condition is met to obtain a noise reduction model.

[0076] The label audio data is noise-free or standard audio corresponding to the sample song information, serving as a reference standard for model training; the sample mixed audio data is audio formed by adding noise audio to the label audio data, simulating the actual collected noise-containing audio; the initial noise reduction model is a basic model architecture without training, having basic audio processing capability; the loss value is a quantitative index for measuring the difference between the model output result and the label data, and the greater the difference, the higher the loss value; the training stop condition is a preset model training termination standard, such as the loss value reaching a preset range or the training times reaching an upper limit.

[0077] Obtaining a plurality of sets of training data, by adding noise audio of different types and intensities to the label audio data, sample mixed audio close to real scenes is constructed, providing rich training materials for the model, so that the model can learn processing methods in various noise environments. Inputting the label audio data and the sample mixed audio data into the initial noise reduction model, the model will try to separate noise from the sample mixed audio, outputting sample audio data after noise reduction, in which the model learns noise reduction rules through internal parameter adjustment.

[0078] The loss value is calculated according to the label audio data and the denoised sample audio data, and the loss value reflects the gap between the current denoising effect of the model and the ideal denoising effect, providing a basis for model parameter optimization. Finally, the parameters of the initial denoising model are adjusted according to the loss value, and iterative training is continuously performed until the loss value is reduced to a preset range or other training stopping conditions are met. At this time, the model can better complete the denoising task and obtain a usable denoising model.

[0079] The trained denoising model can be used for denoising processing of mixed audio data collected on site. The on-site audio x(t) is collected, that is, the mixed audio data sung by the singer is collected, and x(t) contains the singing of the singer, accompaniment, on-site background noise, crowd noise, sound reverberation and other interference. Then, the current song is identified as the first song information ck through rapid matching of the song library and other means, and the corresponding reference audio data r(t) is obtained.

[0080] Then, a small network model is deployed, and the reference audio data r(t) and the noisy mixed audio data x(t) are used as model inputs to train the model using supervised learning. The noise-free "pure audio" y(t) is used as the label audio data, and the model learns how to suppress background noise and other interference components from x(t) with r(t) as the target source, and finally outputs the extracted "pure audio" y(t). That is, the mixed audio data and the reference audio data corresponding to the first song information are input into the denoising model, the mixed audio data is denoised, and the first audio data is output, achieving the goal of maximizing the preservation of the real voices of the singer and the accompaniment and suppressing interference.

[0081] Thus, by constructing diverse training data, the model can adapt to different noise scenes and enhance its generalization ability; the iterative training process based on the loss value continuously optimizes the denoising ability of the model and gradually improves the processing accuracy; and the final denoising model can effectively identify and remove noise in the mixed audio, providing high-quality audio data for subsequent processing and ensuring stable output of clear target audio in actual application.

[0082] The label audio data includes at least one of the following:

[0083] The audio data obtained by randomly performing frequency domain masking on the label audio data, and the audio data obtained by randomly performing time domain masking on the label audio data;

[0084] The frequency domain masking is to set the energy of the label audio data in a specific frequency band to zero, and the time domain masking is to set the energy of the label audio data in a specific time period to zero.

[0085] Frequency masking refers to processing in the frequency dimension of audio, setting the energy of a specific frequency band in the tagged audio data to zero, which is equivalent to "covering" the sound signal in that frequency range. Time masking refers to processing in the time dimension of audio, setting the energy of a specific time period in the tagged audio data to zero, which is equivalent to "covering" the sound signal in that time range. The audio data obtained after these two processes still belongs to the category of tagged audio data, and is used to enrich the diversity of training data.

[0086] By randomly performing frequency masking or time masking on the tagged audio data, audio samples with partial frequency loss or partial time segment loss are artificially created. Frequency masking simulates specific frequency noise interference that may occur in actual scenarios, or attenuation of specific frequency bands during sound propagation. Time masking simulates temporary interruptions and sudden noise coverages during sound transmission. The audio data obtained after masking is included in the tagged audio data, and is used in combination with sample mixed audio data for model training, so that the initial noise reduction model can be exposed to more diverse training data.

[0087] To improve the generalization ability of the model, especially when facing different live acoustic conditions and singer performance styles, the algorithm needs to introduce a "reference track condition masking" mechanism. In the training phase, r(t) is randomly masked in the frequency band, time period, or rhythm level, forcing the model to learn more robust feature mapping relationships and avoiding overfitting caused by simply memorizing the reference track content. Given a reference track r(t), one or a combination of the following two perturbation methods is randomly applied in the training phase:

[0088] Frequency masking (Frequency Masking), randomly selecting a frequency range [f1, f2] of the frequency spectrum and setting it to zero, denoted as:

[0089]

[0090] Time masking (Time Masking), randomly masking the reference track in the time domain [t1, t2]:

[0091]

[0092] The loss function can be a standard signal-to-noise ratio loss:

[0093]

[0094] f θ The noise reduction model obtained by the above method is:

[0095] y(t)=f θ (x(t),r(t))

[0096] Thus, the richness of the label audio data is increased by random frequency domain and time domain covering, so that the model not only learns the characteristics of standard noise-free audio during the training process, but also adapts to various locally missing audio characteristics, enhancing the tolerance and recovery ability of the model to different types of audio damage. When the model faces mixed audio that may have frequency missing or time segment missing in actual application, it can more accurately identify and retain effective sound components, improve the robustness of the noise reduction model, and ensure that high-quality first audio data can be output even in complex sound environments.

[0097] In a possible embodiment, after step 130, the following steps can also be included:

[0098] Removing the first audio data from the mixed audio data to obtain noise audio data;

[0099] Determining non-human audio data in the mixed audio data;

[0100] Removing the noise audio data from the non-human audio data to obtain accompaniment audio data.

[0101] The noise audio data refers to the sound components remaining after the first audio data is removed from the mixed audio data, mainly including environmental noise and other interference sounds; the non-human audio data is all sounds in the mixed audio except human voice, covering accompaniment, instrument sound and noise; and the accompaniment audio data is obtained by removing the noise audio data from the non-human audio data, mainly being the accompaniment or instrument sound of the song.

[0102] By utilizing the characteristics of the mixed audio containing human voice and noise, the noise part not belonging to the target human voice is separated through subtraction operation, because the first audio data is the clear human voice remaining after noise reduction processing, corresponding to the original human voice component in the mixed audio, and the remaining part after removal is noise. The non-human audio data in the mixed audio data is determined, which is based on the human audio data separated in the previous step, and is achieved by subtracting the human audio data from the mixed audio, thereby obtaining the non-human audio part containing accompaniment and noise. Finally, the accompaniment audio data is obtained by removing the noise audio data from the non-human audio data, and the noise audio data obtained is used as a reference to remove this part of noise from the non-human audio data, and the remaining part is the original accompaniment or instrument sound of the song.

[0103] Thus, by separating the noise audio data, the noise component in the mixed audio can be clearly determined, providing materials for subsequent noise analysis or model optimization; determining the non-human audio data lays the foundation for extracting the accompaniment; and the final accompaniment audio data can be used for re-synthesis with the first audio data to form a complete singing audio with better sound quality, or can be used alone as accompaniment material, improving the integrity and practicality of audio processing and meeting the diversified needs of users for clear human voice and pure accompaniment.

[0104] In the embodiments of the present application, by collecting mixed audio data sung by a singer, first song information sung by the singer is identified according to the mixed audio data, and after determining the first song information by analyzing the characteristics in the mixed audio data, reference audio data corresponding to the first song information is determined, so that the noise reduction process has a clear reference standard. The mixed audio data and the reference audio data corresponding to the first song information are input into a noise reduction model, and the reference audio data can help the noise reduction model to distinguish between effective signals and noise in the mixed audio data, such as the singer's voice and the corresponding accompaniment, and noise such as environmental noise. The noise reduction model can more accurately locate and suppress noise components while retaining clear singer's voice, and finally output first audio data with higher purity, which can effectively improve the audio quality.

[0105] The following will be described in combination with Figure 3 An audio processing process provided by the embodiments of the present application is described:

[0106] 301, collecting mixed audio data sung by a singer:

[0107] Using an audio collection device, real-time capture of mixed sound signals containing human voice, accompaniment, environmental noise, etc. during the singer's singing process is converted into processable digital audio data.

[0108] 302, matching lyrics text obtained based on human voice audio data in the mixed audio data with an index library:

[0109] First, the human voice audio is extracted from the mixed audio by a human voice separation technology, and then the lyrics text is generated by voice recognition; the text is compared with the pre-constructed song library index to find the matching song record.

[0110] 303, whether the matching is successful:

[0111] The result of the text matching in step 302 is determined, if the corresponding song information is found, the matching is successful, and the next step is entered; if not found, the matching fails, and the subsequent judgment or retry is entered according to the process.

[0112] 304, matching based on audio features of the mixed audio data and the index library:

[0113] The audio features of the mixed audio are extracted; the features are compared in detail with the audio features of the corresponding candidate songs in the index library to further confirm the song matching condition.

[0114] 305, whether the matching is successful:

[0115] The result of the audio feature matching in the checking step 304 is checked. If the matching is successful, the song is determined, and the reference audio related process is entered. If the matching fails, it is determined whether to retry or switch strategies according to the process.

[0116] 306, determining reference audio data:

[0117] When the text and audio feature matching are both successful, the pure audio data of the matched song is called from an index library or a related data source, as a reference basis for subsequent noise reduction.

[0118] 307, reducing noise in mixed audio data based on reference audio data:

[0119] The mixed audio data and the reference audio data are input into a noise reduction model. The pure features of the reference audio are used to distinguish the target sound and the interference noise in the mixed audio. The noise is suppressed through model operation, and the clear audio after noise reduction is output.

[0120] 308, the number of matching failures exceeds a preset number threshold:

[0121] The number of failures in the text matching, audio feature matching and other links is counted and compared with a preset number threshold, such as 3 times, 5 times, etc. It is determined whether to switch to a general noise reduction alternative scheme due to multiple unsuccessful matching.

[0122] 309, starting general noise reduction:

[0123] If the number of matching failures exceeds the limit, a general noise reduction algorithm that does not depend on specific reference audio is enabled to perform basic noise reduction processing on the mixed audio, and the audio quality is improved as much as possible.

[0124] 310, whether the shooting is ended:

[0125] It is determined whether the audio acquisition process is terminated. If it is terminated, the process is ended. If it is not terminated, the acquisition, matching, noise reduction and other operations are continuously executed in a loop to continuously process the singing audio.

[0126] The audio processing method provided in the embodiments of the present application can be executed by an audio processing device. In the embodiments of the present application, the audio processing device is taken as an example to illustrate the audio processing device provided in the embodiments of the present application.

[0127] Figure 4 is a block diagram of an audio processing device provided in the embodiments of the present application. The device 400 includes:

[0128] The acquisition module 410 is configured to acquire mixed audio data sung by a singer.

[0129] The identification module 420 is configured to identify first song information sung by the singer according to the mixed audio data.

[0130] The input module 430 is configured to input the mixed audio data and reference audio data corresponding to the first song information into a noise reduction model, perform noise reduction processing on the mixed audio data, and output first audio data.

[0131] In a possible implementation, the identification module 420 is specifically configured to:

[0132] identify human voice audio data and non-human voice audio data in the mixed audio data;

[0133] identify the human voice audio data to obtain first lyric text;

[0134] determine the first song information from an index library according to the first lyric text, the index library including a plurality of groups of corresponding lyric text and song information.

[0135] In a possible implementation, the identification module 420 is specifically configured to:

[0136] perform similarity calculation on the first lyric text and lyric text of each song in the index library to obtain a plurality of similarity scores;

[0137] determine the first song information from the index library according to the plurality of similarity scores.

[0138] In a possible implementation, the identification module 420 is specifically configured to:

[0139] determine at least one song information with a similarity score higher than a preset threshold as candidate song information from the plurality of similarity scores;

[0140] determine the first song information from the candidate song information according to the non-human voice audio data.

[0141] In a possible implementation, the apparatus 400 can further include:

[0142] The obtaining module is configured to obtain a plurality of groups of training data, the training data including label audio data corresponding to sample song information and sample mixed audio data, the sample mixed audio data being audio data obtained by adding noise audio data to the label audio data;

[0143] The input module 430 is further configured to input the label audio data and the sample mixed audio data into an initial noise reduction model, and output sample audio data after noise reduction.

[0144] The loss value determination module is configured to determine a loss value according to the label audio data and the sample audio data after noise reduction.

[0145] The training module is configured to train the initial noise reduction model according to the loss value until a training stop condition is met, and obtain the noise reduction model.

[0146] In a possible embodiment, the label audio data comprises at least one of the following:

[0147] audio data obtained by randomly performing frequency domain masking on the label audio data, and audio data obtained by randomly performing time domain masking on the label audio data;

[0148] The frequency domain masking is to set energy of the label audio data in a specific frequency band to zero, and the time domain masking is to set energy of the label audio data in a specific time period to zero.

[0149] In a possible embodiment, the apparatus 400 can further comprise:

[0150] The removing module is configured to remove the first audio data from the mixed audio data, and obtain noise audio data.

[0151] The determining module is configured to determine non-human voice audio data in the mixed audio data.

[0152] The removing module is further configured to remove noise audio data from the non-human voice audio data, and obtain accompaniment audio data.

[0153] In the embodiments of the present application, by collecting mixed audio data sung by a singer, first song information sung by the singer is identified according to the mixed audio data, and after the first song information is determined by analyzing the features in the mixed audio data, reference audio data corresponding to the first song information is determined, so that the noise reduction process has a clear reference standard. The mixed audio data and the reference audio data corresponding to the first song information are input into the noise reduction model, the reference audio data can help the noise reduction model to distinguish effective signals and noise in the mixed audio data, the effective signals are, for example, the singer's voice and the corresponding accompaniment, and the noise is, for example, environmental noise, the noise reduction model can more accurately locate and suppress the noise component by comparing the feature differences between the mixed audio data and the reference audio data, while retaining the clear singer's voice, and finally the first audio data has higher purity and can effectively improve the audio quality.

[0154] The audio processing apparatus in the embodiments of the present application can be an electronic device, or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or other devices than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The electronic device can also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a cash register, or a self-service machine, etc. The embodiments of the present application are not limited in this regard.

[0155] The audio processing apparatus in the embodiments of the present application can be a device with a motion system. The motion system can be an Android motion system, an iOS motion system, or other possible motion systems. The embodiments of the present application are not limited in this regard.

[0156] The audio processing apparatus provided in the embodiments of the present application can implement each process implemented by the method embodiments. To avoid repetition, details are not described herein.

[0157] Optionally, as shown in Figure 5 The embodiments of the present application also provide an electronic device 510, which includes a processor 511, a memory 512, and a program or instruction stored in the memory 512 and executable in the processor 511. When the program or instruction is executed by the processor 511, each step of any of the above audio processing methods is implemented, and the same technical effects are achieved. To avoid repetition, details are not described herein.

[0158] It should be noted that the electronic device in the embodiments of the present application includes the above mobile electronic device and non-mobile electronic device.

[0159] Figure 6 To implement the hardware structure of an electronic device in the embodiments of the present application.

[0160] The electronic device 600 includes, but is not limited to, a radio frequency unit 601, a network module 602, an audio output unit 603, an input unit 604, a sensor 605, a display unit 606, a user input unit 607, an interface unit 608, a memory 609, and a processor 610, etc.

[0161] Those skilled in the art can understand that the electronic device 600 can further include a power supply (such as a battery) for supplying power to each component, and the power supply can be logically connected to the processor 610 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. Figure 6 The electronic device structure shown in the figure does not constitute a limitation on the electronic device, and the electronic device can include more or fewer components than the figure, or combine certain components, or different component arrangements, which are not described here.

[0162] The processor 610 is configured to collect mixed audio data sung by a singer.

[0163] The processor 610 is further configured to identify first song information sung by the singer according to the mixed audio data.

[0164] The processor 610 is further configured to input the mixed audio data and reference audio data corresponding to the first song information into a noise reduction model, perform noise reduction processing on the mixed audio data, and output first audio data.

[0165] Optionally, the processor 610 is further configured to identify human voice audio data and non-human voice audio data in the mixed audio data.

[0166] The processor 610 is further configured to identify the human voice audio data to obtain first lyric text.

[0167] The processor 610 is further configured to determine the first song information from an index library according to the first lyric text, the index library including a plurality of sets of corresponding lyric text and song information.

[0168] Optionally, the processor 610 is further configured to perform similarity calculation on the first lyric text and lyric text of each song in the index library to obtain a plurality of similarity scores.

[0169] The processor 610 is further configured to determine the first song information from the index library according to a plurality of the similarity scores.

[0170] Optionally, the processor 610 is further configured to determine at least one song information with a plurality of the similarity scores exceeding a preset threshold as candidate song information.

[0171] The processor 610 is further configured to determine the first song information from the candidate song information according to the non-human voice audio data.

[0172] Optionally, the processor 610 is further configured to obtain a plurality of sets of training data, the training data including label audio data corresponding to sample song information and sample mixed audio data, the sample mixed audio data being audio data obtained by adding noise audio data to the label audio data.

[0173] The processor 610 is further configured to input the label audio data and the sample mixed audio data into an initial noise reduction model, and output sample audio data after noise reduction.

[0174] The processor 610 is further configured to determine a loss value according to the label audio data and the sample audio data after noise reduction.

[0175] The processor 610 is further configured to train the initial noise reduction model according to the loss value until a training stop condition is met, and obtain the noise reduction model.

[0176] Optionally, the label audio data includes at least one of the following:

[0177] audio data obtained by randomly performing frequency domain covering on the label audio data, and audio data obtained by randomly performing time domain covering on the label audio data.

[0178] The frequency domain covering is to set energy of the label audio data on a specific frequency band to zero, and the time domain covering is to set energy of the label audio data in a specific time period to zero.

[0179] Optionally, the processor 610 is further configured to remove the first audio data from the mixed audio data to obtain noise audio data.

[0180] The processor 610 is further configured to determine non-human voice audio data in the mixed audio data.

[0181] The processor 610 is further configured to remove noise audio data from the non-human voice audio data to obtain accompaniment audio data.

[0182] In the embodiments of the present application, by collecting mixed audio data sung by the singer, the first song information sung by the singer is identified according to the mixed audio data, and after determining the first song information by analyzing the characteristics in the mixed audio data, the reference audio data corresponding to the first song information is determined, so that the noise reduction process has a clear reference standard. The mixed audio data and the reference audio data corresponding to the first song information are input into the noise reduction model, and the reference audio data can help the noise reduction model to distinguish the effective signal in the mixed audio data from the noise, such as the singer's voice and the corresponding accompaniment, and the noise such as the environmental noise. The noise reduction model can more accurately locate and suppress the noise component by comparing the feature differences between the mixed audio data and the reference audio data, while retaining the clear singer's voice. The first audio data output finally has higher purity, which can effectively improve the audio quality.

[0183] It should be understood that in the embodiments of the present application, the input unit 604 can include a graphics processor (GPU) 6041 and a microphone 6042. The graphics processor 6041 processes image data of a still picture or a video image obtained by an image capture device (such as a camera) in a video image capture mode or an image capture mode. The display unit 606 can include a display panel 6061, which can be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 607 includes at least one of a touch panel 6071 and other input devices 6072. The touch panel 6071 is also called a touch screen. The touch panel 6071 can include a touch detection device and a touch controller. The other input devices 6072 can include, but are not limited to, a physical keyboard, function keys (such as volume control keys, on-off keys, etc.), trackballs, mice, motion sticks, etc., which will not be described here. The memory 609 can be used to store software programs and various data, including but not limited to application programs and action systems. The processor 610 can integrate an application processor and a modem processor, wherein the application processor mainly processes action systems, user pages and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 610.

[0184] The memory 609 can be used to store software programs and various data. The memory 609 can mainly include a first storage area storing programs or instructions and a second storage area storing data, wherein the first storage area can store an operating system, application programs or instructions required by at least one function (such as a sound playing function, an image playing function, etc.), and the like. In addition, the memory 609 can include a volatile memory or a non-volatile memory, or the memory 609 can include both a volatile memory and a non-volatile memory. The non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 609 in the embodiments of the present application includes but is not limited to these and any other suitable types of memories.

[0185] The processor 610 can include one or more processing units; optionally, the processor 610 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and an application program, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 610.

[0186] The embodiments of the present application also provide a readable storage medium, the readable storage medium stores programs or instructions, the programs or instructions are executed by a processor to realize various processes of the above-mentioned audio processing method embodiments, and the same technical effects can be achieved. To avoid repetition, details are not described here.

[0187] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes a computer readable storage medium, such as a computer readable only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0188] The embodiment of the present application further provides a chip, which comprises a processor and a communication interface, the communication interface is coupled with the processor, the processor is used for running programs or instructions to realize the processes of the above audio processing method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.

[0189] It should be understood that the chip mentioned in the embodiment of the present application can also be referred to as a system level chip, a system chip, a chip system or a system on chip, etc.

[0190] The embodiment of the present application provides a computer program product, which is stored in a storage medium, and is executed by at least one processor to realize the processes of the above audio processing method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.

[0191] It should be noted that in this document, the term "comprising" or "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the method and device in the embodiment of the present application is not limited to the order of performing the functions as shown or discussed, but can also include performing the functions in a substantially simultaneous manner or in the opposite order, for example, the described method can be performed in an order different from that described, and various steps can also be added, omitted or combined. In addition, the features described with reference to some examples can be combined in other examples.

[0192] Through the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned example methods can be realized by means of software and a necessary general hardware platform, and of course, can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a computer software product in essence or in the form of a part that contributes to the prior art, which is stored in a storage medium (such as a ROM / RAM, a magnetic disk, or an optical disk) and includes a plurality of instructions for causing a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application.

[0193] The embodiments of the present application are described above in combination with the drawings, but the present application is not limited to the above-mentioned specific embodiments, and the above-mentioned specific embodiments are only illustrative and not restrictive. Those skilled in the art can make many forms under the inspiration of the present application without departing from the scope of the present application and the scope protected by the claims.

Claims

1. An audio processing method, characterized in that, The method includes: Collect mixed audio data of the singer's performance; Based on the mixed audio data, identify the information of the first song performed by the singer; The mixed audio data and the reference audio data corresponding to the first song information are input into the noise reduction model to perform noise reduction processing on the mixed audio data and output the first audio data.

2. The method according to claim 1, characterized in that, The step of identifying the first song information sung by the singer based on the mixed audio data includes: Identify human voice audio data and non-human voice audio data in the mixed audio data; The human voice audio data is identified to obtain the first lyrics text; Based on the first lyrics text, the first song information is determined from the index library, which includes multiple sets of corresponding lyrics text and song information.

3. The method according to claim 2, characterized in that, The step of determining the first song information from the index based on the lyrics text includes: The similarity between the first lyrics text and the lyrics text of each song in the index is calculated to obtain multiple similarity scores; The first song information is determined from the index based on multiple similarity scores.

4. The method according to claim 3, characterized in that, The step of determining the first song information from the index based on multiple similarity scores includes: At least one song with a similarity score exceeding a preset threshold is identified as a candidate song. The first song information is determined from the candidate song information based on the non-human voice audio data.

5. The method according to claim 1, characterized in that, Before inputting the mixed audio data and the reference audio data corresponding to the first song information into the noise reduction model, performing noise reduction processing on the mixed audio data, and outputting the first audio data, the method further includes: Multiple sets of training data are acquired, including: labeled audio data corresponding to sample song information and sample mixed audio data, wherein the sample mixed audio data is audio data obtained by adding noise audio data to the labeled audio data; The labeled audio data and the sample mixed audio data are input into the initial noise reduction model, and the noise-reduced sample audio data is output. The loss value is determined based on the tagged audio data and the noise-reduced sample audio data; Based on the loss value, the initial denoising model is trained until the training stopping condition is met, thus obtaining the denoising model.

6. The method according to claim 5, characterized in that, The tagged audio data includes at least one of the following: Audio data obtained by randomly applying frequency domain masking to the tag audio data, and audio data obtained by randomly applying time domain masking to the tag audio data; The frequency domain masking is to set the energy of the tag audio data to zero in a specific frequency band, and the time domain masking is to set the energy of the tag audio data to zero in a specific time period.

7. The method according to claim 1, characterized in that, After inputting the mixed audio data and the reference audio data corresponding to the first song information into the noise reduction model, performing noise reduction processing on the mixed audio data, and outputting the first audio data, the method further includes: Remove the first audio data from the mixed audio data to obtain noisy audio data; Identify the non-human voice audio data in the mixed audio data; Noise audio data is removed from the non-human voice audio data to obtain accompaniment audio data.

8. An audio processing apparatus, characterized in that, The device includes: The acquisition module is used to acquire mixed audio data from the singer's performance; The recognition module is used to identify the first song information sung by the singer based on the mixed audio data; The input module is used to input the mixed audio data and the reference audio data corresponding to the first song information into the noise reduction model, perform noise reduction processing on the mixed audio data, and output the first audio data.

9. The apparatus according to claim 8, characterized in that, The identification module is specifically used for: Identify human voice audio data and non-human voice audio data in the mixed audio data; The human voice audio data is identified to obtain the first lyrics text; Based on the first lyrics text, the first song information is determined from the index library, which includes multiple sets of corresponding lyrics text and song information.

10. The apparatus according to claim 9, characterized in that, The identification module is specifically used for: The similarity between the first lyrics text and the lyrics text of each song in the index is calculated to obtain multiple similarity scores; The first song information is determined from the index based on multiple similarity scores.