Audio noise reduction method, device and computer readable storage medium

By performing speech recognition on audio data, determining the first human voice category and scene label, and generating target noise reduction parameters, the problem of fixed noise reduction parameters being unable to match different audio files is solved, achieving efficient noise reduction processing of audio streams and improving sound quality.

CN113823303BActive Publication Date: 2026-05-19TENCENT TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECH (BEIJING) CO LTD
Filing Date
2021-06-11
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

In existing technologies, fixed noise reduction parameters cannot be matched to different audio files, resulting in unsatisfactory noise reduction effects.

Method used

By performing speech recognition on the audio data, the first human voice category and scene label are determined, and target noise reduction parameters are generated for the current audio data, and the noise reduction processing is adjusted in real time.

Benefits of technology

It enables automatic matching of target noise reduction parameters based on real-time changes in audio data, improving noise reduction effect and efficiency, and enhancing the sound quality of the audio stream.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113823303B_ABST
    Figure CN113823303B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides an audio noise reduction method and device and a computer readable storage medium, and relates to the technical field of speech processing. The method comprises the following steps: obtaining current audio data and a preset scene label at a current moment from an audio stream; performing speech recognition on the current audio data to determine a first human voice category of the current audio data; generating a target noise reduction parameter for the current audio data based on the first human voice category and the scene label; and performing noise reduction processing on the audio stream based on the target noise reduction parameter. The embodiment of the application matches the corresponding target noise reduction parameter by performing speech recognition on the current audio data and combining the scene label, so that the technical effect of improving sound quality is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing technology, and more specifically, to an audio noise reduction method, apparatus, and computer-readable storage medium. Background Technology

[0002] In the digital age, once sound is recorded, whether it's speaking, singing, musical instruments, or even noise, it can be processed by digital music software. In order to pursue excellent sound quality, people often need to further process the audio files to reduce noise interference from external noise.

[0003] In existing technologies, fixed noise reduction parameters are typically configured on the server side. For example, in live streaming scenarios, noise reduction is added during recording to eliminate background noise introduced during the broadcast to improve the broadcaster's voice quality. However, fixed noise reduction parameters cannot be matched to different audio files, resulting in unsatisfactory noise reduction effects. Summary of the Invention

[0004] This application provides an audio noise reduction method, apparatus, and computer-readable storage medium to solve the technical problem of unsatisfactory noise reduction effect.

[0005] Firstly, an audio noise reduction method is provided, the method comprising:

[0006] Retrieve the current audio data and preset scene tags from the audio stream;

[0007] Perform speech recognition on the current audio data to determine the first human voice category of the current audio data;

[0008] Based on the first human voice category and scene label, generate target noise reduction parameters for the current audio data;

[0009] Noise reduction processing is performed on the audio stream based on the target noise reduction parameters.

[0010] In one possible implementation, target noise reduction parameters for the current audio data are generated based on the first voice category and scene label, including:

[0011] Obtain the second voice category corresponding to the audio data from the previous moment;

[0012] If the first voice category does not match the second voice category, then target noise reduction parameters for the current audio data are generated based on the first voice category and the scene label.

[0013] In one possible implementation, speech recognition is performed on the current audio data to determine the first human voice category of the current audio data, including:

[0014] Perform speech detection on the current audio data and extract at least one human voice segment;

[0015] Obtain the audio features of each human voice segment;

[0016] The first voice category of the current audio data is determined based on audio features; wherein, the first voice category includes speaking voices and singing voices.

[0017] In another possible implementation, based on the first voice category and scene label, target noise reduction parameters for the current audio data are generated, including:

[0018] The noise reduction parameters corresponding to the first voice category and the scene noise reduction parameters corresponding to the scene label are weighted to obtain the target noise reduction parameters. Among them, the noise reduction parameters corresponding to the first voice category being speech are greater than the noise reduction parameters corresponding to the first voice category being singing.

[0019] In another possible implementation, based on the first voice category and scene label, target noise reduction parameters for the current audio data are generated, including:

[0020] Determine the acquisition path for audio data;

[0021] Based on the first human voice category, scene label, and acquisition path, target noise reduction parameters are generated for the current audio data.

[0022] In another possible implementation, based on the first voice category, scene label, and acquisition path, target noise reduction parameters for the current audio data are generated, including:

[0023] If no matching noise reduction attribute exists in the acquisition path, then target noise reduction parameters for the current audio data are generated based on the first human voice category, scene label, and acquisition path.

[0024] In another possible implementation, based on the first voice category, scene label, and acquisition path, target noise reduction parameters for the current audio data are generated, including:

[0025] The target noise reduction parameters are obtained by weighting the human voice noise reduction parameters corresponding to the first human voice category, the scene noise reduction parameters corresponding to the scene label, and the path noise reduction parameters corresponding to the acquisition path.

[0026] In another possible implementation, the voice noise reduction parameters corresponding to the first voice category, the scene noise reduction parameters corresponding to the scene label, and the path noise reduction parameters corresponding to the acquisition path are weighted to obtain the target noise reduction parameters, including:

[0027] Determine the first weight of the human voice noise reduction parameters, the second weight of the scene noise reduction parameters, and the third weight of the path noise reduction parameters;

[0028] Based on the first weight, the second weight, and the third weight, the voice noise reduction parameters, scene noise reduction parameters, and path noise reduction parameters are weighted and summed to obtain the target noise reduction parameters; where the first weight is greater than either the second weight or the third weight.

[0029] Secondly, an audio noise reduction device is provided, the device comprising:

[0030] The acquisition module is used to obtain the current audio data and preset scene tags from the audio stream at the current moment;

[0031] The recognition module is used to perform speech recognition on the current audio data and determine the first human voice category of the current audio data.

[0032] The generation module is used to generate target noise reduction parameters for the current audio data based on the first human voice category and scene label;

[0033] The noise reduction module is used to perform noise reduction processing on the audio stream based on the target noise reduction parameters.

[0034] In one possible implementation, the above-mentioned generation module is specifically used for:

[0035] Obtain the second voice category corresponding to the audio data from the previous moment;

[0036] If the first voice category does not match the second voice category, then target noise reduction parameters for the current audio data are generated based on the first voice category and the scene label.

[0037] In one possible implementation, the aforementioned identification module is specifically used for:

[0038] Perform speech detection on the current audio data and extract at least one human voice segment;

[0039] Obtain the audio features of each human voice segment;

[0040] The first voice category of the current audio data is determined based on audio features, where the first voice category includes speaking voices and singing voices.

[0041] In another possible implementation, the above-mentioned generation module is specifically used for:

[0042] The noise reduction parameters corresponding to the first voice category and the scene noise reduction parameters corresponding to the scene label are weighted to obtain the target noise reduction parameters. Among them, the noise reduction parameters corresponding to the first voice category being speech are greater than the noise reduction parameters corresponding to the first voice category being singing.

[0043] In another possible implementation, the above-mentioned generation module specifically includes:

[0044] The determining unit is used to determine the acquisition path for audio data.

[0045] The generation unit is used to generate target noise reduction parameters for the current audio data based on the first human voice category, scene label, and acquisition path.

[0046] In yet another possible implementation, the aforementioned generation unit is specifically used for:

[0047] If no matching noise reduction attribute exists in the acquisition path, then target noise reduction parameters for the current audio data are generated based on the first human voice category, scene label, and acquisition path.

[0048] In yet another possible implementation, the aforementioned generating unit is further used for:

[0049] The target noise reduction parameters are obtained by weighting the human voice noise reduction parameters corresponding to the first human voice category, the scene noise reduction parameters corresponding to the scene label, and the path noise reduction parameters corresponding to the acquisition path.

[0050] In yet another possible implementation, the aforementioned generating unit is further used for:

[0051] Determine the first weight of the human voice noise reduction parameters, the second weight of the scene noise reduction parameters, and the third weight of the path noise reduction parameters;

[0052] Based on the first weight, the second weight, and the third weight, the voice noise reduction parameters, scene noise reduction parameters, and path noise reduction parameters are weighted and summed to obtain the target noise reduction parameters; where the first weight is greater than either the second weight or the third weight.

[0053] Thirdly, an electronic device is provided, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the audio noise reduction method shown in the first aspect of this application.

[0054] Fourthly, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the audio noise reduction method shown in the first aspect of this application.

[0055] Fifthly, embodiments of this application provide a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to implement the methods provided in the first or second aspect embodiments.

[0056] The beneficial effects of the technical solution provided in this application are:

[0057] This application performs speech recognition on the current audio data and combines it with preset scene labels to determine the target noise reduction parameters for the current audio data, thereby achieving effective noise reduction processing of the audio stream. Compared with the existing technology that uses server-side configuration of fixed noise reduction parameters for noise reduction, the noise reduction scheme of this application can configure the target noise reduction parameters in real time based on the current audio data. When the voice category of the audio stream changes, even if the scene label does not change in time, the target noise reduction parameters can still match the speech recognition result of the current audio data in real time, thereby effectively improving the noise reduction effect of the audio stream and improving the noise reduction efficiency. Attached Figure Description

[0058] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.

[0059] Figure 1a This is an application scenario diagram of an audio noise reduction method provided in an embodiment of this application;

[0060] Figure 1b This diagram illustrates another application scenario of an audio noise reduction method provided in this application embodiment.

[0061] Figure 2 A flowchart illustrating an audio noise reduction method provided in an embodiment of this application;

[0062] Figure 3 A configuration diagram of a live streaming start page provided in an embodiment of this application;

[0063] Figure 4 A flowchart illustrating a speech recognition scheme provided in an embodiment of this application;

[0064] Figure 5 This application provides a target noise reduction parameter configuration table for scene labels.

[0065] Figure 6 A target noise reduction parameter configuration table for a first human voice category is provided in the embodiments of this application;

[0066] Figure 7 This application provides a target noise reduction parameter configuration table for the acquisition path;

[0067] Figure 8 A flowchart illustrating an audio noise reduction method in an example provided in this application embodiment;

[0068] Figure 9A flowchart illustrating a live streaming process in an example provided for an embodiment of this application;

[0069] Figure 10 This is a schematic diagram of the structure of an audio noise reduction device provided in an embodiment of this application;

[0070] Figure 11 This is a schematic diagram of the structure of an electronic device for audio noise reduction provided in an embodiment of this application. Detailed Implementation

[0071] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting the invention.

[0072] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.

[0073] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0074] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0075] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0076] Key technologies in speech technology include Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Voiceprint Recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech emerging as one of the most promising methods.

[0077] The audio noise reduction method provided in this application can match the noise reduction requirements of the current audio data in real time, thereby effectively improving the sound quality of the audio stream.

[0078] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close relationship with linguistic research. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.

[0079] Blockchain is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and cryptographic algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying platform, a platform product and service layer, and an application service layer.

[0080] The underlying blockchain platform can include modules for user management, basic services, smart contracts, and operational monitoring. The user management module is responsible for managing the identity information of all blockchain participants, including maintaining public and private key generation (account management), key management, and maintaining the correspondence between user identities and blockchain addresses (access management). Under authorization, it also monitors and audits transactions of certain real identities and provides risk control rule configuration (risk control audit). The basic services module is deployed on all blockchain node devices to verify the validity of business requests. After consensus is reached on valid requests, they are recorded in storage. For a new business request, the basic services first perform interface adaptation parsing and authentication (interface adaptation), and then encrypt the business information using a consensus algorithm (consensus management). The encryption process involves transmitting the encrypted data to the shared ledger (network communication) and storing it in a consistent manner. The smart contract module is responsible for contract registration, issuance, triggering, and execution. Developers can define contract logic using a programming language and publish it to the blockchain (contract registration). Based on the contract terms, the module calls keys or other events to trigger execution and complete the contract logic. It also provides functions for contract upgrades and cancellations. The operation and monitoring module is mainly responsible for deployment, configuration modification, contract settings, cloud adaptation, and real-time status visualization during product launch, such as alarms, network status monitoring, and node device health status monitoring.

[0081] In the audio noise reduction scheme provided in this application embodiment, the scene label of the audio stream, the first human voice category of the audio data, and the corresponding target noise reduction parameters can be stored in the blockchain. When the server or terminal performing audio noise reduction performs audio noise reduction, it can first query whether there are target noise reduction parameters in the blockchain that correspond to the current scene label and the first human voice category, and obtain the target noise reduction parameters from the blockchain in order to perform noise reduction processing on the audio data.

[0082] The solutions provided in this application relate to speech processing technology in natural language processing, which will be specifically illustrated through the following embodiments.

[0083] In the digital age, with the continuous advancement of voice processing technology, recording technology plays an irreplaceable role in promoting social development. For example, the use of smartphones and the development of the film and music industries all require the support and guarantee of recording technology. Once sound is recorded, whether it's speech, singing, musical instruments, or even noise, it can be processed through digital music software. However, in pursuit of excellent sound quality, people often need to further process audio files to reduce noise and minimize interference from external noise.

[0084] In existing technologies, noise reduction parameters are typically configured on the server side or the user manually adjusts the noise reduction level. For example, during live streaming, noise reduction is added during recording to eliminate background noise introduced during the voice capture process, in order to improve the broadcaster's voice quality. In chat scenarios, fixed noise reduction configurations can be used, but the requirements for these configurations differ for singing, dancing, and outdoor lighting conditions. For instance, when singing, noise reduction parameters need to be lowered or even turned off; when dancing, noise reduction reduces the effect of dance music; and outdoors, background wind noise and other environmental noise can affect the clarity of speech.

[0085] Generally, noise reduction parameters are configured on the server side or the broadcaster manually adjusts the noise reduction level. This noise reduction method has the following drawbacks: when the audio data of the audio stream changes in real time, the noise reduction requirements of the audio stream are different, and a fixed configuration cannot match the real-time changing audio stream; when the live broadcast scene changes, it is inconvenient for the broadcaster to adjust the scene label in time, and the noise reduction parameters cannot be automatically configured according to the scene changes.

[0086] The audio noise reduction method provided in this application enables real-time configuration of target noise reduction parameters based on current audio data, which can meet the noise reduction requirements of real-time changing audio streams. Compared with existing technologies, it improves the noise reduction effect and efficiency of audio streams, achieves the goal of improving sound quality, and effectively enhances the user experience.

[0087] The audio noise reduction method, apparatus, and computer-readable storage medium provided in this application are intended to solve the above-mentioned technical problems of the prior art.

[0088] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0089] like Figure 1a As shown, the audio noise reduction method of this application can be applied to... Figure 1a In the scenario shown, specifically, after the server 102 obtains the audio stream 101 to be processed, it performs speech recognition on the current audio data 101 in the audio stream and obtains the preset scene label. Based on the first human voice category and scene label obtained by recognition, it determines the target noise reduction parameters for the current audio data, and then performs noise reduction processing on the audio stream 101 according to the above target noise reduction parameters to obtain the noise-reduced audio stream 103.

[0090] like Figure 1b As shown, the audio noise reduction method of this application can also be applied to Figure 1bIn the scenario shown, specifically, terminal 104 can acquire audio stream 101 and send it to the server. Server 102 identifies the current audio data in the audio stream and then determines the target noise reduction parameters based on the identified first human voice category and preset scene labels. The target noise reduction parameters are then sent to the noise reduction APP (Application) on terminal 104 for processing. In other scenarios, the terminal can also acquire audio data, determine the target noise reduction parameters, and process the audio based on those parameters.

[0091] Those skilled in the art will understand that the “terminal” used herein can be a mobile phone, tablet computer, PDA (Personal Digital Assistant), MID (Mobile Internet Device), etc.; and the “server” can be implemented using a standalone server or a server cluster composed of multiple servers.

[0092] This application provides an audio noise reduction method, such as... Figure 2 As shown, this method can be applied to Figure 1a and Figure 1b The server shown can also be applied to a terminal, and may include the following steps:

[0093] S201, retrieve the current audio data and preset scene labels from the audio stream.

[0094] Among them, the server or terminal used for audio noise reduction processing can use the audio data acquired in real time from the audio stream as the current audio data. For example, it can record the audio stream signal through an audio acquisition device, such as a microphone, to obtain the current audio data. Alternatively, it can process the existing audio stream to obtain the current audio data, such as using audio processing software to capture and extract sound from the audio stream, extracting sound from the audio stream in the video, or extracting a segment of sound from the audio stream as the current audio data.

[0095] Specifically, preset scene tags can be obtained from the terminal or server; the scene tags can represent the scene or activity category of the signal source of the audio stream, or the scene or activity category corresponding to the content of the audio stream; the scene tags can be preset by the user, or determined by the system based on the user's historical usage data.

[0096] Taking live streaming as an example, the audio data during a live stream can be categorized into scenarios such as meetings, lectures, singing, chatting, outdoor activities, and dancing. Figure 3The above is a configuration diagram of the live streaming start page. T1 is the avatar or cover area, displayed as a recommendation page; T2 is the title area, used to display the user-set live streaming title; T3 is the tag area, where tags are differentiated without specific scene categories, such as meeting, singing, lecturing, chatting, outdoor activities, dancing, dating, etc.; T4 is the function area, where the streamer can set the front and rear camera settings and select appropriate beauty filters, etc., when starting the live stream. Users can set the required scene tags in the tag area according to the actual live streaming content and scene.

[0097] S202, Perform speech recognition on the current audio data to determine the first human voice category of the current audio data.

[0098] The first voice category can be a category of audio features in the audio data, such as voice categories including speaking voices and singing voices.

[0099] In some implementations, the server or terminal used for audio noise reduction processing can first identify the text content of the current audio data, and then classify the text content to obtain the first human voice category; specifically, speech recognition can be performed on the audio data based on a speech recognition network to obtain the text data corresponding to the content of the audio data, thereby realizing the conversion of audio to text; then, the text data can be classified based on a pre-trained classification network to determine the first human voice category of the current audio data.

[0100] In other embodiments, the server or terminal used for audio noise reduction processing can extract audio features from the current audio data and then determine the first human voice category of the current audio data by performing digital signal processing on the audio features. The specific recognition process based on audio features will be described in detail below.

[0101] S203 generates target noise reduction parameters for the current audio data based on the first human voice category and scene label.

[0102] The target noise reduction parameter can range from 0 to 1, and it essentially represents the percentage of the total noise reduction capability of the noise reduction software or device used. For example, when the audio stream data has high requirements for noise reduction processing, the target noise reduction parameter can range from 0.6 to 1, meaning that for that audio stream, 60% to 100% of the total noise reduction capability of the noise reduction software or device is used for noise reduction processing.

[0103] In some implementations, the server or terminal used for audio noise reduction processing can be pre-set with the correspondence between different classification results, different scene labels and different noise reduction parameters. The noise reduction parameters for the audio data can be directly queried based on the first human voice category and scene label of the audio data. Alternatively, the functional relationship between the classification results, scene labels and noise reduction parameters can be set, and the noise reduction parameters can be calculated based on the first human voice category and scene label.

[0104] In other implementations, the classification results and scene categories can be combined with other parameters, such as the noise reduction parameters corresponding to the audio data acquisition path, to determine the noise reduction parameters for the current audio data. The specific process of determining the noise reduction parameters will be described in detail below.

[0105] S204 performs noise reduction processing on the audio stream based on the target noise reduction parameters.

[0106] Specifically, the server or terminal used for audio noise reduction processing can configure preset noise reduction software based on target noise reduction parameters, and then use the configured noise reduction software to perform noise reduction processing on the audio stream. Before performing noise reduction processing on the audio stream, it is also necessary to acquire noisy audio, and then configure the noise reduction software based on the noisy audio to accurately identify the noise in order to perform noise reduction processing on the audio stream.

[0107] In one implementation, taking the application scenario of game audio acquisition as an example, in general large-scale games, due to the large number of background sounds and sound effects, and the need for multiple players to maintain real-time voice communication, it is necessary to perform real-time noise reduction according to different game scenarios in order to ensure both the game effect and the voice communication effect of the players. Before audio acquisition, the user pre-sets a scene label such as fighting, music, shooting, etc. according to the needs of the game scene. Then, based on the current game audio data, speech recognition is performed to obtain the first human voice category of the game audio data, such as speaking and singing. Then, target noise reduction parameters are generated based on the scene label and the first human voice category to perform noise reduction processing on the game audio stream.

[0108] This application performs speech recognition on the current audio data and combines it with preset scene labels to determine the target noise reduction parameters for the current audio data, thereby achieving effective noise reduction processing of the audio stream. Compared with the existing technology that uses server-side configuration of fixed noise reduction parameters for noise reduction, the noise reduction scheme of this application can configure the target noise reduction parameters in real time based on the current audio data. When the voice category of the audio stream changes, even if the scene label does not change in time, the target noise reduction parameters can still match the speech recognition result of the current audio data in real time, thereby effectively improving the noise reduction effect of the audio stream and improving the noise reduction efficiency.

[0109] This application embodiment provides a possible implementation method in which the above-mentioned step S202, performing speech recognition on the current audio data to determine the first human voice category of the current audio data, may include:

[0110] (1) Perform speech detection on the current audio data and extract at least one human voice segment.

[0111] Specifically, the fluctuation of the temporal signal of the current audio data can be detected based on the VAD (Voice Activity Detection) algorithm, thereby identifying the human voice and non-human voice parts in the current audio data, and extracting human voice segments.

[0112] The human voice segment includes both voiced and unvoiced sounds. In phonetics, sounds produced by vocal cord vibration are called voiced sounds, while sounds produced without vocal cord vibration are called unvoiced sounds. Short-time energy is more suitable for detecting voiced sounds, while short-time zero-crossing rate is more suitable for detecting unvoiced sounds. The VAD algorithm employs a dual-threshold endpoint detection method, which combines short-time zero-crossing rate and short-time energy as judgment indicators. The short-time average zero-crossing rate refers to the number of times a frame of speech signal crosses the horizontal axis (zero level); short-time energy is the energy of a frame of speech signal, with the energy in the human voice segment typically being lower than that in the non-human voice segment, and the energy in the unvoiced segment being lower than that in the voiced segment.

[0113] The specific VAD detection steps are as follows: First, the current speech data is divided into frames. Then, the short-time energy and short-time zero-crossing rate of each frame of speech signal are calculated. Next, the start frame and end frame of the human voice segment are determined based on the preset upper and lower thresholds of short-time energy or short-time zero-crossing rate, thereby realizing the extraction of the human voice segment.

[0114] (2) Obtain the audio features of each human voice segment.

[0115] Specifically, feature extraction can be performed on each voice segment to obtain a speech feature sequence that changes over time, i.e., audio features. These audio features can be acoustic features such as LPC (Linear Predictive Coding), MFCC (MeI-Freguency CeptraI Coefficients), or CEP (Cepstrum).

[0116] Specifically, taking MFCC feature extraction as an example: the pre-processed human voice segment after VAD is pre-emphasized through a high-pass filter; then the pre-emphasized audio is divided into frames with a time length of 20ms, and each frame signal is windowed to reduce the spectral leakage of the audio signal; then the audio signal is subjected to discrete Fourier transform to obtain the frequency domain signal, and filtered through a Mel scale filter bank to obtain the Mel spectrum; finally, cepstral analysis is performed on the Mel spectrum to obtain the MFCC coefficients.

[0117] Audio features can represent the characteristic information of human voice segments from multiple dimensions. Identifying the human voice category based on audio features can effectively improve the accuracy of identification.

[0118] (3) Determine the first voice category of the current audio data based on audio features, wherein the first voice category includes speaking voice and singing voice.

[0119] In some implementations, digital signal processing can be used to classify audio features to determine the first human voice category.

[0120] The main classification method in digital signal processing described above involves determining the frequency of the fundamental frequency change. By acquiring the fundamental frequency change rate of audio features, it's possible to determine whether a vocal segment belongs to singing or speaking. In musical tones, each note in singing remains constant, while the fundamental frequency of speaking constantly changes. Specific classification methods include: using Matlab (a high-tech computing environment primarily for scientific computing, visualization, and interactive programming) and the findpeaks function to calculate the fundamental frequency value of each audio feature, and then detecting fundamental frequency changes within a preset time period, such as one second. If the fundamental frequency changes from 200Hz to 400Hz and then back to 200Hz within one second, a total of 400 changes, then the vocal segment corresponding to this audio feature can be identified as speaking. If the fundamental frequency changes around 300Hz within one second, with a difference of less than 10Hz, then the vocal segment corresponding to this audio feature can be identified as singing.

[0121] In other implementations, the first voice category may be determined by classifying the text content of the audio data as determined by a speech recognition network.

[0122] Specifically, such as Figure 4 As shown, the process of obtaining the first voice category through a speech recognition network includes:

[0123] First, based on the trained acoustic model, the probability of each frame's audio features being generated by each phoneme in the preset training set is calculated, thereby determining the phoneme sequence with the highest probability, realizing the conversion of audio features into phoneme sequences. The acoustic model can be GMM (Gaussian Mixture Model) or HMM (Hidden Markov Model), etc.

[0124] Then, based on the trained language model, the text data is determined to maximize the probability of the phoneme sequence being converted into the text data, thus realizing the conversion from phoneme sequence to text data. The language model can be used to calculate the probability that the phoneme sequence forms each complete text; the language model can be a statistical N-gram (N-gram language model), a neural network language model, or a model based on the Transformer (data converter) architecture.

[0125] Finally, the text data is classified based on a pre-trained classification network to obtain the first human voice category. The classification network can be based on a random forest model, an SVM (Support Vector Machine) classification model, or a neural network classification model. Among them, the neural network classification model can be a text classification network such as TextCNN (Text Convolutional Neural Networks) or LSTM (Long Short-Term Memory).

[0126] The embodiments of this application obtain the first human voice category through digital signal processing or speech recognition network, which improves the accuracy of human voice recognition in audio data and provides a reliable guarantee for subsequent noise reduction schemes to meet the noise reduction requirements of real-time changing audio stream signals.

[0127] This application embodiment provides a possible implementation method. The step S203 above, which generates target noise reduction parameters for the current audio data based on the first human voice category and scene label, may include:

[0128] (1) Obtain the second voice category corresponding to the audio data of the previous moment.

[0129] Specifically, the second voice category corresponding to the audio data from the previous moment can be retrieved from the database used to store the classification results. The method for identifying the second voice category is the same as that for identifying the first voice category.

[0130] (2) If the first voice category does not match the second voice category, then the target noise reduction parameters for the current audio data are generated based on the first voice category and the scene label.

[0131] Specifically, when the first voice category matches the second voice category, it means that the scene and activity category or voice category corresponding to the current audio data is the same as the scene and activity category or voice category corresponding to the previous audio data. In this case, the target noise reduction parameter is the noise reduction parameter corresponding to the audio data at the previous moment.

[0132] If the first voice category does not match the second voice category, it means that the voice category corresponding to the current audio data is different from the voice category corresponding to the previous audio data. In this case, the target noise reduction parameter is a new noise reduction parameter generated in real time based on the first voice category and the scene label.

[0133] Based on the changing audio stream signal, the embodiments of this application generate target noise reduction parameters for the current audio data in real time according to the first human voice category and scene label, which can effectively improve the noise reduction effect of the audio stream.

[0134] This application embodiment provides another possible implementation method. The step S203 above, which generates target noise reduction parameters for the current audio data based on the first human voice category and scene label, may include:

[0135] The noise reduction parameters corresponding to the first voice category and the scene noise reduction parameters corresponding to the scene label are weighted to obtain the target noise reduction parameters. Among them, the noise reduction parameters corresponding to the first voice category being speech are greater than the noise reduction parameters corresponding to the first voice category being singing.

[0136] Specifically, in some implementations, the voice noise reduction parameters and scene noise reduction parameters can be determined based on the pre-set correspondence between the first voice category and the voice noise reduction parameters, and the correspondence between the scene label and the scene noise reduction parameters.

[0137] Specifically, taking the online live streaming scenario as an example, pre-setting the correspondence between different scene labels and the range of scene noise reduction parameters can be done as follows: Figure 5 As shown, when the scene label is "meeting," "outdoor," or "dating," these scene labels have high requirements for background noise reduction, and their corresponding noise reduction weight is set to "high." When the scene is "singing," "music," or "dancing," these scenes have high requirements for sound fidelity and reproduction, but low requirements for background noise reduction, and their corresponding noise reduction weight is set to "low." When the scene is "lecturing" or "chatting," these scenes need to consider both sound fidelity and reproduction and background noise reduction, and their corresponding noise reduction weight is set to "medium." The different levels of noise reduction weights correspond to different ranges of scene parameter values.

[0138] Pre-setting the comparison relationship between different classification results and the range of human voice noise reduction parameters can be as follows: Figure 6As shown, when the first voice category is speaking, noise other than speaking needs to be reduced, and the corresponding noise reduction weight is set to high; when the first voice category is singing, the corresponding noise reduction weight is set to low because background music needs to be preserved.

[0139] In other implementations, a functional relationship can be established between the first voice category, voice noise reduction parameters, scene label, and scene noise reduction parameters, and the voice noise reduction parameters and scene noise reduction parameters can be calculated based on this functional relationship.

[0140] When weighting the human voice noise reduction parameters corresponding to the first human voice category and the scene noise reduction parameters corresponding to the scene label, the human voice noise reduction parameters have a greater weight than the scene noise reduction parameters because the results are generated based on real-time speech recognition. This avoids the problem of mismatch between the target noise reduction parameters and the current audio data when the scene or activity corresponding to the audio stream changes and the user cannot change the scene label settings in time, thus improving the noise reduction effect and efficiency.

[0141] This application embodiment provides another possible implementation method. The step S203 above, which generates target noise reduction parameters for the current audio data based on the first human voice category and scene label, may include:

[0142] (1) Determine the acquisition path for audio data.

[0143] Specifically, audio data acquisition can be the recording process of recording software on a computer or mobile terminal through a microphone. The acquisition path of the audio data can be determined from the acquisition interface of the terminal. The acquisition path can be a headphone microphone (MIC), a mobile phone microphone, or a sound card. Different acquisition paths will also affect the noise reduction requirements of the audio data. For example, the ambient noise of the audio data acquired by the mobile phone microphone is relatively high, requiring greater noise reduction, while the ambient noise of the audio data acquired by the headphone microphone is relatively low, requiring less noise reduction.

[0144] (2) Based on the first human voice category, scene label and acquisition path, generate target noise reduction parameters for the current audio data.

[0145] In some implementations, a functional relationship can be set between the first voice category, scene label, acquisition path, and target noise reduction parameters, and then the target noise reduction parameters for the current audio data can be calculated based on the first voice category, scene label, and acquisition path of the audio data.

[0146] In other implementations, different correspondences between different voice categories, different scene labels, different acquisition paths, and different target noise reduction parameters can be preset, and then the target noise reduction parameters for the audio data can be determined based on the above correspondences.

[0147] Specifically, taking a live streaming scenario as an example, the relationship between different audio acquisition paths and the value ranges of path noise reduction parameters is pre-set, such as... Figure 7 As shown, when the mobile phone microphone is used as the acquisition path, the background noise is relatively large, so its corresponding noise reduction weight is set to high; when the headphone microphone is used as the acquisition path, the background noise and echo problems are relatively small, so its corresponding noise reduction weight is set to low; when the sound card and Bluetooth headphones are used as the acquisition path, the background noise is moderate, so its corresponding noise reduction weight is medium; different levels of noise reduction weight correspond to different ranges of path parameter values.

[0148] This application embodiment comprehensively considers the impact of audio data classification results, scene labels, and acquisition paths on the required noise reduction intensity, so that the generated target noise reduction parameters are further matched with the audio data. It can not only meet personalized scene labels but also adapt to multiple audio acquisition paths and match the audio data classification results in real time, effectively improving the noise reduction effect and achieving the goal of improving sound quality.

[0149] This application provides yet another possible implementation, in which the above-mentioned generation of target noise reduction parameters for the current audio data based on the first human voice category, scene label, and acquisition path may include:

[0150] If no matching noise reduction attribute exists in the acquisition path, then target noise reduction parameters for the current audio data are generated based on the first human voice category, scene label, and acquisition path.

[0151] Specifically, the server or terminal used for audio noise reduction processing first detects the current acquisition path. When it detects that the acquisition path has matching noise reduction attributes, such as the acquisition path being a headphone microphone with noise reduction function, it generates target noise reduction parameters for the current audio data based on the first human voice category and scene label.

[0152] When it is detected that there is no matching noise reduction attribute in the acquisition path, the target noise reduction parameters for the current audio data are generated based on the first human voice category, scene label and acquisition path.

[0153] The audio noise reduction method provided in this invention embodiment, compared with the prior art, does not require any additional interactive operations for audio noise reduction parameter settings. The noise reduction settings and noise reduction processing are transparent and imperceptible to the user, effectively improving the user experience.

[0154] This application provides another possible implementation, in which the target noise reduction parameters for the current audio data are generated based on the first voice category, scene label, and acquisition path, including:

[0155] The target noise reduction parameters are obtained by weighting the human voice noise reduction parameters corresponding to the first human voice category, the scene noise reduction parameters corresponding to the scene label, and the path noise reduction parameters corresponding to the acquisition path.

[0156] Specifically, the weights of human voice noise reduction parameters, scene noise reduction parameters, and path noise reduction parameters can be determined by category. Based on these weights, the human voice noise reduction parameters, scene noise reduction parameters, and path noise reduction parameters are weighted and summed to obtain the target noise reduction parameters.

[0157] In determining the noise reduction parameters, this invention comprehensively considers three factors: the scene label set by the user, the first voice category identified by speech recognition, and the audio data acquisition path. By taking into account the user's subjective judgment and combining the real-time voice classification results and acquisition path of the audio data, the sound quality of the audio data is further improved, effectively enhancing the user experience.

[0158] This application provides yet another possible implementation, in which the above-mentioned weighted processing of the human voice noise reduction parameters corresponding to the first human voice category, the scene noise reduction parameters corresponding to the scene label, and the path noise reduction parameters corresponding to the acquisition path is performed to obtain the target noise reduction parameters, which may include:

[0159] (1) Determine the first weight of the human voice noise reduction parameter, the second weight of the scene noise reduction parameter, and the third weight of the path noise reduction parameter.

[0160] The first, second, and third weights can be obtained from the terminal or server, or they can be generated based on data statistics calculations according to actual engineering applications.

[0161] (2) Based on the first weight, the second weight and the third weight, the human voice noise reduction parameter, the scene noise reduction parameter and the path noise reduction parameter are weighted and summed to obtain the target noise reduction parameter; wherein the first weight is greater than either the second weight or the third weight.

[0162] Specifically, the first weight can be multiplied by the human voice noise reduction parameter, the second weight by the scene noise reduction parameter, and the third weight by the path noise reduction parameter. The three products are then added together to obtain the target noise reduction parameter.

[0163] The sum of the first, second, and third weights is 1. Since the human voice noise reduction parameters are determined based on real-time speech recognition, the first weight has the largest value among the three weights. When affected by external factors, such as changes in the scene or activity corresponding to the audio stream, when users cannot change the scene label settings in time, or when the noise reduction function of the acquisition channel malfunctions, the human voice noise reduction parameters generated in real time based on speech recognition can have the largest weight. This avoids the problem of mismatch between the target noise reduction parameters and the current audio data caused by the above-mentioned external factors, and further improves the noise reduction effect.

[0164] This embodiment uses live streaming as an example for specific explanation. The audio noise reduction parameters in the live streaming scenario are determined based on three parts: the speech analysis unit, the tag configuration analysis unit, and the noise reduction unit.

[0165] The speech analysis unit determines the voice noise reduction parameters and the path noise reduction parameters based on the first identified human voice category and the acquisition path, respectively.

[0166] The tag configuration analysis unit determines the scene noise reduction parameters based on the scene tags set by the user;

[0167] The noise reduction unit generates target noise reduction parameters based on human voice noise reduction parameters, scene noise reduction parameters, and path noise reduction parameters.

[0168] Different classification results, scene labels, or acquisition paths have different noise reduction weights. Based on these different noise reduction weights, the value ranges of human voice noise reduction parameters, scene noise reduction parameters, and path noise reduction parameters can be determined. Then, based on these value ranges, the actual values ​​of human voice noise reduction parameters, scene noise reduction parameters, and path noise reduction parameters can be determined according to the actual engineering application.

[0169] Specifically, taking the application of live streaming as an example, when the broadcast starts, the host will set a scene tag on the configuration page. When there is a scene change during the broadcast, such as the host moving from indoor to outdoor broadcast, the host may not change the scene tag setting in time. In order to reduce the error caused by human setting, the second weight of the scene noise reduction parameter can be set to 10%. Secondly, since the audio acquisition path depends on the host's usage habits, it generally will not change much, and the existing audio acquisition technology is relatively mature, so it has little impact on the overall noise reduction parameter. The third weight of the path parameter can be set to 20%. Thus, the first weight of the human voice noise reduction parameter determined by real-time audio stream speech recognition is set to 70%. Therefore, the final target noise reduction parameter r of the system can be calculated by the following formula (1):

[0170] (1)

[0171] In the formula, r is the target noise reduction parameter. These are the noise reduction parameters for the scene. Parameters for human voice noise reduction These are the noise reduction parameters for the path.

[0172] The audio target noise reduction parameters provided in this application embodiment are calculated and determined based on multiple parameters, combining user-preset scene tags of the current audio data, speech recognition results, and acquisition path conditions. This can improve the problem that fixed parameter configurations cannot adapt to diverse voice activities and scenarios, and does not require manual adjustment by the user. It effectively improves sound quality while enhancing the user experience. At the same time, the calculation and generation of target noise reduction parameters do not require high system processing performance and occupy less system memory, further ensuring the efficiency of audio noise reduction.

[0173] To better understand the above audio noise reduction methods, such as Figure 8 As shown below, an example of an audio noise reduction method of this application is described in detail:

[0174] S801 retrieves the current audio data and preset scene labels from the audio stream.

[0175] S802 performs speech recognition on the current audio data to determine the first human voice category of the current audio data.

[0176] S803, retrieve the audio data corresponding to the previous moment.

[0177] S804, if the first voice category does not match the second voice category, determine the audio data acquisition path.

[0178] S805 If there is no matching noise reduction attribute for the acquisition path, then determine the first weight of the human voice noise reduction parameter, the second weight of the scene noise reduction parameter, and the third weight of the path noise reduction parameter.

[0179] S806, based on the first weight, the second weight and the third weight, weights and sums the human voice noise reduction parameters, scene noise reduction parameters and path noise reduction parameters to obtain the target noise reduction parameters; wherein the first weight is greater than either the second weight or the third weight.

[0180] S807 performs noise reduction processing on audio streams based on target noise reduction parameters.

[0181] To better understand the above audio noise reduction method, an example from this application is described in detail below, taking a live streaming application as an example, such as... Figure 9 The online live streaming process using the above-mentioned audio noise reduction method may include the following steps:

[0182] (1) The audio acquisition module 901 acquires audio data through a mobile phone MIC, headphone MIC, or microphone MIC; different acquisition devices need to be configured with different parameters. Common parameters include: number of channels (single and dual channels), sampling rate (44100 / 48000), number of bits (8 / 16), etc.

[0183] (2) The scene recognition module 902 acquires the above audio data from the corresponding acquisition channel; wherein, the scene recognition module 902 includes a tag configuration analysis unit 9021 and a voice analysis unit 9022;

[0184] (3) The tag configuration analysis unit 9021 determines the scene noise reduction parameters according to the scene tags set by the user, and the voice analysis unit 9022 performs voice recognition on the audio data, and determines the voice noise reduction parameters and the channel noise reduction parameters according to the first human voice category and the acquisition path.

[0185] (4) The scene recognition module 902 also includes a noise reduction unit 9023, which generates target noise reduction parameters based on scene noise reduction parameters, human voice noise reduction parameters and path noise reduction parameters;

[0186] (5) The audio preprocessing module 903 receives the above target noise reduction parameters and performs preprocessing operations such as noise reduction, echo cancellation, automatic gain processing and sampling rate conversion on the audio data;

[0187] (6) The audio encoding module 904 compresses the preprocessed audio data through the encoder to save storage space and transmission bandwidth. Common encoding standards include: AAC (Advanced Audio Coding), MP3 (Moving Picture Experts Group Audio Layer-3), Opus, etc.; among them, Opus is a lossy audio encoding format that can handle various audio applications. It can be extended from low bit rate narrowband speech to very high-definition stereo music.

[0188] (7) The camera acquisition module 905 acquires images through the mobile phone camera to obtain video data; among which, different camera devices provide data streams of different sizes and frame rates according to their different capabilities;

[0189] (8) The video data is preprocessed by the video preprocessing module 906 before encoding, including cropping, aligning boundaries and converting rotation / color space, etc.

[0190] (9) The video encoding module 907 is used to compress the preprocessed video data, saving storage space and improving transmission efficiency. Common encoding standards include: MPEG2 (Moving Picture Experts Group), H.264, H.265, etc. Among them, H.264 is a highly compressed digital video codec standard jointly proposed by the International Organization for Standardization and the International Telecommunication Union, and H.265 is a video encoding standard improved based on the existing video encoding standard H.264.

[0191] (10) After receiving the compressed audio and video data, the audio and video encapsulation module 908 calibrates and aligns the compressed audio and video data according to the PTS (Presentation Time Stamp), obtains audio and video interleaved data according to different encapsulator specification formats, stores the interleaved data and sends it to the streaming server module 909.

[0192] (11) The streaming server module 909 distributes the above-mentioned interleaved frequency data to the audience's terminal 910, i.e., CDN (Content Delivery Network), through the internal network so that the audience can obtain the live data stream from the CDN server and decode and play it through the player.

[0193] This application performs speech recognition on the current audio data and combines it with preset scene labels to determine the target noise reduction parameters for the current audio data, thereby achieving effective noise reduction processing of the audio stream. Compared with the existing technology that uses server-side configuration of fixed noise reduction parameters for noise reduction, the noise reduction scheme of this application can configure the target noise reduction parameters in real time based on the current audio data. When the voice category of the audio stream changes, even if the scene label does not change in time, the target noise reduction parameters can still match the speech recognition result of the current audio data in real time, thereby effectively improving the noise reduction effect of the audio stream and improving the noise reduction efficiency.

[0194] This application provides an audio noise reduction device, such as... Figure 10 As shown, the audio device 110 may include: an acquisition module 1101, a recognition module 1102, a generation module 1103, and a noise reduction module 1104. The acquisition module 1101 is used to acquire the current audio data and preset scene tags at the current moment from the audio stream.

[0195] The recognition module 1102 is used to perform speech recognition on the current audio data and determine the first human voice category of the current audio data;

[0196] The generation module 1103 is used to generate target noise reduction parameters for the current audio data based on the first human voice category and scene label;

[0197] The noise reduction module 1104 is used to perform noise reduction processing on the audio stream based on the target noise reduction parameters.

[0198] This application embodiment provides a possible implementation, wherein the above-mentioned generation module 1103 is specifically used for:

[0199] Obtain the second voice category corresponding to the audio data from the previous moment;

[0200] If the first voice category does not match the second voice category, then target noise reduction parameters for the current audio data are generated based on the first voice category and the scene label.

[0201] This application embodiment provides a possible implementation, wherein the aforementioned identification module 1102 is specifically used for:

[0202] Perform speech detection on the current audio data and extract at least one human voice segment;

[0203] Obtain the audio features of each human voice segment;

[0204] The first voice category of the current audio data is determined based on audio features, where the first voice category includes speaking voices and singing voices.

[0205] This application embodiment provides a possible implementation, wherein the above-mentioned generation module 1103 is specifically used for:

[0206] The noise reduction parameters corresponding to the first voice category and the scene noise reduction parameters corresponding to the scene label are weighted to obtain the target noise reduction parameters. Among them, the noise reduction parameters corresponding to the first voice category being speech are greater than the noise reduction parameters corresponding to the first voice category being singing.

[0207] This application embodiment provides a possible implementation method, wherein the above-mentioned generation module 1103 specifically includes:

[0208] The determining unit is used to determine the acquisition path for audio data.

[0209] The generation unit is used to generate target noise reduction parameters for the current audio data based on the first human voice category, scene label, and acquisition path.

[0210] This application embodiment provides yet another possible implementation, wherein the above-mentioned generating unit is specifically used for:

[0211] If no matching noise reduction attribute exists in the acquisition path, then target noise reduction parameters for the current audio data are generated based on the first human voice category, scene label, and acquisition path.

[0212] This application provides yet another possible implementation, wherein the above-mentioned generation unit is further configured to:

[0213] The target noise reduction parameters are obtained by weighting the human voice noise reduction parameters corresponding to the first human voice category, the scene noise reduction parameters corresponding to the scene label, and the path noise reduction parameters corresponding to the acquisition path.

[0214] This application embodiment provides a possible implementation, wherein the above-mentioned generation unit is further used for:

[0215] Determine the first weight of the human voice noise reduction parameters, the second weight of the scene noise reduction parameters, and the third weight of the path noise reduction parameters;

[0216] Based on the first weight, the second weight, and the third weight, the voice noise reduction parameters, scene noise reduction parameters, and path noise reduction parameters are weighted and summed to obtain the target noise reduction parameters; where the first weight is greater than either the second weight or the third weight.

[0217] This application performs speech recognition on the current audio data and combines it with preset scene labels to determine the target noise reduction parameters for the current audio data, thereby achieving effective noise reduction processing of the audio stream. Compared with the existing technology that uses server-side configuration of fixed noise reduction parameters for noise reduction, the noise reduction scheme of this application can configure the target noise reduction parameters in real time based on the current audio data. When the voice category of the audio stream changes, even if the scene label does not change in time, the target noise reduction parameters can still match the speech recognition result of the current audio data in real time, thereby effectively improving the noise reduction effect of the audio stream and improving the noise reduction efficiency.

[0218] This application provides an electronic device, which includes a memory and a processor; at least one program stored in the memory, used by the processor to implement the corresponding content as described in the foregoing method embodiments when the program is executed. Compared with the prior art, this application achieves effective noise reduction processing of the audio stream by performing speech recognition on the current audio data and combining it with preset scene labels to determine the target noise reduction parameters for the current audio data. Compared with the prior art, which uses server-side configuration of fixed noise reduction parameters for noise reduction, the noise reduction scheme of this application can configure the target noise reduction parameters in real time based on the current audio data. When the voice category of the audio stream changes, even if the scene label is not changed in time, the target noise reduction parameters can still match the speech recognition result of the current audio data in real time, thereby effectively improving the noise reduction effect of the audio stream and improving the noise reduction efficiency.

[0219] In one alternative embodiment, an electronic device is provided, such as Figure 11 As shown, Figure 11The illustrated electronic device 4000 includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of this application.

[0220] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0221] Bus 4002 may include a pathway for transmitting information between the aforementioned components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 4002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 11 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0222] The memory 4003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.

[0223] The memory 4003 stores application code that executes the scheme of this application, and its execution is controlled by the processor 4001. The processor 4001 executes the application code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.

[0224] Among them, electronic devices include, but are not limited to: mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (such as in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 11 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0225] This application provides a computer-readable storage medium storing a computer program that, when run on a computer, enables the computer to execute the corresponding content in the aforementioned method embodiments.

[0226] This application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the following actions:

[0227] Obtain the current audio data and preset scene labels from the audio stream; perform speech recognition on the current audio data to determine the first human voice category of the current audio data; generate target noise reduction parameters for the current audio data based on the first human voice category and scene labels; and perform noise reduction processing on the audio stream based on the target noise reduction parameters.

[0228] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0229] The above description is only a partial embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. An audio noise reduction method, characterized in that, include: Obtain the current audio data and preset scene tags from the audio stream. The scene tags represent the activity category corresponding to the signal source or content of the audio stream. Speech recognition is performed on the current audio data to determine the first human voice category of the current audio data as either speaking or singing; The acquisition path for the audio data is determined, and the acquisition path is determined based on the type of audio acquisition device; Based on the first human voice category, the scene label, and the acquisition path, target noise reduction parameters are generated for the current audio data; The audio stream is denoised based on the target denoising parameters. The step of generating target noise reduction parameters for the current audio data based on the first voice category, the scene label, and the acquisition path includes: The target noise reduction parameters are obtained by weighting the human voice noise reduction parameters corresponding to the first human voice category, the scene noise reduction parameters corresponding to the scene label, and the channel noise reduction parameters corresponding to the acquisition channel. The step of weighting the voice noise reduction parameters corresponding to the first voice category, the scene noise reduction parameters corresponding to the scene label, and the channel noise reduction parameters corresponding to the acquisition channel to obtain the target noise reduction parameters includes: Determine the first weight of the human voice noise reduction parameters, the second weight of the scene noise reduction parameters, and the third weight of the path noise reduction parameters; Based on the first weight, the second weight, and the third weight, the human voice noise reduction parameter, the scene noise reduction parameter, and the path noise reduction parameter are weighted and summed to obtain the target noise reduction parameter; The step of weighted summing of the human voice noise reduction parameters, the scene noise reduction parameters, and the path noise reduction parameters based on the first weight, the second weight, and the third weight to obtain the target noise reduction parameters includes: The target noise reduction parameters are calculated based on the following formula: in, Let be the target noise reduction parameters. The noise reduction parameters for the scene are as follows. The human voice noise reduction parameters are as follows. For noise reduction parameters of the path, This is the second weight. For the first weight, This is the third weight; The first weight is greater than either the second weight or the third weight; The noise reduction parameter corresponding to the first voice category being speaking is greater than the noise reduction parameter corresponding to the first voice category being singing.

2. The audio noise reduction method according to claim 1, characterized in that, The step of generating target noise reduction parameters for the current audio data based on the first voice category and the scene label includes: Obtain the second voice category corresponding to the audio data from the previous moment; If the first voice category does not match the second voice category, then target noise reduction parameters for the current audio data are generated based on the first voice category and the scene label.

3. The audio noise reduction method according to claim 1, characterized in that, The step of performing speech recognition on the current audio data to determine whether the first human voice category of the current audio data is speaking or singing includes: Perform speech detection on the current audio data and extract at least one human voice segment; Obtain the audio features of each of the aforementioned voice segments; Based on the audio features, the first human voice category of the current audio data is determined to be either speaking or singing.

4. The audio noise reduction method according to claim 1, characterized in that, The step of generating target noise reduction parameters for the current audio data based on the first voice category, the scene label, and the acquisition path includes: If the acquisition path does not have a matching noise reduction attribute, then a target noise reduction parameter for the current audio data is generated based on the first voice category, the scene label, and the acquisition path.

5. An audio noise reduction device, characterized in that, include: The acquisition module is used to acquire the current audio data and preset scene tags from the audio stream at the current moment. The scene tags represent the activity category corresponding to the signal source or content of the audio stream. The recognition module is used to perform speech recognition on the current audio data and determine the first voice category of the current audio data, wherein the first voice category includes speaking voice and singing voice; A determining unit is used to determine the acquisition path of the audio data, wherein the acquisition path is determined based on the type of audio acquisition device; The generation module is used to generate target noise reduction parameters for the current audio data based on the first human voice category, the scene label, and the acquisition path; A noise reduction module is used to perform noise reduction processing on the audio stream based on the target noise reduction parameters; When the generation module generates target noise reduction parameters for the current audio data based on the first voice category, the scene label, and the acquisition path, it is specifically used for: The target noise reduction parameters are obtained by weighting the human voice noise reduction parameters corresponding to the first human voice category, the scene noise reduction parameters corresponding to the scene label, and the channel noise reduction parameters corresponding to the acquisition channel. When the generation module performs weighted processing on the voice noise reduction parameters corresponding to the first voice category, the scene noise reduction parameters corresponding to the scene label, and the channel noise reduction parameters corresponding to the acquisition channel to obtain the target noise reduction parameters, it is specifically used for: Determine the first weight of the human voice noise reduction parameters, the second weight of the scene noise reduction parameters, and the third weight of the path noise reduction parameters; Based on the first weight, the second weight, and the third weight, the human voice noise reduction parameter, the scene noise reduction parameter, and the path noise reduction parameter are weighted and summed to obtain the target noise reduction parameter; When the generation module is used to perform a weighted summation of the human voice noise reduction parameters, the scene noise reduction parameters, and the path noise reduction parameters based on the first weight, the second weight, and the third weight, it is specifically used to include: The target noise reduction parameters are calculated based on the following formula: in, Let be the target noise reduction parameters. The noise reduction parameters for the scene are as follows. The human voice noise reduction parameters are as follows. For noise reduction parameters of the path, This is the second weight. For the first weight, This is the third weight; The first weight is greater than either the second weight or the third weight; The noise reduction parameter corresponding to the first voice category being speaking is greater than the noise reduction parameter corresponding to the first voice category being singing.

6. The audio noise reduction device according to claim 5, characterized in that, Based on the first voice category and the scene label, the generation module, when generating target noise reduction parameters for the current audio data, specifically performs the following: Obtain the second voice category corresponding to the audio data from the previous moment; If the first voice category does not match the second voice category, then target noise reduction parameters for the current audio data are generated based on the first voice category and the scene label.

7. The audio noise reduction device according to claim 5, characterized in that, When the recognition module performs speech recognition on the current audio data and determines that the first human voice category of the current audio data is speaking or singing, it is specifically used for: Perform speech detection on the current audio data and extract at least one human voice segment; Obtain the audio features of each of the aforementioned voice segments; Based on the audio features, the first human voice category of the current audio data is determined to be either speaking or singing.

8. The audio noise reduction device according to claim 5, characterized in that, When the generation module generates target noise reduction parameters for the current audio data based on the first voice category, the scene label, and the acquisition path, it is specifically used for: If the acquisition path does not have a matching noise reduction attribute, then a target noise reduction parameter for the current audio data is generated based on the first voice category, the scene label, and the acquisition path.

9. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the audio noise reduction method according to any one of claims 1-4.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the audio noise reduction method according to any one of claims 1-4.

11. A computer program product, characterized in that, The computer program product includes computer instructions, and the processor executes the computer instructions to implement the audio noise reduction method according to any one of claims 1-4.