Audio noise reduction method, apparatus, device, and storage medium

CN116962940BActive Publication Date: 2026-09-11CHINA MOBILE GROUP SHANDONG +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211005741.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-22
Publication Date
2026-09-11
Estimated Expiration
2042-08-22

AI Technical Summary

Technical Problem

[0003]现有的降噪手段多是适用于普遍的各种噪声类型,在一定程度上可以解决一些用户使用过程中的降噪问题,但特定的噪声处理单元或滤波器组只对特定类型的噪声有明显效果,对于其他类型的环境噪声降噪效果会有明显的退化,在不同的噪声场景和噪声强度下效果存在差异

Benefits of technology

[0057]应当理解的是,本发明实施例的第二~四方面与本发明实施例的第一方面的技术方案一致,各方面及对应的可行实施方式所取得的有益效果相似,不再赘述。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116962940B_ABST
    Figure CN116962940B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an audio noise reduction method, device and equipment, and a storage medium, relating to the technical field of audio processing; natural condition information associated with a real-time scene is fused with an environmental sound recognition result to determine more accurate environmental information, so as to achieve the purpose of reducing noise of audio collected in a specific environment. The method comprises: detecting an environmental sound of collected audio signals to obtain an environment category in which a target time of collecting the audio signals is located; obtaining natural condition information at the time of collecting the audio signals; correcting the environment category by using the natural condition information to obtain a corrected environment category; and selecting a noise reduction algorithm corresponding to the corrected environment category to perform noise reduction processing on the audio signals.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the field of audio processing technology, and in particular to an audio noise reduction method, apparatus, device, and storage medium. [Background Technology]

[0002] In recent years, communication technologies and smart devices have rapidly become widespread, changing people's lifestyles in various ways, enabling them to make calls, hold video conferences, and stream live anytime, anywhere. Intelligent noise cancellation, as a key feature of smartphones, TWS earphones, and other devices, plays a crucial role in many practical application scenarios such as phone calls and earphone listening.

[0003] Existing noise reduction methods are mostly applicable to a wide range of noise types and can solve some noise reduction problems for users to a certain extent. However, specific noise processing units or filter banks are only effective for specific types of noise, and their noise reduction effect will be significantly degraded for other types of environmental noise. The effect varies under different noise scenarios and noise intensities. [Summary of the Invention]

[0004] This invention provides an audio noise reduction method, apparatus, device, and storage medium that fuses real-time scene-related natural condition information with environmental sound recognition results from multiple sources to determine more accurate environmental information, thereby achieving the purpose of noise reduction for audio collected in a specific environment.

[0005] In a first aspect, embodiments of the present invention provide an audio noise reduction method applied to an electronic device. The method includes: performing ambient sound detection on a collected audio signal to obtain the local environment category at the target time of collecting the audio signal; acquiring natural condition information at the time of collecting the audio signal; correcting the environment category using the natural condition information to obtain a corrected environment category; and selecting a noise reduction algorithm corresponding to the corrected environment category to perform noise reduction processing on the audio signal.

[0006] The aforementioned audio noise reduction method targets the real-time scene of audio acquisition, collects relevant natural condition information of the real-time scene, fuses the natural condition information with the environmental sound recognition results, corrects the environmental sound recognition results, and makes a more accurate judgment on the environment in which the audio was acquired; selects a noise reduction algorithm that integrates the environmental sound recognition results with the natural condition information, and performs noise reduction on the acquired audio; and specifically suppresses the noise in the user's current environment, improves the noise reduction effect, and improves the user's call or listening experience.

[0007] In one possible implementation, the environment category includes scene category information; acquiring the natural condition information at the target time includes:

[0008] The positioning sensor is invoked to obtain the device's positioning information at the target time;

[0009] Correcting the environmental category using the aforementioned natural condition information includes:

[0010] The scene category information of the environment category is corrected using the location information.

[0011] In one possible implementation, the environment category includes noise category information; acquiring the natural condition information at the target time includes:

[0012] Retrieve locally stored weather information;

[0013] The initial identification result of the environmental category is corrected using the natural condition information, including:

[0014] The noise category information of the environmental category is corrected using the weather information.

[0015] In one possible implementation, the natural condition information comprises multiple elements, and the environment category is corrected using the natural condition information to obtain a corrected environment category, including:

[0016] Obtain the confidence level of each of the natural condition information and the confidence level of the environmental category;

[0017] The confidence levels of each natural condition information and the environmental category are combined using a pre-set multi-source information processing model to obtain the environmental category with the highest confidence level as the corrected environmental category.

[0018] In one possible implementation, the environment category includes noise category information and scene category information; the natural condition information includes different natural condition information corresponding to the noise category information and the scene information, respectively.

[0019] The environmental category is corrected using the natural condition information to obtain a corrected environmental category, including:

[0020] The noise category information of the environment category is corrected by using the natural condition information corresponding to the noise category information to obtain a first corrected environment category;

[0021] The scene category information of the environment category is corrected by using the natural condition information corresponding to the scene category information to obtain a second corrected environment category;

[0022] Selecting a noise reduction algorithm corresponding to the corrected environment category, and performing noise reduction processing on the audio signal, including:

[0023] The first modified environment category and the second modified environment category are fused using the Dempster combination rule to obtain the fused modified environment category;

[0024] Select the noise reduction algorithm corresponding to the fused corrected environment category to perform noise reduction processing on the audio signal.

[0025] In one possible implementation, the noise category information of the environment category is corrected using natural condition information corresponding to the noise category information to obtain a first corrected environment category, including:

[0026] Calculate the mass function of the natural condition information corresponding to the noise category information and the noise category information of the environmental category to obtain the first corrected environmental category;

[0027] The scene category information of the environment category is corrected using the natural condition information corresponding to the scene category information to obtain a second corrected environment category, including:

[0028] Calculate the mass function of the natural condition information corresponding to the scene category information and the scene category information of the environment category to obtain the second modified environment category;

[0029] The Dempster combination rule is used to fuse the first modified environment category and the second modified environment category, including:

[0030] Substitute the first modified environment category and the second modified environment category into formula (1) to calculate the fused modified environment category:

[0031]

[0032] Where m1(B) represents the first modified environment category, m2(C) represents the second modified environment category; k is the normalization coefficient; B represents the source information corresponding to the first modified environment category, and C represents the source information corresponding to the second modified environment category.

[0033] Secondly, embodiments of the present invention provide an audio noise reduction device, disposed in an electronic device, the device comprising:

[0034] The detection module is used to perform environmental sound detection on the collected audio signal to obtain the local environment category at the target time when the audio signal was collected;

[0035] The acquisition module is used to acquire information about the natural conditions when acquiring the audio signal;

[0036] The correction module is used to correct the environment category using the natural condition information to obtain the corrected environment category;

[0037] The noise reduction module is used to select a noise reduction algorithm corresponding to the corrected environment category and perform noise reduction processing on the audio signal.

[0038] In one possible implementation, the environment category includes scene category information; the acquisition module is specifically used to call the positioning sensor to obtain the device's positioning information at the target time;

[0039] The correction module is specifically used to correct the scene category information of the environment category using the positioning information.

[0040] In one possible implementation, the environment category includes noise category information; the acquisition module is specifically used to retrieve locally stored weather information; and the correction module is specifically used to correct the noise category information of the environment category using the weather information.

[0041] In one possible implementation, the natural condition information is multiple, and the correction module includes:

[0042] The confidence level acquisition submodule is used to obtain the confidence level of each of the natural condition information and the confidence level of the environment category;

[0043] The combination submodule is used to combine the confidence level of each natural condition information and the confidence level of the environmental category using a pre-set multi-source information processing model, and obtain the environmental category with the highest confidence level as the corrected environmental category.

[0044] In one possible implementation, the environment category includes noise category information and scene category information; the natural condition information includes different natural condition information corresponding to the noise category information and the scene information, respectively; the correction module includes:

[0045] The first correction submodule is used to correct the noise category information of the environment category using the natural condition information corresponding to the noise category information, so as to obtain the first corrected environment category;

[0046] The second correction submodule is used to correct the scene category information of the environment category using the natural condition information corresponding to the scene category information, so as to obtain the second corrected environment category;

[0047] The noise reduction module includes:

[0048] The fusion submodule is used to fuse the first modified environment category and the second modified environment category using the Dempster combination rule to obtain the fused modified environment category;

[0049] The noise reduction submodule is used to select a noise reduction algorithm corresponding to the fused corrected environment category and perform noise reduction processing on the audio signal.

[0050] In one possible implementation, the first correction submodule is specifically used to calculate the mass function of the natural condition information corresponding to the noise category information and the noise category information of the environment category to obtain the first corrected environment category;

[0051] The second correction submodule is specifically used to calculate the mass function of the natural condition information corresponding to the scene category information and the scene category information of the environment category to obtain the second corrected environment category;

[0052] The fusion submodule is specifically used to substitute the first modified environment category and the second modified environment category into formula (1) to calculate the fused modified environment category:

[0053]

[0054] Where m1(B) represents the first modified environment category, m2(C) represents the second modified environment category; k is the normalization coefficient; B represents the source information corresponding to the first modified environment category, and C represents the source information corresponding to the second modified environment category.

[0055] Thirdly, embodiments of the present invention provide an apparatus, comprising: at least one processor; and at least one memory communicatively connected to the processor, wherein: the memory stores program instructions executable by the processor, and the processor can execute the method provided in the first aspect by invoking the program instructions.

[0056] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer instructions that cause the computer to perform the method provided in the first aspect.

[0057] It should be understood that the second to fourth aspects of the embodiments of the present invention are consistent with the technical solutions of the first aspect of the embodiments of the present invention, and the beneficial effects achieved by each aspect and the corresponding feasible implementation are similar, and will not be described again. [Attached Image Description]

[0058] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0059] Figure 1 This is a schematic diagram illustrating an application scenario of an example of the audio noise reduction method of the present invention;

[0060] Figure 2This is a flowchart of the audio noise reduction method proposed in the embodiments of the present invention;

[0061] Figure 3 This is a flowchart of another embodiment of the audio noise reduction method of the present invention;

[0062] Figure 4 This is a functional block diagram of the audio noise reduction device proposed in the embodiments of the present invention;

[0063] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.

Detailed Implementation Methods

[0064] To better understand the technical solutions in this specification, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0065] It should be understood that the described embodiments are merely some, not all, of the embodiments in this specification. All other embodiments obtained by those skilled in the art based on the embodiments in this specification without inventive effort are within the scope of protection of this specification.

[0066] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of this specification. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0067] In the daily use of terminal devices (such as smartphones), it is inevitable to encounter scenarios with very noisy environments, such as trains, airplanes, and farmers' markets, or noise scenarios caused by natural environments or weather conditions, such as strong winds (wind noise) and heavy rain. These are also many factors that affect user experience. To address these typical noise scenarios, the following methods based on traditional algorithms or deep learning network models are proposed to solve the problem of noise interference, including:

[0068] This method collects external wind noise information in real time and generates an inverse filter signal for the wind noise signal using an FIR filter or Butterworth filter. After filtering the original wind noise signal, a low-noise or clean speech signal is generated. The main problem with this method is that it ignores the significant differences between different environmental noises in different frequency bands. Specific noise processing units or filter banks are only effective for specific types (frequency bands) of noise, and the noise reduction effect will significantly degrade for other types of environmental noise.

[0069] Key speech signals are pre-saved. When the acquired audio matches the key speech signals, noise reduction parameters are calculated using the matching key speech signals. These parameters are then used to reduce noise in the voice signal during the call. However, this method cannot effectively reduce noise in scenarios where no key speech signal is matched.

[0070] As can be seen from the above description of the existing technology, how to accurately perform noise reduction processing on audio collected in different scenarios remains an urgent problem to be solved.

[0071] In view of the above problems, this invention proposes an audio noise reduction method that can be applied to electronic devices, such as terminals, computers, wearable devices, etc.

[0072] Figure 1 This is a schematic diagram illustrating an application scenario of an example audio noise reduction method according to the present invention, such as... Figure 1 As shown, the audio noise reduction method can be applied to wearable devices such as headphones 11 and mobile terminals 12 carried by users.

[0073] Both the wearable device headphones and the mobile terminal are equipped with at least one processor and at least one memory communicatively connected to the processor. The memory stores program instructions that can be executed by the processor, and the processor can execute the audio noise reduction method proposed in the embodiments of the present invention by calling the program instructions.

[0074] In this example, the wearable device headset receives a user command and begins to call and execute the program instructions associated with the audio noise reduction method; or, the wearable device headset or mobile terminal obtains relevant signals from other software and hardware, such as the wearable device headset or mobile terminal obtaining a communication software startup signal, and begins to call and execute the program instructions associated with the audio noise reduction method.

[0075] For example, when a user clicks the button to turn on "noise reduction mode", the wearable device headphones receive the signal from the user's button trigger, activate the audio sensor, and simultaneously call and execute the relevant program of the audio noise reduction method proposed in this embodiment of the invention, such as the relevant sensor for collecting natural condition information, the calculation system for detecting ambient sound, etc.; the wearable device headphones can process the surrounding ambient noise to ensure that the played audio is not interfered with by noise.

[0076] Figure 2 This is a flowchart of the audio noise reduction method proposed in the embodiments of the present invention, as follows: Figure 2 As shown, the steps include:

[0077] S201: Perform ambient sound detection on the acquired audio signal to obtain the local environment category at the target time of acquiring the audio signal.

[0078] The acquired audio signal can be call audio, ambient audio captured by a microphone, etc. The target time is the moment the audio signal is acquired.

[0079] Environmental sound detection of collected audio signals can determine the environment in which the electronic device collecting the audio is located by recognizing the audio.

[0080] One way to detect ambient sound in the collected audio signal is to use a pre-set ambient sound recognition model to monitor the collected audio signal in real time and output the environment category.

[0081] The pre-set ambient sound recognition model can be trained based on a basic model built using algorithms such as matrix factorization, dictionary learning, wavelet transform, or deep neural network.

[0082] In one example of this invention, the environmental category output for the environmental sound detection of the acquired audio signal can be outdoor natural noise, indoor mechanical noise, etc.

[0083] S202: Obtain information about the natural conditions when the audio signal is acquired.

[0084] Natural condition information is objective information that exists when audio signals are collected, and it can provide a relatively accurate reference for further environmental judgment.

[0085] S203: Use the natural condition information to correct the environmental category and obtain the corrected environmental category.

[0086] S204: Select the noise reduction algorithm corresponding to the corrected environment category and perform noise reduction processing on the audio signal.

[0087] The noise reduction algorithm can be based on the Wiener filtering algorithm of traditional filters or the Ideal RatioMask algorithm based on NN network. The corresponding processing algorithm can be adaptively selected according to the noise intensity or noise type.

[0088] This invention utilizes objective natural condition information during audio signal acquisition to correct the ambient sound detection results, providing more references for judging the environment in which the electronic device is located when the audio is acquired. Therefore, by using natural condition information to correct the environment category, a more accurate environment category judgment result can be obtained. By selecting a noise reduction algorithm corresponding to the corrected environment category, noise in the user's current environment can be suppressed in a targeted manner, improving the noise reduction effect and enhancing the user's call or listening experience.

[0089] This invention also proposes that natural condition information may include location information and weather information.

[0090] Location information can provide accurate location and correct the environmental category output of the audio signal for ambient sound detection; weather information can define the noise range from various aspects, and the environmental category output of the audio signal for ambient sound detection can be corrected using weather information. The intersection of weather information and environmental category is calculated so that the corrected environmental category can provide more accurate noise information.

[0091] The environment category includes scene category information; acquiring the natural condition information at the target time includes:

[0092] The positioning sensor is invoked to obtain the device's positioning information at the target time;

[0093] For example, when a mobile terminal collects audio, it can call GPS, BeiDou positioning system, etc. to obtain location information.

[0094] Correcting the environmental category using the aforementioned natural condition information includes:

[0095] The scene category information of the environment category is corrected using the location information.

[0096] In one example of the present invention, the environmental category output by the ambient sound detection of the collected audio signal is: outdoor natural noise; the location information is YY Park on XX Street, and the scene category information of the environmental category is corrected using the location information to obtain the corrected environmental category as: park natural noise.

[0097] The environmental category includes noise category information; acquiring the natural condition information at the target time includes:

[0098] Retrieve locally stored weather information;

[0099] You can call the API interface to read weather information from local weather programs and corresponding files in your browser.

[0100] The initial identification result of the environmental category is corrected using the natural condition information, including:

[0101] The noise category information of the environmental category is corrected using the weather information.

[0102] In one example of the present invention, the environmental category output by the ambient sound detection of the collected audio signal is: outdoor natural noise; the obtained weather information is cloudy and windy, and the scene category information of the environmental category is corrected by using the weather information to obtain the corrected environmental category as: outdoor wind noise.

[0103] In this embodiment of the invention, S202 can acquire multiple natural condition information; the applicant finds that multiple natural condition information and the environmental category output by environmental sound detection can all be understood as one of the information sources in multi-source information, and the environmental category can be corrected by using the natural condition information through the multi-source information processing method of evidence reasoning.

[0104] The environmental category is corrected using the natural condition information, which may be multiple types, to obtain the corrected environmental category, including:

[0105] Obtain the confidence level of each of the natural condition information and the confidence level of the environmental category;

[0106] The confidence levels of each natural condition information and the environmental category are combined using a pre-set multi-source information processing model to obtain the environmental category with the highest confidence level as the corrected environmental category.

[0107] In one example of this invention, the pre-set recognition framework of the multi-source information processing model includes situations such as park wind sounds, park animal calls, park rain sounds, shopping mall voices, and station announcements. The acquired natural condition information includes location information and weather information. For each of the weather and location information, the confidence level of the above situations in the corresponding recognition framework is obtained. For example, the confidence level of weather information-park wind sounds is 40%, the confidence level of weather information-park wind sounds is 10%, the confidence level of location information-park wind sounds is 50%, the confidence level of location information-shopping mall voices is 5%, the confidence level of environment category-park wind sounds is 35%, and the confidence level of environment category-park animal calls is 15%, etc. The confidence levels corresponding to the above different judgment information sources are combined to calculate a combined mass function. The calculated combined mass function result is a confidence level of 85% for park wind sounds. This realizes the correction of the environmental sound detection result by natural condition information, correcting the environmental sound detection result: environment category - outdoor natural sound to park wind sounds, thereby improving the accuracy of judging the audio acquisition environment.

[0108] Based on environmental perception, this invention integrates various objective natural condition information such as location information and weather information to more accurately restore the real environment in which the electronic device collects audio. It uses the corrected environmental category that accurately represents the real environment as prior information to suppress noise in the user's current environment, improve noise reduction effect, and improve the user's call or listening experience.

[0109] Another embodiment of the present invention also proposes an implementation method for correcting the environmental category using the natural condition information, wherein the environmental category includes noise category information and scene category information; for example, the environmental category output by the ambient sound detection can be indoor living noise or outdoor industrial noise; wherein, indoor or outdoor is the scene category information in the environmental category, and living noise or industrial noise is the noise category information in the environmental category.

[0110] To specifically correct the noise category information and scene category information in the environment category, natural condition information for the corresponding noise category information and natural condition information for the corresponding scene category information are obtained separately. For example, the environment category output by the ambient sound detection can be indoor living noise. Natural condition information for the corresponding noise category information, such as weather information, is obtained to further infer the accurate information of living noise. Natural condition information for the corresponding scene category information, such as location information, is obtained to further infer the accurate information of indoor environments.

[0111] Figure 3 This is a flowchart of another embodiment of the audio noise reduction method of the present invention, as follows: Figure 3 As shown, different natural condition information is first used to correct the environmental category, and then the different correction results are fused.

[0112] K31: Perform ambient sound detection on the acquired audio signal to obtain the local environment category at the target time when the audio signal was acquired.

[0113] K32: Obtain natural condition information corresponding to the noise category information when acquiring the audio signal.

[0114] K33: Obtain natural condition information corresponding to the scene category information when the audio signal is collected.

[0115] K34: The noise category information of the environmental category is corrected using the natural condition information corresponding to the noise category information to obtain the first corrected environmental category.

[0116] For example, the ambient sound detection of the audio signal outputs an ambient category of outdoor water noise, and the natural condition information corresponding to the noise category is obtained as moderate rain. The outdoor water noise is then corrected to obtain a first corrected ambient category of outdoor rain sound.

[0117] The noise category information of the environment category can be corrected by using the natural condition information corresponding to the noise category information in the following way:

[0118] The mass function of the natural condition information corresponding to the noise category information and the noise category information of the environmental category is calculated to obtain the first modified environmental category.

[0119] K35: The scene category information of the environment category is corrected by using the natural condition information corresponding to the scene category information to obtain a second corrected environment category.

[0120] For example, the ambient sound detection of the audio signal outputs an environment category of outdoor water noise, and the natural condition information corresponding to the scene category is obtained as mountain forest. The outdoor water noise is then corrected to obtain a second corrected environment category of mountain forest water noise.

[0121] The scene category information of the environment category can be corrected by using the natural condition information corresponding to the scene category information in the following way:

[0122] The second modified environment category is obtained by calculating the mass function of the natural condition information corresponding to the scene category information and the scene category information of the environment category.

[0123] K36: The first modified environment category and the second modified environment category are fused using the Dempster combination rule to obtain the fused modified environment category.

[0124] The Dempster combination rule is used to fuse the first modified environment category and the second modified environment category, including:

[0125] Substitute the first modified environment category and the second modified environment category into formula (1) to calculate the fused modified environment category:

[0126]

[0127] Where m1(B) represents the first corrected environment category, m2(C) represents the second corrected environment category; k is the normalization coefficient; B represents the source information corresponding to the first corrected environment category, and C represents the source information corresponding to the second corrected environment category. m(A) represents the fused corrected environment category.

[0128] For example, B represents the natural condition information and environmental category information corresponding to the noise category information. That is, B is the water noise in the outdoor water noise environmental category and the rain in the weather information. m1(B) represents the first corrected environmental category obtained after calculating the mass function of the outdoor water noise and the rain in the weather information: outdoor moderate rain. C represents the natural condition information and environmental category information corresponding to the scene category information. That is, C is the outdoor and the mountain / forest location information in the outdoor water noise environmental category. m2(C) represents the second corrected environmental category obtained after calculating the mass function of the outdoor water noise and the mountain / forest location information: mountain / forest water noise.

[0129] According to equation (1), the fusion of m1(B) and m2(C) yields m(A), which can be represented as follows: The first corrected environmental category: outdoor moderate rain and the second corrected environmental category: mountain forest water noise are fused to obtain the fused corrected environmental category: mountain forest rain sound.

[0130] K37: Select the noise reduction algorithm corresponding to the fused corrected environment category and perform noise reduction processing on the audio signal.

[0131] This invention integrates weather, location, and ambient sound recognition results from multiple sources to provide more accurate environmental information. This provides the noise reduction module with accurate environmental noise information, improves the noise reduction effect, significantly enhances the user experience of the noise reduction function, and also improves the call quality of the terminal device.

[0132] The electronic device in the embodiments of the present invention may be a mobile terminal or a wearable device. The process of the mobile terminal or wearable device performing the audio noise reduction method is illustrated through different embodiments.

[0133] In one example of the present invention, the application scenario is the process of audio noise reduction through a wearable device (earphone) held by a user. The noise reduction process includes:

[0134] B111: When the headphones are connected to the phone via Bluetooth, the headphones' microphone begins to collect audio signals and detects whether the user has enabled the headphones' noise cancellation mode.

[0135] B112: Noise cancellation mode detected. Connect via Bluetooth and request the phone to obtain location information, weather information, and start the ambient sound detection program.

[0136] B113: Sends the collected audio signal to the mobile phone. The mobile phone's built-in LSTM network performs environmental sound recognition on the audio signal collected by B111 and outputs the environmental category, such as indoor social noise, outdoor natural noise, etc.

[0137] B114: By using the location and weather functions in the mobile phone, obtain the mobile phone's current real-time location (such as outdoors, subway, shopping mall, etc.) and weather (such as wind speed, rainfall), and use the DS evidence reasoning method to fuse the location, weather and ambient sound recognition results, and output the corrected environment category: outdoor wind noise, indoor human noise, etc.

[0138] B115: The corrected environmental category from B114 is transmitted from the mobile phone to the headset via Bluetooth as input. The noise reduction module selects the corresponding noise reduction algorithm according to the noise category, such as the CRNN noise reduction model, and controls the noise reduction amplitude according to the noise intensity information.

[0139] B116: Based on the noise reduction algorithm and noise reduction amplitude in B115, the input speech signal is denoised in real time, and the denoised clean speech signal is output.

[0140] In one example of this invention, the application scenario is the process of audio noise reduction using a user-held mobile terminal. The noise reduction process includes:

[0141] C111: When the phone enters call mode, it actively enters noise reduction mode. The phone's microphone begins to collect call voice signals, and at the same time, the phone's location, weather, and ambient sound recognition modules are activated.

[0142] C112: The ambient sound recognition model LSTM network in the mobile phone is used to perform ambient sound recognition on the call voice signal collected in step one, and output the environment category, such as indoor social noise, outdoor natural noise, etc.

[0143] C113: By using the location and weather functions in the mobile phone, the current real-time location and weather of the mobile phone are obtained. The location, weather and ambient sound recognition results are fused using the DS evidence reasoning method, and the corrected environmental category is output: outdoor wind noise, indoor human noise, etc.

[0144] C114: Taking the corrected environment category in C113 as input, the call noise reduction module selects the corresponding noise reduction algorithm and an appropriate noise reduction amplitude based on the noise category and noise intensity information.

[0145] C115: Performs real-time noise reduction on the input call voice signal based on the noise reduction algorithm and noise reduction amplitude in C114, and outputs a clean call voice signal after noise reduction.

[0146] Figure 4 This is a functional block diagram of the audio noise reduction device proposed in an embodiment of the present invention. The audio noise reduction device is installed in an electronic device, such as... Figure 4 As shown, the device includes:

[0147] Detection module 41 is used to perform environmental sound detection on the collected audio signal to obtain the local environment category at the target time of collecting the audio signal;

[0148] Acquisition module 42 is used to acquire natural condition information when acquiring the audio signal;

[0149] Correction module 43 is used to correct the environment category using the natural condition information to obtain the corrected environment category;

[0150] The noise reduction module 44 is used to select a noise reduction algorithm corresponding to the corrected environment category and perform noise reduction processing on the audio signal.

[0151] Figure 4The audio noise reduction device provided in the illustrated embodiment can be used to perform the functions described in this specification. Figures 1 to 3 The implementation principle and technical effects of the method embodiment shown can be further referred to the relevant description in the method embodiment.

[0152] Optionally, the environment category includes scene category information; the acquisition module is specifically used to call the positioning sensor to obtain the device's positioning information at the target time;

[0153] The correction module is specifically used to correct the scene category information of the environment category using the positioning information.

[0154] Optionally, the acquisition module is specifically used to retrieve locally stored weather information; the correction module is specifically used to use the weather information to correct the noise category information of the environmental category.

[0155] Optionally, the natural condition information may be multiple, and the correction module includes:

[0156] The confidence level acquisition submodule is used to obtain the confidence level of each of the natural condition information and the confidence level of the environment category;

[0157] The combination submodule is used to combine the confidence level of each natural condition information and the confidence level of the environmental category using a pre-set multi-source information processing model, and obtain the environmental category with the highest confidence level as the corrected environmental category.

[0158] Optionally, the environment category includes noise category information and scene category information; the natural condition information includes different natural condition information corresponding to the noise category information and scene category information, respectively; the correction module includes:

[0159] The first correction submodule is used to correct the noise category information of the environment category using the natural condition information corresponding to the noise category information, so as to obtain the first corrected environment category;

[0160] The second correction submodule is used to correct the scene category information of the environment category using the natural condition information corresponding to the scene category information, so as to obtain the second corrected environment category;

[0161] The noise reduction module includes:

[0162] The fusion submodule is used to fuse the first modified environment category and the second modified environment category using the Dempster combination rule to obtain the fused modified environment category;

[0163] The noise reduction submodule is used to select a noise reduction algorithm corresponding to the fused corrected environment category and perform noise reduction processing on the audio signal.

[0164] Optionally, the first correction submodule is specifically used to calculate the mass function of the natural condition information corresponding to the noise category information and the noise category information of the environment category to obtain the first corrected environment category;

[0165] The second correction submodule is specifically used to calculate the mass function of the natural condition information corresponding to the scene category information and the scene category information of the environment category to obtain the second corrected environment category;

[0166] The fusion submodule is specifically used to substitute the first modified environment category and the second modified environment category into formula (1) to calculate the fused modified environment category:

[0167]

[0168] Where m1(B) represents the first modified environment category, m2(C) represents the second modified environment category; k is the normalization coefficient; B represents the source information corresponding to the first modified environment category, and C represents the source information corresponding to the second modified environment category.

[0169] The apparatus provided in the above embodiments is used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effects can be further referred to the relevant descriptions in the method embodiments, and will not be repeated here.

[0170] The apparatus provided in the above embodiments may be, for example, a chip or a chip module. The apparatus provided in the above embodiments is used to execute the technical solutions of the above-described method embodiments. Its implementation principles and technical effects can be further referred to the relevant descriptions in the method embodiments, and will not be repeated here.

[0171] Regarding the modules / units included in the various devices described in the above embodiments, they can be software modules / units, hardware modules / units, or a combination of both. For example, for devices applied to or integrated into a chip, all modules / units can be implemented using hardware methods such as circuits, or at least some modules / units can be implemented using software programs running on a processor integrated within the chip, while the remaining modules / units can be implemented using hardware methods such as circuits. For devices applied to or integrated into a chip module, all modules / units can be implemented using hardware methods such as circuits. Different modules / units can be located in the same component (e.g., chip, circuit module, etc.) or different components of the chip module, or at least some modules / units can be implemented using software programs. The software program runs on the processor integrated inside the chip module, and the remaining modules / units can be implemented using hardware methods such as circuits. For each device applied to or integrated into an electronic terminal device, each of its modules / units can be implemented using hardware methods such as circuits. Different modules / units can be located in the same component (e.g., chip, circuit module, etc.) or different components within the electronic terminal device. Alternatively, at least some modules / units can be implemented using software programs that run on the processor integrated inside the electronic terminal device, and the remaining (if any) modules / units can be implemented using hardware methods such as circuits.

[0172] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. The electronic device 500 includes a processor 510, a memory 511, and a computer program stored in the memory 511 and executable on the processor 510. When the processor 510 executes the program, it implements the steps in the aforementioned method embodiment. The electronic device provided in this embodiment can be used to execute the technical solution of the method embodiment shown above. Its implementation principle and technical effect can be further referred to the relevant description in the method embodiment, which will not be repeated here.

[0173] This invention provides a computer-readable storage medium storing computer instructions that cause a computer to execute the present specification. Figures 1-3 The illustrated embodiment provides an audio noise reduction method. A computer-readable storage medium may refer to a non-volatile computer storage medium.

[0174] The aforementioned computer-readable storage medium may be any combination of one or more computer-readable media. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium that contains or stores a program that may be used by or in connection with an instruction execution system, apparatus, or device.

[0175] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including—but not limited to—electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of transmitting, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.

[0176] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including—but not limited to—wireless, wire, optical fiber, radio frequency (RF), etc., or any suitable combination thereof.

[0177] Computer program code for performing the operations described herein can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as "C" or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0178] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0179] In the description of the embodiments of the present invention, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this specification. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0180] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this specification, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0181] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this specification includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which the embodiments of this specification pertain.

[0182] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."

[0183] It should be noted that the terminals involved in the embodiments of the present invention may include, but are not limited to, personal computers (PCs), personal digital assistants (PDAs), wireless handheld devices, tablet computers, mobile phones, MP3 players, MP4 players, etc.

[0184] In the several embodiments provided in this specification, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0185] Furthermore, the functional units in the various embodiments of this specification can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in a combination of hardware and software functional units.

[0186] The integrated units implemented as software functional units described above can be stored in a computer-readable storage medium. These software functional units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods described in the various embodiments of this specification. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0187] The above description is merely a preferred embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of protection of this specification.

Claims

1. An audio noise reduction method, characterized by, The method includes: Environmental sound detection is performed on the collected audio signal to obtain the local environmental category at the target time when the audio signal was collected; Acquire information about the natural conditions at the time the audio signal was acquired; The environmental category is corrected using the natural condition information to obtain a corrected environmental category; Select the noise reduction algorithm corresponding to the corrected environment category and perform noise reduction processing on the audio signal; The environmental category includes independent noise category information and scene category information; the natural condition information includes weather information used to correct the noise category information and location information used to correct the scene category information. The environmental category is corrected using the natural condition information to obtain a corrected environmental category, including: The noise category information of the environmental category is corrected using the weather information to obtain a first corrected environmental category; The scene category information of the environment category is corrected using the location information to obtain a second corrected environment category; Selecting a noise reduction algorithm corresponding to the corrected environment category, and performing noise reduction processing on the audio signal, including: The Dempster combination rule is used to fuse the first modified environment category and the second modified environment category to obtain the fused modified environment category. The first modified environment category and the second modified environment category belong to different information dimensions. Select the noise reduction algorithm corresponding to the fused corrected environment category to perform noise reduction processing on the audio signal.

2. The method according to claim 1, characterized in that, The natural condition information includes multiple elements. The environmental category is corrected using this natural condition information to obtain a corrected environmental category, including: Obtain the confidence level of each of the natural condition information and the confidence level of the environmental category; The confidence levels of each natural condition information and the environmental category are combined using a pre-set multi-source information processing model to obtain the environmental category with the highest confidence level as the corrected environmental category.

3. The method according to claim 1, characterized in that, The noise category information of the environmental category is corrected using the weather information to obtain a first corrected environmental category, including: Calculate the mass function of the noise category information of the weather information and the environmental category to obtain the first corrected environmental category; The scene category information of the environment category is corrected using the location information to obtain a second corrected environment category, including: Calculate the mass function of the location information and the scene category information of the environment category to obtain the second corrected environment category; The Dempster combination rule is used to fuse the first modified environment category and the second modified environment category, including: Substitute the first modified environment category and the second modified environment category into formula (1) to calculate the fused modified environment category: (1); in, Indicates the first modified environment category. denoted as the second modified environment category; k is the normalization coefficient; B represents the source information corresponding to the first modified environment category, and C represents the source information corresponding to the second modified environment category.

4. An audio noise reduction device, comprising: The device includes: The detection module is used to perform environmental sound detection on the collected audio signal to obtain the local environment category at the target time when the audio signal was collected; The acquisition module is used to acquire information about the natural conditions when acquiring the audio signal; The correction module is used to correct the environment category using the natural condition information to obtain the corrected environment category; The noise reduction module is used to select a noise reduction algorithm corresponding to the corrected environment category and perform noise reduction processing on the audio signal; The environmental category includes independent noise category information and scene category information; the natural condition information includes weather information used to correct the noise category information and location information used to correct the scene category information. The correction module is specifically used for: The noise category information of the environmental category is corrected using the weather information to obtain a first corrected environmental category; The scene category information of the environment category is corrected using the location information to obtain a second corrected environment category; The noise reduction module is specifically used for: The Dempster combination rule is used to fuse the first modified environment category and the second modified environment category to obtain the fused modified environment category. The first modified environment category and the second modified environment category belong to different information dimensions. Select the noise reduction algorithm corresponding to the fused corrected environment category to perform noise reduction processing on the audio signal.

5. An apparatus comprising: At least one processor; as well as At least one memory communicatively connected to the processor, characterized in that, The memory stores program instructions that can be executed by the processor, and the processor can execute the method as described in any one of claims 1 to 3 by calling the program instructions.

6. A computer-readable storage medium storing computer instructions, wherein, The computer instructions cause the computer to perform the method as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • D-S evidence theory-based high-voltage circuit breaker mechanical fault diagnosis method

    CN110119713A

  • Noise processing method, intelligent terminal and storage medium

    CN114095617A

  • Electronic device and audio noise reduction method and medium therefor

    WO2022022585A1