Noise suppression method, construction method of noise suppression strategy and related equipment
By modeling and gain control of multi-speaker systems, the problem of mutual influence of noise between speakers is solved, and the sound quality and volume of speakers are improved.
Patent Information
- Application Number
- CN202411499030.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-24
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2044-10-24
AI Technical Summary
In electronic devices, the multi-speaker system has not been isolated in the cavity structure design, causing noises to affect each other between speakers, and the prior art is difficult to effectively suppress this noise phenomenon.
By modeling the multi-speaker system, the equivalent rectangular bandwidth energy of the input signal is calculated in real time, gain control is performed using the ideal control strategy of joint correction, and noise suppression is performed for each speaker separately, taking into account separate and co-sounding scenarios.
It effectively suppresses the mutual influence of noise in the multi-speaker system, ensures the sound quality and high volume output of the speaker, and improves the user's listening experience.
Smart Images

Figure CN120475094A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application belong to the field of electronic equipment and relate to audio processing technology, and in particular to a noise suppression method, a method for constructing a noise suppression strategy, and related equipment. Background Art
[0002] With the development of terminal technology, the functions of electronic devices are becoming more and more abundant, especially the sound playback function has gradually become a more commonly used function of electronic devices.
[0003] Electronic devices can be equipped with dual or more speakers to enhance their stereo audio playback. Without proper isolation between multiple speakers, noise can easily interfere with each other. For example, excessive vibration from one speaker can cause resonance in other speakers, generating noise. Summary of the Invention
[0004] In view of the above, it is necessary to provide a noise suppression method, a method for constructing a noise suppression strategy and related equipment, which can suppress the problem of mutual influence of noise when electronic equipment plays sound through multiple speakers.
[0005] In a first aspect, the present application provides a noise suppression method, which is applied to an electronic device, wherein the electronic device includes a first speaker and a second speaker, and the cavity where the first speaker is located is connected to the cavity where the second speaker is located. The noise suppression method includes: in response to a first operation of a user, obtaining a first audio and a second audio; performing frame processing on the first audio to determine a first audio data frame, and performing frame processing on the second audio to determine a second audio data frame; if the tonality of the first audio data frame is a preset tonality, using a first suppression strategy to process the first audio data frame to obtain a third audio data frame, the first suppression strategy being determined according to a first control strategy of the first speaker and a second control strategy of the second speaker when the first speaker and the second speaker make sounds together; if the tonality of the second audio data frame is a preset tonality, using the first suppression strategy to process the second audio data frame to obtain a fourth audio data frame; the first speaker plays the third audio data frame, and the second speaker plays the fourth audio data frame.
[0006] By adopting the above technical solution, the preset tonality may refer to a tonality that is more likely to cause the speaker to produce noise, that is, for audio data frames that are more likely to cause the speaker to produce noise, the first suppression strategy obtained by modeling based on the joint sound emission of multiple speakers is adopted for processing. This can effectively suppress the problem of mutual influence of noise when electronic equipment plays sound through multiple speakers, and the noise suppression of each speaker is performed independently without affecting other speakers without noise, thereby ensuring the high volume and sound quality of the speaker.
[0007] In one possible implementation, the first suppression strategy is determined based on a first control strategy, a second control strategy, a third control strategy, and a fourth control strategy. The third control strategy is a control strategy for the first speaker when the first speaker makes a sound alone. The fourth control strategy is a control strategy for the second speaker when the second speaker makes a sound alone.
[0008] By adopting the above technical solution, not only is the modeling performed based on the joint sounding of multiple speakers, but the sounding situation of a single speaker is also taken into consideration, so that the accuracy of the first suppression strategy obtained is higher, and the effect of suppressing the problem of mutual influence of noise when electronic devices play sound through multiple speakers is further improved, and the first suppression strategy can be applied to scenarios where the speakers sound alone or together.
[0009] In one possible implementation, the first suppression strategy includes a first ideal control strategy corresponding to the first speaker, the first ideal control strategy includes a first equivalent rectangular bandwidth (ERB) energy threshold set, a first control frequency point set, a first gain control value set and a first quality factor set. The first suppression strategy is used to process the first audio data frame, including: obtaining the first ERB energy of the first audio data frame input to the first speaker; based on the first control frequency point set, comparing the first ERB energy with the corresponding ERB energy threshold in the first ERB energy threshold set; if the first ERB energy is greater than the corresponding ERB energy threshold, performing gain adjustment processing on the first audio data frame based on the first gain control value set and the first quality factor set.
[0010] Using the above technical solution, if the ERB energy of the first audio data frame at certain control frequencies exceeds the energy threshold, it indicates that the first audio data frame is more likely to generate noise at these control frequencies. If the ERB energy of the first audio data frame at certain control frequencies is greater than the corresponding ERB energy threshold, the first audio data frame is gain-adjusted to reduce the ERB energy of the first audio data frame at these control frequencies, thereby reducing the possibility of noise being generated by the first speaker when playing the first audio data frame.
[0011] In one possible implementation, gain adjustment processing is performed on the first audio data frame based on the first gain control value set and the first quality factor set, including: determining an audio bandwidth to be gain adjusted in the first audio data frame based on the first quality factor set; and gain adjusting audio data within the audio bandwidth in the first audio data frame based on the first gain control value set.
[0012] With the above technical solution, due to the continuity of the audio signal, when the calculated ERB energy of the first audio data frame at a certain control frequency point is greater than the corresponding ERB energy threshold, it indicates that there is a high probability that multiple frequency points near the control frequency point will also have ERB energies greater than the corresponding ERB energy threshold. The audio bandwidth requiring gain adjustment at each control frequency point of the first audio data frame is determined by the control frequency point and the quality factor corresponding to the control frequency point, thereby reducing the ERB energy of all audio data in the first audio data frame at these control frequencies and within the audio bandwidth corresponding to these control frequencies, thereby further reducing the possibility of noise generated when the first speaker plays the first audio data frame.
[0013] In a possible implementation, the noise suppression method further includes: if the first ERB energy is less than or equal to a corresponding ERB energy threshold, not performing gain adjustment processing on the first audio data frame.
[0014] Using the above technical solution, if the ERB energy of the first audio data frame at certain control frequencies is less than or equal to the energy threshold, it indicates that the first audio data frame is not likely to / will not generate noise at certain control frequencies. In this case, gain adjustment processing is not required, saving software and hardware resources of the electronic device.
[0015] In one possible implementation, the first suppression strategy includes a second ideal control strategy corresponding to the second speaker, the second ideal control strategy includes a second ERB energy threshold set, a second control frequency set, a second gain control value set and a second quality factor set. The first suppression strategy is used to process the second audio data frame, including: obtaining the second ERB energy of the second audio data frame input to the second speaker; based on the second control frequency set, comparing the second ERB energy with the corresponding ERB energy threshold in the ERB energy threshold set; if the second ERB energy is greater than the corresponding ERB energy threshold, performing gain adjustment processing on the second audio data frame based on the second gain control value set and the second quality factor set.
[0016] Using the above technical solution, if the ERB energy of the second audio data frame at certain control frequencies exceeds the energy threshold, it indicates that the second audio data frame is more likely to generate noise at these control frequencies. If the ERB energy of the second audio data frame at certain control frequencies is greater than the corresponding ERB energy threshold, the ERB energy of the second audio data frame at these control frequencies is reduced by performing gain adjustment processing on the second audio data frame, thereby reducing the possibility of noise being generated by the second speaker when playing the second audio data frame.
[0017] In one possible implementation, gain adjustment processing is performed on the second audio data frame based on the second gain control value set and the second quality factor set, including: determining an audio bandwidth to be gain adjusted in the second audio data frame based on the second quality factor set; and gain adjusting audio data within the audio bandwidth in the second audio data frame based on the second gain control value set.
[0018] With the above technical solution, due to the continuity of the audio signal, when the calculated ERB energy of the second audio data frame at a certain control frequency point is greater than the corresponding ERB energy threshold, it indicates that there is a high probability that the ERB energy of multiple frequency points near the control frequency point will be greater than the corresponding ERB energy threshold. The audio bandwidth for which gain adjustment is required at each control frequency point of the second audio data frame is determined by the control frequency point and the quality factor corresponding to the control frequency point, thereby reducing the ERB energy of all audio data in the second audio data frame at these control frequencies and within the audio bandwidth corresponding to these control frequencies, thereby further reducing the possibility of noise generated when the second speaker plays the second audio data frame.
[0019] In a possible implementation, the noise suppression method further includes: if the second ERB energy is less than or equal to a corresponding ERB energy threshold, not performing gain adjustment processing on the second audio data frame.
[0020] Using the above technical solution, if the ERB energy of the second audio data frame at certain control frequencies is less than or equal to the energy threshold, it indicates that the second audio data frame is not likely to / will not generate noise at certain control frequencies. In this case, gain adjustment processing is not required, saving software and hardware resources of the electronic device.
[0021] In one possible implementation, the noise suppression method also includes: if the tonality of the first audio data frame is not the preset tonality, adopting a second suppression strategy to process the first audio data frame; if the tonality of the second audio data frame is not the preset tonality, adopting a second suppression strategy to process the second audio data frame, the second suppression strategy including adjusting the first audio data frame and the second audio data frame by a preset gain size, or not adjusting the gain of the first audio data frame and the second audio data frame.
[0022] Using the above technical solution, if the tonality of the first audio data frame / the second audio data is not the preset tonality, indicating that the first audio data frame / the second audio data bit is not likely to / will not cause the speaker to produce noise, in this case, a smaller fixed-size gain adjustment can be performed on the first audio data frame / the second audio data, or no gain adjustment can be performed, thereby reducing the data processing capacity of the electronic device and the amount of software and hardware resources occupied.
[0023] In second aspect, the present application provides a method for constructing a noise suppression strategy, including: when the first speaker and the second speaker of an electronic device play a preset audio signal together and there is no noise, obtaining a first control strategy for the first speaker and a second control strategy for the second speaker, the first control strategy being determined based on the sound working condition of the first speaker and a first influence factor of the second speaker on the first speaker, and the second control strategy being determined based on the sound working condition of the second speaker and the second influence factor of the first speaker on the second speaker; based on the first control strategy, determining a first ideal control strategy for the first speaker, and based on the second control strategy, determining a second ideal control strategy for the second speaker, the first ideal control strategy being used to suppress noise for the audio data played by the first speaker, and the second ideal control strategy being used to suppress noise for the audio data played by the second speaker.
[0024] By adopting the above technical solution, an ideal control strategy corresponding to each speaker is constructed based on the working conditions of multiple speakers emitting sound together and the influence factors between them. This can effectively suppress the problem of mutual influence of noise when electronic equipment plays sound through multiple speakers. Since each speaker has a corresponding control strategy, the noise suppression processing is performed independently based on its own control strategy, which can further ensure the high volume and sound quality of the speaker's external sound.
[0025] In one possible implementation, based on the first control strategy, a first ideal control strategy for the first speaker is determined, including: obtaining the first control strategies of multiple first speakers of multiple electronic devices; determining the first ideal control strategy of the first speaker based on the multiple first control strategies; based on the second control strategy, a second ideal control strategy for the second speaker is determined, including: obtaining the second control strategies of multiple second speakers of multiple electronic devices; determining the second ideal control strategy of the second speaker based on the multiple second control strategies.
[0026] By adopting the above technical solution, there are also slight differences in consistency among the devices of the same model of electronic equipment. By modeling the ideal control strategy for the joint sound generation of multiple speakers of multiple electronic devices, the robustness of the obtained ideal control strategy can be improved, so that users can obtain a better listening experience.
[0027] In one possible implementation, based on the first control strategy, a first ideal control strategy for the first speaker is determined, including: obtaining a third control strategy for the first speaker of the electronic device when the preset audio signal is played alone and there is no noise; based on the first control strategy and the third control strategy, the first ideal control strategy for the first speaker is determined; based on the second control strategy, a second ideal control strategy for the second speaker is determined, including: obtaining a fourth control strategy for the second speaker of the electronic device when the preset audio signal is played alone and there is no noise; based on the second control strategy and the fourth control strategy, the second ideal control strategy for the second speaker is determined.
[0028] By adopting the above technical solution, the ideal control strategy is not only modeled based on the joint sound of multiple speakers, but also comprehensively considers the sound of a single speaker, so that the obtained ideal control strategy is more accurate, further improving the effect of suppressing the problem of mutual influence of noise when electronic equipment plays sound through multiple speakers, and making the ideal control strategy applicable to scenarios where speakers sound alone or together.
[0029] In one possible implementation, the first control strategy includes a first ERB energy threshold set and a first control frequency set, the first control frequency set includes multiple first control frequencies, and the first ERB energy threshold set includes multiple first ERB energy thresholds corresponding one-to-one to the multiple first control frequencies. The method also includes: obtaining the resonant frequency of the first speaker, and determining the first control frequency set based on the resonant frequency of the first speaker; at any first control frequency, obtaining the first ERB energy of the first speaker playing the preset audio signal and the second ERB energy of the second speaker playing the preset audio signal; and determining the first ERB energy threshold corresponding to any first control frequency based on the first ERB energy, the second ERB energy and the first influencing factor.
[0030] By adopting the above technical solution, based on the resonant frequency of the first speaker, it is possible to accurately determine at which control frequencies the first speaker is more likely to produce noise. Based on the first ERB energy of the preset audio signal played by the first speaker, the second ERB energy of the preset audio signal played by the second speaker, and the influence factor of the second speaker on the first speaker, it is possible to accurately determine the ERB energy threshold corresponding to the first speaker at any control frequency when multiple speakers are emitting sound together, so as to facilitate the subsequent determination of whether it is necessary to perform gain adjustment on the first audio data frame based on the ERB energy threshold.
[0031] In one possible implementation, determining a first ERB energy threshold corresponding to any first control frequency point based on the first ERB energy, the second ERB energy, and the first influencing factor includes: determining the first ERB energy threshold corresponding to any first control frequency point based on the following formula: ERB A,AB =ERBA1 +α*ERB B1 , ERB A1 =sum(s A,AB (t) 2 )*f s / BW / L,ERB B1 =sum(s B,AB (t) 2 )*f s / BW / L, BW=1.019*24.7*(4.37*f c1 / 1000+1),
[0032] Among them, ERB A,AB is the first ERB energy threshold, ERB A1 For the first ERB energy, ERB B1 is the second ERB energy, α is the first impact factor, s A,AB (t) is a signal of audio data obtained by downsampling and filtering the preset audio data input to the first speaker, s B,AB (t) is a signal of audio data obtained by downsampling and filtering the preset audio data input to the second speaker, and f s is the sampling rate of the preset audio data, BW is the filter bandwidth, L is the frame length of the preset audio data after downsampling, f c1 is any first control frequency point.
[0033] By adopting the above technical solution, the preset audio data input to the first speaker is downsampled and filtered to realize the Gammatone filter model, which can accurately characterize the ERB energy threshold corresponding to the first speaker at any control frequency point when multiple speakers play the preset audio data together, so as to facilitate the subsequent determination of whether the first audio data frame needs to be gain adjusted based on the ERB energy threshold.
[0034] In a possible implementation, each of the multiple first control frequency points corresponds to a first impact factor.
[0035] By adopting the above technical solution, each control frequency point corresponds to a first impact factor, so that the ERB energy threshold corresponding to the first speaker at any control frequency point when multiple speakers are emitting sound together can be accurately obtained through the above formula.
[0036] In a possible implementation, when the second speaker emits sound alone, the first impact factor is obtained based on vibration displacement information of the second speaker and vibration displacement information of the first speaker.
[0037] By adopting the above technical solution, the first influence factor of the second speaker on the first speaker is characterized as the vibration displacement of the first speaker caused by the vibration displacement of the second speaker when the second speaker makes a sound alone, thereby accurately determining the first influence factor.
[0038] In one possible implementation, the first impact factor is determined based on the following formula:
[0039]
[0040] Among them, x BtoA (f c1 , n) is the first vibration displacement information of the first speaker recorded when the second speaker sounds alone, x B (f c1 , n) is the second vibration displacement information of the second speaker recorded when the second speaker sounds alone, average(|x BtoA (f c1 , n)|) is the average value of the absolute values of the multiple first vibration displacement information recorded, average(|x B (f c1 , n)|) is the average value of the absolute values of the multiple second vibration displacement information recorded, f c1 is any first control frequency point, and n is the sampling point.
[0041] By adopting the above technical solution, by recording the average of the multiple vibration displacements of the second speaker at each control frequency point and the average of the multiple vibration displacements of the first speaker, the first influence factor of the second speaker on the first speaker at each control frequency point can be accurately characterized.
[0042] In one possible implementation, the second control strategy includes a second ERB energy threshold set and a second control frequency set, the second control frequency set includes multiple second control frequencies, and the second ERB energy threshold set includes multiple second ERB energy thresholds corresponding one-to-one to the multiple second control frequencies. The method also includes: obtaining the resonant frequency of the second speaker, and determining the second control frequency set based on the resonant frequency of the second speaker; at any second control frequency, obtaining the third ERB energy of the first speaker playing the preset audio signal, and the fourth ERB energy of the second speaker playing the preset audio signal; and determining the second ERB energy threshold corresponding to any second control frequency based on the third ERB energy, the fourth ERB energy and the second influencing factor.
[0043] By adopting the above technical solution, based on the resonant frequency of the second speaker, it is possible to accurately determine at which control frequency points the second speaker is more likely to produce noise. Based on the first ERB energy of the preset audio signal played by the second speaker, the second ERB energy of the preset audio signal played by the first speaker, and the influence factor of the first speaker on the second speaker, it is possible to accurately determine the ERB energy threshold corresponding to the second speaker at any control frequency point when multiple speakers are emitting sound together, which is convenient for subsequent determination based on the ERB energy threshold whether it is necessary to perform gain adjustment on the second audio data frame.
[0044] In one possible implementation, determining the second ERB energy threshold corresponding to any second control frequency point based on the third ERB energy, the fourth ERB energy, and the second influencing factor includes: determining the second ERB energy threshold corresponding to any second control frequency point based on the following formula: ERB B,AB =ERB B2 +β*ERB A2 , ERB A2 =sum(s A,AB (t) 2 )*f s / BW / L,ERB B2 =sum(s B,AB (t) 2 )*f s / BW / L, BW=1.019*24.7*(4.37*f c2 / 1000+1),
[0045] Among them, ERB B,AB is the second ERB energy threshold, ERB A2 For the third ERB energy, ERB B2 is the fourth ERB energy, β is the second impact factor, s A,AB (t) is a signal of audio data obtained by downsampling and filtering the preset audio data input to the first speaker, s B,AB (t) is a signal of audio data obtained by downsampling and filtering the preset audio data input to the second speaker, and f s is the sampling rate of the preset audio data, BW is the filter bandwidth, L is the frame length of the preset audio data after downsampling, f c2 is any second control frequency point.
[0046] By adopting the above technical solution, the preset audio data input to the second speaker is downsampled and filtered to achieve Gammaton eThe filter model can accurately characterize the ERB energy threshold corresponding to the second speaker at any control frequency point when multiple speakers play the audio data together, so as to facilitate the subsequent determination of whether the gain adjustment of the second audio data frame is required based on the ERB energy threshold.
[0047] In a possible implementation, each of the multiple second control frequency points corresponds to a second impact factor.
[0048] By adopting the above technical solution, each control frequency point corresponds to a second impact factor, so that the ERB energy threshold corresponding to the second speaker at any control frequency point when multiple speakers are emitting sound together can be accurately obtained through the above formula.
[0049] In a possible implementation, when the first speaker emits sound alone, the second impact factor is obtained based on vibration displacement information of the first speaker and vibration displacement information of the second speaker.
[0050] By adopting the above technical solution, the second influence factor of the first speaker on the second speaker is characterized as the vibration displacement of the second speaker caused by the vibration displacement of the first speaker when the first speaker makes a sound alone, thereby accurately determining the second influence factor.
[0051] In one possible implementation, the second impact factor is determined based on the following formula:
[0052]
[0053] Among them, x A (f c2 , n) is the third vibration displacement information of the first speaker recorded when the first speaker sounds alone, x AtoB (f c2 , n) is the fourth vibration displacement information of the second speaker recorded when the first speaker sounds alone, average(|x A (f c2 , n)|) is the average value of the absolute values of the third vibration displacement information recorded, average(|x AtoB (f c2 , n)|) is the average value of the absolute values of the plurality of fourth vibration displacement information recorded, f c2 is any second control frequency point, and n is the sampling point.
[0054] By adopting the above technical solution, by recording the average of the multiple vibration displacements of the first speaker at each control frequency point and the average of the multiple vibration displacements of the second speaker, the second influence factor of the first speaker on the second speaker at each control frequency point can be accurately characterized.
[0055] In a third aspect, the present application provides an electronic device, which includes a first speaker, a second speaker, a memory and a processor; the first speaker, the second speaker and the memory are all coupled to the processor; the memory is used to store program instructions; the processor is used to read the program instructions stored in the memory to implement the noise suppression method of the above-mentioned first aspect and its possible implementation methods, or to implement the method for constructing the noise suppression strategy of the above-mentioned second aspect and its possible implementation methods.
[0056] In a fourth aspect, the present application provides a computer-readable storage medium storing computer-readable instructions. When the computer-readable instructions are executed by a processor, the noise suppression method of the first aspect and its possible implementation methods are implemented, or the method for constructing a noise suppression strategy of the second aspect and its possible implementation methods are implemented.
[0057] In a fifth aspect, the present application provides a computer program product, which includes computer-readable instructions. When the computer-readable instructions are executed by a processor, the noise suppression method of the above-mentioned first aspect and its possible implementation methods are implemented, or the method for constructing a noise suppression strategy of the above-mentioned second aspect and its possible implementation methods are implemented.
[0058] In the sixth aspect, the present application provides a chip system, which is coupled to a memory, and the chip system is used to read and execute a computer program stored in the memory to implement the noise suppression method of the above-mentioned first aspect and its possible implementation methods, or to implement the noise suppression strategy construction method of the above-mentioned second aspect and its possible implementation methods.
[0059] In addition, the technical effects brought about by the second to fifth aspects can be found in the descriptions of the methods of each design in the above method section, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 A schematic structural diagram of a possible sound outlet duct of a speaker of an electronic device provided in one embodiment of the present application;
[0061] Figure 2 This is a schematic diagram of a possible installation position of a speaker of an electronic device provided in an embodiment of the present application;
[0062] Figure 3 This is a schematic diagram of an application scenario of the noise suppression method provided in one embodiment of the present application;
[0063] Figure 4 This is a schematic diagram of an application scenario of a noise suppression method provided by another embodiment of the present application;
[0064] Figure 5A schematic diagram comparing an ideal ERB energy curve and an actual ERB energy curve of a loudspeaker provided in one embodiment of the present application;
[0065] Figure 6A A flowchart of noise suppression for a speaker by an electronic device provided in one embodiment of the present application;
[0066] Figure 6B A schematic diagram of ERB energy changes of an audio data frame before and after noise suppression provided in one embodiment of the present application;
[0067] Figure 6C A schematic diagram of the spectrum changes of an audio data frame before and after noise suppression provided in one embodiment of the present application;
[0068] Figure 7 A flowchart for constructing an ideal control strategy for a loudspeaker provided in one embodiment of the present application;
[0069] Figure 8 A flowchart for constructing an ideal control strategy for a loudspeaker provided in another embodiment of the present application;
[0070] Figure 9 A schematic diagram of a structure for measuring an influence factor between two speakers provided in an embodiment of the present application;
[0071] Figure 10 A schematic diagram of a waveform of a test signal for measuring an influence factor between two speakers provided in an embodiment of the present application;
[0072] Figure 11 A flowchart of noise suppression for a speaker by an electronic device provided in another embodiment of the present application;
[0073] Figure 12 This is a diagram of the software and hardware architecture of an electronic device provided in one embodiment of the present application;
[0074] Figure 13 This is an interactive flow chart of audio data processing implemented by internal software and hardware modules of an electronic device provided by an embodiment of the present application;
[0075] Figure 14 This is a hardware architecture diagram of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0076] The following will describe the technical solutions in the embodiments of this application in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0077] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, words such as "exemplary", "or", and "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary", "or", and "for example" is intended to present related concepts in a concrete way.
[0078] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of this application are for the purpose of describing specific embodiments only and are not intended to limit this application. It should be understood that, unless otherwise specified in this application, " / " means or. For example, A / B can mean A or B. "And / or" in this application is merely a way to describe the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. "At least one" means one or more. "Multiple" means two or more than two. For example, at least one of a, b or c can mean: a, b, c, a and b, a and c, b and c, a, b and c. It should be understood that the order of the steps shown in the flowcharts herein can be changed, and some can be omitted.
[0079] To facilitate understanding of the various embodiments of this application, first, the technical terms involved in this application are introduced:
[0080] User interface (UI): It is the media interface for interaction and information exchange between applications or operating systems and users. It realizes the conversion between the internal form of information and the form acceptable to users. The user interface is source code written in a specific computer language such as Java and extensible markup language (XML). The interface source code is parsed and rendered on the electronic device, and ultimately presented as content that the user can recognize. The most common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operations that is displayed graphically. It can be visual interface elements such as text, icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, widgets, etc. displayed on the display screen of an electronic device.
[0081] Application (APP): A software program that can perform one or more specific functions. For example, a calling application, an instant messaging application, a video playback application, an audio playback application, etc.
[0082] Currently, most electronic devices are equipped with dual or more speakers to enhance stereo audio playback. In external sound playback scenarios, since the speakers on electronic devices are generally small in size and are affected by factors such as the speaker's sound outlet pipe and cavity structure, it is more likely that the speakers will produce noise when playing low-frequency, high-amplitude audio data. This is especially true when the fundamental frequency of the audio data to be played coincides with the speaker's resonant frequency, which further exacerbates the speaker's noise problem.
[0083] like Figure 1 As shown, two different sound outlet pipe designs are shown. Figure 1 The sound outlet pipe of the speaker shown in (a) is a diagonal pipe. Figure 1 The speaker shown in (b) has a zigzag sound outlet. Both types of outlets are long and narrow. Low-frequency, high-amplitude audio data corresponds to longer wavelengths and larger amplitudes, which are more likely to cause resonance in the speaker, resulting in more noise when playing low-frequency, high-amplitude audio data.
[0084] The noise of the speaker is reduced by adopting a noise suppression algorithm based on single-speaker modeling. Single-speaker modeling refers to modeling based on the sound conditions of a single speaker to obtain a corrected ideal gain curve, which can suppress the phenomenon of the speaker generating noise when playing low-frequency, large-amplitude audio data. For electronic devices that use multiple speakers, the cavity structure design of the speakers of most electronic devices does not achieve mutual non-interference between the speakers, and the ideal gain curve constructed by single-speaker modeling does not take into account the factors that will affect each other when the speakers sound, which leads to poor noise suppression effect of the speaker.
[0085] like Figure 2 As shown, the electronic device 100 includes two speakers as an example for description. The installation positions of the two speakers can be selected according to actual product design requirements, and this embodiment of the application does not limit this. Figure 2 As shown in (a) in the figure, if the two speakers each have two independent cavities, and the front cavity and the rear cavity are independent of each other (for example, the rear cavity is a closed cavity relative to the front cavity), the two speakers will not affect each other. In this case, single speaker modeling can be performed for each speaker to effectively reduce the noise of each speaker. Figure 2As shown in (b), if the two cavities of two speakers are interconnected, the sound emitted by any speaker can be transmitted to the other speaker through the interconnected cavity, causing the vibration of the diaphragm of the other speaker, resulting in mutual influence between the two speakers. In this case, if only single-speaker modeling is performed for each speaker, the noise of each speaker cannot be effectively reduced. Because the noise of the two speakers will affect each other, for example, if the diaphragm amplitude of one speaker is too large, it will cause the diaphragm of the other speaker to resonate through the interconnected cavity, which will also cause noise problems in the other speaker. As a result, the noise suppression algorithm using single-speaker modeling cannot solve the noise problem of multiple speakers with mutual influence.
[0086] Based on this, an embodiment of the present application provides a noise suppression method, which models the joint sound emission of multiple speakers to obtain an ideal control strategy for multiple speakers. In actual use, the signal energy input to each speaker (for example, the equivalent rectangular bandwidth (ERB) energy) is calculated in real time, and gain control is performed according to the ideal control strategy to suppress the noise problem of multiple speakers that affect each other.
[0087] An embodiment of the present application also provides a noise suppression method, which not only models the sound emitted by multiple speakers together, but also models the sound emitted by a single speaker, obtains an ideal control strategy for the multiple speakers after joint correction, and performs gain control based on the joint corrected ideal control strategy to suppress the noise problem of multiple speakers that affect each other.
[0088] like Figure 3 FIG2 is a schematic diagram of a scenario in which the noise suppression method provided in an embodiment of the present application is applicable. The scenario may be a voice call speaker scenario, and the scenario includes an electronic device 100. The electronic device 100 may be a device with a sound speaker function, for example, the electronic device 100 includes two or more speakers. The embodiment of the present application does not limit the device type of the electronic device. For example, the electronic device 100 may be a mobile phone, a tablet computer, a laptop computer, etc.
[0089] For example, electronic device 100 is a smartphone, and an application for making calls is installed on electronic device 100. Electronic device 100 can initiate a request to establish a call connection with another electronic device in response to a user operation. The call connection can be implemented over the Internet or a carrier network, and this embodiment of the application is not limited thereto.
[0090] When the electronic device 100 displays the call interface 101 , the electronic device 100 may also use a speaker to play a call prompt tone or the other party's voice in response to the user's operation of the hands-free control 102 in the call interface 101 .
[0091] The electronic device 100 can effectively reduce noise from each speaker by calculating the signal energy input to each speaker in real time and performing gain control according to a jointly modified ideal control strategy (e.g., the first ideal control strategy and the second ideal control strategy described below), thereby improving the user's call experience. The specific implementation of noise suppression will be described in the following embodiments.
[0092] See also Figure 4 , which is a schematic diagram of another scenario to which the noise suppression method provided in an embodiment of the present application is applicable. The scenario may be an audio playback scenario. For example, the electronic device 100 may be installed with an application that can play video or audio. Take the example of an electronic device 100 being installed with an audio application that can play audio. The electronic device 100 may, in response to a user operation, open the playback interface 103 of the audio application to play specified music.
[0093] When the electronic device 100 is not connected to an external audio playback device (e.g., headphones, speakers, etc.), the electronic device 100 uses the speakers to play music. The electronic device 100 can calculate the signal energy input to each speaker in real time and perform gain control based on the joint modified ideal control strategy to effectively reduce the noise of each speaker and enhance the user's music listening experience.
[0094] like Figure 5 As shown, it is assumed that curve S11 is an ideal ERB energy curve of audio data played by a speaker in the electronic device 100, and curve S12 is an ERB energy curve of audio data played by the speaker in an actual external speaker scenario. For the speaker, there is an ideal energy threshold at certain frequency points or frequency bands. Once the ideal energy threshold is exceeded, the speaker is more likely to generate noise, for example, Figure 5 The illustrated frequency points f1-f3 each have a corresponding ideal energy threshold. To reduce speaker noise, if the ERB energy of the audio data calculated in real time at these frequencies exceeds the corresponding ideal energy threshold, gain adjustment is required to eliminate the speaker's noise at these frequencies. For example, the electronic device can adjust the gain of the audio data based on the ERB energy difference between curves S11 and S12 for frequency points f1-f3, respectively.
[0095] For another example, if the ERB energy of the audio data calculated in real time is greater than the corresponding ideal energy threshold, gain adjustment is performed on the audio data (the gain of the audio data is reduced) to reduce the ERB energy of the audio data, thereby reducing the noise generated by the speaker when playing the audio data. If the ERB energy of the audio data calculated in real time is less than or equal to the corresponding ideal energy threshold, indicating that the speaker is less likely to generate noise when playing the audio data, in this case, gain adjustment of the audio data is not required and the gain adjustment of the audio data can be omitted.
[0096] like Figure 6A As shown, the specific process of suppressing noise on a speaker based on the noise suppression method of the electronic device provided in the embodiment of the present application is introduced.
[0097] In the embodiment of the present application, an electronic device including a first speaker and a second speaker is taken as an example for explanation. The installation positions of the first speaker and the second speaker can be selected according to actual product design requirements. The embodiment of the present application does not limit the number of speakers included in the electronic device and the installation positions of each speaker.
[0098] S601, in response to an instruction to play audio using a first speaker and a second speaker, first audio data input to the first speaker is framed to obtain a first audio data frame, and second audio data input to the second speaker is framed to obtain a second audio data frame.
[0099] In some embodiments, the electronic device may generate an instruction to play audio using the first speaker and the second speaker in response to a preset trigger event. The preset trigger event may be a user clicking the electronic device, or other events that automatically trigger the electronic device to play audio (e.g., an alarm event, an incoming call event, etc.).
[0100] For example, during a voice or video call, the electronic device may generate an instruction to play audio using the first and second speakers in response to a user clicking a hands-free control in the call interface. For another example, when the electronic device is not connected to an external audio playback device, the electronic device may generate an instruction to play audio using the first and second speakers in response to a user's operation to play a specified audio / video.
[0101] In some embodiments, the first audio data / the second audio data may be framed using a relevant audio framing algorithm, which is not limited in the present embodiment. For example, the first audio data / the second audio data may be framed using a fixed frame length framing algorithm, an endpoint detection framing algorithm, an overlay window framing algorithm, or the like.
[0102] In some embodiments, by framing the first audio data, one or more first audio data frames can be obtained, and by framing the second audio data, one or more second audio data frames can be obtained, thereby achieving tonality detection and processing of the first audio data and the second audio data in units of frames (processed using the first suppression strategy or the second suppression strategy), which can improve the tonality detection accuracy and processing accuracy of the first audio data / second audio data.
[0103] S602: Perform tonality detection on the first audio data frame and the second audio data frame respectively.
[0104] In some embodiments, the tonality detection algorithm may be used to perform tonality detection on the first audio data frame, and the tonality detection algorithm may be used to perform tonality detection on the second audio data frame, which is not limited in the present embodiment. By performing tonality detection on the first audio data frame and the second audio data frame, it is possible to determine whether the first audio data frame and the second audio data frame are audio data frames that are more likely to cause the speaker to produce noise. The first suppression strategy is used to process the audio data frame that is more likely to cause the speaker to produce noise, and the second suppression strategy is used to process the audio data frame that is not likely to cause the speaker to produce noise.
[0105] By performing tonality detection on the first audio data frame and the second audio data frame respectively, it is possible to determine whether the first audio data frame and the second audio data frame are audio data frames that are more likely to cause the speaker to produce noise. Different suppression strategies can then be used to process different types of audio data frames, so as to maximize the utilization of the software and hardware resources of the electronic device while suppressing the noise produced by the speaker.
[0106] S603: If the tonality of the first audio data frame is the preset tonality, a first suppression strategy is adopted for processing the first audio data frame.
[0107] In some embodiments, the preset tonality may refer to a tonality that is more likely to cause a speaker to produce noise. The preset tonality may be set according to the actual playback working conditions of the speaker, and the embodiments of the present application are not limited to this. The working conditions involved in the embodiments of the present application may refer to the working state of the speaker during audio playback. For example, the preset tonality characterizes that the audio data frame is a piano sound data frame or a piano-like sound data frame, that is, the preset tonality can be used to indicate that the audio type of the audio data frame is a piano sound data frame or a piano-like sound data frame.
[0108] For another example, if the tonality of the first audio data frame is greater than a preset threshold, the tonality of the first audio data frame is considered to be the preset tonality, and the first audio data frame is an audio data frame that is more likely to cause a speaker to produce noise. A first suppression strategy can be applied to the first audio data frame to suppress the problem of the speaker being more likely to produce noise when playing the first audio data frame. If the tonality of the first audio data frame is less than or equal to the preset threshold, the tonality of the first audio data frame is considered to be not the preset tonality, and the first audio data frame is an audio data frame that is less likely to cause a speaker to produce noise. A second suppression strategy can be applied to the first audio data frame.
[0109] In some embodiments, the first suppression strategy may be determined according to a strategy for suppressing noise on the first speaker and a strategy for suppressing noise on the second speaker when the first speaker and the second speaker make sounds together.
[0110] In some embodiments, the first suppression strategy can also be determined based on the strategy for suppressing noise for the first speaker and the strategy for suppressing noise for the second speaker when the first speaker and the second speaker make sounds together, as well as the strategy for suppressing noise for the first speaker when the first speaker makes sounds alone and the strategy for suppressing noise for the second speaker when the second speaker makes sounds alone.
[0111] In some embodiments, the first suppression strategy may include a first ideal control strategy corresponding to the first speaker and a second ideal control strategy corresponding to the second speaker.
[0112] In some embodiments, adopting a first suppression strategy to process the first audio data frame may include: a1. calculating the first ERB energy of the first audio data frame input to the first speaker; b1. obtaining a first ideal control strategy for the first speaker; c1. processing the first audio data frame based on the calculated first ERB energy and the first ideal control strategy to obtain a third audio data frame, and the third audio data frame is played by the first speaker.
[0113] S604: If the tonality of the second audio data frame is the preset tonality, the first suppression strategy is adopted for processing the second audio data frame.
[0114] Similarly, for the second audio data frame, whether the tonality of the second audio data frame is greater than a preset threshold can be determined to determine whether the tonality of the second audio data frame is the preset tonality. If the tonality of the second audio data frame is the preset tonality, the second audio data frame is considered to be an audio data frame that is more likely to cause the speaker to produce noise. The first suppression strategy is applied to the second audio data frame to suppress the problem of the speaker being more likely to produce noise when playing the second audio data frame.
[0115] In some embodiments, adopting the first suppression strategy to process the second audio data frame may include: a2. calculating the second ERB energy of the second audio data frame input to the second speaker; b2. obtaining the first ideal control strategy of the second speaker; c2. processing the second audio data frame based on the calculated second ERB energy and the second ideal control strategy to obtain a fourth audio data frame, and the fourth audio data frame is played by the second speaker.
[0116] In some embodiments, the first ERB energy of the first audio data frame input to the first speaker and the second ERB energy of the second audio data frame input to the second speaker can both be calculated using relevant ERB energy calculation algorithms, which are not limited in this embodiment of the present application.
[0117] In some embodiments, the first ideal control strategy for the first speaker and the second ideal control strategy for the second speaker can both be constructed before the electronic device leaves the factory and stored in the electronic device.
[0118] S605: If the tonality of the first audio data frame is not the preset tonality, adopt a second suppression strategy to process the first audio data frame.
[0119] In some embodiments, if the tonality of the first audio data frame is not a preset tonality, indicating that the first audio data frame is an audio data frame that is not likely to / will not cause the first speaker to generate noise, in this case, a smaller fixed-size gain adjustment can be performed on the first audio data frame, or no gain adjustment can be performed, which can reduce the data processing capacity of the electronic device and the amount of software and hardware resources occupied.
[0120] For example, the second suppression strategy processing may be not performing gain adjustment processing, or performing suppression with a smaller gain according to preset rules. For another example, for the first audio data frame, not performing gain adjustment processing may mean not changing the gain processing method of the first audio data frame, and using the default gain rule to perform gain processing on the first audio data frame. Performing suppression with a smaller gain according to preset rules may mean reducing the gain of the first audio data frame by a preset value, which can avoid the problem of the first speaker playing the first audio data frame with fluctuating volume. The preset value can be set according to actual needs, and the embodiment of the present application does not limit this.
[0121] S606: Output the processed first audio data frame.
[0122] In some embodiments, the processed first audio data frame may be the first audio data frame processed using the first suppression strategy, or the first audio data frame processed using the second suppression strategy. For example, the processed first audio data frame may be output to the first speaker A and played by the first speaker A, thereby achieving playback of the first audio data frame.
[0123] S607: If the tonality of the second audio data frame is not the preset tonality, adopt a second suppression strategy to process the second audio data frame.
[0124] In some embodiments, if the tonality of the second audio data frame is not a preset tonality, it indicates that the second audio data frame is an audio data frame that is not likely to / will not cause the second speaker to produce noise. In this case, a smaller fixed-size gain adjustment can be performed on the second audio data frame, or no gain adjustment can be performed, which can reduce the data processing capacity of the electronic device and the usage of software and hardware resources.
[0125] In some embodiments, by performing a small fixed gain adjustment on the second audio data frame, it is also possible to avoid the problem of fluctuating volume when the second speaker plays the second audio data frame.
[0126] S608: Output the processed second audio data frame.
[0127] In some embodiments, the processed second audio data frame may be the second audio data frame processed using the first suppression strategy, or the second audio data frame processed using the second suppression strategy. For example, the processed second audio data frame may be output to the second speaker B and played by the second speaker B, thereby playing the second audio data frame.
[0128] like Figure 6B , which is a comparison diagram of ERB energy of multiple first audio data frames before and after noise suppression. Curve S21 is the ERB energy curve of multiple first audio data frames before noise suppression, and curve S22 is the ERB energy curve of multiple first audio data frames after noise suppression.
[0129] For example, Figure 6B The tonality of the first audio data frame from the 102nd frame to the 1119th frame is the preset tonality. Figure 6B It can be seen that if the tonality of the first audio data frame is the preset tonality, by adopting the first suppression strategy to process the first audio data frame, the ERB energy of the first audio data frame changes greatly before and after noise suppression, and the ERB energy before noise suppression is significantly greater than the ERB energy after noise suppression, thereby suppressing the problem that the first speaker is more likely to generate noise when playing the first audio data frame. If the tonality of the first audio data frame is not the preset tonality, by adopting the second suppression strategy to process the first audio data frame, the ERB energy changes relatively little before and after noise suppression, which can reduce the data processing capacity of the electronic device and the occupation of software and hardware resources, and can avoid the problem of the first speaker playing the first audio data frame with fluctuating volume.
[0130] like Figure 6CThe figure shows the spectrum change of the first audio data frame with a preset tonality before and after noise suppression. Curve S31 is the spectrum curve of the first audio data frame before noise suppression, and curve S32 is the spectrum curve of the first audio data frame after noise suppression.
[0131] from Figure 6C It can be seen that by adopting the first suppression strategy for the first audio data frame, the amplitude of the first audio data frame after noise suppression is smaller than the amplitude before noise suppression, that is, gain suppression is performed on the first audio data frame, which can suppress the problem that the first speaker is more likely to generate noise when playing the first audio data frame.
[0132] The following combination Figure 7 and Figure 8 , introduces the construction process of the first ideal control strategy and the second ideal control strategy.
[0133] like Figure 7 As shown, the construction process includes at least the following steps:
[0134] S701: Construct a first control strategy for a first speaker and a second control strategy for a second speaker of each of a plurality of electronic devices.
[0135] In some embodiments, the multiple electronic devices may be electronic devices of the same brand and model, and each electronic device is equipped with the same speaker and the speakers are installed in the same position. The embodiment of the present application is described by taking each electronic device as an example in which a first speaker and a second speaker are installed. The number of the multiple electronic devices can be set according to the actual construction requirements of the control strategy, and the embodiment of the present application does not limit this.
[0136] In some embodiments, the construction method of the first control strategy of the first speaker is basically the same as the construction method of the second control strategy of the second speaker. The following takes the construction of the first control strategy of the first speaker in a certain electronic device as an example to illustrate: i. Perform a separate sound test on the first speaker (the second speaker does not make a sound) to obtain the third control strategy of the first speaker in a noise-free state; ii. Perform a joint sound test on the first speaker and the second speaker to obtain the fourth control strategy of the first speaker when both the first speaker and the second speaker are noise-free; iii. Based on the third control strategy and the fourth control strategy, the first control strategy of the first speaker is obtained. Similarly, the fifth control strategy and the sixth control strategy of the second speaker can be obtained by performing a separate sound test and a joint sound test on the second speaker, and then the second control strategy of the second speaker can be obtained based on the fifth control strategy and the sixth control strategy of the second speaker.
[0137] Similarly, for other electronic devices, the above method can be referred to to obtain the first control strategy of the first speaker and the second control strategy of the second speaker installed therein.
[0138] S702: Determine a first ideal control strategy for the first speakers based on the first control strategies for the multiple first speakers of the multiple electronic devices, and determine a second ideal control strategy for the second speakers based on the second control strategies for the multiple second speakers of the multiple electronic devices.
[0139] In some embodiments, after obtaining a first control strategy for a first speaker of each of the plurality of electronic devices, a first ideal control strategy for the first speaker can be determined by averaging the plurality of first control strategies. Similarly, a second ideal control strategy for the second speaker can also be determined by averaging the plurality of second control strategies.
[0140] like Figure 8 As shown, the construction process of the first ideal control strategy and the second ideal control strategy is introduced in detail.
[0141] S801, controlling a first speaker in an electronic device to play preset audio data alone, and obtaining a third control strategy for the first speaker in a noise-free state.
[0142] In some embodiments, the preset audio data may be audio data that is more likely to cause the first speaker to generate noise. For example, the preset audio data may be audio data such as sweep tone, white noise, piano sound, etc.
[0143] In some embodiments, for the first speaker A, a sound collector (eg, a microphone or a mic) may be arranged near the first speaker A to monitor the sound signal y emitted by the first speaker A during the process of playing the preset audio data. A (t), and then calculate the ERB energy of the sound signal monitored by the sound collector in real time. A The ERB energy of (t) can be calculated based on the signal value of the sound signal monitored by the sound collector. A The ERB energy of (t) may be subsequently compared with an ERB energy threshold in an ERB energy threshold set to set a gain control value.
[0144] In order to better reflect the auditory perception of the human ear to the preset audio data played by the first speaker, the preset audio data can be first downsampled and then filtered through a finite impulse response (FIR) filter to obtain a signal s of the filtered audio data. A (t), where s A (t) = yA (t)⊙h(t), ⊙ is the convolution operation, h(t) is the equivalent of the FIR filter. Gammaton can be achieved by filtering the preset audio data using the FIR filter. e Filter Model (Gammaton e The filter model is a set of filter models designed to simulate the frequency decomposition characteristics of the cochlea, thereby better reflecting the human ear's auditory perception of the preset audio data played by the first speaker. It is understood that other filters that can simulate the frequency decomposition characteristics of the cochlea can also be used to filter the audio data.
[0145] Assume that the sampling rate of the preset audio data after downsampling is f d , the sampling rate of the preset audio data before downsampling is f s , the FIR filter can be equivalent to the following formula 1:
[0146]
[0147] Where t is the sampling time, c is the preset proportional coefficient, n is the order of the FIR filter, b is the time attenuation coefficient, f0 is the center frequency of the FIR filter, is the phase of the FIR filter.
[0148] In some embodiments, the ERB energy of the first speaker A when playing the preset audio data can be calculated using the following formula 2:
[0149] ERB A =sum(s A (t) 2 )*f s / BW / L
[0150] ERB A,dB =10*log 10 (ERB A )--Formula 2;
[0151] Among them, ERB A The ERB energy of the first speaker when playing the preset audio data, ERB A,d B is for ERB A The decibel value, f s is the sampling rate of the preset audio data before downsampling, BW is the bandwidth of the FIR filter, and L is the frame length of the preset audio data after downsampling. The value of bandwidth BW can be obtained by the following formula: BW = 1.019*24.7*(4.37*f c / 1000+1), f C For example, f CIt can be any one of the following control frequency points (1 / 2f r ,f1,f2,f3,...,f i 、3f r ), that is, different control frequency points correspond to different bandwidths BW, and different ERB energy thresholds can be calculated based on Formula 2.
[0152] In some embodiments, the third control strategy of the first speaker A in the noise-free state may include: the ERB energy threshold set U A1 , control frequency set F A1 , Gain control value set Gain A1 and the quality factor set Q A1 .
[0153] In some embodiments, the first speaker A can be placed in a noise-free state by adjusting a multi-band equalizer (EQ), and the ERB energy threshold set U of the first speaker A can be obtained. A1 , control frequency set F A1 , Gain control value set Gain A1 and the quality factor set Q A1 .
[0154] Assume that the resonance frequency of the first speaker A is measured in advance to be f r , the control frequency band of the first speaker A can be set to 1 / 2f r -3f r , indicating that the first speaker A is more likely to produce noise in the control frequency band. The control frequency point may refer to a plurality of frequency points specified in the control frequency band that are more likely to cause the first speaker A to produce noise, that is, at these control frequency points, if the calculated ERB energy of the audio data exceeds the ERB energy threshold, the first speaker is more likely to produce noise. It can be understood that different types of speakers may be provided with different control frequency bands and control frequency points. The control frequency bands and control frequency points are set according to the actual use of the speaker, and the embodiment of the present application does not limit this. Assume that in the control frequency band 1 / 2f r -3f r The control frequency points set from small to large include: 1 / 2f r ,f1,f2,f3,...,f i 、3f r , where i is a positive integer, and the ERB energy threshold set U A 1. Control frequency set FA1, gain control value set Gain A1 and the quality factor set Q A1 It can be expressed by the following equations 3 to 6:
[0155]
[0156] in, They are respectively the control frequency 1 / 2f r ,f1,f2,f3,...,f i 、3f r The corresponding ERB energy threshold, It can be calculated using Formula 2 and the bandwidth BW corresponding to each control frequency point.
[0157] F A1 =[1 / 2f r ,f1,f2,f3,…,f i 、3f r ]--Formula 4;
[0158]
[0159] in, Each control frequency is 1 / 2f r ,f1,f2,f3,…,f i 、3f r The corresponding gain control value, The value of makes the first speaker A in a noise-free state at each control frequency point.
[0160] For example, the ERB energy of the preset audio data input to the first speaker can be calculated using a relevant ERB energy calculation algorithm, thereby generating an ERB energy curve for the preset audio data. If the real-time calculated ERB energy exceeds a corresponding ERB energy threshold, the gain control value can be set to the decibel value corresponding to the difference in ERB energy between the two. The set gain control value can also be fine-tuned during the multi-band EQ adjustment stage to ensure that the first speaker A is in a noise-free state at each control frequency. The fine-tuned gain control value can then be added to the gain control value set.
[0161]
[0162] in, To control the frequency 1 / 2f r ,f1,f2,f3,…,f i 、3f r The corresponding bandwidth adjustment range is: This can be set based on historical speaker noise measurement and adjustment experience. For example, when the calculated ERB energy of audio data at a certain control frequency point is greater than the corresponding ERB energy threshold, there is a high probability that multiple frequency points near the control frequency point will also have ERB energies greater than the corresponding ERB energy threshold. The multiple frequency points near the control frequency point are characterized by the quality factor.
[0163] In some embodiments, It can also be set according to the ERB energy curve and ERB energy threshold set of preset audio data so that after gain adjustment, the ERB energy of the control frequency point and multiple frequency points near the control frequency point is less than the ERB energy threshold corresponding to the control frequency point.
[0164] For example, when the calculated ERB energy of the audio data at the control frequency point f1 is greater than the corresponding ERB energy threshold According to The frequency band to be adjusted is determined by f1, and the audio data of the frequency band is adjusted according to The gain is adjusted so that the first speaker is less likely to generate noise in this frequency band.
[0165] S802: Control the second speaker in the electronic device to play preset audio data alone, and obtain a fifth control strategy for the second speaker in a noise-free state.
[0166] In some embodiments, for the second speaker B, the fifth control strategy for the second speaker in the noise-free state may include: an ERB energy threshold set U B1 , control frequency set F B1 , Gain control value set Gain B1 and the quality factor set Q B1 . ERB energy threshold set U B1 , control frequency set F B1 , Gain control value set Gain B1 and the quality factor set Q B1 The acquisition method and ERB energy threshold set U A1 , control frequency set F A1 , Gain control value set Gain A1 and the quality factor set Q A1 The method of obtaining is similar, so in order to avoid repetition, it will not be described here.
[0167] S803, controlling the first speaker and the second speaker in the electronic device to play the preset audio data together, and obtaining the fourth control strategy of the first speaker and the sixth control strategy of the second speaker when both the first speaker and the second speaker are in a state without noise.
[0168] In the case where the first speaker A and the second speaker B in the control electronic device jointly play preset audio data, a first sound collector can also be arranged near the first speaker A to monitor the sound signal emitted by the first speaker A during the playback of the preset audio data, and a second sound collector can be arranged near the second speaker B to monitor the sound signal emitted by the second speaker B during the playback of the preset audio data. Then, based on the sound signal monitored by the first sound collector, the signal of the audio data obtained after the preset audio data input to the first speaker A is filtered by the FIR filter can be determined, and based on the sound signal monitored by the second sound collector, the signal of the audio data obtained after the preset audio data input to the second speaker B is filtered by the FIR filter can be determined. Subsequently, the ERB energy of the preset audio data played by the first speaker A can be calculated based on the influence factor of the second speaker B on the first speaker A (hereinafter referred to as the first influence factor), and the ERB energy of the preset audio data played by the second speaker B can be calculated based on the influence factor of the first speaker A on the second speaker B (hereinafter referred to as the second influence factor).
[0169] For example, the ERB energy of the preset audio data played by the first speaker A may be based on: the signal s of the audio data obtained by downsampling and filtering the preset audio data input to the first speaker A; A,AB (t), a signal s of audio data obtained by downsampling and filtering the preset audio data input to the second speaker B B,AB (t) and the first impact factor are obtained.
[0170] For another example, when a first speaker A and a second speaker B in an electronic device are controlled to play preset audio data together, the ERB energy of the first speaker A playing the preset audio data can be calculated using the following formula 7:
[0171] ERB A,AB =sum(s A,AB (t) 2 )*f s / BW / L+α*sum(s B,AB (t) 2 )*f s / BW / L
[0172] ERB A,AB,dB =10*log 10 (ERB A,AB )--Formula 7;
[0173] Among them, ERB A,AB is the ERB energy of the first speaker A playing the preset audio data when the first speaker A and the second speaker B play the preset audio data together, and ERB A,AB,dB For ERBA,AB The decibel value is α, and α is the first influencing factor.
[0174] Similarly, the ERB energy of the preset audio data played by the second speaker B can be based on: the signal s of the audio data obtained by downsampling and filtering the preset audio data input to the first speaker A A,AB (t), a signal s of audio data obtained by downsampling and filtering the preset audio data input to the second speaker B B,AB (t) and the second impact factor are obtained.
[0175] For example, when a first speaker A and a second speaker B in an electronic device are controlled to play preset audio data together, the ERB energy of the preset audio data played by the second speaker B can be calculated using the following formula 8:
[0176] ERB B,AB =sum(s B,AB (t) 2 )*f s / B / L+β*sum(s A,AB (t) 2 )*f s / BW / L
[0177] ERB B,AB,dB =10*log 10 (ERB B,AB )--Formula 8;
[0178] Among them, ERB B,AB In the case where the first speaker A and the second speaker B play the preset audio data together, the ERB energy of the second speaker B playing the preset audio data, ERB B,AB,dB For ERB B,AB The decibel value of the noise level is , and β is the second influencing factor.
[0179] In some embodiments, when controlling the first speaker A and the second speaker B in an electronic device to play preset audio data together, the above-mentioned multi-band EQ adjustment method can also be used to make the first speaker A and the second speaker B both in a noise-free state, thereby obtaining the fourth control strategy of the first speaker A and the sixth control strategy of the second speaker B.
[0180] In some embodiments, the fourth control strategy for the first speaker A may include: an ERB energy threshold set U A2 , control frequency set F A2 , Gain control value set Gain A2 and the quality factor set Q A2 For example, the control frequency set F A2Each control frequency point in the set F can be set based on the resonance frequency of the first speaker A. For example, the control frequency point set F A2 Can be combined with the control frequency set F A1 Same. ERB energy threshold set U A2 Each ERB energy threshold in the control frequency set F A2 The bandwidth BW corresponding to each control frequency point in is determined by formula 7. Gain control value set Gain A2 is the control frequency set F A2 The gain control value corresponding to each control frequency point, the gain control value set Gain A2 Each gain control value in is used to make the first speaker A and the second speaker B at the control frequency set F A2 All the control frequency points in are in a noise-free state (when the first speaker A and the second speaker B play the preset audio data together). A2 is the control frequency set F A2 The bandwidth adjustment range corresponding to each control frequency point.
[0181] In some embodiments, the gain control value set Gain A2 The gain control values and quality factor set Q in A2 The various quality factors can be obtained during the multi-band EQ adjustment process.
[0182] In some embodiments, the quality factor set Q A2 The quality factors in the above equations can also be calculated based on the ERB energy curve of the preset audio data and the ERB energy threshold set U A2 Make settings.
[0183] The sixth control strategy of the second speaker B may include: an ERB energy threshold set UB2, a control frequency point set F B2 , Gain control value set Gain B2 and the quality factor set Q B2 For example, the control frequency set F B2 Each control frequency point in the set F can be set based on the resonance frequency of the second speaker B. For example, the control frequency point set F B2 and the control frequency set F B1 Each ERB energy threshold in the ERB energy threshold set UB2 can be based on the control frequency point set F B2 The bandwidth BW corresponding to each control frequency point in is determined by formula 8. Gain control value set Gain B2 is the control frequency set F B2 The gain control value corresponding to each control frequency point, the gain control value set Gain B2Each gain control value in the control frequency set F makes the first speaker A and the second speaker B B2 All the control frequency points are in a noise-free state (when the first speaker A and the second speaker B play the preset audio data together). B2 is the control frequency set F B2 The bandwidth adjustment range corresponding to each control frequency point.
[0184] In some embodiments, the gain control value set Gain B2 The gain control values and quality factor set Q in B2 The various quality factors can be obtained during the multi-band EQ adjustment process.
[0185] In some embodiments, the quality factor set Q B2 The quality factors in the above equations can also be calculated based on the ERB energy curve of the preset audio data and the ERB energy threshold set U B2 Make settings.
[0186] In some embodiments, the values of the first impact factor α and the second impact factor β may vary with the frequency. For example, the frequency band f that causes the loudspeaker noise may be set to a ~f b , assign values to the first influence factor α and the second influence factor β, and set the first influence factor α and the second influence factor β in other frequency bands to zero, to indicate that the two speakers (the first speaker A and the second speaker B) will not affect each other and produce noise in other frequency bands.
[0187] In some embodiments, in the frequency band f causing the loudspeaker noise a ~f b The values assigned to the first influencing factor α and the second influencing factor β may correspond to different values at different frequencies.
[0188] like Figure 9 As shown, the electronic device 100 includes a first speaker A and a second speaker B. The first speaker A is the upper speaker of the electronic device 100, and the second speaker B is the lower speaker of the electronic device 100. The first impact factor α and the second impact factor β can be obtained by testing using a laser vibrometer. While the second speaker B plays a preset test signal alone, the first vibration displacement information x of the first speaker A is recorded using the first laser vibrometer. BtoA (f c , n), and simultaneously use the second laser vibrometer to record the second vibration displacement information x of the second speaker B B (f c , n), where n is the sampling point, f c is the control frequency. c Can be 1 / 2fr ,f1,f2,f3,...,f i 、3f r Any one of .
[0189] For example, while the second speaker B is playing a preset test signal alone, the first laser vibrometer is used to emit laser light to the diaphragm of the first speaker A, and the second laser vibrometer is used to emit laser light to the diaphragm of the second speaker B, so that the vibration displacement information of the diaphragm of the first speaker A and the vibration displacement information of the diaphragm of the second speaker B can be recorded respectively. That is, the first vibration displacement information x is recorded by the first laser vibrometer BtoA (f c , n), the second vibration displacement information x is obtained by recording with a second laser vibrometer B (f c , n).
[0190] After obtaining the first vibration displacement information x BtoA (f c , n) and the second vibration displacement information x B (f c , n), based on the first vibration displacement information x BtoA (f c , n) and the second vibration displacement information x B (f c , n) determine the first impact factor α.
[0191] For example, for any control frequency point f c , multiple first vibration displacement information x can be recorded BtoA (f c , n) and a plurality of second vibration displacement information x B (f c , n), at the control frequency f c The first influence factor of the second speaker B on the first speaker A α It can be expressed as:
[0192]
[0193] Among them, average(|x BtoA (f c , n)|) is a plurality of first vibration displacement information x BtoA (f c , the average of the absolute values of n), average(|x B (f c , n)|) is a plurality of second vibration displacement information x B (f c , n) is the average value of the absolute value. That is, at different control frequency points 1 / 2fr ,f1,f2,f3,...,f i 、3f r The corresponding first impact factor α can be calculated and substituted into Formula 7 for calculation.
[0194] Similarly, when the first speaker A plays the preset test signal alone, the first laser vibrometer is used to record the third vibration displacement information x of the first speaker A. A (f c , n), and at the same time, the fourth vibration displacement information x of the second speaker B is recorded by the second laser vibrometer AtoB (f c , n).
[0195] For example, while the first speaker A plays a preset test signal alone, the first laser vibrometer emits laser light to the diaphragm of the first speaker A, and the second laser vibrometer emits laser light to the diaphragm of the second speaker B, thereby respectively recording the vibration displacement information of the diaphragm of the first speaker A and the vibration displacement information of the diaphragm of the second speaker B. That is, the third vibration displacement information x is recorded by the first laser vibrometer. A (f c , n), the fourth vibration displacement information x is obtained by recording with the second laser vibrometer AtoB (f c , n).
[0196] After obtaining the third vibration displacement information x A (f c , n) and the fourth vibration displacement information x AtoB (f c , n), based on the third vibration displacement information x A (f c , n) and the fourth vibration displacement information x AtoB (f c , n) determine the second influencing factor β.
[0197] For example, for any control frequency point f c , multiple third vibration displacement information x can be recorded A (f c , n) and a plurality of fourth vibration displacement information x AtoB (f c , n), at the control frequency f c The second influence factor β of the first loudspeaker A on the second loudspeaker B can be expressed as:
[0198]
[0199] Among them, average(ox A(f c , n)o) are multiple third vibration displacement information x A (f c , the average of the absolute values of n), average(ox AtoB (f c , n)o) are multiple fourth vibration displacement information x AtoB (f c , the average of the absolute values of n).
[0200] like Figure 10 As shown, the preset test signal can be a superposition signal including pink noise V1 and signals V2_1 to V2_j with different bandwidths, where j is a positive integer greater than 1. The value of j can be set according to actual test requirements, and the embodiment of the present application does not limit this. For example, the preset test signal includes multiple segments of signals V2_1 to V2_j with different bandwidths, and the pink noise V1 is located between two adjacent signals. Figure 10 As shown in the figure, the preset test signal includes a superposition signal of pink noise V1 and seven signals with different bandwidths V2_1 to V27. The first signal is at a frequency of 1 / 2f r The first segment of the signal is a time domain signal with frequency f1 as the center and bandwidth B1, and the second segment of the signal is a time domain signal with frequency f1 as the center and bandwidth B2.
[0201] S804: Determine a first control strategy for the first speaker based on the third control strategy and the fourth control strategy of the first speaker, and determine a second control strategy for the second speaker based on the fifth control strategy and the sixth control strategy of the second speaker.
[0202] For the first speaker A, after determining the third control strategy and the fourth control strategy, the first control strategy of the first speaker A can be determined based on the third control strategy and the fourth control strategy, so that the first control strategy not only considers the working condition of the first speaker A making sound independently, but also considers the working condition of the first speaker A and the second speaker B making sound together.
[0203] For example, the first control strategy for the first speaker A may include: the ERB energy threshold set U A3 , control frequency set F A3 , Gain control value set Gain A3 and the quality factor set Q A3 The first control strategy based on the third control strategy and the fourth control strategy may be: for the ERB energy threshold set, compare U A1 with U A2 , retain the lower threshold value and get the ERB energy threshold set U A3By retaining the lower value of the threshold, a lower gain adjustment trigger condition is achieved, thereby making it possible to adjust the gain of audio data at more frequency points. The possibility of generating noise when the first speaker A plays the adjusted audio data is lower. For the gain control value set, compare Gain A1 and Gain A2 , retain the higher gain value and get the gain control value set Gain A3 By retaining a higher gain value, the gain adjustment range is larger, and the possibility of the first speaker A producing noise after playing the adjusted audio data is also lower. A1 With F A2 The way to find the union, or other methods (for example, F A2 or F A1 As F A3 ), get the control frequency set F A3 , for the quality factor set, we can A1 With Q A2 The way to find the union, or other methods (for example, Q A2 or Q A1 As Q A3 ), and obtain the quality factor set Q A3 Similarly, for the second speaker B, after determining the fifth control strategy and the sixth control strategy, the second control strategy of the second speaker B can be determined based on the fifth control strategy and the sixth control strategy, so that the second control strategy not only considers the working condition of the second speaker B making sound independently, but also considers the working condition of the first speaker A and the second speaker B making sound together.
[0204] For example, the second control strategy for the second speaker B may include: the ERB energy threshold set U B3 , control frequency set F B3 , Gain control value set Gain B3 and the quality factor set Q B3 The second control strategy based on the fifth control strategy and the sixth control strategy may be: for the ERB energy threshold set, compare U B1 with U B2 , retain the lower threshold value and get the ERB energy threshold set U B3 For a set of gain control values, compare Gain B1 and Gain B2 , retain the higher gain value and get the gain control value set Gain B3 For the control frequency set, the F B1 With F B2 The way to find the union, or other methods (for example, F B2 or F B1As F B3 ), get the control frequency set F B3 , for the quality factor set, we can B1 With Q B2 The way to find the union, or other methods (for example, Q B2 or Q B1 As Q B3 ), and obtain the quality factor set Q B3 .
[0205] S805 : Acquire a first control strategy for a first speaker and a second control strategy for a second speaker of each of the plurality of electronic devices.
[0206] In some embodiments, the first control strategy for the first speaker and the second control strategy for the second speaker of each of the multiple electronic devices can be obtained by repeatedly executing steps S801 to S804. For example, n electronic devices can be randomly selected, and steps S801 to S804 can be executed for each electronic device to obtain the first control strategy for the first speaker and the second control strategy for the second speaker of each electronic device. n is a positive integer, and the value of n can be set according to the actual noise suppression requirements, which is not limited in the embodiments of the present application.
[0207] For example, for the first electronic device, the first control strategy of the first speaker includes: ERB energy threshold set U A31 , control frequency set F A31 , Gain control value set Gain A31 and the quality factor set Q A31 , the second control strategy for the second speaker includes: ERB energy threshold set U B31 , control frequency set F B31 , Gain control value set Gain B31 and the quality factor set Q B31 For the nth electronic device, the first control strategy of the first speaker includes: ERB energy threshold set U A3n , control frequency set F A3n , Gain control value set Gain A3n and the quality factor set Q A3n , the second control strategy for the second speaker includes: ERB energy threshold set U B3n , control frequency set F B3n , Gain control value set Gain B3n and the quality factor set Q B3n .
[0208] S806 , preprocessing the first control strategy for the multiple first speakers of the multiple electronic devices, and preprocessing the second control strategy for the multiple second speakers of the multiple electronic devices, wherein the preprocessing includes outlier elimination processing.
[0209] In some embodiments, in order to improve the accuracy of the ideal control strategy, the multiple first control strategies and the multiple second control strategies may be pre-processed. The pre-processing may include, but is not limited to, outlier removal processing to remove values that are obviously abnormal in the multiple first control strategies and / or the multiple second control strategies. For example, if the difference between the parameter value in a certain first control strategy and the mean or median of all parameter values is greater than a preset value, it can be considered to be a value that is obviously abnormal. For another example, each strategy factor (ERB energy threshold, control frequency, gain control value, and quality factor) in the multiple first control strategies can be analyzed by cluster analysis to remove abnormal values.
[0210] In some embodiments, if the probability of an outlier existing in the plurality of first control strategies and the plurality of second control strategies is low, or if no outlier exists, step S806 may be omitted.
[0211] S807, determining a first ideal control strategy for the first speaker based on the preprocessed first control strategy for the multiple first speakers of the multiple electronic devices, and determining a second ideal control strategy for the second speaker based on the preprocessed second control strategy for the multiple second speakers of the multiple electronic devices.
[0212] In some embodiments, the first ideal control strategy for the first speaker A can be determined by averaging the first control strategies for the first speakers A of the multiple electronic devices, and the second ideal control strategy for the second speaker B can be determined by averaging the second control strategies for the second speakers B of the multiple electronic devices. For example, the multiple first control strategies for the multiple electronic devices include the ERB energy threshold set U A31 ~U A3n , by averaging the ERB energy thresholds, we get the final modified ERB energy threshold set U Afinal The first control strategies of the plurality of electronic devices further include a gain control value set Gain A31 ~Gain A3n , by averaging the gain control values, determine the final corrected gain control value set Gain Afinal The first control strategies of the plurality of electronic devices further include a control frequency set F A31 ~F A3n , by averaging the control frequency points, the final corrected control frequency point set F is determined Aifinal , or set the control frequency point F A31 ~FA3n Any control frequency set in is used as the final modified control frequency set F Aifinal The first control strategies of the plurality of electronic devices further include a quality factor set Q A31 ~Q A3n , by averaging the quality factors, the final corrected quality factor set Q is determined Afinal , or the quality factor set Q A31 ~Q A3n Any quality factor set in is used as the final modified quality factor set Q Afinal That is, the first ideal control strategy for the first loudspeaker A includes: the ERB energy threshold set U Afinal , Gain control value set Gain Afinal , control frequency set F Afinal , and the quality factor set Q Afinal .
[0213] In some optional embodiments, the ERB energy threshold set U in the first ideal control strategy of the first speaker A is A3 , control frequency set F A3 , Gain control value set Gain A3 and the quality factor set Q A3 It may also be determined based on the first control strategy of only one electronic device, or based on the four control strategies of multiple first speakers of multiple electronic devices.
[0214] Similarly, the second ideal control strategy for the second speaker B can be determined: the ERB energy threshold set U Bfinal , control frequency set F Bfinal , Gain control value set Gain Bfinal and the quality factor set Q Bfinal .
[0215] In some optional embodiments, the ERB energy threshold set U in the second ideal control strategy for the second speaker B is Bfinal , control frequency set F Bfinal , Gain control value set Gain Bfinal and the quality factor set Q Bfinal It may also be determined based on the second control strategy of only one electronic device, or based on the sixth control strategy of multiple second speakers of multiple electronic devices.
[0216] See also Figure 11 , introduces the specific process of suppressing noise on a speaker based on a noise suppression method of an electronic device provided in another embodiment of the present application.
[0217] In the embodiment of the present application, an electronic device including a first speaker and a second speaker is used as an example for explanation. The installation positions of the first speaker and the second speaker can be selected according to actual product design requirements. The embodiment of the present application does not limit the number of speakers included in the electronic device and the installation position of each speaker:
[0218] S1101, in response to an instruction to play audio using a first speaker and a second speaker, first audio data input to the first speaker is framed to obtain a first audio data frame, and second audio data input to the second speaker is framed to obtain a second audio data frame.
[0219] Step S1101 of the embodiment of the present application is similar to step S601 of the aforementioned embodiment, and will not be described again here to avoid repetition.
[0220] By framing the first audio data, one or more first audio data frames can be obtained, and by framing the second audio data, one or more second audio data frames can be obtained, thereby achieving tonality detection and processing of the first audio data and the second audio data in units of frames (using the first suppression strategy for processing, or using the second suppression strategy for processing), thereby improving the tonality detection accuracy and processing accuracy of the first audio data / second audio data.
[0221] S1102 : Perform tonality detection on the first audio data frame and the second audio data frame respectively.
[0222] Step S1102 of the embodiment of the present application is similar to step S602 of the aforementioned embodiment, and will not be described again here to avoid repetition.
[0223] By performing tonality detection on the first audio data frame and the second audio data frame respectively, it is possible to determine whether the first audio data frame and the second audio data frame are audio data frames that are more likely to cause the speaker to produce noise. Different suppression strategies can then be used to process different types of audio data frames, so as to maximize the utilization of the software and hardware resources of the electronic device while suppressing the noise produced by the speaker.
[0224] S1103: If the tonality of the first audio data frame is a preset tonality, calculate a first ERB energy of the first audio data frame input to the first speaker.
[0225] In some embodiments, if the tonality of the first audio data frame is a preset tonality, indicating that the first audio data frame is more likely to cause the speaker to generate noise, a first suppression strategy needs to be adopted to process the first audio data frame, that is, the first audio data frame needs to be processed based on the first ideal control strategy corresponding to the first speaker. By calculating the first ERB energy of the first audio data frame to the first speaker, it is convenient to subsequently compare it with the ERB energy threshold set U in the first ideal control strategy. Afinal A comparison is performed, and then based on the comparison result, it can be accurately determined whether the first audio data frame will cause the speaker to generate noise.
[0226] The first ideal control strategy includes the ERB energy threshold set UAfinal, the gain control value set Gain Afinal , control frequency set F Afinal , and the quality factor set Q Afinal The first ERB energy of the first audio data frame input to the first speaker can be calculated, and the first ERB energy and the ERB energy threshold set U Afinal The corresponding ERB energy threshold is compared to determine whether gain adjustment needs to be performed on the first audio data frame.
[0227] S1104, in the control frequency set F Aifinal Each control frequency point in the first ERB energy and the ERB energy threshold set U Afinal The corresponding ERB energy thresholds were compared.
[0228] In some embodiments, by controlling the frequency set F Afinal Each control frequency point in the first ERB energy and the ERB energy threshold set U Afinal By comparing the corresponding ERB energy threshold in the first audio data frame, it can be determined whether the first audio data frame will cause the speaker to generate noise at each control frequency point.
[0229] S1105: If the first ERB energy is greater than the corresponding ERB energy threshold, based on the gain control value set Gain Afinal and the quality factor set Q Afinal Perform gain adjustment on the first audio data frame.
[0230] In some embodiments, if the first ERB energy is greater than the ERB energy threshold corresponding to a certain control frequency point, it indicates that the first audio data frame will cause the speaker to produce noise at the control frequency point. And when the calculated ERB energy of the first audio data frame at a certain control frequency point is greater than the corresponding ERB energy threshold, there is a high possibility that the ERB energy of multiple frequency points near the control frequency point is greater than the corresponding ERB energy threshold.Afinal and the quality factor set Q Afinal By performing gain suppression (reducing the gain) on the first audio data frame, the ERB energy of the first audio data frame at the control frequency and multiple frequency points near the control frequency can be adjusted to be less than or equal to the corresponding ERB energy threshold, so that the first audio data frame at the control frequency and multiple frequency points near the control frequency will not cause noise in the speaker.
[0231] S1106: If the first ERB energy is less than or equal to the corresponding ERB energy threshold, do not perform gain adjustment on the first audio data frame.
[0232] In some embodiments, if the first ERB energy is less than or equal to the ERB energy threshold corresponding to a certain control frequency point, it indicates that the first audio data frame will not cause noise in the speaker at the control frequency point. In this case, there is no need to perform gain adjustment on the first audio data frame, and the first audio data frame does not need to be gain adjusted, thereby saving software and hardware resources of the electronic device.
[0233] S1107: If the tonality of the first audio data frame is not the preset tonality, adopt a second suppression strategy to process the first audio data frame.
[0234] Step S1107 of the embodiment of the present application is similar to step S605 of the aforementioned embodiment, and will not be described again here to avoid repetition.
[0235] In some embodiments, if the tonality of the first audio data frame is not a preset tonality, indicating that the first audio data frame is an audio data frame that is not likely to / will not cause the first speaker to generate noise, in this case, a smaller fixed-size gain adjustment can be performed on the first audio data frame, or no gain adjustment can be performed, which can reduce the data processing capacity of the electronic device and the amount of software and hardware resources occupied.
[0236] In some embodiments, by performing a small fixed gain adjustment on the first audio data frame, it is also possible to avoid the problem of fluctuating volume when the first speaker plays the first audio data frame.
[0237] S1108: Output the processed first audio data frame.
[0238] In some embodiments, the processed first audio data frame may be the first audio data frame processed by the first suppression strategy, or the first audio data frame processed by the second suppression strategy. The first audio data frame processed by the first suppression strategy may be the first audio data frame processed by step S1105 or step S1106. The processed first audio data frame may be output to the first speaker and played by the first speaker, thereby achieving playback of the first audio data frame.
[0239] S1109: If the tonality of the second audio data frame is the preset tonality, calculate a second ERB energy of the second audio data frame input to the second speaker.
[0240] In some embodiments, if the tonality of the second audio data frame is a preset tonality, indicating that the second audio data frame is an audio data frame that is more likely to cause the speaker to generate noise, the first suppression strategy needs to be adopted to process the second audio data frame, that is, the first audio data frame needs to be processed based on the second ideal control strategy corresponding to the second speaker. By calculating the second ERB energy of the second audio data frame to the second speaker, it is convenient to subsequently compare it with the ERB energy threshold set U in the second ideal control strategy. Bfinal A comparison is performed, and then based on the comparison result, it can be accurately determined whether the second audio data frame will cause the speaker to generate noise.
[0241] The second ideal control strategy includes the ERB energy threshold set U Bfinal , control frequency set F Bfinal , Gain control value set Gain Bfinal and the quality factor set Q Bfinal The second ERB energy of the second audio data frame input to the second speaker can be calculated, and the second ERB energy and the ERB energy threshold set U Bfinal The corresponding ERB energy threshold is compared to determine whether gain adjustment needs to be performed on the second audio data frame.
[0242] S1110, in the control frequency set F Bfinal Each control frequency point in the second ERB energy and the ERB energy threshold set U Bfinal The corresponding ERB energy thresholds were compared.
[0243] In some embodiments, by controlling the frequency set F Bfinal Each control frequency point in the second ERB energy and the ERB energy threshold set U Bfinal By comparing the corresponding ERB energy threshold in the second audio data frame, it can be determined whether the second audio data frame will cause the speaker to generate noise at each control frequency point.
[0244] S1111: If the second ERB energy is greater than the corresponding ERB energy threshold, based on the gain control value set Gain Bfinal and the quality factor set Q Bfinal Perform gain adjustment on the second audio data frame.
[0245] In some embodiments, if the second ERB energy is greater than the ERB energy threshold corresponding to a certain control frequency point, it indicates that the second audio data frame will cause the speaker to produce noise at the control frequency point. And when the calculated ERB energy of the second audio data frame at a certain control frequency point is greater than the corresponding ERB energy threshold, there is a high possibility that the ERB energy of multiple frequency points near the control frequency point is greater than the corresponding ERB energy threshold. Bfinal and the quality factor set Q Bfinal By performing gain suppression (reducing the gain) on the second audio data frame, the ERB energy of the second audio data frame at the control frequency and multiple frequency points near the control frequency can be adjusted to be less than or equal to the corresponding ERB energy threshold, so that the second audio data frame at the control frequency and multiple frequency points near the control frequency will not cause noise in the speaker.
[0246] S1112: If the second ERB energy is less than or equal to the corresponding ERB energy threshold, no gain adjustment is performed on the second audio data frame.
[0247] In some embodiments, if the second ERB energy is less than or equal to the ERB energy threshold corresponding to a certain control frequency point, it indicates that the second audio data frame will not cause noise in the speaker at the control frequency point. In this case, there is no need to perform gain adjustment on the second audio data frame, and the second audio data frame does not need to be gain adjusted, thereby saving software and hardware resources of the electronic device.
[0248] S1113: If the tonality of the second audio data frame is not the preset tonality, adopt a second suppression strategy to process the second audio data frame.
[0249] Step S1113 of the embodiment of the present application is similar to step S607 of the aforementioned embodiment, and will not be described again here to avoid repetition.
[0250] In some embodiments, if the tonality of the second audio data frame is not a preset tonality, it indicates that the second audio data frame is an audio data frame that is not likely to / will not cause the second speaker to produce noise. In this case, a smaller fixed-size gain adjustment can be performed on the second audio data frame, or no gain adjustment can be performed, which can reduce the data processing capacity of the electronic device and the usage of software and hardware resources.
[0251] In some embodiments, by performing a small fixed gain adjustment on the second audio data frame, it is also possible to avoid the problem of fluctuating volume when the second speaker plays the second audio data frame.
[0252] S1114: Output the processed second audio data frame.
[0253] In some embodiments, the processed second audio data frame may be the second audio data frame processed by the first suppression strategy, or the second audio data frame processed by the second suppression strategy. The second audio data frame processed by the first suppression strategy may be the second audio data frame obtained by processing in step S1111 or step S1112. The processed second audio data frame may be output to the second speaker and played by the second speaker, thereby playing the second audio data frame.
[0254] The noise suppression method for suppressing noise of a loudspeaker according to the embodiment of the present application is compared with Figure 6A The noise suppression method for suppressing loudspeaker noise shown in the figure further calculates the ERB energy of the audio data frame when the tonality of the audio data frame is determined to be a preset tonality and compares it with the ERB energy threshold in the corresponding ideal control strategy. Based on the comparison result, it can accurately determine whether the audio data frame will actually cause the loudspeaker to generate noise. Different processing strategies can then be adopted to process the audio data frame, maximizing the utilization of the software and hardware resources of the electronic device while suppressing the noise generated by the loudspeaker playing the audio data frame. For example, if the calculated ERB energy is greater than the ERB energy threshold corresponding to a certain control frequency point, it indicates that the audio data frame will cause the loudspeaker to generate noise at the control frequency point. In this case, the audio data frame is gain-suppressed (gain reduced) based on the gain control value set and quality factor set in the corresponding ideal control strategy. This can adjust the ERB energy of the audio data frame at the control frequency point and multiple frequency points near the control frequency point to be less than or equal to the corresponding ERB energy threshold, thereby ensuring that the audio data frame does not cause the loudspeaker to generate noise at the control frequency point and multiple frequency points near the control frequency point. If the calculated ERB energy is less than or equal to the ERB energy threshold corresponding to a certain control frequency point, it indicates that the audio data frame will not cause noise in the speaker at the control frequency point. In this case, there is no need to adjust the gain of the audio data frame, saving the software and hardware resources of the electronic equipment.
[0255] In order to more clearly understand the implementation details of the above noise suppression method on electronic equipment, Figure 12 and Figure 13 The following describes how the various software and hardware components in electronic devices work together to implement the aforementioned noise suppression method. The details are as follows:
[0256] The operating system of the electronic device can adopt a layered architecture, event-driven architecture, micro-kernel architecture, micro-service architecture, or cloud architecture. The embodiment of the present application takes the Android system of layered architecture as an example to illustrate the software structure of the electronic device. Figure 12As shown, a layered architecture divides software into several layers, each with distinct roles and divisions of labor. Layers communicate with each other through software interfaces. Taking the Android system as an example, in some embodiments, the Android system is divided into four layers: the application layer (APK), the application framework layer (Framework), the hardware abstraction layer (HAL), and the kernel layer (Kernel).
[0257] The application layer can include a series of application packages. For example, an application package may include a calling application, an audio playback application, and a video playback application. Calling applications may include applications that support voice calls and video calls. Calling applications, audio playback applications, and video playback applications all support the ability to play audio through the speaker.
[0258] The application framework layer provides an application programming interface (API) and a programming framework for the applications in the application layer. For example, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, and the like.
[0259] Among them, the window manager is used to manage window programs. The window manager can obtain the size of the display screen, determine whether there is a status bar, lock the screen, take screenshots, etc. The content provider is used to store and obtain data and make this data accessible to applications. The data may include video, images, audio, calls made and received, browsing history and bookmarks, phone books, etc. The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build applications. The display interface can be composed of one or more views. For example, a display interface including a text message notification icon may include a view for displaying text and a view for displaying pictures. The phone manager is used to provide communication functions for electronic devices. For example, the management of call status (including answering, hanging up, etc.). The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, etc.
[0260] The hardware abstraction layer includes an audio algorithm module. The audio algorithm module can be used to execute the noise suppression method provided in the embodiment of the present application to achieve noise suppression for multiple speakers.
[0261] For example, the audio algorithm module may include an audio framing module, a tonality detection module, a first suppression strategy processing module, a second suppression strategy processing module, and an audio output module. The audio framing module is used to frame the audio data to obtain multiple audio data frames. The tonality detection module is used to perform tonality detection on the audio data frames. Different tonality detection results are processed by the first suppression strategy processing module or by the second suppression strategy processing module. The first suppression strategy processing module is used to process the audio data frames using the first suppression strategy. The second suppression strategy processing module is used to process the audio data frames using the second suppression strategy. The audio output module is used to output the audio data frames processed by the first suppression strategy processing module or the second suppression strategy processing module.
[0262] In some embodiments, the first suppression strategy processing module may include multiple ERB energy calculation and gain control units, and the number of ERB energy calculation and gain control units may match the number of speakers, that is, one ERB energy calculation and gain control unit is used to correspond to one speaker. Each ERB energy calculation and gain control unit is used to calculate the ERB energy of the audio data frame input to the corresponding speaker in real time, and compare the calculated ERB energy with the ERB energy threshold set corresponding to the speaker based on the control frequency point set corresponding to the speaker. If the calculated ERB energy value is greater than the corresponding ERB energy threshold, the audio data frame is gain-adjusted based on the gain control value set and quality factor set corresponding to the speaker. If the calculated ERB energy value is less than or equal to the corresponding ERB energy threshold, the audio data frame may not be gain-adjusted. The audio output module may send the audio data frame processed by the ERB energy calculation and gain control unit to the speaker corresponding to the ERB energy calculation and gain control unit for playback. Assuming that the electronic device includes m speakers, the first suppression strategy processing module may include m ERB energy calculation and gain control units. m is a positive integer greater than or equal to 2. As Figure 12 The first ERB energy calculation and gain control unit shown corresponds to the first loudspeaker, and the mth ERB energy calculation and gain control unit corresponds to the mth loudspeaker.
[0263] The embodiment of the present application performs ERB energy threshold comparison and gain adjustment on each speaker individually, thereby accurately suppressing the noise of each speaker and ensuring a high volume of audio playback.
[0264] The kernel layer may include audio drivers, display drivers, sensor drivers, and more. It also includes programs closely related to hardware, such as interrupt handlers and device drivers. It also includes basic, common, and high-frequency modules, such as clock management and process scheduling, as well as critical data structures. The kernel layer can be installed in the processor or stored in internal memory.
[0265] The hardware layer may include a display screen, multiple speakers, etc. The multiple speakers are used to play audio data.
[0266] It can be understood that the above software structure is only exemplary and does not constitute a limitation on the software structure of the electronic device. In other embodiments, the electronic device may have more or fewer structures, and this application does not impose any limitation on this.
[0267] In order to more intuitively understand the process of the above software modules cooperating to achieve the noise suppression of the speaker, Figure 11 The interactive diagram shown is used as an example to introduce the noise suppression process of an electronic device when a speaker plays audio data.
[0268] like Figure 13 FIG. 1 shows the interaction process of the software and hardware modules. This embodiment is described by taking an electronic device including a first speaker and a second speaker as an example.
[0269] 1301: In response to an audio playing operation, the first application sends first audio data and second audio data to an audio framing module.
[0270] In some embodiments, the first application may be an application installed on the electronic device that can implement an audio playback function, for example, the first application may be a call application, an audio playback application, a video playback application, etc. The first audio data corresponds to the first speaker, and the second audio data corresponds to the second speaker.
[0271] 1302: The audio framing module frames the first audio data to determine a first audio data frame, and frames the second audio data to determine a second audio data frame.
[0272] In some embodiments, performing tonality detection and processing on audio data in frames (using the first suppression strategy for processing or the second suppression strategy for processing) can improve the tonality detection accuracy and processing efficiency of audio data. By framing the first audio data, multiple first audio data frames can be obtained, which facilitates subsequent tonality detection of each first audio data frame to select the first suppression strategy or the second suppression strategy for processing. By framing the second audio data, multiple second audio data frames can be obtained, which facilitates subsequent tonality detection of each second audio data frame to select the first suppression strategy or the second suppression strategy for processing.
[0273] 1303: The audio framing module sends the first audio data frame and the second audio data frame to the tonality detection module.
[0274] 1304: The tonality detection module performs tonality detection on the first audio data frame and the second audio data frame respectively.
[0275] By performing tonality detection on each of the first and second audio data frames, it is possible to determine whether the first and second audio data frames are likely to cause a speaker to produce noise. For audio data frames that are likely to cause a speaker to produce noise, a first suppression strategy is used to suppress the noise that is likely to be produced when the speaker plays these audio data frames. For audio data frames that are less likely to cause a speaker to produce noise, a second suppression strategy is used to conserve software and hardware resources in the electronic device.
[0276] 1305: If the tonality of the first audio data frame is the preset tonality, the tonality detection module sends the first audio data frame to the first suppression strategy processing module.
[0277] 1306: The first suppression strategy processing module processes the first audio data frame using the first suppression strategy.
[0278] If the tonality of the first audio data frame is the preset tonality, indicating that the first audio data frame is likely to cause noise to be generated by the speaker, the first audio data frame is processed by adopting a first suppression strategy to suppress the problem that the first audio data frame is likely to generate noise when played by the first speaker.
[0279] 1307: The first suppression strategy processing module sends the processed first audio data frame to the audio output module.
[0280] 1308: If the tonality of the first audio data frame is not the preset tonality, the tonality detection module sends the first audio data frame to the second suppression strategy processing module.
[0281] 1309: The second suppression strategy processing module processes the first audio data frame using the second suppression strategy.
[0282] If the tonality of the first audio data frame is not the preset tonality, indicating that the first audio data frame is unlikely to cause noise in the speaker, the first audio data frame is processed using the second suppression strategy, resulting in a smaller amount of data processing and saving software and hardware resources of the electronic device.
[0283] 1310: The second suppression strategy processing module sends the processed first audio data frame to the audio output module.
[0284] 1311: The audio output module sends the processed first audio data frame to the first speaker.
[0285] By sending the processed first audio data frame to the first speaker, the first audio data frame can be played, and the first speaker plays the first audio data frame without noise.
[0286] 1312: If the tonality of the second audio data frame is the preset tonality, the tonality detection module sends the second audio data frame to the first suppression strategy processing module.
[0287] 1313: The first suppression strategy processing module processes the second audio data frame using the first suppression strategy.
[0288] If the tonality of the second audio data frame is the preset tonality, indicating that the second audio data frame is likely to cause noise to be generated by the speaker, the second audio data frame is processed by adopting the first suppression strategy to suppress the problem that the second audio data frame is likely to generate noise when played by the second speaker.
[0289] 1314: The first suppression strategy processing module sends the processed second audio data frame to the audio output module.
[0290] 1315: If the tonality of the second audio data frame is not the preset tonality, the tonality detection module sends the second audio data frame to the second suppression strategy processing module.
[0291] 1316: The second suppression strategy processing module processes the second audio data frame using the second suppression strategy.
[0292] If the tonality of the second audio data frame is not the preset tonality, indicating that the second audio data frame is not likely to cause noise in the speaker, the second audio data frame is processed using the second suppression strategy, resulting in a relatively small amount of data processing, thereby saving software and hardware resources of the electronic device.
[0293] 1317: The second suppression strategy processing module sends the processed second audio data frame to the audio output module.
[0294] 1318: The audio output module sends the processed second audio data frame to the second speaker.
[0295] By sending the processed second audio data frame to the second speaker, the second audio data frame can be played, and the second speaker plays the second audio data frame without noise.
[0296] See Figure 14 As shown, the electronic device 100 involved in the embodiment of the present application is introduced below. The electronic device 100 in the embodiment of the present application can be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a large screen, a smart TV, a netbook, as well as a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) or virtual reality (VR) device, etc., which includes a microphone and a display screen. The embodiment of the present application does not impose any special restrictions on the specific form of the electronic device. Please refer to Figure 14 , Figure 14 1 is a schematic structural diagram of an electronic device 100 provided in an embodiment of the present application.
[0297] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.
[0298] The number of speakers 170A can be multiple. For example, the speakers 170A include a first speaker and a second speaker. For example, the first speaker is used to play the first audio processed by the first suppression strategy or the second suppression strategy, and the second speaker is used to play the second audio processed by the first suppression strategy or the second suppression strategy.
[0299] It should be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0300] In addition, operating systems run on the above components, such as the iOS operating system developed by Apple, the Android open source operating system developed by Google, and the Windows operating system developed by Microsoft.
[0301] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0302] For example, the processor 110 may invoke the first suppression strategy or the second suppression strategy to process the first audio played by the first speaker, and the processor 110 may also invoke the first suppression strategy or the second suppression strategy to process the second audio played by the second speaker.
[0303] For another example, the processor 110 may perform frame segmentation and tonality detection on the first audio, and based on the tonality detection result, select to invoke the first suppression strategy or the second suppression strategy to process the first audio. The processor 110 may also perform frame segmentation and tonality detection on the second audio, and based on the tonality detection result, select to invoke the first suppression strategy or the second suppression strategy to process the second audio.
[0304] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use these instructions or data again, it can directly access them from the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.
[0305] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.
[0306] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.
[0307] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0308] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.
[0309] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc., which are applied to the electronic device 100. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.
[0310] In some embodiments, the antenna 1 of the electronic device 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the electronic device 100 can communicate with the network and other devices through wireless communication technology. The wireless communication technology may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include a global positioning system (GPS), a global navigation satellite system (GLONASS), a Beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS) and / or a satellite based augmentation system (SBAS).
[0311] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.
[0312] The display screen 194 is used to display images, videos, etc. The display screen 194 can also be used to display an audio playback interface, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), Miniled, MicroLed, Micro-oled, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include 1 or N display screens 194, where N is a positive integer greater than 1. Among them, the display screen 194 in the embodiment of the present application can be a touch screen. That is, the touch sensor 180K is integrated into the display screen 194.
[0313] The internal memory 121 may include one or more random access memories (RAMs) and one or more non-volatile memories (NVMs). RAMs may include static random access memories (SRAMs), dynamic random access memories (DRAMs), synchronous dynamic random access memories (SDRAMs), and double data rate synchronous dynamic random access memories (DDR SDRAMs, such as the fifth generation DDR SDRAM, commonly referred to as DDR5 SDRAMs). NVMs may include disk storage devices and flash memories.
[0314] Flash memory can be divided into NOR FLASH, NAND FLASH, 3D NAND FLASH, etc. according to the operating principle; single-level cell (SLC), multi-level cell (MLC), triple-level cell (TLC), quad-level cell (QLC), etc. according to the storage cell potential level; universal flash storage (UFS), embedded multi media card (eMMC), etc. according to the storage specification.
[0315] The random access memory can be directly read and written by the processor 110, and can be used to store executable programs (such as machine instructions) of the operating system or other running programs, and can also be used to store user and application data.
[0316] The non-volatile memory may also store executable programs and user and application data, etc., and may be loaded into the random access memory in advance for direct reading and writing by the processor 110 .
[0317] The external memory interface 120 can be used to connect to an external non-volatile memory to expand the storage capacity of the electronic device 100. The external non-volatile memory communicates with the processor 110 through the external memory interface 120 to implement a data storage function.
[0318] The noise suppression methods in the above embodiments can all be implemented in the electronic device 100 having the above hardware structure.
[0319] This embodiment also provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed on the electronic device 100, the electronic device 100 executes the above-mentioned related method steps to implement the noise suppression method in the above-mentioned embodiment, or the method for constructing an ideal control strategy.
[0320] This embodiment further provides a computer program product. When the computer program product runs on a computer, it enables the computer to execute the above-mentioned related steps to implement the noise suppression method or the method for constructing an ideal control strategy in the above-mentioned embodiment.
[0321] This embodiment also provides a chip system, which is coupled to a memory and is used to read and execute a computer program stored in the memory to implement the noise suppression method in the above embodiment, or the method for constructing an ideal control strategy.
[0322] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0323] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0324] The units described as separate components may or may not be physically separate, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0325] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0326] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0327] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A noise suppression method, applied to electronic equipment, characterized in that: The electronic device includes a first speaker and a second speaker, a cavity where the first speaker is located is connected to a cavity where the second speaker is located, and the method includes: In response to a first operation of the user, obtaining a first audio and a second audio; Performing frame processing on the first audio to determine a first audio data frame, and performing frame processing on the second audio to determine a second audio data frame; If the tonality of the first audio data frame is a preset tonality, processing the first audio data frame using a first suppression strategy to obtain a third audio data frame, wherein the first suppression strategy is determined based on a first control strategy of the first speaker and a second control strategy of the second speaker when the first speaker and the second speaker produce sound together; If the tonality of the second audio data frame is the preset tonality, processing the second audio data frame using the first suppression strategy to obtain a fourth audio data frame; The first speaker plays the third audio data frame, and the second speaker plays the fourth audio data frame.
2. The method according to claim 1, wherein The first suppression strategy is determined based on the first control strategy, the second control strategy, the third control strategy and the fourth control strategy. The third control strategy is the control strategy for the first speaker when the first speaker makes a sound alone. The fourth control strategy is the control strategy for the second speaker when the second speaker makes a sound alone.
3. The method according to claim 1 or 2, wherein: The first suppression strategy includes a first ideal control strategy corresponding to the first speaker, the first ideal control strategy including a first equivalent rectangular bandwidth (ERB) energy threshold set, a first control frequency point set, a first gain control value set, and a first quality factor set. Processing the first audio data frame using the first suppression strategy includes: Obtaining a first ERB energy of a first audio data frame input to the first speaker; comparing the first ERB energy with a corresponding ERB energy threshold in the first ERB energy threshold set based on the first control frequency point set; If the first ERB energy is greater than the corresponding ERB energy threshold, gain adjustment processing is performed on the first audio data frame based on the first gain control value set and the first quality factor set.
4. The method according to claim 3, wherein The performing gain adjustment processing on the first audio data frame based on the first gain control value set and the first quality factor set includes: determining an audio bandwidth to be gain adjusted in the first audio data frame based on the first quality factor set; The audio data in the first audio data frame that is within the audio bandwidth is gain-adjusted based on the first gain control value set.
5. The method according to claim 3, wherein The method further comprises: If the first ERB energy is less than or equal to the corresponding ERB energy threshold, no gain adjustment processing is performed on the first audio data frame.
6. The method according to claim 1 or 2, wherein: The first suppression strategy includes a second ideal control strategy corresponding to the second speaker, the second ideal control strategy includes a second ERB energy threshold set, a second control frequency point set, a second gain control value set, and a second quality factor set. Processing the second audio data frame using the first suppression strategy includes: Obtaining a second ERB energy of a second audio data frame input to the second speaker; comparing the second ERB energy with a corresponding ERB energy threshold in the second ERB energy threshold set based on the second control frequency point set; If the second ERB energy is greater than the corresponding ERB energy threshold, gain adjustment processing is performed on the second audio data frame based on the second gain control value set and the second quality factor set.
7. The method according to claim 6, wherein The performing gain adjustment processing on the second audio data frame based on the second gain control value set and the second quality factor set includes: determining an audio bandwidth to be gain adjusted in the second audio data frame based on the second quality factor set; The audio data in the second audio data frame that is within the audio bandwidth is gain-adjusted based on the second gain control value set.
8. The method according to claim 6, wherein The method further comprises: If the second ERB energy is less than or equal to the corresponding ERB energy threshold, no gain adjustment processing is performed on the second audio data frame.
9. The method according to any one of claims 1 to 8, characterized in that The method further comprises: If the tonality of the first audio data frame is not the preset tonality, processing the first audio data frame using a second suppression strategy; If the tonality of the second audio data frame is not the preset tonality, the second audio data frame is processed using the second suppression strategy, and the second suppression strategy includes adjusting the first audio data frame and the second audio data frame by a preset gain size, or not adjusting the gain of the first audio data frame and the second audio data frame.
10. A method for constructing a noise suppression strategy, characterized in that: The method comprises: Under the condition that a first speaker and a second speaker of an electronic device jointly play a preset audio signal and both are free of noise, obtaining a first control strategy for the first speaker and a second control strategy for the second speaker, the first control strategy being determined based on a sound emission condition of the first speaker and a first influence factor of the second speaker on the first speaker, and the second control strategy being determined based on a sound emission condition of the second speaker and a second influence factor of the first speaker on the second speaker; Based on the first control strategy, a first ideal control strategy for the first speaker is determined, and based on the second control strategy, a second ideal control strategy for the second speaker is determined. The first ideal control strategy is used to suppress noise in the audio data played by the first speaker, and the second ideal control strategy is used to suppress noise in the audio data played by the second speaker.
11. The method according to claim 10, wherein The determining, based on the first control strategy, a first ideal control strategy for the first speaker includes: Acquire the first control strategy of the plurality of the first speakers of the plurality of the electronic devices; determining a first ideal control strategy for the first speaker based on a plurality of the first control strategies; The determining, based on the second control strategy, a second ideal control strategy for the second speaker includes: Acquire the second control strategy of the plurality of second speakers of the plurality of electronic devices; A second ideal control strategy for the second speaker is determined based on a plurality of the second control strategies.
12. The method according to claim 10, wherein The determining, based on the first control strategy, a first ideal control strategy for the first speaker includes: Acquire a third control strategy for the first speaker of the electronic device when the first speaker plays the preset audio signal alone without any noise; determining the first ideal control strategy for the first speaker based on the first control strategy and the third control strategy; The determining, based on the second control strategy, a second ideal control strategy for the second speaker includes: Acquire a fourth control strategy for the second speaker of the electronic device when the second speaker plays the preset audio signal alone without any noise; The second ideal control strategy for the second speaker is determined based on the second control strategy and the fourth control strategy.
13. The method according to claim 10, wherein The first control strategy includes a first ERB energy threshold set and a first control frequency set, the first control frequency set includes multiple first control frequencies, the first ERB energy threshold set includes multiple first ERB energy thresholds corresponding one-to-one to the multiple first control frequencies, and the method further includes: Acquiring a resonant frequency of the first speaker, and determining the first control frequency point set based on the resonant frequency of the first speaker; At any first control frequency point, obtaining a first ERB energy of the preset audio signal played by the first speaker and a second ERB energy of the preset audio signal played by the second speaker; A first ERB energy threshold corresponding to any one of the first control frequency points is determined based on the first ERB energy, the second ERB energy, and the first influencing factor.
14. The method according to claim 13, wherein The determining, based on the first ERB energy, the second ERB energy, and the first influencing factor, a first ERB energy threshold corresponding to any one of the first control frequency points includes: The first ERB energy threshold corresponding to any one of the first control frequency points is determined based on the following formula: ERB A,AB =ERB A1 +α*ERB B1 , ERB A1 =sum(s A,AB (t) 2 )*f s / BW / L, ERB B1 =sum(s B,AB (t) 2 )*f s / BW / L, BW=1.019*24.7*(4.37*f c1 / 1000+1), Among them, ERB A,AB is the first ERB energy threshold, ERB A1 For the first ERB energy, ERB B1 is the second ERB energy, α is the first impact factor, s A,AB (t) is a signal of audio data obtained by downsampling and filtering the preset audio data input to the first speaker, s B,AB (t) is a signal of audio data obtained by downsampling and filtering the preset audio data input to the second speaker, and f s is the sampling rate of the preset audio data, BW is the filter bandwidth, L is the frame length of the preset audio data after downsampling, f c1 is any first control frequency point.
15. The method according to claim 13, wherein Each of the multiple first control frequency points corresponds to a first impact factor.
16. The method according to claim 15, wherein In the case where the second speaker emits sound alone, the first impact factor is obtained based on vibration displacement information of the second speaker and vibration displacement information of the first speaker.
17. The method according to claim 16, wherein The first impact factor is determined based on the following formula: Among them, x BtoA (f c1 , n) is the first vibration displacement information of the first speaker recorded when the second speaker sounds alone, x B (f c1 ,n) is the second vibration displacement information of the second speaker recorded when the second speaker sounds alone, average(|x BtoA (f c1 ,n)|) is the average value of the absolute values of the multiple first vibration displacement information recorded, average(|x B (f c1 ,n)|) is the average value of the absolute values of the multiple second vibration displacement information recorded, f c1 is any first control frequency point, and n is the sampling point.
18. The method according to claim 10, wherein The second control strategy includes a second ERB energy threshold set and a second control frequency set, the second control frequency set includes multiple second control frequencies, the second ERB energy threshold set includes multiple second ERB energy thresholds corresponding one-to-one to the multiple second control frequencies, and the method further includes: Acquiring a resonant frequency of the second speaker, and determining the second control frequency point set based on the resonant frequency of the second speaker; At any second control frequency point, obtaining a third ERB energy of the preset audio signal played by the first speaker and a fourth ERB energy of the preset audio signal played by the second speaker; A second ERB energy threshold corresponding to any one of the second control frequency points is determined based on the third ERB energy, the fourth ERB energy, and the second influencing factor.
19. The method according to claim 18, wherein The determining, based on the third ERB energy, the fourth ERB energy, and the second influencing factor, a second ERB energy threshold corresponding to any second control frequency point includes: The second ERB energy threshold corresponding to any second control frequency point is determined based on the following formula: ERB B,AB =ERB B2 +β*ERB A2 , ERB A2 =sum(s A,AB (t) 2 )*f s / BW / L, ERB B2 =sum(s B,AB (t) 2 )*f s / BW / L, BW=1.019*24.7*(4.37*f c2 / 1000+1), Among them, ERB B,AB is the second ERB energy threshold, ERB A2 For the third ERB energy, ERB B2 is the fourth ERB energy, β is the second impact factor, s A,AB (t) is a signal of audio data obtained by downsampling and filtering the preset audio data input to the first speaker, s B,AB (t) is a signal of audio data obtained by downsampling and filtering the preset audio data input to the second speaker, and f s is the sampling rate of the preset audio data, BW is the filter bandwidth, L is the frame length of the preset audio data after downsampling, f c2 is any second control frequency point.
20. The method of claim 18, wherein Each of the plurality of second control frequency points corresponds to a second impact factor.
21. The method according to claim 20, wherein In the case where the first speaker emits sound alone, the second impact factor is obtained based on the vibration displacement information of the first speaker and the vibration displacement information of the second speaker.
22. The method according to claim 21, wherein The second impact factor is determined based on the following formula: Among them, x A (f c2 , n) is the third vibration displacement information of the first speaker recorded when the first speaker makes a sound alone, x AtoB (f c2 ,n) is the fourth vibration displacement information of the second speaker recorded when the first speaker sounds alone, average(|x A (f c2 ,n)|) is the average value of the absolute values of the third vibration displacement information recorded, average(|x AtoB (f c2 ,n)|) is the average value of the absolute values of the plurality of fourth vibration displacement information recorded, f c2 is any second control frequency point, and n is the sampling point.
23. An electronic device, characterized in that: The electronic device includes a first speaker, a second speaker, a memory and a processor; The first speaker, the second speaker, and the memory are all coupled to the processor; The memory is used to store program instructions; The processor is configured to read the program instructions stored in the memory, execute the noise suppression method according to any one of claims 1 to 9, or execute the method for constructing a noise suppression strategy according to any one of claims 10 to 22.
24. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the noise suppression method according to any one of claims 1 to 9, or implement the method for constructing a noise suppression strategy according to any one of claims 10 to 22.
25. A computer program product comprising computer-readable instructions, characterized in that: When the computer-readable instructions are executed by a processor, the noise suppression method according to any one of claims 1 to 9 is implemented, or the method for constructing a noise suppression strategy according to any one of claims 10 to 22 is implemented.
Citation Information
Patent Citations
Speaker device
CN108471579A
Noise detection method and device and electronic equipment
CN112163117A
Noise suppression method and electronic equipment
CN116546126A
Sound reproducing system
JP2009118366A
Audio signal processing method and related electronic device
WO2023000778A1