Method of processing wind noise on electronic device
By using low-power processing devices to detect wind noise in electronic devices such as smart glasses, and waking up the high-power processing device to perform noise reduction processing when wind noise is detected, the problem of wind noise affecting the wake-up signal is solved, and lower power consumption and more effective wake-up signal transmission is achieved.
Patent Information
- Application Number
- CN202510168309.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-16
AI Technical Summary
Electronic devices such as smart glasses are easily affected by wind noise when used outdoors, resulting in the wake-up sound being overwhelmed by wind noise, which in turn leads to the wake-up failure.
A low-power consumption processing device is used for wind noise detection, and a high-power consumption processing device is awakened for wind noise reduction processing when wind noise is detected. The specific method includes collecting audio data from the dual microphone, and the low-power processing device calculates the frequency domain data of the audio frame and calculates the correlation value, and wakes up the high-power processing device for noise reduction when the correlation value is less than the threshold.
It effectively reduces resource consumption for wind noise detection, reduces power consumption, and ensures effective communication of wake-up signals.
Smart Images

Figure CN120015008A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method for processing wind noise on an electronic device, a computing device, a storage medium, a program product, etc. Background Art
[0002] In outdoor usage scenarios where there is wind noise (for example, when walking on the road or riding a bicycle), electronic devices such as smart glasses may receive relatively strong wind noise. At this time, the wind noise energy may drown out the user's wake-up sound, resulting in wake-up failure.
[0003] It is therefore desirable to propose some solutions for dealing with wind noise on electronic devices. Summary of the invention
[0004] Embodiments of the present disclosure provide methods for processing wind noise on electronic devices, as well as corresponding electronic devices, non-transitory machine-readable storage media, and computer program products for executing these methods.
[0005] According to a first aspect of an embodiment of the present disclosure, a method for processing wind noise on an electronic device is provided, wherein the electronic device comprises a first microphone, a second microphone, a low-power processing device and a high-power processing device, the power consumption of the low-power processing device is less than the power consumption of the high-power processing device, and the method comprises: collecting a first channel of audio by the first microphone, and collecting a second channel of audio by the second microphone; performing wind noise detection on only one frame of audio and skipping another frame of audio within two consecutive frames, thereby completing wind noise detection by frame skipping, comprising: within a first time period, the low-power processing device calculates and obtains frequency domain data of a first frame to be processed of the first channel of audio; within a second time period adjacent to the first time period, the low-power processing device calculates and obtains The frequency domain data of the first frame to be processed of the second audio channel is obtained, the first frame to be processed of the first audio channel and the first frame to be processed of the second audio channel are time-aligned, the duration of the first time period is equal to the duration of the second time period, and is equal to the frame length of any frame to be processed; the low power consumption processing device calculates the correlation value between the frequency domain data of the first frame to be processed of the first audio channel and the frequency domain data of the first frame to be processed of the second audio channel; in response to the correlation value between the frequency domain data of the first frame to be processed of the first audio channel and the frequency domain data of the first frame to be processed of the second audio channel being less than the first correlation value threshold, the low power consumption processing device wakes up the high power consumption processing device; the high power consumption processing device performs wind noise reduction processing on the first audio channel and the second audio channel.
[0006] Optionally, the method also includes: within the first time period, the low-power processing device calculates the energy value of the first frame to be processed of the first audio within a predetermined frequency range based on the frequency domain data of the first frame to be processed of the first audio; and before the low-power processing device obtains the frequency domain data of the first frame to be processed of the second audio, the method also includes: determining that the energy value is greater than an energy threshold value.
[0007] Optionally, a preset number of frames is spaced between a first frame to be processed of the first audio channel and a previous processed frame of the first audio channel, and a preset number of frames is spaced between a first frame to be processed of the second audio channel and a previous processed frame of the second audio channel.
[0008] Optionally, after the low power consumption processing device wakes up the high power consumption processing device, the method further comprises: the low power consumption processing device enters a sleep state.
[0009] Optionally, after the low power consumption processing device wakes up the high power consumption processing device, the method further includes: the high power consumption processing device calculates and obtains the correlation value between the frequency domain data of the second frame to be processed of the first audio channel and the frequency domain data of the second frame to be processed of the second audio channel, and the second frame to be processed of the first audio channel and the second frame to be processed of the second audio channel are time-aligned; in response to the correlation value between the frequency domain data of the second frame to be processed of the first audio channel and the frequency domain data of the second frame to be processed of the second audio channel being not less than a second correlation value threshold, the high power consumption processing device wakes up the low power consumption processing device; and the high power consumption processing device enters a sleep state.
[0010] Optionally, the first correlation value threshold is smaller than the second correlation value threshold.
[0011] Optionally, the high power consumption processing device calculates and obtains the correlation value between the frequency domain data of the second frame to be processed of the first audio channel and the frequency domain data of the second frame to be processed of the second audio channel, including: within a third time period, the high power consumption processing device calculates and obtains the correlation value between the frequency domain data of the second frame to be processed of the first audio channel and the frequency domain data of the second frame to be processed of the second audio channel, and the length of the third time period is equal to the frame length of any frame to be processed.
[0012] Optionally, before the high power consumption processing device wakes up the low power consumption processing device, the method further includes: the high power consumption processing device starts a timer; during the timing of the timer, the high power consumption processing device continues to calculate and obtain the correlation value between the frequency domain data of other frames to be processed in the first audio channel and the second audio channel; and determines that the correlation values between the frequency domain data of other frames to be processed in the first audio channel and the second audio channel during the timing of the timer are not less than the second correlation value threshold.
[0013] Optionally, after the low power consumption processing device wakes up the high power consumption processing device, the method further includes: the high power consumption processing device obtains the frequency domain data of the second frame to be processed of the first audio or the second audio and calculates the energy value of the second frame to be processed within a predetermined frequency range based on the frequency domain data of the second frame to be processed; in response to the energy value being not greater than the energy threshold value, the high power consumption processing device wakes up the low power consumption processing device; and the high power consumption processing device enters a sleep state.
[0014] According to a second aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a first microphone; a second microphone; a low power consumption processing device and a high power consumption processing device, for running program instructions so that the method described in any scheme of the first aspect above is executed.
[0015] According to a third aspect of an embodiment of the present disclosure, a non-temporary machine-readable storage medium is provided, on which an executable code is stored. When the executable code is executed by a processor of an electronic device, the processor executes a method as described in any of the schemes in the first aspect above.
[0016] According to a fourth aspect of an embodiment of the present disclosure, a computer program product is provided, comprising an executable code. When the executable code is executed by a processor of an electronic device, the processor is caused to execute a method as described in any one of the schemes in the first aspect above. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The above and other objects, features and advantages of the present disclosure will become more apparent through a more detailed description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, wherein like reference numerals generally represent like components in the exemplary embodiments of the present disclosure.
[0018] Figure 1 A schematic structural diagram of smart glasses as an example of an electronic device according to at least one embodiment of the present disclosure is exemplarily shown.
[0019] Figure 2 A schematic flow chart of a method for processing wind noise on an electronic device according to at least one embodiment of the present disclosure is exemplarily shown.
[0020] Figure 3 A schematic timing diagram exemplarily illustrates a method for processing wind noise on an electronic device according to at least one embodiment of the present disclosure.
[0021] Figure 4 A schematic flow chart of a method for processing wind noise on an electronic device according to at least one embodiment of the present disclosure is exemplarily shown.
[0022] Figure 5 A structural schematic diagram of an electronic device according to at least one embodiment of the present disclosure is exemplarily shown. DETAILED DESCRIPTION
[0023] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0024] Smart glasses are a type of wearable smart product, which can include XR (Extended Reality) glasses, etc. XR glasses can include AR (Augmented Reality) glasses, VR (Virtual Reality) glasses, MR (Mixed Reality) glasses, etc.
[0025] like Figure 1 As shown, in some examples, the smart glasses can embed some hardware modules in their frames (including temples), including two microphones installed at different positions, namely a first microphone and a second microphone, and two different processing devices, namely a low-power processing device and a high-power processing device, etc. The first and second microphones can be used to collect the first audio and the second audio, respectively, and then the low-power processing device and the high-power processing device can execute relevant program instructions to process the two audios, thereby performing wind noise detection and wind noise reduction processing. Those skilled in the art should understand that Figure 1 The above is merely exemplary and not intended to limit the structure of the smart glasses of the present disclosure; the smart glasses may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently; or the components in the smart glasses may be deployed in other different locations. For example, the smart glasses may include more microphones and / or more processing devices.
[0026] For example, the low-power processing device and the high-power processing device can be two different computing chips, such as a low-power chip and a main chip. The low-power chip has weaker computing power but lower power consumption, and the main chip has stronger computing power than the low-power chip, but its power consumption is also relatively higher.
[0027] The low power consumption processing device and the high power consumption processing device may both include a processor and a memory, and the processors and memories of the two may have the same or different architectures.
[0028] For example, the processor may include one or more processing units, such as a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processor (NPU), etc., wherein different processing units may be independent devices or integrated into one or more processors. The memory may be used to store computer executable program codes, and the executable program codes include instructions. The memory may include a program storage area and a data storage area. The program storage area may store an operating system, an application required for at least one function, etc. The data storage area may store data (such as audio data, etc.) created during the use of the smart glasses.
[0029] In one example, the processor and memory may be integrated into a microcontroller unit. In one example, the two microphones and / or two processing devices may be attached to the glasses instead of being embedded therein, and so on.
[0030] As mentioned above, in some usage scenarios, there may be strong wind noise. Therefore, it is expected to detect wind noise and perform wind noise reduction when wind noise is detected. However, since it is impossible to predict when wind noise will occur, wind noise detection needs to be kept running all the time, which consumes a lot of resources of smart glasses and consumes a lot of power.
[0031] Therefore, the embodiment of the present disclosure proposes a solution for processing wind noise on an electronic device, which uses a processing device with lower power consumption in the electronic device to detect wind noise, and wakes up a processing device with higher power consumption to perform wind noise reduction processing when wind noise is detected, thereby reducing the resources consumed by wind noise detection. In addition, the solution of the embodiment of the present disclosure uses a frame skipping method to detect wind noise, which can reduce the running time of wind noise detection, and can allow the use of a processing device with very low computing power (correspondingly, low power consumption) to detect wind noise, thereby further saving power consumption.
[0032] In addition, in some embodiments, the scheme of the present disclosure can also use the energy value of one audio channel to pre-judge the wind noise. Only when the energy value is greater than a threshold will the frequency domain conversion processing of the other audio channel and the correlation value calculation of the two audio channels be performed to determine the wind noise situation, thereby further reducing the running time of the wind noise detection operation and saving power consumption.
[0033] In addition, in some embodiments, the solution of the present disclosure may also preset the number of frames between two wind noise detections to reduce the detection frequency, thereby further saving power consumption.
[0034] It should be understood that the method for processing wind noise on an electronic device provided by the embodiments of the present disclosure can be applied not only to the aforementioned smart glasses, but also to various other types of handheld devices (such as mobile phones, personal digital assistants (PDA), etc.), various types of computers (such as tablets, notebook computers, ultra-mobile personal computers (UMPC), netbooks, laptop computers, etc.), wearable devices, vehicle-mounted devices, AR / VR devices and other electronic devices. The embodiments of the present disclosure do not impose any restrictions on the specific types of application devices such as electronic devices, as long as the application device is equipped with two microphones and has the need to process wind noise.
[0035] As an example but not limitation, when the electronic device is a wearable device, the wearable device can also be a general term for wearable devices that are developed by applying wearable technology to intelligently design daily wearables, such as gloves, watches, AR head-mounted display devices, VR head-mounted display devices, or MR head-mounted display devices equipped with far-field communication modules and / or near-field communication modules.
[0036] Figure 2 A schematic flow chart of a method for processing wind noise on an electronic device according to at least one embodiment of the present disclosure is exemplarily shown. In the present disclosure, "multiple" or similar expressions refer to two or more. The electronic device includes a first microphone, a second microphone, a low-power processing device, and a high-power processing device, wherein the power consumption of the low-power processing device is less than the power consumption of the high-power processing device.
[0037] like Figure 2 As shown, in step S210, a first channel of audio is collected by a first microphone, and a second channel of audio is collected by a second microphone.
[0038] In some implementations, the first microphone and the second microphone may continuously collect the first channel of audio and the second channel of audio, and send them to the low power consumption processing device or the high power consumption processing device.
[0039] The first audio channel and the second audio channel may have the same frame length, for example, 10, 15, 20, or 30 milliseconds (ms), and their sampling rate may be 16000 Hz and word width may be 16 bits.
[0040] Next, in step S220, wind noise detection is performed on only one frame of audio within two consecutive frames and the other frame of audio is skipped, thereby completing wind noise detection by skipping frames.
[0041] Wind noise is the sound produced by flowing air hitting the surface of an object, and its sound source comes from the contact surface between the air and the object. In an electronic device with two microphones (or more microphones), since the positions of the multiple microphones are different, the sound sources of the wind noise collected by each microphone are different, and the sound correlation is very small. When the sound emitted by the same sound source, such as human voice, is collected by multiple microphones, the audio collected by these multiple microphones is related. Although there may be differences in arrival time, phase shift, amplitude change, etc., the correlation is very high.
[0042] Due to this characteristic of wind noise, the embodiment of the present disclosure can use the correlation between two channels of audio collected by dual microphones to perform wind noise detection.
[0043] The correlation between the two audio channels is usually calculated in the frequency domain, and the first audio channel and the second audio channel collected above are usually time domain signals. Therefore, the first audio channel and the second audio channel need to be transformed from the time domain to the frequency domain. For example, fast Fourier transform (FFT) can be used to obtain the frequency domain data of the two audio channels.
[0044] FFT has a large amount of calculation, and for each wind noise detection, it is necessary to perform FFT calculations on the same frame (i.e., time-aligned frame) in the two audio channels, that is, two FFT calculations are required. In order to reduce the computing power requirements for low-power processing devices, thereby reducing the power consumption of low-power processing devices, a wind noise detection can be divided into two steps, which are completed in two consecutive frames. Therefore, wind noise detection can be performed by skipping frames. For example, only one frame of audio is tested for wind noise every 2 frames, and the other frame is skipped. This will basically not affect the wind noise detection results. For example, an FFT calculation of one audio frame can be performed in the first frame, and an FFT calculation of the same frame of another audio channel can be performed in the second frame, as well as correlation calculation and threshold judgment of the two signals.
[0045] like Figure 2As shown, step S220 includes sub-steps S221 to S225. In step S221, in a first time period, the low power consumption processing device calculates and obtains the frequency domain data of the first frame to be processed of the first audio channel; then, in step S222, in a second time period adjacent to the first time period, the low power consumption processing device calculates and obtains the frequency domain data of the first frame to be processed of the second audio channel, the first frame to be processed of the first audio channel and the first frame to be processed of the second audio channel are time-aligned, the duration of the first time period is equal to the duration of the second time period, and is equal to the frame length of any frame to be processed; and, in step S223, the low power consumption processing device calculates the correlation value between the frequency domain data of the first frame to be processed of the first audio channel and the frequency domain data of the first frame to be processed of the second audio channel.
[0046] In some implementations, the correlation value calculation in step S223 is also completed in the second time period. In addition, the comparison and judgment between the correlation value and the first correlation value threshold can also be completed in the second time period.
[0047] When the correlation value is less than the first correlation value threshold, it indicates that there is wind noise, and the process goes to step S224, the low power consumption processing device wakes up the high power consumption processing device, and in step S225, the high power consumption processing device performs wind noise reduction processing on the first audio channel and the second audio channel.
[0048] In some implementations, when the correlation value is not less than the first correlation value threshold, it indicates that there is no wind noise at present, and the wind noise detection process of the next frame to be processed can be continued in the next time period, and the cycle continues until wind noise is detected.
[0049] The first time period mentioned above may correspond to a frame after the first frame to be processed, with no interval or one or more frames interval therebetween; the second time period corresponds to the frame immediately after the first time period.
[0050] In addition, since the audio energy of wind noise is mainly concentrated in the low-frequency part, the low-frequency energy can be used to assist in determining whether there is wind noise, that is, before executing the above step S222, it is first determined whether the energy value of the first frame to be processed of the first audio channel within the predetermined frequency range is greater than a preset threshold value; if not, it means that there is no wind noise at this time, and there is no need to perform subsequent calculations and judgments related to wind noise detection. This can save the amount of frequency domain conversion (such as FFT) calculations and further save power consumption.
[0051] For example, in some embodiments, within the above-mentioned first time period, after step S221, the low-power processing device can also calculate the energy value of the first frame to be processed of the first audio channel within a predetermined frequency range based on the frequency domain data of the first frame to be processed of the first audio channel; and when it is determined that the energy value is greater than the energy threshold value, perform the operation of step S222, otherwise, terminate the wind noise detection processing of the first frame to be processed, that is, do not perform step S222 and its subsequent operations, and wait until the next time period after the second time period to start the wind noise detection processing of the next frame to be processed.
[0052] In some implementations, the frequency domain data of the first to-be-processed frames of the first and second audio channels are both complex values of N frequency points obtained by FFT.
[0053] For example, the frequency domain data obtained by FFT contains 512 complex values, which means that the entire frequency domain is divided into 512 frequency points in equal proportion. Each complex value corresponds to one of the frequency points, and its amplitude (modulus) can be used to represent the energy at these 512 frequency points.
[0054] The frequency domain ranges from 0 to K, where K is determined by the sampling rate. For example, when the sampling rate is 16000 Hz, according to the Nyquist theorem, only 8 kHz signals can be sampled, that is, K = 8000. Therefore, the frequency domain of 0 to 8000 Hz is evenly divided into 512 frequency points, with an interval of 8000 / 512 = 15.625. For example, the 0th frequency point represents 0 frequency, the 1st frequency point represents 15.625 frequency, the 2nd frequency point represents 31.25 frequency, and so on.
[0055] As mentioned above, the energy of wind noise is mainly concentrated in the low frequency, and the low frequency range is a series of frequency points starting from 0. Therefore, for example, the 2nd to 15th frequency points, that is, the low frequency range of 31.25 to 234.375 Hz, can be selected as the above-mentioned predetermined frequency range. The energy value within the predetermined frequency range can be represented by the amplitude (modulus value) of the FFT value corresponding to the selected frequency points.
[0056] For example, the energy value within the predetermined frequency range may be the sum of the moduli or the sum of the squares of the moduli of the complex values of M consecutive frequency points in the first half of the N frequency points, where M <N / 2。
[0057] In addition, assume that:
[0058] The FFT value of the kth frequency point of the first frame to be processed of the first audio channel is x(k)=a1+b1i.
[0059] The FFT value of the kth frequency point of the first frame to be processed of the second audio channel is y(k)=a2+b2i.
[0060] Then the correlation value between the two audio channels at the kth frequency point is:
[0061]
[0062] The correlation value at each frequency point may be calculated as above, and then the correlation values at all frequency points may be summed or averaged to obtain a final correlation value.
[0063] That is, the correlation value between the frequency domain data of the first frame to be processed of the first audio channel and the frequency domain data of the first frame to be processed of the second audio channel in the above step S223 can be obtained as follows:
[0064] For each of the N frequency points, respectively calculate the square of the modulus of the product of the corresponding complex value of the first audio channel and the transpose of the complex value of the second audio channel, and then divide it by the square of the modulus of the corresponding complex value of the first audio channel and the square of the modulus of the complex value of the second audio channel to obtain a correlation value of each frequency point;
[0065] The correlation values of the N frequency points are summed or averaged to obtain a correlation value between the frequency domain data of the first frame to be processed of the first audio channel and the frequency domain data of the first frame to be processed of the second audio channel.
[0066] In addition, in some implementations, the number of frames between two wind noise detections may be preset, that is, the number of frames between two adjacent frames to be processed may be preset, thereby reducing the detection frequency and further saving power consumption.
[0067] For example, the first frame to be processed of the first audio channel is spaced apart from the last processed frame of the first audio channel by a preset number of frames, and the first frame to be processed of the second audio channel is spaced apart from the last processed frame of the second audio channel by the preset number of frames. The preset number of frames may be, for example, 3, 5, or a greater value.
[0068] In some embodiments, after the low power processing device wakes up the high power processing device in the aforementioned step S224, in addition to performing the wind noise reduction processing of step S225, the low power processing device and the high power processing device may also perform other processing, which may be executed in parallel or partially in parallel with step S225 according to actual conditions.
[0069] For example, after the low power consumption processing device wakes up the high power consumption processing device, the low power consumption processing device may enter a sleep state.
[0070] For example, after the low power processing device wakes up the high power processing device, the high power processing device can take over the low power processing device to continue similar wind noise detection until no wind noise is detected. The high power processing device can wake up the low power processing device and enter a sleep state.
[0071] For example, the high power consumption processing device calculates and obtains the correlation value between the frequency domain data of the second frame to be processed of the first audio channel and the frequency domain data of the second frame to be processed of the second audio channel, and the second frame to be processed of the first audio channel and the second frame to be processed of the second audio channel are time-aligned; in response to the correlation value between the frequency domain data of the second frame to be processed of the first audio channel and the frequency domain data of the second frame to be processed of the second audio channel being not less than the second correlation value threshold, the high power consumption processing device wakes up the low power consumption processing device; and the high power consumption processing device enters a sleep state.
[0072] In some examples, the second correlation value threshold is greater than the first correlation value threshold. The first correlation value threshold can be regarded as a detection threshold value for entering a state with wind noise from a state without wind noise, and the second correlation value threshold can be regarded as a detection threshold value for entering a state with wind noise from a state with wind noise, so setting the second correlation value threshold to be greater than the first correlation value threshold can prevent the current wind noise detection value from entering a ping-pong switching state when it is near the threshold value.
[0073] In some examples, since the high-power processing device has strong computing power, the frequency domain conversion calculation of the same frame of the two audio channels can be completed within the duration of one frame. For example, the high-power processing device can calculate the correlation value between the frequency domain data of the second frame to be processed of the first audio channel and the frequency domain data of the second frame to be processed of the second audio channel in a third time period, and the duration of the third time period is equal to the duration of the first and second time periods, and is equal to the frame length of any frame to be processed.
[0074] In addition, in some examples, the high power consumption processing device may also assist in determining the wind noise state through the energy value similar to the aforementioned low power consumption processing device.
[0075] For example, the high power consumption processing device calculates and obtains the frequency domain data of the second frame to be processed of the first audio or the second audio and calculates the energy value of the second frame to be processed within a predetermined frequency range based on the frequency domain data; in response to the energy value being not greater than the energy threshold value, the high power consumption processing device wakes up the low power consumption processing device; and the high power consumption processing device enters a sleep state.
[0076] The operations related to wind noise detection, such as correlation value calculation, energy value calculation, energy threshold value calculation, etc., of the high power consumption processing device may be the same as those of the aforementioned low power consumption processing device.
[0077] In addition, in some examples, a buffer time can be set before the high-power processing device wakes up the low-power processing device and enters a dormant state, thereby improving the user experience. For example, the high-power processing device can start a timer, and during the timing of the timer, the high-power processing device continues to calculate and obtain the correlation value between the frequency domain data of other frames to be processed in the first and second audio channels; determine that the correlation value between the frequency domain data of other frames to be processed in the first and second audio channels during the timing of the timer is not less than the second correlation value threshold. When the timing times out, the high-power processing device wakes up the low-power processing device and enters a dormant state.
[0078] The following will be combined Figure 3 to Figure 4 Some embodiments of the present disclosure are described in detail by taking smart glasses as an example of an electronic device.
[0079] The smart glasses contain two computing chips, namely a low-power chip and a main chip, which correspond to the aforementioned low-power processing device and high-power processing device respectively. The low-power chip has weak computing power and low power consumption. When the smart glasses are dormant, the main chip is dormant while the low-power chip is always running, which is equivalent to standby. The low-power chip runs wind noise detection. The computing power of the main chip is relatively strong compared to the low-power chip, but the power consumption is also relatively high. The main chip can run more algorithms, such as wind noise reduction processing.
[0080] like Figure 3 As shown, when the smart glasses are in a sleep state, the main chip is in sleep mode, and the low-power chip keeps running, and wind noise detection is continuously performed in step S310 until wind noise is detected.
[0081] The low power chip wind noise detection state machine can be set as shown in the following table:
[0082]
[0083]
[0084] In some examples, in order to prevent frequent switching at the critical point between wind noise and no wind noise, the results of multiple detections can be smoothed. For example, wind noise is considered to be present only after three consecutive detections, and wind noise is considered to have disappeared only after eight consecutive detections when no wind noise is detected.
[0085] For example, the wind noise detection state machine in the low-power chip may be switched from other states (Default or no-wind-noise state) to the wind-noise state only when the wind noise is determined to be present after several consecutive wind noise detection processes.
[0086] Similarly, the wind noise detection state machine in the main chip can also be set as shown in the following table:
[0087]
[0088] For example, the wind noise detection state machine in the main chip may be switched from other states (Default or wind noise state) to the no-wind noise state only when it is determined that there is no wind noise after several consecutive wind noise detection processes.
[0089] Figure 3 A specific example of wind noise detection of a low power chip in step S310 is shown in FIG. Figure 4 in the flow chart. Figure 4 The details of each step and alternative implementation methods can refer to the previous combination Figure 1 and Figure 2 What has been described will not be repeated here.
[0090] like Figure 4 As shown, in step S1, the low-power chip performs FFT on one frame (currently processed frame) of the received first audio channel to obtain frequency domain data, and accordingly calculates the energy value in the low-frequency range.
[0091] Then, in step S2, a judgment is made to determine whether the calculated energy value is greater than the energy threshold; in the case of "no", it indicates that there is no wind noise at present, so the wind noise detection of the current frame to be processed is terminated, and the process returns to step S1 to start the wind noise detection process of the next frame to be processed; in the case of "yes", the process proceeds to step S3 to continue the wind noise detection. The above steps S1 and S2 can be completed within 1 frame, and step S3 is started in the next frame.
[0092] In step S3, FFT is performed on the same frame (current frame to be processed) of the received second audio channel to obtain frequency domain data, and a correlation value between the frequency domain data obtained in step S1 is calculated accordingly.
[0093] Then, in step S4, a judgment is made to see whether the calculated correlation value is less than the first threshold; in the case of "no", it means that there is no wind noise at present, so the process returns to step S1 and starts the wind noise detection process for the next frame to be processed; in the case of "yes", the process proceeds to step S5. The above steps S3 and S4 can be completed within 1 frame time. Therefore, the wind noise detection for a frame to be processed can be completed within 2 frames.
[0094] In some cases, it may be arranged that the process proceeds to step S5 only when the result of step S4 in several consecutive wind noise detections is “yes”.
[0095] Therefore, if Figure 4 As shown, the low power consumption chip continues to perform wind noise detection in steps S1 to S4 until wind noise is detected.
[0096] When wind noise is detected, Figure 3 Step S320 and Figure 4As shown in step S5, the low-power chip starts the main chip and sends a message that there is wind noise to the main chip, so that the main chip takes over the work, and the low-power chip can prepare to enter sleep.
[0097] Then, if Figure 3 Steps S330, S340 and Figure 4 As shown in step S6, the main chip starts, prohibits itself from sleeping, instructs the low-power chip to sleep, and the main chip starts wind noise detection and noise reduction processing.
[0098] like Figure 3 As shown, as the main chip is awakened, the low-power chip enters a sleep state, and the smart glasses enter a non-sleep state.
[0099] That is, when the smart glasses are in a non-sleep state, the low-power chip is in sleep mode, while the main chip keeps running, and continues to perform wind noise detection and noise reduction processing in step S350 until no wind noise is detected.
[0100] Figure 3 A specific example of wind noise detection of the main chip in step S350 is shown in Figure 4 In the flowchart, it is similar to the wind noise detection process of the low-power chip.
[0101] like Figure 4 As shown, in step S7, the main chip performs FFT on one frame (current frame to be processed) of at least one of the first and second audio channels received to obtain frequency domain data, and calculates the energy value in the low frequency range accordingly. It should be understood that this is only for the sake of clarity. Figure 4 In step S7, the first and second audio channels as input are not drawn as in step S1, but the main chip also needs to obtain the first and second audio channels collected by the first and second microphones.
[0102] Since the main chip has strong computing power, it can complete the FFT of two audio frames within one frame, so the wind noise detection of each frame to be processed can be completed within one frame, that is, Figure 4 Steps S7 to S10 in the above method can be completed within one frame. Therefore, it is possible to select to perform FFT of one audio channel in step S7 and then perform FFT of another audio channel in step S9 as needed; or perform FFT of two audio channels in step S7 and only perform correlation value calculation in step S9. Accordingly, the energy value calculated in step S7 can be the energy value of only one audio channel, or can be the energy value of two audio channels.
[0103] Then, in step S8, a judgment is made to determine whether the calculated energy value is greater than the energy threshold; in the case of "No", it indicates that there is no wind noise at present, so the wind noise detection is terminated and the process proceeds to step S11; in the case of "Yes", the process proceeds to step S9 to continue the wind noise detection.
[0104] In step S9, the correlation value between the frequency domain data of the current frames to be processed of the two audio channels is calculated; if only the FFT of one audio channel is performed in step S7, the FFT of the other audio channel is also performed in step S9.
[0105] Then, in step S10, a judgment is made to determine whether the calculated correlation value is less than the second threshold; in the case of "yes", it indicates that there is still wind noise, so return to step S7 and start the wind noise detection process for the next frame to be processed; in the case of "no", it indicates that there is no wind noise, and proceed to step S11.
[0106] In some cases, it may be arranged that the process proceeds to step S11 only when the result of S8 is “No” or the result of S10 is “No” in several consecutive wind noise detections.
[0107] Therefore, if Figure 4 As shown, the main chip continues to perform wind noise detection from step S7 to step S10 until no wind noise is detected. This indicates that the wind noise has disappeared, and the smart glasses can enter a dormant state, that is, the low-power chip works and the main chip sleeps, thereby reducing power consumption.
[0108] When no wind noise is detected, Figure 4 As shown in step S11, the main chip starts a sleep timer, such as a 15ms timer, and enters step S12 when it times out. During this timing period, the main chip can still perform the aforementioned wind noise detection operation on the real-time collected audio; if wind noise is detected during this period, the timer and subsequent operations can be canceled, and the loop operation of wind noise detection from S7 to S10 can be returned.
[0109] like Figure 3 Steps S360, S370 and Figure 4 As shown in step S12, the main chip starts the low-power chip, and the main chip enters the sleep state. After the low-power chip is started, it enters step S380 to start wind noise detection, and then returns to step S400. Figure 3 Step S310, and Figure 4 Step S1. Thus, the smart glasses can detect wind noise in real time with low power consumption, and switch to the main chip with high computing power to process the wind noise in time after the wind noise is detected, and switch back to the low power consumption mode to detect the wind noise after the wind noise disappears.
[0110] Figure 5 A schematic structural diagram of an electronic device that can be used to implement the above-mentioned method for processing wind noise on an electronic device according to at least one embodiment of the present disclosure is shown.
[0111] See also Figure 5, the computing device 500 includes a first microphone 510, a second microphone 520, a low-power processing device 530 and a high-power processing device 540, wherein the low-power processing device 530 includes a first memory 531 and a first processor 532, and the high-power processing device 540 includes a second memory 541 and a second processor 542. The power consumption of the low-power processing device 530 is less than that of the high-power processing device 540, so that the operating frequency and / or computing power of the first processor 532 can be lower than that of the second processor 542, and the two can adopt the same or different architectures. The first memory 531 and the second memory 541 can also adopt the same or different architectures.
[0112] The first and second microphones may be any available microphones, and may respectively collect the first and second audio channels in the method according to the embodiment of the present disclosure.
[0113] The first or second processor may be a multi-core processor or may include multiple processors. In some embodiments, the first or second processor may include a general-purpose main processor and one or more special coprocessors, such as a graphics processing unit (GPU), a digital signal processor (DSP), etc. In some embodiments, the first or second processor may be implemented using a customized circuit, such as an application-specific integrated circuit (ASIC) or a field programmable gate array (FPGA).
[0114] The first or second memory may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage. Among them, ROM can store static data or instructions required by the processor or other modules of the processing device. The permanent storage device may be a readable and writable storage device. The permanent storage device may be a non-volatile storage device that does not lose the stored instructions and data even after the processing device is powered off. In some embodiments, the permanent storage device uses a large-capacity storage device (such as a magnetic or optical disk, flash memory) as a permanent storage device. In some other embodiments, the permanent storage device may be a removable storage device (such as a floppy disk, optical drive). The system memory may be a readable and writable storage device or a volatile readable and writable storage device, such as a dynamic random access memory. The system memory may store some or all instructions and data required by the processor at run time. In addition, the first or second memory may include any combination of computer-readable storage media, including various types of semiconductor memory chips (DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and disks and / or optical disks may also be used. In some embodiments, the first or second memory may include a readable and / or writable removable storage device, such as a laser disc (CD), a read-only digital versatile disc (such as a DVD-ROM, a double-layer DVD-ROM), a read-only Blu-ray disc, an ultra-density optical disc, a flash memory card (such as an SD card, a mini SD card, a Micro-SD card, etc.), a magnetic floppy disk, etc. The computer-readable storage medium does not include carrier waves and transient electronic signals transmitted wirelessly or wired.
[0115] The first and second memories store executable codes, and when the executable codes are processed by the first and second processors respectively, the first and second processors can execute the above-mentioned method for processing wind noise on an electronic device.
[0116] In addition, the method according to the present disclosure may also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing the above steps defined in the above method of the present disclosure.
[0117] Alternatively, the present disclosure may also be implemented as a non-temporary machine-readable storage medium (or computer-readable storage medium, or machine-readable storage medium) on which executable code (or computer program, or computer instruction code) is stored. When the executable code (or computer program, or computer instruction code) is executed by a processor of an electronic device (or computing device, server, etc.), the processor executes the various steps of the above-mentioned method according to the present disclosure.
[0118] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein may be implemented as electronic hardware, computer software, or a combination of both.
[0119] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the system and method according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0120] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A method for processing wind noise on an electronic device, wherein the electronic device comprises a first microphone, a second microphone, a low power consumption processing device and a high power consumption processing device, the power consumption of the low power consumption processing device is less than the power consumption of the high power consumption processing device, the method comprising: The first microphone collects a first channel of audio, and the second microphone collects a second channel of audio; Wind noise detection is performed on only one frame of audio within two consecutive frames and the other frame of audio is skipped, thereby completing wind noise detection by skipping frames, including: In a first time period, the low-power processing device calculates and obtains frequency domain data of a first frame to be processed of the first audio channel; In a second time period adjacent to the first time period, the low-power processing device calculates and obtains frequency domain data of a first frame to be processed of the second audio channel, the first frame to be processed of the first audio channel and the first frame to be processed of the second audio channel are time-aligned, and the length of the first time period is equal to the length of the second time period, and is equal to the frame length of any frame to be processed; The low power consumption processing device calculates a correlation value between the frequency domain data of the first frame to be processed of the first audio channel and the frequency domain data of the first frame to be processed of the second audio channel; In response to a correlation value between the frequency domain data of the first frame to be processed of the first audio channel and the frequency domain data of the first frame to be processed of the second audio channel being less than a first correlation value threshold, the low power consumption processing device waking up the high power consumption processing device; The high power consumption processing device performs wind noise reduction processing on the first audio channel and the second audio channel.
2. The method according to claim 1, further comprising: In the first time period, the low power consumption processing device calculates the energy value of the first frame to be processed of the first audio channel within a predetermined frequency range according to the frequency domain data of the first frame to be processed of the first audio channel; and Before the low-power processing device obtains the frequency domain data of the first frame to be processed of the second audio channel, the method further includes: It is determined that the energy value is greater than an energy threshold value.
3. The method according to claim 1, wherein: The first to-be-processed frame of the first audio channel is spaced apart from a previously processed frame of the first audio channel by a preset number of frames, and the first to-be-processed frame of the second audio channel is spaced apart from a previously processed frame of the second audio channel by the preset number of frames.
4. The method according to claim 1, after the low power consumption processing device wakes up the high power consumption processing device, further comprising: The low power consumption processing device enters a sleep state.
5. The method according to claim 1, after the low power consumption processing device wakes up the high power consumption processing device, further comprising: The high power consumption processing device calculates and obtains a correlation value between frequency domain data of a second frame to be processed of the first audio channel and frequency domain data of a second frame to be processed of the second audio channel, and the second frame to be processed of the first audio channel and the second frame to be processed of the second audio channel are time-aligned; In response to a correlation value between the frequency domain data of the second to-be-processed frame of the first audio channel and the frequency domain data of the second to-be-processed frame of the second audio channel being not less than a second correlation value threshold, the high power consumption processing device waking up the low power consumption processing device; The high power consumption processing device enters a sleep state.
6. The method according to claim 5, wherein: The first correlation value threshold is smaller than the second correlation value threshold.
7. The method according to claim 5, wherein: The high power consumption processing device calculates and obtains a correlation value between the frequency domain data of the second frame to be processed of the first audio channel and the frequency domain data of the second frame to be processed of the second audio channel, including: In a third time period, the high power consumption processing device calculates and obtains a correlation value between frequency domain data of a second frame to be processed of the first audio channel and frequency domain data of a second frame to be processed of the second audio channel, and the length of the third time period is equal to the frame length of any frame to be processed.
8. The method according to claim 5, before the high power consumption processing device wakes up the low power consumption processing device, further comprising: The high power consumption processing device starts a timer; During the timing of the timer, the high power consumption processing device continues to calculate and obtain the correlation value between the frequency domain data of other frames to be processed in the first audio channel and the second audio channel; Determine that the correlation values between the frequency domain data of other to-be-processed frames in the first audio channel and the second audio channel during the timing of the timer are not less than the second correlation value threshold.
9. The method according to claim 1, after the low power consumption processing device wakes up the high power consumption processing device, further comprising: The high power consumption processing device acquires frequency domain data of a second frame to be processed of the first audio channel or the second audio channel and calculates an energy value of the second frame to be processed within a predetermined frequency range according to the frequency domain data of the second frame to be processed; In response to the energy value being not greater than the energy threshold value, the high power consumption processing device wakes up the low power consumption processing device; The high power consumption processing device enters a sleep state.
10. An electronic device, comprising: First Microphone; Second microphone; The low power consumption processing device and the high power consumption processing device are used to run program instructions so that the method according to any one of claims 1 to 9 is executed.
Citation Information
Patent Citations
Audio frequency verification method and device, storage medium and electronic device
CN110021307A
Wind noise detection method and device, earphone and storage medium
CN115914971A
Wind noise detection method, electronic equipment and storage medium
CN117221805A
Wind noise detection method and device, wearable equipment and readable storage medium
CN117636893A
Dynamic wind detection for adaptive noise cancellation (ANC)
US20240205596A1