Single-microphone long-distance pickup method and device

A single microphone system uses analog and digital algorithms to enhance distant sound capture by dynamically adjusting signal levels and reducing noise, addressing interference issues in video monitoring systems.

CN120321534AActive Publication Date: 2025-07-15SHANGHAI WEIJING SEMICONDUCTOR CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510556046.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-15
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

In video surveillance scenarios, it is difficult to effectively deal with long-distance sound sources when picking up sounds on a single microphone. The existing technologies such as microphone arrays and analog gain control have hardware limitations and poor results.

Method used

Using a combination of analog domain and digital domain, the overall level of the audio signal is automatically adjusted through automatic level control algorithm and voice enhancement algorithm, and noise reduction and gain adjustment are performed in the digital domain to improve the sound pick-up distance and signal-to-noise ratio of a single microphone.

Benefits of technology

The effect of long-distance sound pickup under single microphone conditions is achieved, the audio signal quality is optimized, noise interference is reduced, and signal strength and signal-to-noise ratio are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321534A_ABST
    Figure CN120321534A_ABST
Patent Text Reader

Abstract

The invention discloses a single-microphone long-distance pickup method and device, and the method comprises the steps: determining a microphone bias voltage according to the microphone type, circuit and characteristics of an employed single microphone, enabling the single microphone to directly face a sound source for pickup, and obtaining a to-be-processed audio signal; automatically adjusting the overall level of the audio signal to be processed by adopting an automatic level control algorithm; converting the audio signal to be processed after the level adjustment into a digital signal; a voice enhancement algorithm and an automatic gain control algorithm are adopted to carry out noise reduction, voice enhancement and gain adjustment processing on the digital signal in sequence, that is, a method of combining an analog domain and a digital domain is used to increase the pickup distance; the quality of audio signals acquired by a single microphone is optimized by using an automatic level control (ALC) technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio and video, and specifically, to a single microphone long-distance sound pickup method and device, an electronic device, and a storage medium. Background Art

[0002] In the video surveillance application scenario, the sound source and the camera are usually at a relatively long distance, and the present invention aims to solve the problem of sound pickup of a microphone at a long distance.

[0003] The prior art usually uses a microphone array, that is, by using multiple microphones and adding an array algorithm to achieve the effect of long-distance sound pickup. However, limited by the hardware size and product cost, the microphone array scheme is not always applicable.

[0004] In the single microphone scenario, the analog gain can be increased to amplify the smaller signal, that is, the signal collected from the distant sound source. However, if there is a large signal (a closer sound source) at this time, the collected sound will be clipped.

[0005] There are also solutions that use automatic gain control technology in the digital domain. However, since the analog gain cannot be increased as much as possible, the effect is not good. Summary of the Invention

[0006] One of the purposes of the embodiments of the present invention is to provide a single microphone long-distance sound pickup method and device, an electronic device, and a storage medium, so as to provide a single microphone long-distance sound pickup scheme that combines the analog domain and the digital domain for the deficiencies of the prior art.

[0007] To solve the above technical problems, in a first aspect, a single microphone long-distance sound pickup method provided by an embodiment of the present invention includes:

[0008] Determine the microphone bias voltage according to the microphone type, circuit, and characteristics of the single microphone used, and pick up sound with the single microphone facing the sound source to obtain a to-be-processed audio signal;

[0009] Use an automatic level control algorithm to automatically adjust the overall level of the to-be-processed audio signal;

[0010] Convert the to-be-processed audio signal with the adjusted level into a digital signal;

[0011] Use a voice enhancement algorithm and an automatic gain control algorithm to perform noise reduction, voice enhancement, and gain adjustment processing on the digital signal in sequence.

[0012] Preferably, the step of using an automatic level control algorithm to automatically adjust the overall level of the to-be-processed audio signal specifically includes:

[0013] Obtain the input signal level by sampling and measuring the amplitude of the input to-be-processed audio signal;

[0014] Calculate the required target output signal level according to the preset target level;

[0015] Calculate the required gain value according to the difference between the input signal level and the target output signal level, and adjust the output signal level according to the calculated gain value;

[0016] Detect and adjust the input audio signal to be processed according to the output signal level.

[0017] Preferably, the conversion of the audio signal to be processed with the adjusted level into a digital signal specifically includes:

[0018] At the sampling period T s Take values of the audio signal to be processed x(t) to obtain a series of discrete sample values x(nT s ), where n = 0, 1, …;

[0019] Compare the amplitude A of each sampled sample value with the preset quantization levels to obtain quantization values, and the quantization levels are represented by q0, q1, q2, …, q L-1 denoted, where L is the number of quantization levels. If the amplitude A is between the nth and (n + 1)th quantization levels and A is closer to the nth quantization level, the quantization value is q n-1 , and if it is closer to the (n + 1)th quantization level, the quantization value is q n ;

[0020] Convert each quantized quantization value q i into the corresponding binary digital signal x(n) according to the preset coding rules.

[0021] Preferably, the use of the voice enhancement algorithm and the automatic gain control algorithm to sequentially perform noise reduction, voice enhancement, and gain adjustment processing on the digital signal specifically includes:

[0022] Perform frame division and windowing processing on the digital signal, and perform short-time Fourier transform on each frame of the signal to obtain the spectral characteristics of the audio signal. Among them, the formula for Fourier transform is:

[0023]

[0024] Calculate the noise power spectrum according to the spectral characteristics;

[0025] Calculate the voice presence probability according to the noise power spectrum, and update the noise power spectrum estimate according to the voice presence probability.

[0026] Preferably, the frame division processing specifically includes:

[0027] When the audio signal is x(n), where n = 0, 1, …, N - 1 and N is the signal length, the frame length is M, the frame shift is S, and the signal of the i-th frame x i (m) = x(iS + m), where m = 0, 1, …, M - 1;

[0028] The windowing process specifically includes: expressing the Hamming window function w(m) as Then the windowed signal y i (m) = x i (m)w(m) = x(iS + m)w(m), where m = 0, 1, …, M - 1.

[0029] Preferably, according to the spectral characteristics, calculate the noise power spectrum;

[0030] Take the square of the complex spectral amplitude of the audio signal of the 0-th frame as the initial noise power spectrum, that is, P(k, 0) = |Y(k, 0)| 2 , where P(k, 0) is the noise power spectrum of the k-th frequency point of the 0-th frame, and Y(k, 0) is the complex spectrum of the k-th frequency point of the 0-th frame;

[0031] If in the frequency domain, the power spectrum value after smoothing the noisy speech spectrum of each subsequent frame is b(i) represents the normalized window function, the length of the window function is 2ω + 1, and Y(k - i, l) represents the amplitude value of the short-time Fourier transform of the noisy speech in the time-frequency domain. Then, use recursive smoothing to calculate the power spectrum of the current frame, that is, P(k, l) = α P P(k, l - 1) + (1 - α P )|Y(k, l)| 2 , where α P is the smoothing factor, taking 0.9 - 0.95, P(k, l) is the noise power spectrum of the k-th frequency point of the l-th frame, P(k, l - 1) is the noise power spectrum of the k-th frequency point of the (l - 1)-th frame, and Y(k, l) is the complex spectrum of the k-th frequency point of the l-th frame.

[0032] Preferably, calculating the speech presence probability according to the noise power spectrum specifically includes:

[0033] In the starting stage, when the speech signal does not exist or the speech signal intensity is less than the preset signal intensity threshold, directly take the power spectrum of the collected audio signal as the initial estimate of the noise power spectrum;

[0034] For each subsequent frame of the audio signal, obtain the minimum value of the spectral amplitude of the current frame. When both speech and noise follow a Gaussian distribution, calculate the speech presence probability by comparing the spectral amplitude of the audio signal in the current frame with the noise power spectrum estimate.

[0035] In a second aspect, an embodiment of the present invention further provides a single microphone long-distance sound pickup device, and the device includes:

[0036] An audio acquisition unit, configured to determine a microphone bias voltage according to the microphone type, circuit and characteristics of the adopted single microphone, pick up sound with the single microphone facing the sound source, and obtain a to-be-processed audio signal;

[0037] A level control unit, configured to automatically adjust the overall level of the to-be-processed audio signal by using an automatic level control algorithm;

[0038] An analog-to-digital conversion unit, configured to convert the to-be-processed audio signal with adjusted level into a digital signal;

[0039] A digital domain signal processing unit, which adopts a voice enhancement algorithm and an automatic gain control algorithm to perform noise reduction, voice enhancement and gain adjustment processing on the digital signal in sequence.

[0040] In a third aspect, an embodiment of the present invention further provides an electronic device, which includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is configured to call the program instructions to execute the method as described above.

[0041] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor or a calculator, the processor is caused to execute the method as described above.

[0042] Compared with the prior art, a single microphone long-distance sound pickup method and device provided by an embodiment of the present invention at least have the following beneficial effects:

[0043] In the embodiment of the present invention, a microphone bias voltage is determined according to the microphone type, circuit and characteristics of the adopted single microphone, sound is picked up with the single microphone facing the sound source to obtain a to-be-processed audio signal; an automatic level control algorithm is used to automatically adjust the overall level of the to-be-processed audio signal; the to-be-processed audio signal with adjusted level is converted into a digital signal; a voice enhancement algorithm and an automatic gain control algorithm are adopted to perform noise reduction, voice enhancement and gain adjustment processing on the digital signal in sequence, that is, a method combining the analog domain and the digital domain is used to improve the sound pickup distance; the automatic level control (ALC) technology is used to optimize the quality of the audio signal collected by the single microphone. Description of the Drawings

[0044] The above characteristics, technical features, advantages and their implementation manners of the present invention will be further described below in a clear and understandable manner in combination with the drawings in the preferred embodiments.

[0045] Figure 1 Schematic diagram of the process of a single - microphone long - distance sound pickup method according to an embodiment of the present invention;

[0046] Figure 2 Schematic diagram of a single - microphone long - distance sound pickup device according to an embodiment of the present invention;

[0047] Figure 3 Schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed implementation manners

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will describe the specific implementation manners of the present invention with reference to the accompanying drawings. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings, and other implementation manners can also be obtained.

[0049] To make the drawings concise, only the parts related to the invention are schematically shown in each drawing, and they do not represent the actual structure of the product. In addition, to make the drawings concise and easy to understand, in some drawings, components with the same structure or function are only schematically shown for one of them, or only one of them is marked. In this article, "one" not only means "only this one", but also means "more than one" situation.

[0050] The following mainly takes some specific embodiments as examples to detail the implementation manners of the technical solutions of the present invention.

[0051] As Figure 1 shown, in order to achieve the invention purpose of the present invention, a single - microphone long - distance sound pickup method provided by an embodiment of the present invention includes:

[0052] Determine the microphone bias voltage according to the microphone type, circuit and characteristics of the single microphone, and pick up sound with the single microphone facing the sound source to obtain the audio signal to be processed;

[0053] Adopt an automatic level control algorithm to automatically adjust the overall level of the audio signal to be processed;

[0054] Convert the audio signal to be processed with the adjusted level into a digital signal;

[0055] Adopt a voice enhancement algorithm and an automatic gain control algorithm to perform noise reduction, voice enhancement and gain adjustment processing on the digital signal in sequence.

[0056] The embodiment of the present invention adopts a method combining the analog domain and the digital domain to improve the sound collection distance.

[0057] The analog domain introduces automatic level control to automatically and dynamically adjust the signal amplitude. When the signal is large, a smaller gain is given to prevent clipping; when the signal is small, a larger gain is given to amplify the small signal. Eventually, the output level of the analog domain reaches a relatively balanced level.

[0058] The digital domain introduces noise reduction and automatic gain control.

[0059] Since the output of the analog domain amplifies small signals, but the problem is that noise is also amplified. Therefore, a noise reduction module is first introduced in the digital domain to reduce the noise amplification problem caused by automatic level control.

[0060] Finally, automatic gain control is added. The reason is that the amplification factor of the above-mentioned automatic level control is usually limited. Adding automatic gain control with voice detection further increases the signal amplitude and improves the signal-to-noise ratio.

[0061] The specific implementation steps of the embodiments of the present invention are as follows:

[0062] Pickup step:

[0063] Use a single microphone and determine an appropriate microphone bias voltage according to the type, circuit and characteristics of the microphone. Generally speaking, the bias voltage of an electret microphone is usually between 2V and 5V. The specific value can refer to the microphone's specification manual. Point the single microphone at the sound source to pick up the audio signal to be processed;

[0064] Analog domain processing step:

[0065] Adopt the automatic level control (ALC) algorithm to automatically adjust the overall level of the audio signal so that it remains constant within a certain range to prevent signal clipping and distortion.

[0066] The analog domain processing includes the following steps:

[0067] Detect the input audio signal level: By sampling and measuring the amplitude of the input signal, obtain the current level of the audio signal;

[0068] Output signal level calculation: Calculate the required output signal level according to the set target level.

[0069] Gain control: Calculate the required gain value according to the difference between the input signal level and the target output signal level, and adjust the output signal level according to the calculated gain value.

[0070] Feedback control: Detect and adjust the input signal according to the level of the output signal.

[0071] Analog-to-digital conversion step: mainly includes three steps of sampling, quantization and encoding.

[0072] Sampling: At regular time intervals Ts (Sampling period) samples the analog signal x(t) to obtain a series of discrete sampling values x(nT s ), where n = 0, 1, ….

[0073] Quantization: Compares the amplitude A of each sampled sample value with a pre - defined quantization level. The quantization levels are usually represented by q0, q1, q2, …, q L-1 , where L is the number of quantization levels. If the amplitude A is between the nth and (n + 1)th quantization levels, and if A is closer to the nth quantization level, it is quantized to q n-1 , and if it is closer to the (n + 1)th quantization level, it is quantized to q n .

[0074] Coding: Converts each quantized value q i into the corresponding binary code according to a certain coding rule. For example, if the quantized value is 12345, its 16 - bit binary conversion is 0011000000111001.

[0075] Digital domain signal processing steps:

[0076] Adopts a voice enhancement algorithm and an automatic gain control algorithm to perform noise reduction, voice enhancement and gain adjustment on the collected audio signal. The specific steps are as follows:

[0077] Signal pre - processing: Performs frame division and windowing on the collected audio signal, and then performs a short - time Fourier transform on each frame of the signal to obtain the spectral characteristics of the audio signal.

[0078] The specific steps of frame division are as follows: Assume the audio signal is x(n), n = 0, 1, …, N - 1, where N is the signal length. The frame length is M, the frame shift is S, and the signal of the i - th frame x i (m)=x(iS + m), m = 0, 1, …, M - 1.

[0079] Taking the Hamming window as an example, the specific steps of windowing are as follows: The expression of the Hamming window function w(m) is The windowed signal y i (m)=x i (m)w(m)=x(iS + m)w(m), m = 0, 1, …, M - 1.

[0080] The formula for Fourier transform is:

[0081]

[0082] Calculates the noise power spectrum according to the said spectral characteristics;

[0083] Calculate the voice presence probability according to the noise power spectrum, and update the noise power spectrum estimate according to the voice presence probability.

[0084] Among them, for the noise power spectrum estimate, at the initial stage, assume that the voice signal does not exist or the voice signal strength is very weak, and directly use the power spectrum of the collected audio signal as the initial estimate of the noise power spectrum. For each subsequent frame of the audio signal, calculate the minimum value of the spectral amplitude. Assume that both the voice and the noise follow a Gaussian distribution. Calculate the voice presence probability by comparing the spectral amplitude of the audio signal in the current frame with the noise power spectrum estimate. Finally, update the noise power spectrum estimate according to the voice presence probability. If the voice presence probability is low, use the minimum value to update the noise power spectrum estimate. If the voice presence probability is high, do not update the noise power spectrum estimate.

[0085] Among them, calculating the noise power spectrum according to the spectral characteristics specifically includes:

[0086] Take the square of the complex spectral amplitude of the audio signal in the 0th frame as the initial noise power spectrum, that is, P(k,0) = |Y(k,0)| 2 , where P(k,0) is the noise power spectrum at the kth frequency point in the 0th frame, and Y(k,0) is the complex spectrum at the kth frequency point in the 0th frame;

[0087] If in the frequency domain, the power spectrum value after smoothing the noisy speech spectrum of each subsequent frame is b(i) represents the normalized window function, the length of the window function is 2ω + 1, and Y(k - i, l) represents the amplitude value of the short-time Fourier transform of the noisy speech in the time-frequency domain. Then use recursive smoothing to calculate the power spectrum of the current frame, that is, P(k, l) = α P P(k, l - 1) + (1 - α P )|Y(k, l)| 2 , where α P is the smoothing factor, taking 0.9 - 0.95, P(k, l) is the noise power spectrum at the kth frequency point in the lth frame, P(k, l - 1) is the noise power spectrum at the kth frequency point in the (l - 1)th frame, and Y(k, l) is the complex spectrum at the kth frequency point in the lth frame.

[0088] Among them, for the initial noise power spectrum:

[0089] Take the square of the complex spectral amplitude of the audio signal in the 0th frame as the initial noise power spectrum, that is, P(k,0) = |Y(k,0)| 2 , where P(k,0) is the noise power spectrum at the kth frequency point in the 0th frame, and Y(k,0) is the complex spectrum at the kth frequency point in the 0th frame.

[0090] Noise power spectrum of other frames:

[0091] The power spectral values after smoothing each frame of noisy speech spectrum in the frequency domain are b(i) represents the normalized window function with a window length of 2ω + 1, and Y(k - i, l) represents the amplitude value of the noisy speech after short-time Fourier transform in the time-frequency domain.

[0092] Next, the power spectrum of the current frame is calculated using a recursive smoothing method, i.e., P(k, l) = α P P(k, l - 1) + (1 - α P )|Y(k, l)| 2 , where α P is the smoothing factor, usually taken as 0.9 - 0.95, P(k, l) is the noise power spectrum at the k-th frequency point of the l-th frame, P(k, l - 1) is the noise power spectrum at the k-th frequency point of the (l - 1)-th frame, and Y(k, l) is the complex spectrum at the k-th frequency point of the l-th frame.

[0093] Minimum value tracking: Search for the minimum value within the time window as the noise candidate, N is the length of the search window.

[0094] Define the variable B min = 1.66 is the bias compensation factor for estimating the minimum value of the noise power spectrum.

[0095] Define I(k, l) = 1 indicates the absence of speech, and I(k, l) = 0 indicates the presence of speech;

[0096] where the empirical values are γ0 = 4.6 and ζ0 = 1.67.

[0097] The smoothed power spectral values are obtained by performing quadratic smoothing on the power spectra of different frequencies:

[0098]

[0099] The quadratic smoothed power spectral values obtained by using the first-order recursive smoothing method in the time domain:

[0100]

[0101] Track the minimum value of the quadratic smoothed power spectrum :

[0102]

[0103] Define two more variables:

[0104]

[0105] Calculate the estimated value of the speech presence probability ζ0 = 1.67, γ1 = 3

[0106]

[0107] Update the noise power spectrum estimate according to the voice presence probability, specifically including:

[0108] Perform recursive update of the noise power spectrum estimate by combining the estimated value of the voice presence probability:

[0109]

[0110] where α d is the noise smoothing factor, usually 0.85 - 0.95, is the updated result of the noise power spectrum estimate at the k-th frequency bin of the l-th frame.

[0111] In order to achieve the effect of noise reduction to eliminate noise, the following steps are required:

[0112] Voice enhancement: According to the updated result of the noise power spectrum estimate, calculate the posterior signal-to-noise ratio, then combine the updated noise power spectrum estimate and the posterior signal-to-noise ratio to calculate the prior signal-to-noise ratio, and then calculate the gain function according to the minimum mean square error criterion. Finally, add the logarithmic spectral amplitude of the audio signal and the logarithm of the gain function to obtain the adjusted logarithmic spectral amplitude.

[0113] Among them, the formula for calculating the posterior signal-to-noise ratio:

[0114] The formula for calculating the prior signal-to-noise ratio:

[0115]

[0116] The formula for the gain function when voice is present: where

[0117] Calculate the voice presence value according to the estimated value of the voice presence probability:

[0118] Obtain the spectrum of the enhanced audio signal:

[0119] Signal restoration: Through the inverse short-time Fourier transform, convert the adjusted logarithmic spectral amplitude back to the time-domain signal to obtain the enhanced audio signal. Among them, the inverse short-time Fourier transform is: Assume X(k, l) is the frequency-domain signal after the short-time Fourier transform, where k represents the frequency index, l represents the time-frame index, and the window function is ω(l), then the restored time-domain signal is N is the number of Fourier transform points.

[0120] Next, automatic gain control is performed to stabilize the voice volume, including the following steps:

[0121] First, signal detection and analysis are performed. The enhanced audio signal is monitored and analyzed in real time to obtain the relevant characteristics of the signal. The volume level is judged by calculating the amplitude of the audio signal samples, and the voice activity detection (VAD) algorithm is used to determine whether there is voice present.

[0122] The voice activity detection algorithm is as follows: First, calculate the energy of the audio signal within a certain time window where x(i) is the value of the audio signal at the i-th sampling point, n represents the serial number of the current frame, N is the number of sampling points corresponding to the frame length, n usually starts from N - 1, representing the last sampling point of the first complete frame of data, and subsequent ones increase at intervals of the number of points m corresponding to the frame shift, that is, n = N - 1, N - 1 + m, N - 1 + 2m, …

[0123] Then calculate the zero-crossing rate, ZCR(n) represents the number of zero crossings of the audio signal within a certain time, that is, the zero-crossing rate, sgn(·) is the sign function. If the energy is higher than a certain threshold T E and the zero-crossing rate exceeds the threshold, it can be judged that there is voice activity. Among them, both the energy threshold and the zero-crossing rate threshold need to be determined according to the specific application scenario, audio data characteristics, etc.

[0124] Gain calculation: According to the target volume and the volume of the current signal, calculate the required gain value, and amplify or attenuate the input audio signal.

[0125] Let the input audio signal be x(n), the output audio signal be y(n), and the gain be G(n).

[0126] In the voice segment, calculate the gain according to the target volume T and the current signal energy E(n), such as then y(n) = G(n) × x(n).

[0127] In the non-voice segment, a fixed small gain G can be set min , generally between 0 dB and 6 dB, and the specific value varies according to different application scenarios.

[0128] Such as Figure 2 As shown, an embodiment of the present invention also provides a single microphone long-distance sound pickup device, and the device includes:

[0129] An audio acquisition unit, configured to determine a microphone bias voltage according to the microphone type, circuit and characteristics of the single microphone used, pick up sound from the sound source with the single microphone to obtain an audio signal to be processed; a level control unit, configured to automatically adjust the overall level of the audio signal to be processed by using an automatic level control algorithm; an analog-to-digital conversion unit, configured to convert the audio signal to be processed with adjusted level into a digital signal; a digital domain signal processing unit, which uses a voice enhancement algorithm and an automatic gain control algorithm to perform noise reduction, voice enhancement and gain adjustment processing on the digital signal in sequence.

[0130] Here, the implementation manner of the device is the same as that of the method described above, and will not be elaborated one by one.

[0131] The embodiment of the present invention uses a method combining the analog domain and the digital domain to improve the pickup distance, and uses the automatic level control (ALC) technology to optimize the quality of the audio signal collected by the single microphone. Therefore, under sound sources of different intensities, it has obvious advantages over the existing solutions. For example, when comparing two other solutions, the test results are shown in Table 1, indicating that the embodiment of the present invention has obvious advantages.

[0132] It should be noted that for Competitor 2, when the sound source is strong, the noise signal is suppressed to 0 level through noise reduction, so the SNR is the maximum (represented by 100 dB). However, this solution does not sound good and is generally not recommended. When the signal is all 0, -100 dBFS is used to indicate that there is no signal at all.

[0133] (Table 1)

[0134]

[0135] In a third aspect, the embodiment of the present invention further provides an electronic device, which includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor or calculator is configured to call the program instructions to execute the method as described above.

[0136] In a fourth aspect, the embodiment of the present invention further provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions, and when the program instructions are executed by a processor or calculator, the processor or calculator is caused to execute the method as described above.

[0137] As Figure 3As shown, an electronic device provided by an embodiment of the present application, the electronic device 1000 includes a processor or calculator (not shown in the figure) 1001 and a memory 1002. The processor or calculator 1001 and the memory 1002 can be interconnected through a communication bus 1003. The communication bus 1003 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus 1003 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the memory 1002 is used to store a computer program, the computer program includes program instructions, and the processor 1001 is configured to call the program instructions. The above program includes those for executing some or all of the steps in the foregoing included methods.

[0138] The processor 1001 can be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the above program.

[0139] The memory 1002 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or can also be an Electrically Erasable Programmable Read-Only Memory (EEPROM), a Compact Disc Read-Only Memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory can exist independently and be connected to the processor through a bus. The memory can also be integrated with the processor.

[0140] The electronic device 1000 can further include a communication module 1004 and a display 1005. The communication module 1004 can be communicatively connected to an optical tracking device. The communication module 1004 can be a wireless communication module (such as a WiFi module, a Bluetooth module, etc.) or a wired communication module.

[0141] In addition, the electronic device 1000 may further include general components such as a communication interface (e.g., a USB interface, a microphone interface, etc.), an antenna, etc., which will not be elaborated herein.

[0142] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps may be adopted in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0143] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not elaborated in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0144] In the several embodiments provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical or other forms.

[0145] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0146] In addition, the functional units in the respective embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software program modules.

[0147] When the integrated unit is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned memory includes various media that can store program codes, such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs.

[0148] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable memory, and the memory can include: flash drives, read-only memories, random access memories, magnetic disks, or optical discs, etc.

[0149] The above has introduced the embodiments of this application in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.

[0150] It should be noted that the above embodiments can be freely combined according to needs. The above are only the preferred implementation manners of the present invention. It should be pointed out that for those of ordinary skill in the technical field, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A single microphone long-distance sound pickup method, characterized in that, The method includes: Determining a microphone bias voltage according to the microphone type, circuit, and characteristics of the single microphone used, and picking up sound from the sound source with the single microphone to obtain an audio signal to be processed; Adopting an automatic level control algorithm to automatically adjust the overall level of the audio signal to be processed; Converting the audio signal to be processed with adjusted level into a digital signal; Adopting a voice enhancement algorithm and an automatic gain control algorithm to perform noise reduction, voice enhancement, and gain adjustment processing on the digital signal in sequence.

2. The single microphone long-distance sound pickup method according to claim 1, wherein The adopting of the automatic level control algorithm to automatically adjust the overall level of the audio signal to be processed specifically includes: Obtaining an input signal level by sampling and measuring the amplitude of the input audio signal to be processed; Calculating a required target output signal level according to a preset target level; Calculating a required gain value according to the difference between the input signal level and the target output signal level, and adjusting the output signal level according to the calculated gain value; Detecting and adjusting the input audio signal to be processed according to the output signal level.

3. The single microphone long-distance sound pickup method according to claim 1, wherein The converting of the audio signal to be processed with adjusted level into a digital signal specifically includes: At sampling period T s Take values of the audio signal x(t) to be processed to obtain a series of discrete sample values x(nT s ), where n = 0, 1, …; Compare the amplitude A of each sampled sample value with a pre-specified quantization level to obtain a quantization value, where the quantization levels are represented by q0, q1, q2, …, q L-1 , where L is the number of quantization levels. If the amplitude A is between the n-th and (n + 1)-th quantization levels and A is closer to the n-th quantization level, the quantization value is q n-1 , and if it is closer to the (n + 1)-th quantization level, the quantization value is q n ; Each quantized value q after quantization i is converted into a corresponding binary digital signal x(n) according to a preset coding rule.

4. The single microphone long-distance sound pickup method according to claim 1, wherein The adopting of the voice enhancement algorithm and the automatic gain control algorithm to perform noise reduction, voice enhancement, and gain adjustment processing on the digital signal in sequence specifically includes: Performing frame division and windowing processing on the digital signal, and performing short-time Fourier transform on each frame of the signal to obtain the spectral characteristics of the audio signal, where the formula for Fourier transform is: Calculating a noise power spectrum according to the spectral characteristics; Calculating a voice presence probability according to the noise power spectrum, and updating the noise power spectrum estimate according to the voice presence probability.

5. The single microphone long-distance sound pickup method according to claim 4, characterized in that The frame division processing specifically includes: When the audio signal is x(n), where n = 0, 1, …, N-1 and N is the signal length, the frame length is M, the frame shift is S, and the signal of the i-th frame x i (m) = x(iS + m), where m = 0, 1, …, M-1; The windowing process specifically includes: expressing the Hamming window function w(m) as Then the windowed signal y i (m) = x i (m)w(m) = x(iS + m)w(m), m = 0, 1, …, M - 1.

6. The single microphone long-distance sound pickup method according to claim 4, characterized in that, The calculating of the noise power spectrum according to the spectral characteristics specifically includes: Take the square of the complex spectral magnitude of the audio signal in the 0th frame as the initial noise power spectrum, i.e., P(k, 0) = |Y(k, 0)| 2 , where P(k, 0) is the noise power spectrum at the kth frequency point in the 0th frame, and Y(k, 0) is the complex spectrum at the kth frequency point in the 0th frame; If in the frequency domain, the power spectrum value after smoothing the noisy speech spectrum of each subsequent frame is b(i) represents the normalized window function with a window length of 2ω + 1, and Y(k - i, l) represents the magnitude value of the short-time Fourier transform of the noisy speech in the time-frequency domain. Then, the power spectrum of the current frame is calculated by recursive smoothing, that is, P(k, l) = α P P(k, l - 1)+(1 - α P )|Y(k, l)| 2 , where α P is the smoothing factor, taking values from 0.9 to 0.

95. P(k, l) is the noise power spectrum at the k-th frequency point of the l-th frame, P(k, l - 1) is the noise power spectrum at the k-th frequency point of the (l - 1)-th frame, and Y(k, l) is the complex spectrum at the k-th frequency point of the l-th frame.

7. The single microphone long-distance sound pickup method according to claim 4, wherein The calculating of the voice presence probability according to the noise power spectrum specifically includes: In the starting stage, when the voice signal does not exist or the voice signal strength is less than a preset signal strength threshold, directly using the power spectrum of the collected audio signal as the initial estimate of the noise power spectrum; For each subsequent frame of the audio signal, obtaining the minimum value of the spectral amplitude of the current frame, and calculating the voice presence probability by comparing the spectral amplitude of the audio signal in the current frame with the noise power spectrum estimate when both the voice and the noise follow a Gaussian distribution.

8. A single microphone long-distance sound pickup device, characterized in that, The device includes: An audio acquisition unit, configured to determine a microphone bias voltage according to the microphone type, circuit, and characteristics of the single microphone used, and pick up sound from the sound source with the single microphone to obtain an audio signal to be processed; A level control unit, configured to adopt an automatic level control algorithm to automatically adjust the overall level of the audio signal to be processed; An analog-to-digital conversion unit, configured to convert the audio signal to be processed with adjusted level into a digital signal; A digital domain signal processing unit, adopting a voice enhancement algorithm and an automatic gain control algorithm to perform noise reduction, voice enhancement, and gain adjustment processing on the digital signal in sequence.

9. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program includes program instructions which, when executed by a processor, cause the processor or calculator to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Dual-microphone directional pickup method and device with adjustable pickup angle range

    CN113660578A

  • Single microphone noise suppression method and device

    CN113870884A

  • Double-microphone speech enhancement method and system

    CN115831145A

  • Noise reduction method and remote directional pickup microphone

    CN118382030A

  • Voice pickup method and device, electronic equipment and computer readable storage medium

    CN118918917A