A method and related device for dynamically adjusting the volume of a microphone

CN120018003BActive Publication Date: 2025-06-27SHENZHEN FENGHUO HONGSHENG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510468184.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-06-27
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The poor real-time adjustment of audio signals in traditional microphone systems leads to insufficient robustness and adaptability of voice acquisition. Especially when sound source displacement or environmental noise fluctuates, overload distortion or signal-to-noise ratio deterioration is prone to problems.

Method used

By continuously obtaining the audio signal received by the microphone within a preset time range, the stability is determined. If the condition is met, the microphone output volume will be dynamically adjusted based on the distance detection result and the gender judgment result; if the condition is not met, the adjustment will be made only based on the gender judgment result. This method uses features such as Mel frequency cepspectral coefficient, spectral center of mass and spectral contrast, and combines a hybrid model of convolutional neural network, residual convolutional network and recurrent neural network to make gender judgments.

Benefits of technology

It realizes rapid and accurate adjustment of microphone volume, adapts to different speaking distances and genders, reduces manual operations, provides clear and consistent audio output, reduces usage costs, and improves the robustness and adaptability of the voice acquisition system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120018003B_ABST
    Figure CN120018003B_ABST
Patent Text Reader

Abstract

This application belongs to a method for automatic volume adjustment. Aiming at the problems of poor real-time adjustment, poor robustness and adaptability of voice acquisition in the traditional microphone system, a method and related device for dynamically adjusting the microphone volume are provided. First, the stationarity of the audio signal within a preset time range is judged. If the audio signal is stationary, the distance between the source of the audio signal and the microphone is determined by detecting the distance of the audio signal. According to the Mel Frequency Cepstral Coefficients, spectral centroid and spectral contrast extracted from the audio signal, a first gender judgment result and a second gender judgment result are obtained. Among them, the second gender judgment result is obtained by combining a hybrid model structure composed of a convolutional neural network, a residual convolutional network and a recurrent neural network, further verifying the first gender judgment result and the second gender judgment result with each other to improve the accuracy of the gender judgment result. If the audio signal is not stationary, only the final gender judgment result is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to a method for automatic volume adjustment, and particularly relates to a method and related device for dynamically adjusting the volume of a microphone. Background Art

[0002] In traditional microphone systems, the acquisition quality of audio signals is often limited by the manual gain adjustment mechanism. Existing technologies usually require users to manually adjust the input sensitivity parameters through physical knobs or software interfaces according to the change of the sound source distance or the fluctuation of environmental noise. This method has significant real-time defects and technical limitations. Specifically, when the sound source moves or the environmental noise level changes dynamically, the fixed gain setting will cause the voice signal to be overloaded and distorted or the signal-to-noise ratio to deteriorate; in the scenario of sudden interference noise, the operation delay will cause the voice dynamic range compression to fail; in addition, due to the lack of considering the physiological differences of different gender speakers in terms of fundamental frequency range and sound pressure level characteristics, the existing system lacks a differential gain compensation mechanism based on bioacoustic characteristics, resulting in a significant reduction in the speech intelligibility of specific user groups, seriously affecting the robustness and adaptability of the voice acquisition system. Summary of the Invention

[0003] Aiming at the technical problems of poor real-time adjustment and poor robustness and adaptability of voice acquisition in the traditional microphone system for adjusting audio signals, this application provides a method and related device for dynamically adjusting the volume of a microphone.

[0004] To achieve the above object, this application adopts the following technical solutions:

[0005] In the first aspect, this application proposes a method for dynamically adjusting the volume of a microphone, including:

[0006] Continuously obtain the audio signal received by the microphone within a preset time range, and judge whether the smoothness of the audio signal meets the preset requirements. If it meets, dynamically adjust the output volume of the microphone according to the distance result and the final gender judgment result; otherwise, dynamically adjust the output volume of the microphone according to the final gender judgment result. The methods for obtaining the distance result and the final gender judgment result include:

[0007] Perform distance detection on the audio signal to determine the distance between the audio signal source and the microphone, and obtain a distance detection result;

[0008] Extract the Mel-frequency cepstral coefficients, spectral centroid, and spectral contrast from the audio signal respectively, and determine the first gender judgment result of the person who emits the audio signal by combining the Mel-frequency cepstral coefficients, spectral centroid, and spectral contrast; input the Mel-frequency cepstral coefficients, spectral centroid, and spectral contrast into the gender judgment model to obtain the second gender judgment result of the person who emits the audio signal; the gender judgment model adopts a hybrid model structure of a convolutional neural network, a residual convolutional network, and a recurrent neural network; if the first gender judgment result and the second gender judgment result are consistent, use the consistent result as the final gender judgment result.

[0009] Further, a method for distance detection of the audio signal includes:

[0010] Apply a window function to the audio signal to obtain a first preprocessed audio signal;

[0011] Perform frame processing on the first preprocessed audio signal to obtain a set of first-sequence frame signals;

[0012] Calculate the power of all the first-sequence frame signals respectively to obtain a set of frame powers correspondingly;

[0013] Calculate the average power of the set of frame powers;

[0014] Convert the average power to decibels and take the absolute value, and determine the distance between the audio signal source and the microphone according to the absolute value.

[0015] Further, the method for extracting the Mel-frequency cepstral coefficients includes:

[0016] Pre-emphasize the audio signal to obtain an emphasized signal;

[0017] Perform frame segmentation on the emphasized signal to obtain segmented signals;

[0018] Apply a window function to the segmented signals to obtain a second preprocessed audio signal;

[0019] Perform frame processing on the second preprocessed audio signal to obtain a set of second-sequence frame signals;

[0020] Perform a fast Fourier transform on the set of second-sequence frame signals to convert the second-sequence frame signals into frequency-domain signals, denoted as first frequency-domain signals;

[0021] Pass the first frequency-domain signals through a Mel filter bank to obtain Mel-frequency domain signals;

[0022] Perform logarithmic compression on the Mel-frequency domain signals to obtain compressed signals;

[0023] Perform a discrete cosine transform on the compressed signals to extract the Mel-frequency cepstral coefficients.

[0024] Further, the calculation method of the spectral centroid includes:

[0025] Applying a window function to the audio signal to obtain a first preprocessed audio signal;

[0026] Performing frame processing on the first preprocessed audio signal to obtain a set of first-sequence frame signals;

[0027] Performing a fast Fourier transform on the set of first-sequence frame signals to convert the first-sequence frame signals into frequency-domain signals, denoted as second frequency-domain signals;

[0028] Calculating the energy-weighted average value of all frequencies in the second frequency-domain signal to obtain the spectral centroid.

[0029] Further, the calculation method of the spectral contrast includes:

[0030] Applying a window function to the audio signal to obtain a first preprocessed audio signal;

[0031] Performing frame processing on the first preprocessed audio signal to obtain a set of first-sequence frame signals;

[0032] Performing a fast Fourier transform on the set of first-sequence frame signals to convert the first-sequence frame signals into frequency-domain signals, denoted as second frequency-domain signals;

[0033] Dividing the second frequency-domain signal into multiple non-overlapping frequency bands to obtain multiple frequency bands;

[0034] Calculating the energy of each frequency band respectively, and obtaining the energy peak value and energy valley value of each frequency band;

[0035] Calculating the spectral contrast of each frequency band according to the energy peak value and energy valley value of each frequency band.

[0036] Further, the method for determining the first gender judgment result of the person who emits the audio signal by combining the Mel frequency cepstral coefficients, the spectral centroid, and the spectral contrast includes:

[0037] Aligning the Mel frequency cepstral coefficients, the spectral centroid, and the spectral contrast to the same number of frames;

[0038] Concatenating the aligned Mel frequency cepstral coefficients, the spectral centroid, and the spectral contrast into a feature matrix, denoted as the first concatenated feature matrix;

[0039] Performing normalization processing on the first concatenated feature matrix;

[0040] Assigning weights to different features in the normalized first concatenated feature matrix to obtain a weighted feature matrix;

[0041] Compare the weighted feature matrix with a preset constant matrix. If it is greater than or equal to the constant matrix, the first gender judgment result is female; otherwise, the first gender judgment result is male.

[0042] Further, the convolutional neural network adopted in the gender judgment model includes a first convolutional neural network and a second convolutional neural network;

[0043] The method of inputting the Mel-frequency cepstral coefficients, spectral centroid, and spectral contrast into the gender judgment model to obtain the second gender judgment result of the person who emits the audio signal includes:

[0044] Concatenate the Mel-frequency cepstral coefficients, spectral centroid, and spectral contrast by frame to form a second concatenated feature matrix;

[0045] Input the second concatenated feature matrix into the gender judgment model. In the gender judgment model, use the first convolutional neural network to extract the first local features from the Mel-frequency cepstral coefficients; use the second convolutional neural network to extract statistical features from the spectral centroid and spectral contrast; use the residual convolutional network to process the speech spectrogram to obtain the second local features; use the recurrent neural network to capture the dependencies in the time series to obtain multiple third local features; the speech spectrogram includes a first frequency-domain signal and a second frequency-domain signal; the time series is obtained by arranging the first local features, statistical features, and second local features in chronological order;

[0046] Perform weighted averaging on the multiple third local features to obtain the second gender judgment result.

[0047] In a second aspect, the present application proposes a microphone volume dynamic adjustment system, including:

[0048] A signal acquisition module, configured to continuously acquire the audio signal received by the microphone within a preset time range, determine whether the smoothness of the audio signal meets the preset requirements. If it meets, dynamically adjust the output volume of the microphone according to the distance result and the final gender judgment result; otherwise, dynamically adjust the output volume of the microphone according to the final gender judgment result; the distance result and the final gender judgment result are obtained through the outputs of the distance detection module and the gender judgment module;

[0049] A distance detection module, configured to perform distance detection on the audio signal to determine the distance between the audio signal source and the microphone, and obtain a distance detection result;

[0050] A gender judgment module is used to separately extract Mel Frequency Cepstral Coefficients (MFCCs), spectral centroid, and spectral contrast from the audio signal, and determine a first gender judgment result of the person who emits the audio signal by combining the MFCCs, spectral centroid, and spectral contrast; input the MFCCs, spectral centroid, and spectral contrast into a gender judgment model to obtain a second gender judgment result of the person who emits the audio signal; the gender judgment model adopts a hybrid model structure of a convolutional neural network, a residual convolutional network, and a recurrent neural network; if the first gender judgment result and the second gender judgment result are consistent, the consistent result is used as the final gender judgment result.

[0051] In a third aspect, the present application proposes an electronic device, including: a memory, and one or more processors; the memory is coupled to the processor; wherein, computer program code is stored in the memory, and the computer program code includes computer instructions, when the computer instructions are executed by the processor, the electronic device executes the steps of the above-mentioned microphone volume dynamic adjustment method.

[0052] In a fourth aspect, the present application proposes a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned microphone volume dynamic adjustment method are implemented.

[0053] Compared with the prior art, the present application has the following beneficial effects:

[0054] The present application proposes a microphone volume dynamic adjustment method. First, the stationarity of the audio signal within a preset time range is judged. If the audio signal is stationary, distance detection is performed on the audio signal to determine the distance between the audio signal source and the microphone. At the same time, according to the MFCCs, spectral centroid, and spectral contrast extracted from the audio signal, a first gender judgment result and a second gender judgment result are obtained, wherein the second gender judgment result is obtained by combining a hybrid model structure composed of a convolutional neural network, a residual convolutional network, and a recurrent neural network, further enabling the first gender judgment result and the second gender judgment result to corroborate each other and improving the accuracy of the gender judgment result. By using the microphone volume dynamic adjustment method of the present application, it is possible to quickly and accurately adjust the volume according to the distance between the person who emits the audio signal and the microphone, and the gender of the person who emits the audio signal, so as to adapt to different speaking distances and genders of people, reduce manual operations, provide clear and consistent audio output, cancel the traditional distance sensor, and effectively reduce the use cost. If the audio signal is not stationary, the output volume of the microphone is dynamically adjusted according to the final gender judgment result, and the above effects can also be achieved.

[0055] The present application also proposes a microphone volume dynamic adjustment system, an electronic device, and a computer storage medium, which have all the advantages of the above-mentioned microphone volume dynamic adjustment method. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following accompanying drawings only show some embodiments of the present application, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related accompanying drawings can also be obtained based on these accompanying drawings.

[0057] Figure 1 It is the first flowchart of the method for dynamically adjusting the microphone volume of the present application;

[0058] Figure 2 It is the second flowchart of the method for dynamically adjusting the microphone volume of the present application;

[0059] Figure 3 It is a schematic diagram of the system for dynamically adjusting the microphone volume of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Usually, the components of the embodiments of the present application described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0061] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application to be protected, but merely represents the selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0062] It should be noted that: similar reference numerals and letters denote similar items in the following accompanying drawings. Therefore, once an item is defined in one accompanying drawing, it does not need to be further defined and explained in subsequent accompanying drawings.

[0063] In the description of the embodiments of the present application, it should be noted that if terms such as "upper", "lower", "horizontal", "inner", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship in which the inventive product is usually placed during use, it is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation of the present application. In addition, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.

[0064] In a microphone system, the adjustment of the microphone volume is a crucial step that directly affects the clarity and intelligibility of the audio signal. In various scenarios such as meetings, speeches, teaching, recording, etc., as the front-end device for sound collection, whether the volume setting of the microphone is appropriate directly relates to the effect of subsequent audio processing and the ultimate user experience. However, in practical applications, due to the continuous change in the distance between the speaker and the microphone, as well as the fluctuation of the ambient noise level, the microphone volume needs to be frequently adjusted to adapt to these changes. This manual adjustment method is not only cumbersome but also difficult to maintain the stability of the audio signal quality in a dynamically changing environment. In addition, due to the differences in voice characteristics (such as pitch, volume, timbre, etc.) of speakers of different genders, the requirements for the microphone volume are also different. For example, male speakers usually have a lower voice and may require a higher microphone volume to ensure the clarity of the voice; while female speakers usually have a higher voice and may require a relatively lower volume to avoid voice distortion or overload. Therefore, in the occasion of multiple people taking turns to speak, the adjustment of the microphone volume is more complex and meticulous.

[0065] Traditional methods for adjusting the microphone volume are that users or operators manually adjust the microphone volume according to the actual situation. This method is simple and direct, but it requires manual intervention and cannot adapt to environmental changes and speaker differences in real time. At the same time, there are also automatic gain control methods, noise threshold control methods, and intelligent volume control algorithms, etc., for adjusting the microphone volume. Specifically, the automatic gain control method is a technology that automatically adjusts the gain of the audio signal through an electronic circuit to keep the amplitude of the output signal within a certain range, and can automatically adjust the volume according to the size of the input signal, but it will produce unstable audio effects in a rapidly changing environment, such as the sound being sometimes loud and sometimes soft. The noise threshold control method is a method of automatically adjusting the volume by setting a noise threshold. When the ambient noise is below the threshold, the microphone volume remains at a low level; when the ambient noise exceeds the threshold, the microphone volume automatically increases. This method can reduce the impact of noise on the audio signal to a certain extent, but the setting of the threshold requires rich experience and may need to be adjusted repeatedly in different environments.

[0066] Based on the above situation of the prior art, although the prior art methods can improve the audio signal quality to a certain extent, there are still many defects and challenges. This application proposes a method and related device for dynamically adjusting the microphone volume, which can automatically adjust the output volume according to the distance between the audio signal source and the microphone, as well as the gender of the person emitting the audio signal, to ensure the clarity and consistency of the audio signal, and perform more intelligent, accurate, and real-time adjustment of the microphone volume to meet the needs of different scenarios and speakers.

[0067] Such as Figure 1As shown in the figure, this is the first process schematic diagram of the method for dynamically adjusting the microphone volume in this application, which may include:

[0068] S101. Continuously obtain the audio signal received by the microphone within a preset time range, and determine whether the smoothness of the audio signal meets the preset requirements. If it meets, dynamically adjust the output volume of the microphone according to the distance result and the final gender judgment result; otherwise, dynamically adjust the output volume of the microphone according to the final gender judgment result.

[0069] The microphone can capture sound fluctuations, convert them into electrical signals, and then digitize the electrical signals to provide basic data for subsequent processing in real time for dynamically adjusting the microphone volume. To ensure the accuracy of subsequent distance detection, first judge the smoothness of the audio signal. In practical applications, a certain threshold can be set according to the detection standard as the judgment basis. If the audio signal is smooth, the distance detection result can be ensured to be accurate. If the audio signal is not smooth, the distance detection result is not used as the basis for subsequent dynamic adjustment.

[0070] In practical applications, the microphone can be connected to an audio chip to obtain the audio signal received by the microphone, and convert the voice signal of the person emitting the audio signal into a digital signal. This process usually involves sampling and quantization of analog signals. Specifically, sampling is the process of discretizing an analog signal that changes continuously in time at a certain time interval. This time interval is called the sampling period, and its reciprocal is the sampling frequency. The audio chip usually contains a sample-and-hold circuit for capturing the voltage value of the analog signal at the sampling moment. The higher the sampling frequency, the closer the sampling result is to the original analog signal, but the amount of data will also increase accordingly, and it can be adjusted according to actual needs. Quantization is the process of discretizing the sampled voltage value and converting it into a digital signal. During the quantization process, the voltage value is mapped to a set of discrete values. The audio chip usually contains a quantizer for comparing the sampled voltage value with a set of predefined quantization levels and outputting the corresponding digital signal. The number of quantization levels determines the quantization accuracy. The more quantization levels, the higher the quantization accuracy, but the amount of data will also increase accordingly, and it can also be adjusted according to the requirements of actual applications.

[0071] The method for obtaining the distance result and the final gender judgment result includes:

[0072] S102. Perform distance detection on the audio signal to determine the distance between the audio signal source and the microphone, and obtain the distance detection result.

[0073] Distance detection provides a distance basis for volume adjustment, ensuring that the audio output remains clearly audible at different distances and improving the user experience, especially in multi-person meetings or public speaking scenarios. In practical applications, the distance can be estimated by using the intensity of the audio signal, the time difference of arrival, or the characteristics of sound wave propagation. It is also possible to calculate the relative distance between the audio signal source and the microphone by analyzing the attenuation, echo, and other characteristics of the audio signal through algorithms.

[0074] S103, extract the Mel-frequency cepstral coefficients, spectral centroid, and spectral contrast from the audio signal respectively, and determine the first gender judgment result of the person who emits the audio signal by combining the Mel-frequency cepstral coefficients, spectral centroid, and spectral contrast; input the Mel-frequency cepstral coefficients, spectral centroid, and spectral contrast into the gender judgment model to obtain the second gender judgment result of the person who emits the audio signal; the gender judgment model adopts a hybrid model structure of a convolutional neural network, a residual convolutional network, and a recurrent neural network; if the first gender judgment result and the second gender judgment result are consistent, then use the consistent result as the final gender judgment result.

[0075] It should be noted that if the first gender judgment result and the second gender judgment result are inconsistent, no adjustment will be made temporarily, and the judgment will be re-made and then adjusted.

[0076] In this application, a dual judgment method is adopted for gender judgment. On the one hand, the judgment is directly made based on the extraction results of the Mel-frequency cepstral coefficients, spectral centroid, and spectral contrast. On the other hand, the judgment is combined with the gender judgment model. The dual judgment method improves the judgment accuracy and enables this application to provide more refined personalized settings for volume adjustment, reducing the misjudgment rate and enhancing the user experience.

[0077] In practical applications, the gender judgment model adopts a hybrid model structure of a convolutional neural network, a residual convolutional network, and a recurrent neural network. The convolutional neural network extracts local spatial features from the feature representation that conforms to the human ear perception characteristics. The convolutional neural network of this application can realize the automatic learning and abstraction of features through the combination of convolutional layers, pooling layers, and fully connected layers. The residual convolutional network alleviates the gradient vanishing problem in deep networks by introducing residual connections, improving the training effect and generalization ability of the gender judgment model. The recurrent neural network is used to process the temporal information in the audio signal and capture the time-dependent relationship in the audio signal through the recurrent layer.

[0078] In practical applications, the Mel-frequency cepstral coefficients reflect the spectral characteristics of the audio signal, the spectral centroid represents the spectral center position of the audio signal, reflecting the "brightness" of the sound, and the spectral contrast is used to measure the energy difference between different frequency bands in the spectrum, which helps to distinguish different sound characteristics.

[0079] In practical applications, the microphone gain or output volume can be adjusted according to the distance judgment result to ensure that the volume increases at a long distance and decreases at a short distance. Combining with the gender judgment result, the volume can be finely adjusted to adapt to the characteristics of voices of different genders. For example, female voices are usually higher in frequency, and the volume of the high-frequency part may need to be appropriately increased, etc. Furthermore, intelligent volume control can be achieved to improve the clarity and comfort of audio output. This enables the microphone to adapt to different scenarios and user needs, enhancing the overall user experience. In practical applications, gain compensation can also be added during the microphone output to keep the final output of the microphone stable.

[0080] As Figure 2 shown, the following is the second flow diagram of the method for dynamically adjusting the microphone volume of this application, which is used to further illustrate this application. In this embodiment, taking the example that the smoothness of the input audio signal meets the preset requirements, it may specifically include:

[0081] S201, connect to the microphone through the audio chip to receive the audio signal from the microphone, and convert the audio signal into a digital signal.

[0082] S202, detect the distance between the audio signal source and the microphone. The following methods can be specifically used:

[0083] (1) Window function: To reduce spectral leakage, apply a window function to the converted digital signal to obtain the first preprocessed audio signal. For example, Hamming window and Hanning window can be used, and the Hanning window is used in this embodiment.

[0084] (2) Frame processing: Divide the first preprocessed audio signal into small frames to obtain a group of first-sequence frame signals. Each frame of the first-sequence frame signal usually contains dozens to hundreds of sampling points. In this embodiment, the length of the frame is set to 128.

[0085] (3) Calculate the power of the current frame: For each frame of the first-sequence frame signal, calculate its power respectively, that is, the sum of the squares of all sampling points divided by the number of sampling points. The specific formula is as follows:

[0086]

[0087] Among them, is the power of the first-sequence frame signal, N is the number of sampling points in a frame, is the i th sampling point value.

[0088] (4) Calculate the average power: Calculate the average power of the entire audio signal, and the average of all frame powers can be calculated. The formula is as follows:

[0089]

[0090] Among them, is the average power of the entire audio signal, M is the number of frames, is the j power of the

[0091] (5) Conversion to decibels: Convert the calculated average power to decibels and take the absolute value. The formula is as follows:

[0092]

[0093] Among them, is the absolute value of the converted decibels.

[0094] S203, jointly judge the gender of the person who emits the audio signal through the neural network algorithm combined with voice features.

[0095] Specifically, the following method can be adopted:

[0096] 1. Extraction of voice features: First, features that are helpful for gender judgment can be extracted from the audio signal. The features used in this application include Mel-scale Frequency Cepstral Coefficients (MFCC), spectral centroid, and spectral contrast.

[0097] (1) Extraction of Mel-scale Frequency Cepstral Coefficients.

[0098] ① Pre-emphasis: Pre-emphasize the audio signal to obtain the pre-emphasized signal.

[0099] ② Frame division: Divide the pre-emphasized signal into frames to obtain the framed signal.

[0100] ③ Window function application: Apply the window function to the framed signal to obtain the second preprocessed audio signal;

[0101] ④ Frame processing: Perform frame processing on the second preprocessed audio signal to obtain a set of second-sequence frame signals. Divide a set of second-sequence frame signals into small frames, and each frame usually contains dozens to hundreds of sampling points. In this embodiment, the length of the frame is set to 1024.

[0102] ⑤ Fast Fourier transform: Perform a fast Fourier transform on a set of second-sequence frame signals to convert the second-sequence frame signals into frequency-domain signals, denoted as the first frequency-domain signals.

[0103] ⑥ Mel filter bank: Pass the first frequency-domain signals through the Mel filter bank to obtain the Mel-frequency domain signals. The Mel filter bank simulates the auditory characteristics of the human ear, converting linear frequencies to Mel frequencies, which is a non-linear frequency scale and is more suitable for human auditory perception. The formula for Mel frequency is:

[0104]

[0105] Among them, is the frequency, with the unit of Hz, is the frequency of the first frequency-domain signal in Mel frequency.

[0106] ⑦ Logarithmic compression: Perform logarithmic compression on the Mel frequency-domain signal to obtain a compressed signal, in order to reduce the dynamic range and make the eigenvalue distribution more uniform. Logarithmic compression helps to highlight the important features in the speech signal and reduce the dynamic range difference between different frequency components. The calculation formula is:

[0107]

[0108] Among them, is the compressed signal, x is the Mel frequency-domain signal, that is, the output data of the Mel filter bank, and are both logarithmic compression coefficients, and in this embodiment, a = 2 and b = 0.5 can be set.

[0109] ⑧ Discrete cosine transform: Perform discrete cosine transform on the compressed signal to extract Mel frequency cepstral coefficients. Mel frequency cepstral coefficients can convert strong Mel spectrum features into a series of Mel frequency cepstral coefficient features with less correlation, and these features are very important for the representation of audio signals.

[0110] (2) Extraction of spectral centroid.

[0111] ① Window function: In order to reduce spectral leakage, apply a window function to the audio signal to obtain a first preprocessed audio signal.

[0112] ② Frame processing: Perform frame processing on the first preprocessed audio signal to obtain a set of first-sequence frame signals. Divide the audio signal into small frames, and each frame usually contains dozens to hundreds of sampling points. In this embodiment, the length of the frame is set to 1024.

[0113] ③ Fast Fourier transform: Perform fast Fourier transform on a set of first-sequence frame signals to convert the first-sequence frame signals into frequency-domain signals, denoted as the second frequency-domain signal. Converting the time-domain signal into the frequency-domain signal is the spectrum.

[0114] ④ Calculate spectral centroid: The spectral centroid can be obtained by calculating the energy-weighted average of all frequencies in the second frequency-domain signal:

[0115]

[0116] Among them, is the spectral centroid, is the frequency Spectral energy is the maximum frequency.

[0117] (3)Extraction of spectral contrast.

[0118] ① Window function: Apply a window function to the audio signal to obtain a first preprocessed audio signal.

[0119] ② Frame processing: Perform frame processing on the first preprocessed audio signal to obtain a set of first-sequence frame signals.

[0120] ③ Fast Fourier transform: Perform a fast Fourier transform on a set of first-sequence frame signals to convert the first-sequence frame signals into frequency-domain signals, denoted as second frequency-domain signals. Converting the time-domain signal into the frequency-domain signal is the spectrum.

[0121] ④ Frequency band division: Divide the second frequency-domain signal into multiple non-overlapping frequency bands to obtain multiple non-overlapping frequency bands. In this embodiment, it can be divided into 6 frequency bands.

[0122] ⑤ Calculate the energy of each frequency band: Calculate the energy of each frequency band respectively, and obtain the energy peak and energy valley of each frequency band. For each frequency band, the formula for calculating its energy is:

[0123]

[0124] where is the energy of the k th frequency band, is the complex amplitude of the i th frequency component in the second frequency-domain signal, which can be obtained through Fourier transform, and are the starting frequency index and ending frequency index of the k th frequency band respectively.

[0125] For each frequency band, the peak usually corresponds to the part with the highest energy in the frequency band, while the valley corresponds to the part with the lowest energy.

[0126] ⑥ Calculate the spectral contrast: The spectral contrast is calculated by comparing the energy peak and energy valley of each frequency band. The spectral contrast can be expressed by the formula:

[0127]

[0128] where is the energy peak of the k th frequency band, is the energy valley of the k th frequency band, is the energy of the k th frequency band.

[0129] 2. Gender Judgment Model Construction: The gender judgment model uses a hybrid model structure of CNN (Convolutional Neural Network) + ResCNN (Residual Convolutional Neural Network) + RNN (Recurrent Neural Network), which can combine the architectures of convolutional neural network, residual convolutional neural network, and recurrent neural network. This hybrid model structure can combine the advantages of each single network structure, process the time series and spectral features of audio signals, and is very effective for gender recognition.

[0130] Specifically, the calculation method in the gender judgment model includes:

[0131] (1) Concatenate the Mel Frequency Cepstral Coefficients, spectral centroid, and spectral contrast by frame to form a second concatenated feature matrix.

[0132] In practical applications, the Mel Frequency Cepstral Coefficients, spectral centroid, and spectral contrast can be concatenated by frame to form a second concatenated feature matrix, and the second concatenated feature matrix is normalized using Z-score normalization (Z-value normalization) to eliminate the dimensional difference.

[0133] (2) Input the second concatenated feature matrix into the gender judgment model. In the gender judgment model, the first convolutional neural network is used to extract the first local features from the Mel Frequency Cepstral Coefficients; the second convolutional neural network is used to extract statistical features from the spectral centroid and spectral contrast; the residual convolutional neural network is used to process the speech spectrogram to obtain the second local features; the recurrent neural network is used to capture the dependencies in the time series to obtain multiple third local features; the speech spectrogram includes a first frequency domain signal and a second frequency domain signal; the time series is obtained by arranging the first local features, statistical features, and second local features in chronological order.

[0134] (3) Perform weighted averaging on the multiple third local features to obtain the second gender judgment result.

[0135] It should be noted that by weighted summing the time step features according to the attention weights, a global feature representation can be obtained.

[0136] In practical applications, when training a gender judgment model, the training dataset used can be the publicly available male and female voice datasets on the Internet to train the gender judgment model, and the data results of male voice training and female voice training are saved separately. When making an actual judgment, the real-time voice signal is processed by the gender judgment model, and its results are compared with the data results of males and females respectively. If the correlation with males is greater than 90%, it is considered that the voice signal is male, and the corresponding first gender judgment result is used to further confirm whether it is male; if the correlation with females is greater than 90%, it is considered that the voice signal is female, and the corresponding first gender judgment result is used to further confirm whether it is female.

[0137] S204. jointly determine the final output size of the volume according to the distance detection result and the gender judgment result.

[0138] In practical applications, the microphone volume can be dynamically adjusted according to the following method:

[0139] (1) Set the initial value V1 of the volume output. When the device is just turned on and the gender of the speaker is not determined, and the distance between the person emitting the audio signal and the microphone is not known, play the audio at the volume of this initial value V1.

[0140] (2) Obtain the audio signal received by the microphone, and through the aforementioned method of converting to decibels, obtain the absolute value of the converted decibels corresponding to this audio signal. , and by comparing with the internal threshold K, judge the distance between the person emitting the audio signal and the microphone at this time. In this embodiment, the internal threshold K is set to 30. If < K, it is considered that the person emitting the audio signal is too close to the microphone, and the adjusted volume value V2 is calculated by the following formula:

[0141] .

[0142] If ≥ K, it is considered that the person emitting the audio signal is too far from the microphone, and the adjusted volume value V2 is calculated by the following formula:

[0143] .

[0144] (3) Output different volumes V3 for voices of different genders. If the speaker is male, amplify the data of V2 by 1.2 times and output V3; if the speaker is female, reduce the data of V2 by 0.8 times and output V3.

[0145] It should be noted that in other embodiments of the present application, if the smoothness of the input audio signal does not meet the preset requirements, based on step S204 in the above embodiments, the adjustment of the distance result can be omitted. In addition, the smoothness of the input audio signal does not affect the final gender determination result.

[0146] As Figure 3 shown, it is a schematic diagram of a microphone volume dynamic adjustment system of the present application, which may include:

[0147] A signal acquisition module, configured to continuously acquire the audio signal received by the microphone within a preset time range, determine whether the smoothness of the audio signal meets the preset requirements, if so, dynamically adjust the output volume of the microphone according to the distance result and the final gender determination result, otherwise, dynamically adjust the output volume of the microphone according to the final gender determination result; the distance result and the final gender determination result are obtained through the outputs of a distance detection module and a gender determination module;

[0148] A distance detection module, configured to perform distance detection on the audio signal to determine the distance between the audio signal source and the microphone, and obtain a distance detection result;

[0149] A gender determination module, configured to respectively extract Mel frequency cepstral coefficients, spectral centroid, and spectral contrast from the audio signal, combine the Mel frequency cepstral coefficients, spectral centroid, and spectral contrast to determine a first gender determination result of the person who emits the audio signal; input the Mel frequency cepstral coefficients, spectral centroid, and spectral contrast into a gender determination model to obtain a second gender determination result of the person who emits the audio signal; the gender determination model adopts a hybrid model structure of a convolutional neural network, a residual convolutional network, and a recurrent neural network; if the first gender determination result and the second gender determination result are consistent, the consistent result is used as the final gender determination result.

[0150] It should be noted that in several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of each module is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules can be combined or integrated into another device, or some features can be ignored or not executed. The modules described as separate components may or may not be physically separated, and the components shown as modules may be a physical unit or multiple physical units, that is, they may be located in one place, or may be distributed to multiple different places. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0151] In addition, in each embodiment of the present invention, each module can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0152] The embodiment of the present application further provides an electronic device, which may include one or more processors, a memory, and a communication interface.

[0153] Among them, the memory and the communication interface are coupled to the processor. For example, the memory and the communication interface can be coupled together through a bus.

[0154] Among them, the communication interface is used for data transmission with other devices. The memory stores computer program code. The computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device is caused to execute the steps of the above-mentioned method for dynamically adjusting the microphone volume.

[0155] Among them, the processor can be a processor or a controller. For example, it can be a Central Processing Unit (CPU), a general-purpose processor, a Digital Signal Processor (DSP), an Application-Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in conjunction with the present disclosure. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and so on. The processor can be used to support the electronic device to execute the method steps provided in the above embodiments.

[0156] Among them, the bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The above bus can be divided into an address bus, a data bus, a control bus, etc.

[0157] The embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned method for dynamically adjusting the microphone volume are implemented.

[0158] The computer-readable storage medium involved in the present application includes random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the technical field.

[0159] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A method for dynamically adjusting microphone volume, characterized in that: include: Within a preset time range, continuously obtain the audio signal received by the microphone, determine whether the stability of the audio signal meets the preset requirements, and if so, dynamically adjust the output volume of the microphone according to the distance result and the final gender judgment result; otherwise, dynamically adjust the output volume of the microphone according to the final gender judgment result; The method for obtaining the distance result and the final gender determination result includes: performing distance detection on the audio signal, determining the distance between the source of the audio signal and the microphone, and obtaining a distance detection result; Mel-frequency cepstral coefficients, spectral centroid and spectral contrast are extracted from the audio signal respectively, and a first gender judgment result of the person who sends the audio signal is determined by combining the Mel-frequency cepstral coefficients, spectral centroid and spectral contrast; the Mel-frequency cepstral coefficients, spectral centroid and spectral contrast are input into a gender judgment model to obtain a second gender judgment result of the person who sends the audio signal; the gender judgment model adopts a hybrid model structure of a convolutional neural network, a residual convolutional network and a recurrent neural network; if the first gender judgment result and the second gender judgment result are consistent, the consistent result is used as the final gender judgment result.

2. The method for dynamically adjusting microphone volume according to claim 1, characterized in that: The method for performing distance detection on the audio signal comprises: Applying a window function to the audio signal to obtain a first preprocessed audio signal; Performing frame processing on the first preprocessed audio signal to obtain a set of first sequence number frame signals; Calculate the power of all first sequence number frame signals respectively, and obtain a set of corresponding frame powers; Calculate the average power of a group of frame powers; The average power is converted into decibels and the absolute value is calculated. The distance between the source of the audio signal and the microphone is determined based on the absolute value.

3. The method for dynamically adjusting microphone volume according to claim 1, characterized in that: The method for extracting Mel-frequency cepstral coefficients comprises: Pre-emphasize the audio signal to obtain an emphasized signal; Framing the emphasized signal to obtain a framed signal; Applying a window function to the framed signal to obtain a second preprocessed audio signal; Performing frame processing on the second preprocessed audio signal to obtain a set of second sequence number frame signals; Performing a fast Fourier transform on a set of second sequence number frame signals to convert the second sequence number frame signals into frequency domain signals, recorded as first frequency domain signals; Passing the first frequency domain signal through a Mel filter bank to obtain a Mel frequency domain signal; Logarithmically compress the Mel frequency domain signal to obtain a compressed signal; The compressed signal is subjected to discrete cosine transform to extract Mel-frequency cepstrum coefficients.

4. The method for dynamically adjusting microphone volume according to claim 3, characterized in that: The method for calculating the spectral centroid comprises: Applying a window function to the audio signal to obtain a first preprocessed audio signal; Performing frame processing on the first preprocessed audio signal to obtain a set of first sequence number frame signals; Performing a fast Fourier transform on a group of first sequence number frame signals to convert the first sequence number frame signals into frequency domain signals, recorded as second frequency domain signals; The energy-weighted average of all frequencies in the second frequency domain signal is calculated to obtain a spectrum centroid.

5. The method for dynamically adjusting microphone volume according to claim 4, characterized in that: The calculation method of spectral contrast includes: Applying a window function to the audio signal to obtain a first preprocessed audio signal; Performing frame processing on the first preprocessed audio signal to obtain a set of first sequence number frame signals; Performing a fast Fourier transform on a group of first sequence number frame signals to convert the first sequence number frame signals into frequency domain signals, recorded as second frequency domain signals; Dividing the second frequency domain signal into a plurality of non-overlapping frequency bands to obtain a plurality of frequency bands; Calculate the energy of each frequency band respectively, and obtain the energy peak value and energy valley value of each frequency band; The spectral contrast of each frequency band is calculated based on the energy peak and energy valley of each frequency band.

6. The method for dynamically adjusting microphone volume according to claim 5, characterized in that: The method for determining the first gender judgment result of the person who sends the audio signal by combining the Mel frequency cepstrum coefficient, the spectrum centroid and the spectrum contrast comprises: Align the Mel-frequency cepstral coefficients, spectral centroids, and spectral contrast to the same number of frames; The aligned Mel-frequency cepstrum coefficients, spectral centroids and spectral contrasts are concatenated into a feature matrix, which is recorded as the first concatenated feature matrix; Normalize the first concatenated feature matrix; Assigning weights to different features in the normalized first concatenated feature matrix to obtain a weighted feature matrix; The weighted feature matrix is ​​compared with a preset constant matrix. If it is greater than or equal to the constant matrix, the first gender judgment result is female. Otherwise, the first gender judgment result is male.

7. The method for dynamically adjusting microphone volume according to claim 5, characterized in that: The convolutional neural network used in the gender judgment model includes a first convolutional neural network and a second convolutional neural network; The method of inputting the Mel frequency cepstrum coefficients, the spectral centroid and the spectral contrast into the gender judgment model to obtain the second gender judgment result of the person who sends the audio signal includes: The Mel frequency cepstral coefficients, the spectral centroid and the spectral contrast are spliced ​​into a second spliced ​​feature matrix by frame; The second concatenated feature matrix is ​​input into the gender judgment model. In the gender judgment model, a first convolutional neural network is used to extract a first local feature from the Mel frequency cepstral coefficient; a second convolutional neural network is used to extract statistical features from the spectral centroid and spectral contrast; a residual convolutional network is used to process the speech spectrogram to obtain a second local feature; a recurrent neural network is used to capture the dependency in the time series to obtain a plurality of third local features; the speech spectrogram includes a first frequency domain signal and a second frequency domain signal; the time series is obtained by arranging the first local feature, the statistical feature and the second local feature in chronological order; The weighted average of the multiple third local features is used to obtain the second gender judgment result.

8. A microphone volume dynamic adjustment system, characterized in that: include: A signal acquisition module, used to continuously acquire the audio signal received by the microphone within a preset time range, and determine whether the stability of the audio signal meets the preset requirements. If so, dynamically adjust the output volume of the microphone according to the distance result and the final gender judgment result; otherwise, dynamically adjust the output volume of the microphone according to the final gender judgment result; the distance result and the final gender judgment result are obtained through the output of the distance detection module and the gender judgment module; A distance detection module is used to perform distance detection on the audio signal, determine the distance between the source of the audio signal and the microphone, and obtain a distance detection result; The gender judgment module is used to extract the Mel-frequency cepstral coefficients, spectral centroid and spectral contrast from the audio signal respectively, and determine the first gender judgment result of the person who sends the audio signal by combining the Mel-frequency cepstral coefficients, spectral centroid and spectral contrast; input the Mel-frequency cepstral coefficients, spectral centroid and spectral contrast into the gender judgment model to obtain the second gender judgment result of the person who sends the audio signal; the gender judgment model adopts a hybrid model structure of a convolutional neural network, a residual convolutional network and a recurrent neural network; if the first gender judgment result and the second gender judgment result are consistent, the consistent result is used as the final gender judgment result.

9. An electronic device, characterized in that: include: A memory and one or more processors; the memory is coupled to the processor; wherein the memory stores computer program code, the computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device executes the steps of the microphone volume dynamic adjustment method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for dynamically adjusting the microphone volume as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Howling suppression method based on feedback signal spectrum estimation

    CN102740214A

  • Microphone control method and microphone

    CN107547978A