Dynamic microphone volume adjustment method and related device

By acquiring audio signals in real time and making distance and gender judgments in the microphone system, dynamically adjusting the microphone output volume, the problem of poor real-time adjustment and gender differences in traditional microphone systems is solved, and a higher robustness and adaptability of the voice acquisition system is achieved.

CN120018003AActive Publication Date: 2025-05-16SHENZHEN FENGHUO HONGSHENG TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510468184.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-16
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The real-time adjustment of audio signals in traditional microphone systems leads to overload distortion or signal-to-noise ratio deterioration when the sound source displacement or environmental noise changes, and the physiological differences in the basic frequency range and sound pressure level characteristics of speakers of different genders are not considered, resulting in a decrease in speech intelligibility.

Method used

By continuously obtaining the audio signal received by the microphone within a preset time range, the stability is judged, and the output volume of the microphone is dynamically adjusted according to the distance detection results and gender judgment results. Specific methods include distance detection and gender judgment on audio signals, and gender judgment using a mixed model of a Mel frequency cepspectral coefficient, spectral center of mass and spectral contrast combined with a convolutional neural network, residual convolutional network and recurrent neural network for gender judgment.

Benefits of technology

It realizes rapid and accurate adjustment of microphone volume, adapts to different speaking distances and genders, reduces manual operations, provides clear and consistent audio output, and improves the robustness and adaptability of the voice acquisition system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120018003A_ABST
    Figure CN120018003A_ABST
Patent Text Reader

Abstract

The invention belongs to an automatic volume adjustment method, and provides a microphone volume dynamic adjustment method and a related device aiming at the problems of poor adjustment real-time performance, poor voice acquisition robustness and poor adaptability of an audio signal adjustment method in a traditional microphone system. And if the audio signal is stable, distance detection is performed on the audio signal, and the distance between the source of the audio signal and the microphone is determined. A first gender judgment result and a second gender judgment result are obtained according to the Mel-frequency cepstrum coefficient, the spectral centroid and the spectral contrast extracted from the audio signal, and the second gender judgment result is obtained by combining a hybrid model structure composed of a convolutional neural network, a residual convolutional network and a recurrent neural network. Furthermore, the first gender judgment result and the second gender judgment result are mutually verified, so that the accuracy of the gender judgment result is improved. And if the audio signal is unstable, only a final gender judgment result is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to a method for automatically adjusting volume, and more particularly to a method for dynamically adjusting microphone volume and a related device. Background Art

[0002] In traditional microphone systems, the quality of audio signal acquisition is often limited by the manual gain adjustment mechanism. Existing technologies usually require users to manually adjust the input sensitivity parameters through physical knobs or software interfaces according to changes in the distance of the sound source or fluctuations in ambient noise. This method has significant real-time defects and technical limitations. Specifically, when the sound source is displaced or the ambient noise level changes dynamically, the fixed gain setting will cause the voice signal to be overloaded and distorted or the signal-to-noise ratio to deteriorate; in sudden interference noise scenarios, operation delays will cause the voice dynamic range compression to fail; in addition, because the physiological differences in fundamental frequency range and sound pressure level characteristics of speakers of different genders are not taken into account, the existing system lacks a differentiated gain compensation mechanism based on bioacoustic characteristics, resulting in a significant reduction in the voice intelligibility of specific user groups, which seriously affects the robustness and adaptability of the voice acquisition system. Summary of the invention

[0003] The present application aims at the technical problems of poor real-time adjustment, poor robustness and adaptability of voice acquisition in the adjustment method of audio signals in the traditional microphone system, and provides a microphone volume dynamic adjustment method and related devices.

[0004] In order to achieve the above objectives, this application adopts the following technical solutions: In a first aspect, the present application proposes a method for dynamically adjusting microphone volume, comprising: Within a preset time range, continuously obtain an audio signal received by a microphone, determine whether the stability of the audio signal meets a preset requirement, and if so, dynamically adjust the output volume of the microphone according to the distance result and the final gender judgment result; otherwise, dynamically adjust the output volume of the microphone according to the final gender judgment result; the method for obtaining the distance result and the final gender judgment result includes: Performing distance detection on the audio signal to determine the distance between the source of the audio signal and the microphone, and obtaining a distance detection result; Mel-frequency cepstral coefficients, spectral centroid and spectral contrast are extracted from the audio signal respectively, and a first gender judgment result of the person who sends the audio signal is determined by combining the Mel-frequency cepstral coefficients, spectral centroid and spectral contrast; the Mel-frequency cepstral coefficients, spectral centroid and spectral contrast are input into a gender judgment model to obtain a second gender judgment result of the person who sends the audio signal; the gender judgment model adopts a hybrid model structure of a convolutional neural network, a residual convolutional network and a recurrent neural network; if the first gender judgment result and the second gender judgment result are consistent, the consistent result is used as the final gender judgment result.

[0005] Furthermore, the method for performing distance detection on the audio signal includes: Applying a window function to the audio signal to obtain a first preprocessed audio signal; Performing frame processing on the first preprocessed audio signal to obtain a set of first sequence number frame signals; Calculate the power of all first sequence number frame signals respectively, and obtain a set of corresponding frame powers; Calculate the average power of a group of frame powers; The average power is converted into decibels and the absolute value is calculated. The distance between the source of the audio signal and the microphone is determined based on the absolute value.

[0006] Furthermore, the method for extracting Mel-frequency cepstral coefficients comprises: Pre-emphasize the audio signal to obtain an emphasized signal; Framing the emphasized signal to obtain a framed signal; Applying a window function to the framed signal to obtain a second preprocessed audio signal; Performing frame processing on the second preprocessed audio signal to obtain a set of second sequence number frame signals; Performing a fast Fourier transform on a set of second sequence number frame signals to convert the second sequence number frame signals into frequency domain signals, recorded as first frequency domain signals; Passing the first frequency domain signal through a Mel filter bank to obtain a Mel frequency domain signal; Logarithmically compress the Mel frequency domain signal to obtain a compressed signal; The compressed signal is subjected to discrete cosine transform to extract Mel-frequency cepstrum coefficients.

[0007] Furthermore, the method for calculating the spectral centroid includes: Applying a window function to the audio signal to obtain a first preprocessed audio signal; Performing frame processing on the first preprocessed audio signal to obtain a set of first sequence number frame signals; Performing a fast Fourier transform on a group of first sequence number frame signals to convert the first sequence number frame signals into frequency domain signals, recorded as second frequency domain signals; The energy-weighted average of all frequencies in the second frequency domain signal is calculated to obtain a spectrum centroid.

[0008] Furthermore, the calculation method of the spectral contrast includes: Applying a window function to the audio signal to obtain a first preprocessed audio signal; Performing frame processing on the first preprocessed audio signal to obtain a set of first sequence number frame signals; Performing a fast Fourier transform on a group of first sequence number frame signals to convert the first sequence number frame signals into frequency domain signals, recorded as second frequency domain signals; Dividing the second frequency domain signal into a plurality of non-overlapping frequency bands to obtain a plurality of frequency bands; Calculate the energy of each frequency band respectively, and obtain the energy peak value and energy valley value of each frequency band; The spectral contrast of each frequency band is calculated based on the energy peak and energy valley of each frequency band.

[0009] Furthermore, the method for determining the first gender judgment result of the person who sends the audio signal by combining the Mel-frequency cepstrum coefficient, the spectral centroid and the spectral contrast includes: Align the Mel-frequency cepstral coefficients, spectral centroids, and spectral contrast to the same number of frames; The aligned Mel-frequency cepstrum coefficients, spectral centroids and spectral contrasts are concatenated into a feature matrix, which is recorded as the first concatenated feature matrix; Normalize the first concatenated feature matrix; Assigning weights to different features in the normalized first concatenated feature matrix to obtain a weighted feature matrix; The weighted feature matrix is ​​compared with a preset constant matrix. If it is greater than or equal to the constant matrix, the first gender judgment result is female. Otherwise, the first gender judgment result is male.

[0010] Furthermore, the convolutional neural network used in the gender determination model includes a first convolutional neural network and a second convolutional neural network; The method of inputting the Mel frequency cepstrum coefficients, the spectral centroid and the spectral contrast into the gender judgment model to obtain the second gender judgment result of the person who sends the audio signal includes: The Mel frequency cepstral coefficients, the spectral centroid and the spectral contrast are spliced ​​into a second spliced ​​feature matrix by frame; The second concatenated feature matrix is ​​input into the gender judgment model. In the gender judgment model, a first convolutional neural network is used to extract a first local feature from the Mel frequency cepstral coefficient; a second convolutional neural network is used to extract statistical features from the spectral centroid and spectral contrast; a residual convolutional network is used to process the speech spectrogram to obtain a second local feature; a recurrent neural network is used to capture the dependency in the time series to obtain a plurality of third local features; the speech spectrogram includes a first frequency domain signal and a second frequency domain signal; the time series is obtained by arranging the first local feature, the statistical feature and the second local feature in chronological order; The weighted average of the multiple third local features is used to obtain the second gender judgment result.

[0011] In a second aspect, the present application proposes a microphone volume dynamic adjustment system, comprising: A signal acquisition module, used to continuously acquire the audio signal received by the microphone within a preset time range, and determine whether the stability of the audio signal meets the preset requirements. If so, dynamically adjust the output volume of the microphone according to the distance result and the final gender judgment result; otherwise, dynamically adjust the output volume of the microphone according to the final gender judgment result; the distance result and the final gender judgment result are obtained through the output of the distance detection module and the gender judgment module; A distance detection module is used to perform distance detection on the audio signal, determine the distance between the source of the audio signal and the microphone, and obtain a distance detection result; The gender judgment module is used to extract the Mel-frequency cepstral coefficients, spectral centroid and spectral contrast from the audio signal respectively, and determine the first gender judgment result of the person who sends the audio signal by combining the Mel-frequency cepstral coefficients, spectral centroid and spectral contrast; input the Mel-frequency cepstral coefficients, spectral centroid and spectral contrast into the gender judgment model to obtain the second gender judgment result of the person who sends the audio signal; the gender judgment model adopts a hybrid model structure of a convolutional neural network, a residual convolutional network and a recurrent neural network; if the first gender judgment result and the second gender judgment result are consistent, the consistent result is used as the final gender judgment result.

[0012] In a third aspect, the present application proposes an electronic device, comprising: a memory and one or more processors; the memory is coupled to the processor; wherein computer program code is stored in the memory, and the computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device performs the steps of the above-mentioned microphone volume dynamic adjustment method.

[0013] In a fourth aspect, the present application proposes a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned method for dynamically adjusting the microphone volume are implemented.

[0014] Compared with the prior art, this application has the following beneficial effects: The present application proposes a method for dynamically adjusting the volume of a microphone. First, the stability of an audio signal within a preset time range is judged. If the audio signal is stable, the distance detection of the audio signal is performed to determine the distance between the source of the audio signal and the microphone. At the same time, according to the Mel frequency cepstrum coefficients, spectral centroid and spectral contrast extracted from the audio signal, a first gender judgment result and a second gender judgment result are obtained, wherein the second gender judgment result is obtained by combining a hybrid model structure composed of a convolutional neural network, a residual convolutional network and a recurrent neural network, so that the first gender judgment result and the second gender judgment result are further verified with each other, thereby improving the accuracy of the gender judgment result. The method for dynamically adjusting the volume of a microphone of the present application can quickly and accurately adjust the volume according to the distance between the person who sends the audio signal and the microphone, as well as the gender of the person who sends the audio signal, to adapt to different speaking distances and genders of the person, reduce manual operations, provide clear and consistent audio output, cancel the traditional distance sensor, and effectively reduce the cost of use. If the audio signal is not stable, the output volume of the microphone can be dynamically adjusted according to the final gender judgment result, which can also achieve the above effect.

[0015] The present application also proposes a microphone volume dynamic adjustment system, an electronic device and a computer storage medium, which have all the advantages of the above-mentioned microphone volume dynamic adjustment method. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0017] Figure 1 This is a schematic diagram of the first flow chart of the method for dynamically adjusting microphone volume in this application; Figure 2 This is a second flow chart of the method for dynamically adjusting microphone volume in the present application; Figure 3 A schematic diagram of a microphone volume dynamic adjustment system of the present application. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations.

[0019] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for which protection is sought, but merely represents selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.

[0020] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.

[0021] In the description of the embodiments of the present application, it should be noted that if the terms "upper", "lower", "horizontal", "inner", etc. appear, the orientation or position relationship indicated is based on the orientation or position relationship shown in the drawings, or the orientation or position relationship in which the invented product is usually placed when used. It is only for the convenience of describing the present application and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first", "second", etc. are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.

[0022] In microphone systems, the adjustment of microphone volume is a crucial link, which directly affects the clarity and intelligibility of audio signals. In various scenarios such as meetings, lectures, teaching, and recording, the microphone is used as a front-end device for sound collection. Whether its volume setting is appropriate is directly related to the effect of subsequent audio processing and the final user experience. However, in actual applications, due to the constant change in the distance between the speaker and the microphone and the fluctuation of the ambient noise level, the microphone volume needs to be adjusted frequently to adapt to these changes. This manual adjustment method is not only cumbersome, but also difficult to maintain the stability of audio signal quality in a dynamically changing environment. In addition, speakers of different genders have different requirements for microphone volume due to differences in their voice characteristics (such as pitch, volume, timbre, etc.). For example, male speakers usually have a deeper voice and may require a higher microphone volume to ensure the clarity of the voice; while female speakers usually have a higher voice and may require a relatively lower volume to avoid sound distortion or overload. Therefore, in situations where multiple people speak alternately, the adjustment of microphone volume is more complex and delicate.

[0023] The traditional method of adjusting the microphone volume is for the user or operator to manually adjust the microphone volume according to the actual situation. This method is simple and direct, but requires manual intervention and cannot adapt to environmental changes and speaker differences in real time. There are also automatic gain control methods, noise threshold control methods, and intelligent volume control algorithms for adjusting the microphone volume. Specifically, the automatic gain control method is a technology that automatically adjusts the gain of the audio signal through an electronic circuit to keep the amplitude of the output signal within a certain range. The volume can be automatically adjusted according to the size of the input signal, but in a rapidly changing environment, unstable audio effects will be produced, such as the sound fluctuating. The noise threshold control method is a method that automatically adjusts the volume by setting a noise threshold. When the ambient noise is lower than the threshold, the microphone volume remains at a low level; when the ambient noise exceeds the threshold, the microphone volume automatically increases. This method can reduce the impact of noise on the audio signal to a certain extent, but the setting of the threshold requires rich experience and may need to be adjusted repeatedly in different environments.

[0024] Based on the above-mentioned prior art, although the prior art methods can improve the quality of audio signals to a certain extent, there are still many defects and challenges. This application proposes a method and a related device for dynamically adjusting the volume of a microphone, which can automatically adjust the output volume according to the distance between the source of the audio signal and the microphone, and the gender of the person who sends the audio signal, so as to ensure the clarity and consistency of the audio signal, and adjust the microphone volume more intelligently, accurately and in real time to meet the needs of different scenarios and speakers.

[0025] like Figure 1 As shown, it is a first flow chart of the method for dynamically adjusting the microphone volume of the present application, which may include: S101, continuously obtain the audio signal received by the microphone within a preset time range, and determine whether the smoothness of the audio signal meets the preset requirements. If so, dynamically adjust the output volume of the microphone according to the distance result and the final gender judgment result; otherwise, dynamically adjust the output volume of the microphone according to the final gender judgment result.

[0026] The microphone can capture sound fluctuations and convert them into electrical signals, which are then digitized to provide basic data for subsequent processing in real time for dynamic adjustment of the microphone volume. In order to ensure the accuracy of subsequent distance detection, the stability of the audio signal is first determined. In practical applications, a certain threshold can be set as a basis for judgment based on the detection standard. If the audio signal is stable, the distance detection result can be guaranteed to be accurate. If the audio signal is not stable, the distance detection result will not be used as a basis for subsequent dynamic adjustment.

[0027] In practical applications, a microphone can be connected through an audio chip to obtain the audio signal received by the microphone, and the voice signal of the person who sends the audio signal can be converted into a digital signal. This process usually involves sampling and quantization of analog signals. Specifically, sampling is the process of discretizing an analog signal that changes continuously in time according to a certain time interval. This time interval is called a sampling period, and its reciprocal is the sampling frequency. The audio chip usually contains a sample-and-hold circuit to capture the voltage value of the analog signal at the sampling moment. The higher the sampling frequency, the closer the sampling result is to the original analog signal, but the amount of data will also increase accordingly, which can be adjusted according to actual needs. Quantization is the process of discretizing the sampled voltage value and converting it into a digital signal. During the quantization process, the voltage value is mapped to a set of discrete values. The audio chip usually contains a quantizer to compare the sampled voltage value with a set of predefined quantization levels and output the corresponding digital signal. The number of quantization levels determines the accuracy of quantization. The more quantization levels, the higher the quantization accuracy, but the amount of data will also increase accordingly, which can also be adjusted according to actual application requirements.

[0028] The method for obtaining the distance result and the final gender determination result includes: S102, performing distance detection on the audio signal to determine the distance between the source of the audio signal and the microphone, and obtaining a distance detection result.

[0029] Distance detection provides a distance basis for volume adjustment, ensuring that audio output remains clear and audible at different distances, improving user experience, especially in multi-person meetings or public speeches. In practical applications, distance estimation can be performed using the intensity of the audio signal, arrival time difference, or sound wave propagation characteristics. Algorithms can also be used to analyze the attenuation, echo, and other characteristics of the audio signal to calculate the relative distance between the source of the audio signal and the microphone.

[0030] S103, extracting the Mel-frequency cepstral coefficients, spectral centroid and spectral contrast from the audio signal respectively, and determining a first gender judgment result of the person who sends the audio signal by combining the Mel-frequency cepstral coefficients, spectral centroid and spectral contrast; inputting the Mel-frequency cepstral coefficients, spectral centroid and spectral contrast into a gender judgment model to obtain a second gender judgment result of the person who sends the audio signal; the gender judgment model adopts a hybrid model structure of a convolutional neural network, a residual convolutional network and a recurrent neural network; if the first gender judgment result and the second gender judgment result are consistent, the consistent result is used as the final gender judgment result.

[0031] It should be noted that if the first gender judgment result and the second gender judgment result are inconsistent, no adjustment will be made temporarily, and adjustment will be made after re-judgment.

[0032] In this application, a dual judgment method is adopted for gender judgment. On the one hand, judgment is made directly based on the extraction results of Mel-frequency cepstral coefficients, spectral centroid and spectral contrast. On the other hand, judgment is made in combination with the gender judgment model. The dual judgment method improves the judgment accuracy and also enables the application to provide more refined personalized settings for volume adjustment, reduce the misjudgment rate, and improve user experience.

[0033] In practical applications, the gender judgment model adopts a hybrid model structure of convolutional neural networks, residual convolutional networks and recurrent neural networks. Convolutional neural networks extract local spatial features from feature representations that conform to the perceptual characteristics of the human ear. The convolutional neural network of this application can achieve automatic learning and abstraction of features through a combination of convolutional layers, pooling layers and fully connected layers. The residual convolutional network alleviates the gradient vanishing problem in deep networks by introducing residual connections, thereby improving the training effect and generalization ability of the gender judgment model. Recurrent neural networks are used to process timing information in audio signals and capture time dependencies in audio signals through recurrent layers.

[0034] In practical applications, the Mel-frequency cepstral coefficient reflects the spectral characteristics of the audio signal. The spectral centroid indicates the center position of the spectrum of the audio signal, reflecting the "brightness" of the sound. The spectral contrast is used to measure the energy difference between different frequency bands in the spectrum, which helps to distinguish different sound characteristics.

[0035] In actual applications, the microphone gain or output volume can be adjusted according to the distance judgment result to ensure that the volume increases at a long distance and decreases at a close distance. Combined with the gender judgment result, the volume can be fine-tuned to adapt to the characteristics of voices of different genders. For example, female voices are usually higher-frequency, and the volume of the high-frequency part may need to be appropriately increased. This will achieve intelligent volume control and improve the clarity and comfort of audio output. This allows the microphone to adapt to different scenarios and user needs and improve the overall user experience. In actual applications, gain compensation can also be added to the microphone output to keep the final output of the microphone stable.

[0036] like Figure 2 As shown, it is a second flow chart of the method for dynamically adjusting the microphone volume of the present application, which is used to further illustrate the present application. In this embodiment, the stability of the input audio signal meets the preset requirements as an example, which may specifically include: S201, receiving an audio signal from a microphone through an audio chip connected to the microphone, and converting the audio signal into a digital signal.

[0037] S202, detecting the distance between the source of the audio signal and the microphone. Specifically, the following method may be used: (1) Windowing function: In order to reduce spectrum leakage, a window function is applied to the converted digital signal to obtain a first preprocessed audio signal. For example, a Hamming window or a Hanning window may be used. In this embodiment, a Hanning window is used.

[0038] (2) Frame processing: The first pre-processed audio signal is divided into small frames to obtain a group of first sequence number frame signals, each of which usually contains dozens to hundreds of sampling points. In this embodiment, the frame length is set to 128.

[0039] (3) Calculate the power of the current frame: For the first frame signal of each frame, calculate its power, that is, the sum of the squares of all sampling points divided by the number of sampling points. The specific formula is as follows:

[0040] in, is the power of the first sequence frame signal, N is the number of sampling points in a frame, It is i The value of the sampling point.

[0041] (4) Calculate the average power: To calculate the average power of the entire audio signal, you can average the power of all frames. The formula is as follows:

[0042] in, is the average power of the entire audio signal, M is the number of frames, For the j The power of the frame.

[0043] (5) Conversion to decibels: Convert the calculated average power to decibels and find the absolute value. The formula is as follows:

[0044] in, is the absolute value of decibel after conversion.

[0045] S203, determining the gender of the person who sends the audio signal through a neural network algorithm combined with voice features.

[0046] The following methods can be used: 1. Extraction of speech features: First, features that are helpful for gender determination can be extracted from the audio signal. The features used in this application include Mel-scale Frequency Cepstral Coefficients (MFCC), spectral centroid, and spectral contrast.

[0047] (1) Extraction of Mel-frequency cepstral coefficients.

[0048] ① Pre-emphasis: Pre-emphasize the audio signal to obtain an emphasized signal.

[0049] ② Framing: Framing the emphasized signal to obtain a framed signal.

[0050] ③ Windowing function: applying the windowing function to the framed signal to obtain a second preprocessed audio signal; ④ Frame processing: Frame processing is performed on the second pre-processed audio signal to obtain a set of second sequence frame signals. The set of second sequence frame signals is divided into small frames, each of which usually contains dozens to hundreds of sampling points. In this embodiment, the frame length is set to 1024.

[0051] ⑤ Fast Fourier Transform: Perform fast Fourier transform on a group of second sequence frame signals, convert the second sequence frame signals into frequency domain signals, which are recorded as first frequency domain signals.

[0052] ⑥ Mel filter bank: The first frequency domain signal passes through the Mel filter bank to obtain a Mel frequency frequency domain signal. The Mel filter bank simulates the auditory characteristics of the human ear and converts the linear frequency into the Mel frequency, which is a nonlinear frequency scale that is more suitable for human auditory perception. The calculation formula of the Mel frequency is:

[0053] in, is the frequency in Hz, For frequency The Mel frequencies of the first frequency domain signal.

[0054] ⑦ Logarithmic compression: Logarithmic compression is performed on the Mel frequency domain signal to obtain a compressed signal to reduce the dynamic range and make the eigenvalue distribution more uniform. Logarithmic compression helps to highlight the important features in the speech signal and reduce the dynamic range differences between different frequency components. The calculation formula is:

[0055] in, is the compressed signal, x is the Mel frequency domain signal, that is, the output data of the Mel filter bank, and Both are logarithmic compression coefficients, which can be set to a=2 and b=0.5 in this embodiment.

[0056] ⑧ Discrete cosine transform: Perform discrete cosine transform on the compressed signal to extract the Mel-frequency cepstral coefficients. The Mel-frequency cepstral coefficients can convert the strong Mel-frequency spectrum features into a series of Mel-frequency cepstral coefficient features with less correlation. These features are very important for the representation of audio signals.

[0057] (2) Extraction of spectral centroid.

[0058] ① Windowing function: In order to reduce spectrum leakage, a windowing function is applied to the audio signal to obtain a first preprocessed audio signal.

[0059] ② Frame processing: Frame processing is performed on the first preprocessed audio signal to obtain a set of first sequence number frame signals. The audio signal is divided into small frames, each of which usually contains dozens to hundreds of sampling points. In this embodiment, the frame length is set to 1024.

[0060] ③ Fast Fourier Transform: Perform fast Fourier transform on a group of first-numbered frame signals, convert the first-numbered frame signals into frequency domain signals, recorded as second frequency domain signals, and convert the time domain signals into frequency domain signals, which are the spectrum.

[0061] ④ Calculate the spectral centroid: The spectral centroid can be obtained by calculating the energy-weighted average of all frequencies in the second frequency domain signal:

[0062] in, is the spectral centroid, For frequency The spectral energy of is the maximum frequency.

[0063] (3) Extraction of spectral contrast.

[0064] ① Windowing function: Apply a window function to the audio signal to obtain a first preprocessed audio signal.

[0065] ② Frame processing: Perform frame processing on the first preprocessed audio signal to obtain a set of first sequence number frame signals.

[0066] ③ Fast Fourier Transform: Perform fast Fourier transform on a group of first-numbered frame signals, convert the first-numbered frame signals into frequency domain signals, recorded as second frequency domain signals, and convert the time domain signals into frequency domain signals, which are the spectrum.

[0067] ④ Frequency band division: The second frequency domain signal is divided into a plurality of non-intersecting frequency bands to obtain a plurality of non-intersecting frequency bands. In this embodiment, the frequency bands may be divided into 6 frequency bands.

[0068] ⑤ Calculate the energy of each frequency band: Calculate the energy of each frequency band separately and obtain the energy peak and energy valley of each frequency band. For each frequency band, the formula for calculating its energy is:

[0069] in, For the k The energy of the frequency band, is the first iThe complex amplitude of the frequency component can be obtained by Fourier transform. and Respectively k The starting frequency index and ending frequency index of the frequency band.

[0070] For each frequency band, the peaks generally correspond to the highest energy portions of the band, while the valleys correspond to the lowest energy portions.

[0071] ⑥ Calculate spectral contrast: Spectral contrast is calculated by comparing the energy peak and energy valley of each frequency band. The calculation formula can be expressed as:

[0072] in, It is k The peak energy of the frequency band, It is k The energy valley value of the frequency band, It is k The energy of a frequency band.

[0073] 2. Gender identification model construction: The gender identification model uses a hybrid model structure of CNN (Convolutional Neural Network) + ResCNN (Residual Convolutional Neural Network) + RNN (Recurrent Neural Network), which can combine the architectures of convolutional neural network, residual convolutional network and recurrent neural network. This hybrid model structure can combine the advantages of each single network structure to process the time series and spectrum characteristics of audio signals, which is very effective for gender identification.

[0074] Specifically, the calculation method in the gender determination model includes: (1) The Mel-frequency cepstral coefficients, spectral centroids, and spectral contrasts are concatenated frame by frame into the second concatenated feature matrix.

[0075] In practical applications, the Mel-frequency cepstral coefficients, spectral centroids, and spectral contrasts can be spliced ​​frame by frame to form a second spliced ​​feature matrix, and the second spliced ​​feature matrix can be normalized using Z-score standardization to eliminate dimensional differences.

[0076] (2) The second concatenated feature matrix is ​​input into a gender judgment model. In the gender judgment model, a first convolutional neural network is used to extract a first local feature from the Mel-frequency cepstral coefficient; a second convolutional neural network is used to extract statistical features from the spectral centroid and spectral contrast; a residual convolutional network is used to process the speech spectrogram to obtain a second local feature; a recurrent neural network is used to capture the dependency in the time series to obtain a plurality of third local features; the speech spectrogram includes a first frequency domain signal and a second frequency domain signal; and the time series is obtained by arranging the first local feature, the statistical feature and the second local feature in chronological order.

[0077] (3) Taking a weighted average of multiple third local features, a second gender judgment result is obtained.

[0078] It should be noted that by weighting and summing the time step features according to the attention weights, a global feature representation can be obtained.

[0079] In practical applications, when training the gender judgment model, the training data set used can use the male and female voice data sets publicly available on the Internet to train the gender judgment model, and save the data results of male voice training and female voice training respectively. When making actual judgments, the real-time voice signal is processed by the gender judgment model, and its results are compared with the data results of male and female respectively. If the correlation of male is greater than 90%, the voice signal is considered to be male, and the corresponding first gender judgment result is used to further confirm whether it is male; if the correlation of female is greater than 90%, the voice signal is considered to be female, and the corresponding first gender judgment result is used to further confirm whether it is female.

[0080] S204, determining a final output volume according to the distance detection result and the gender determination result.

[0081] In actual applications, the microphone volume can be dynamically adjusted as follows: (1) Set the initial value V1 of the volume output. When the device is just turned on and the gender of the speaker and the distance between the person sending the audio signal and the microphone are not determined, the audio signal is played at the volume of the initial value V1.

[0082] (2) Obtain the audio signal received by the microphone, and obtain the converted decibel absolute value corresponding to the audio signal through the above-mentioned conversion method to decibels , by comparing with the internal threshold K, the distance between the person who sends the audio signal and the microphone is determined. In this embodiment, the internal threshold K is set to 30. <K, it is considered that the person sending the audio signal is too close to the microphone, and the adjusted volume value V2 is calculated by the following formula: .

[0083] like ≥K, it is considered that the person sending the audio signal is too far away from the microphone, and the adjusted volume value V2 is calculated by the following formula: .

[0084] (3) Output different volumes V3 for voices of different genders. If the speaker is male, the data of V2 is amplified by 1.2 times and output as V3; if the speaker is female, the data of V2 is reduced by 0.8 times and output as V3.

[0085] It should be noted that, in other embodiments of the present application, if the stability of the input audio signal does not meet the preset requirements, the adjustment of the distance result can be omitted based on step S204 in the above embodiment. In addition, the stability of the input audio signal will not affect the final gender determination result.

[0086] like Figure 3 As shown, it is a schematic diagram of a microphone volume dynamic adjustment system of the present application, which may include: A signal acquisition module, used to continuously acquire the audio signal received by the microphone within a preset time range, and determine whether the stability of the audio signal meets the preset requirements. If so, dynamically adjust the output volume of the microphone according to the distance result and the final gender judgment result; otherwise, dynamically adjust the output volume of the microphone according to the final gender judgment result; the distance result and the final gender judgment result are obtained through the output of the distance detection module and the gender judgment module; A distance detection module is used to perform distance detection on the audio signal, determine the distance between the source of the audio signal and the microphone, and obtain a distance detection result; The gender judgment module is used to extract the Mel-frequency cepstral coefficients, spectral centroid and spectral contrast from the audio signal respectively, and determine the first gender judgment result of the person who sends the audio signal by combining the Mel-frequency cepstral coefficients, spectral centroid and spectral contrast; input the Mel-frequency cepstral coefficients, spectral centroid and spectral contrast into the gender judgment model to obtain the second gender judgment result of the person who sends the audio signal; the gender judgment model adopts a hybrid model structure of a convolutional neural network, a residual convolutional network and a recurrent neural network; if the first gender judgment result and the second gender judgment result are consistent, the consistent result is used as the final gender judgment result.

[0087] It should be noted that in the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the system embodiments described above are merely schematic. For example, the division of each module is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules can be combined or integrated into another device, or some features can be ignored or not executed. The module described as a separate component may or may not be physically separated. The component displayed as a module may be a physical unit or multiple physical units, that is, it may be located in one place, or it may be distributed in multiple different places. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0088] In addition, each module in each embodiment of the present invention may be integrated into a processing unit, each module may exist physically separately, or two or more modules may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0089] An embodiment of the present application also provides an electronic device, which may include one or more processors, a memory, and a communication interface.

[0090] The memory, the communication interface and the processor are coupled, for example, the memory, the communication interface and the processor may be coupled together via a bus.

[0091] The communication interface is used for data transmission with other devices. The memory stores computer program code. The computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device executes the steps of the above-mentioned microphone volume dynamic adjustment method.

[0092] Wherein, the processor can be a processor or a controller, for example, a central processing unit (CPU), a general processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the present disclosure. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of DSP and microprocessors, and the like. The processor can be used to support electronic devices to execute the method steps provided in the above embodiments.

[0093] The bus may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The above bus may be divided into an address bus, a data bus, a control bus, etc.

[0094] An embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned method for dynamically adjusting the microphone volume are implemented.

[0095] The computer-readable storage medium involved in the present application includes random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the technical field.

[0096] The above are only preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for dynamically adjusting microphone volume, characterized in that: include: Within a preset time range, continuously obtain the audio signal received by the microphone, determine whether the stability of the audio signal meets the preset requirements, and if so, dynamically adjust the output volume of the microphone according to the distance result and the final gender judgment result; otherwise, dynamically adjust the output volume of the microphone according to the final gender judgment result; The method for obtaining the distance result and the final gender determination result includes: performing distance detection on the audio signal, determining the distance between the source of the audio signal and the microphone, and obtaining a distance detection result; Mel-frequency cepstral coefficients, spectral centroid and spectral contrast are extracted from the audio signal respectively, and a first gender judgment result of the person who sends the audio signal is determined by combining the Mel-frequency cepstral coefficients, spectral centroid and spectral contrast; the Mel-frequency cepstral coefficients, spectral centroid and spectral contrast are input into a gender judgment model to obtain a second gender judgment result of the person who sends the audio signal; the gender judgment model adopts a hybrid model structure of a convolutional neural network, a residual convolutional network and a recurrent neural network; if the first gender judgment result and the second gender judgment result are consistent, the consistent result is used as the final gender judgment result.

2. The method for dynamically adjusting microphone volume according to claim 1, characterized in that: The method for performing distance detection on the audio signal comprises: Applying a window function to the audio signal to obtain a first preprocessed audio signal; Performing frame processing on the first preprocessed audio signal to obtain a set of first sequence number frame signals; Calculate the power of all first sequence number frame signals respectively, and obtain a set of corresponding frame powers; Calculate the average power of a group of frame powers; The average power is converted into decibels and the absolute value is calculated. The distance between the source of the audio signal and the microphone is determined based on the absolute value.

3. The method for dynamically adjusting microphone volume according to claim 1, characterized in that: The method for extracting Mel-frequency cepstral coefficients comprises: Pre-emphasize the audio signal to obtain an emphasized signal; framing the emphasized signal to obtain a framed signal; Applying a window function to the framed signal to obtain a second preprocessed audio signal; Performing frame processing on the second preprocessed audio signal to obtain a set of second sequence number frame signals; Performing a fast Fourier transform on a set of second sequence number frame signals to convert the second sequence number frame signals into frequency domain signals, recorded as first frequency domain signals; Passing the first frequency domain signal through a Mel filter bank to obtain a Mel frequency domain signal; Logarithmically compress the Mel frequency domain signal to obtain a compressed signal; The compressed signal is subjected to discrete cosine transform to extract Mel-frequency cepstrum coefficients.

4. The method for dynamically adjusting microphone volume according to claim 3, characterized in that: The method for calculating the spectral centroid comprises: Applying a window function to the audio signal to obtain a first preprocessed audio signal; Performing frame processing on the first preprocessed audio signal to obtain a set of first sequence number frame signals; Performing a fast Fourier transform on a group of first sequence number frame signals to convert the first sequence number frame signals into frequency domain signals, recorded as second frequency domain signals; The energy-weighted average of all frequencies in the second frequency domain signal is calculated to obtain a spectrum centroid.

5. The method for dynamically adjusting microphone volume according to claim 4, characterized in that: The calculation method of spectral contrast includes: Applying a window function to the audio signal to obtain a first preprocessed audio signal; Performing frame processing on the first preprocessed audio signal to obtain a set of first sequence number frame signals; Performing a fast Fourier transform on a group of first sequence number frame signals to convert the first sequence number frame signals into frequency domain signals, recorded as second frequency domain signals; Dividing the second frequency domain signal into a plurality of non-overlapping frequency bands to obtain a plurality of frequency bands; Calculate the energy of each frequency band respectively, and obtain the energy peak value and energy valley value of each frequency band; The spectral contrast of each frequency band is calculated based on the energy peak and energy valley of each frequency band.

6. The method for dynamically adjusting microphone volume according to claim 5, characterized in that: The method for determining the first gender judgment result of the person who sends the audio signal by combining the Mel frequency cepstrum coefficient, the spectrum centroid and the spectrum contrast comprises: Align the Mel-frequency cepstral coefficients, spectral centroids, and spectral contrast to the same number of frames; The aligned Mel-frequency cepstrum coefficients, spectral centroids and spectral contrasts are concatenated into a feature matrix, which is recorded as the first concatenated feature matrix; Normalize the first concatenated feature matrix; Assigning weights to different features in the normalized first concatenated feature matrix to obtain a weighted feature matrix; The weighted feature matrix is ​​compared with a preset constant matrix. If it is greater than or equal to the constant matrix, the first gender judgment result is female. Otherwise, the first gender judgment result is male.

7. The method for dynamically adjusting microphone volume according to claim 5, characterized in that: The convolutional neural network used in the gender determination model includes a first convolutional neural network and a second convolutional neural network; The method of inputting the Mel frequency cepstrum coefficients, the spectral centroid and the spectral contrast into the gender judgment model to obtain the second gender judgment result of the person who sends the audio signal includes: The Mel frequency cepstral coefficients, the spectral centroid and the spectral contrast are spliced ​​into a second spliced ​​feature matrix by frame; The second concatenated feature matrix is ​​input into the gender judgment model. In the gender judgment model, a first convolutional neural network is used to extract a first local feature from the Mel frequency cepstral coefficient; a second convolutional neural network is used to extract statistical features from the spectral centroid and spectral contrast; a residual convolutional network is used to process the speech spectrogram to obtain a second local feature; a recurrent neural network is used to capture the dependency in the time series to obtain a plurality of third local features; the speech spectrogram includes a first frequency domain signal and a second frequency domain signal; the time series is obtained by arranging the first local feature, the statistical feature and the second local feature in chronological order; The weighted average of multiple third local features is used to obtain the second gender judgment result.

8. A microphone volume dynamic adjustment system, characterized in that: include: A signal acquisition module, used to continuously acquire the audio signal received by the microphone within a preset time range, and determine whether the stability of the audio signal meets the preset requirements. If so, dynamically adjust the output volume of the microphone according to the distance result and the final gender judgment result; otherwise, dynamically adjust the output volume of the microphone according to the final gender judgment result; the distance result and the final gender judgment result are obtained through the output of the distance detection module and the gender judgment module; A distance detection module is used to perform distance detection on the audio signal, determine the distance between the source of the audio signal and the microphone, and obtain a distance detection result; The gender judgment module is used to extract the Mel-frequency cepstral coefficients, spectral centroid and spectral contrast from the audio signal respectively, and determine the first gender judgment result of the person who sends the audio signal by combining the Mel-frequency cepstral coefficients, spectral centroid and spectral contrast; input the Mel-frequency cepstral coefficients, spectral centroid and spectral contrast into the gender judgment model to obtain the second gender judgment result of the person who sends the audio signal; the gender judgment model adopts a hybrid model structure of a convolutional neural network, a residual convolutional network and a recurrent neural network; if the first gender judgment result and the second gender judgment result are consistent, the consistent result is used as the final gender judgment result.

9. An electronic device, characterized in that: include: A memory and one or more processors; the memory is coupled to the processor; wherein the memory stores computer program code, the computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device executes the steps of the microphone volume dynamic adjustment method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for dynamically adjusting the microphone volume as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Howling suppression method based on feedback signal spectrum estimation

    CN102740214A

  • Microphone control method and microphone

    CN107547978A

  • Selecting a microphone based on estimated proximity to sound source

    US20200301651A1