Voice Data Processing Method and Apparatus, Voice Data Processing System, and Electronic Device
By comparing the analog voice signal with gain amplification and reference voltage comparison, combined with the adaptively adjusted gain and voltage threshold judgment, the complex circuit and high power consumption problems in traditional voice endpoint detection are solved, and efficient voice segment and non-voice segment detection is achieved, reducing power consumption and improving detection accuracy.
Patent Information
- Application Number
- CN202110169539.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-07
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-02-07
AI Technical Summary
Traditional voice endpoint detection technology has problems of complex circuits and high power consumption. Especially in voice control devices with high power consumption requirements, it is difficult to achieve efficient precise detection of voice segments and non-voice segments.
By comparing the analog voice signal with the reference voltage Vref, the number ratio and threshold are calculated to judge, the gain and reference voltage are adaptively adjusted to realize the detection of voice segments or non-voice segments, combined with a simple hardware circuit design to reduce power consumption.
It realizes accurate detection of voice segments or non-voice segments under different analog voice signals input, reducing power consumption at the analog end, ensuring detection accuracy and simplifying hardware circuit design.
Smart Images

Figure CN114913879B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech detection, and in particular, to a method and device for processing speech data, a speech data processing system, and an electronic device. Background Art
[0002] The main function of voice activity detection (VAD) is to detect the start point and end point of a speech signal. The voice activity detection at the analog end mainly performs initial detection of the start point and end point of the analog signal, and then sends the detected speech signal to the digital speech processing module for precise digital-end voice activity detection, speech recognition, noise reduction, etc.
[0003] The voice activity detection at the analog end is one of the more important functional modules in speech signal processing, and is widely used in electronic products such as Bluetooth headsets and smart speakers with high power consumption requirements. The voice activity detection at the analog end is generally a bridge connecting the analog speech signal and the digital speech signal, and plays a role in waking up the digital speech processing module by detecting the analog speech signal, which can reduce the continuous use of the digital signal processor at the digital speech processing module end to a certain extent. The voice activity detection at the analog end has a very large effect on reducing the system power consumption.
[0004] Traditionally, a fixed threshold is generally used to judge the energy value sequence of the analog-end speech signal, or the signal-to-noise ratio of the analog signal is calculated and compared with the threshold to classify the analog speech signal into speech segments (such as user speech input) and non-speech segments (such as ambient noise). The implementation of traditional technical solutions usually has the drawbacks of complex circuits and high power consumption. Summary of the Invention
[0005] Based on the above situation, the main object of the present invention is to provide a method and device for processing speech data, a speech data processing system, and an electronic device.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows: A method for processing speech data, comprising: Step S1: receiving an analog speech signal; Step S2: performing gain amplification processing on the analog speech signal; Step S3: comparing the voltage of the speech signal after gain amplification processing with a reference voltage V ref to obtain a first comparison result, where the first comparison result indicates whether the voltage of the speech signal after gain amplification processing is greater than the reference voltage V ref; Step S4: Calculate the quantity ratio of the first comparison results within the t time period, compare the quantity ratio with a preset first threshold P to obtain a second comparison result; calculate the sum num_T of the second comparison results within a preset T time, where the T time includes n equivalent t time periods; compare num_T with a preset second threshold N and determine whether the analog voice signal corresponds to a voice segment or a non-voice segment based on the comparison result. If the judgment result is a voice segment, wake up the digital voice processing module to enter the voice state. If the judgment is a non-voice segment, the digital voice processing module remains in the silent state; and Step S5: Calculate the quantity ratio of the first comparison results within the S_T time, increase or decrease the gain and the reference voltage V according to the comparison result between the quantity ratio and the first threshold P ref for determining whether the analog voice signal within the next S_T time corresponds to a voice segment or a non-voice segment.
[0007] Preferably, the gain of the gain amplification process in Step S2 matches the value of the reference voltage V ref numerically.
[0008] Preferably, in Step S5, determine the adjustment step size of the gain and the reference voltage V according to the difference between the quantity ratio and the first threshold P ref of.
[0009] Preferably, in Step S3, the first comparison result is 0 or 1; Step S4 includes:
[0010] Step S41: Calculate the quantity ratio a0 of 0 and 1 within the t time, the quantity ratio ɑ0 = quantity value 0 / quantity value all, quantity value all is the sum of the quantity values of 0 and 1 within the t time, quantity value 0 is the quantity value number of 0 within the t time, and quantity value 1 is the quantity value number of 1 within the t time; Step S42: Compare the quantity ratio a0 with the first threshold P. If a0 < P, the second comparison result num_temp = 0. If a0 ≥ P, the second comparison result num_temp = 1; Step S43: Calculate num_T = sum(num_temp1 +... num_tempn), the T time includes n equivalent t time periods t1, t1, ······ tn, and num_tempn is the second comparison result obtained by comparing the quantity ratio a0 with the first threshold P within the tn time; and Step S44: Compare num_T with the second threshold N. If num_T > N, then judge that tn is a non-voice segment. If num_T ≤ N, then judge that tn is a voice segment.
[0011] Preferably, the second threshold values adopted in the voice state and the silent state are different.
[0012] Preferably, the second thresholds adopted in the voice state and the silent state are different, and the value of the second threshold in the voice state is greater than the value of the second threshold in the silent state.
[0013] Preferably, step S5 includes: step S51: calculating the quantity ratio a0 of the number of 0s and 1s within the time period S_T, where the quantity ratio ɑ0 = quantity value of 0 / quantity value of all, the quantity value of all is the sum of the quantity values of 0 and 1 within the time period t, the quantity value of 0 is the number of quantity values of 0 within the time period t, and the quantity value of 1 is the number of quantity values of 1 within the time period t; and step S52: comparing the quantity ratio a0 with the first threshold P. If a0 < P, increasing the gain and the reference voltage; if a0 ≥ P, decreasing the gain and the reference voltage.
[0014] Preferably, in the voice data processing method, the gain and the reference voltage V are adjusted in a time period of S_T ref , and the duration of the time period S_T is greater than the duration of the time period T.
[0015] Preferably, the number of voice signal frames included in the time period t is 4 - 10 frames, the number of voice signal frames included in the time period T is 70 - 95 frames, and the number of data frames included in the time period S_T is 100 - 200 frames.
[0016] The present invention also provides a voice data processing device, including: a voice receiving module for receiving an analog voice signal; an amplifying module for performing gain amplification processing on the analog voice signal; a comparing module for comparing the voltage of the voice signal after gain amplification processing with a reference voltage V ref to obtain a first comparison result, where the first comparison result represents whether the voltage of the voice signal after gain amplification processing is greater than the reference voltage V ref ; a processing module for calculating the quantity ratio of the first comparison results within the time period t, comparing the quantity ratio with a preset first threshold P to obtain a second comparison result; calculating the sum num T of the second comparison results within the time period T, where the time period T includes n identical time periods t; comparing num T with a preset second threshold N and judging whether the analog voice signal corresponds to a voice segment or a non - voice segment according to the comparison result. If the judgment result is a voice segment, waking up the digital voice processing module to enter the voice state. If the judgment is a non - voice segment, the digital voice processing module remains in the silent state; and an adjusting module for calculating the quantity ratio of the first comparison results within the time period S_T, and increasing or decreasing the gain and the reference voltage V according to the comparison result of the quantity ratio and the first threshold P ref for judging whether the analog voice signal in the next time period S_T corresponds to a voice segment or a non - voice segment.
[0017] Preferably, the comparison module sets a threshold voltage, which is divided into multiple levels of voltage. The comparison module selects one level of voltage that matches the gain value as the reference voltage.
[0018] Preferably, the adjustment module determines the gain and the reference voltage V according to the difference between the ratio of the number of the first comparison results within the S_T time and the first threshold P. ref of the adjustment step.
[0019] Preferably, the processing module includes: a first calculation module for calculating the ratio a0 of the number of 0s and 1s within the t time, where the ratio ɑ0 = quantity value 0 / quantity value all, quantity value all is the sum of the quantity value 0 and the quantity value 1 within the t time, quantity value 0 is the number of quantity values of 0 within the t time, and quantity value 1 is the number of quantity values of 1 within the t time; a second calculation module for calculating num temp, comparing the ratio a0 with the first threshold P, if a0 < P, the second comparison result num temp = 0, if a0 ≥ P, the second comparison result num temp = 1; a third calculation module for calculating num T = sum(num temp1 +... num tempn), the T time includes several n equal t time periods t1, t1,..., tn, and num tempn is the second comparison result obtained by comparing the ratio a0 with the first threshold P within the tn time; and a fourth calculation module for comparing num T with the second threshold N, if num T > N, it is determined that tn is a non-speech segment, if num T ≤ N, it is determined that tn is a speech segment.
[0020] Preferably, the second threshold values adopted in the speech state and the silent state are different.
[0021] Preferably, the adjustment module includes: a first adjustment calculation module for calculating the ratio a0 of the number of 0s and 1s within the S_T time, where the ratio ɑ0 = quantity value 0 / quantity value all, quantity value all is the sum of the quantity value 0 and the quantity value 1 within the t time, quantity value 0 is the number of quantity values of 0 within the t time, and quantity value 1 is the number of quantity values of 1 within the t time; and a second adjustment calculation module for comparing the ratio a0 with the first threshold P, if a0 < P, increasing the gain and the reference voltage; if a0 ≥ P, decreasing the gain and the reference voltage.
[0022] Preferably, the speech data processing device adjusts the gain and the reference voltage V in a cycle of S_T time. ref The duration of the S_T time is greater than the duration of the t time.
[0023] The present invention further provides a voice data processing system, including the voice data processing device and the digital voice processing module as described above. When the voice data processing device detects a voice segment, it wakes up the digital voice processing module to work.
[0024] The present invention further provides an electronic device, which includes the voice data processing system as described above.
[0025] Preferably, the electronic device is a mobile phone, a speaker, a headset, a sports bracelet, a computer, a recording pen or an electronic toy.
[0026] Compared with the prior art, the voice data processing method provided by the present invention performs a gain amplification process on the analog voice signal and then compares it with the reference voltage V ref to obtain a first comparison result, calculates the quantity ratio of the first comparison results within the time t, compares the quantity ratio with a preset first threshold P to obtain a second comparison result; calculates the sum num_T of the second comparison results within the time T, and the time T includes n time periods of t; compares num_T with a preset second threshold N and determines whether the analog voice signal corresponds to a voice segment or a non-voice segment according to the comparison result. If the judgment result is a voice segment, the digital voice processing module is woken up to enter the voice state. If it is judged as a non-voice segment, the digital voice processing module remains in the silent state. Calculates the quantity ratio of the first comparison results within the time S_T, and increases or decreases the gain and the reference voltage V ref according to the comparison result of the quantity ratio and the first threshold P. It realizes the application of different gains and reference voltages under different inputs of analog voice signals. Specifically, when the quantity ratio within the S_T period of the analog voice signal is greater than the first threshold P, the reference voltage and the corresponding selected gain are increased. When the quantity ratio within the S_T period of the analog voice signal is less than the first threshold P, the reference voltage and the selected corresponding gain are decreased. In this way, the analog voice signal can be better amplified or reduced to more precisely correspond to a certain quantized reference voltage Vref, achieving the effective utilization of efficiency as a whole, reducing the overall power consumption of detecting voice segments or non-voice segments at the analog end, and at the same time ensuring the detection accuracy of voice segments or non-voice segments. The voice data processing method provided by the present invention can be realized by a simple hardware design circuit scheme, such as a combination of a microphone, an amplifier, a comparator and a DSP can implement this scheme. The hardware circuit is simple, which can reduce the power consumption of analog-end voice data processing. In the case of a simple hardware circuit, the method of calculating the noise energy of the analog voice signal through the first comparison result (two-level threshold comparison) further effectively guarantees the detection accuracy of voice segments or non-voice segments.
[0027] In the present invention, the gain matches the value of the reference voltage, and the matching of the two guarantees the detection accuracy of voice segments or non-voice segments.
[0028] In the present invention, by calculating the quantity ratio a0 or a1, and then comparing it with the first threshold to calculate num temp, and then calculating the sum num T of num temp within the T time period, comparing num T with the second threshold, and judging whether the middle time period within the T time period is a speech segment or a non-speech segment according to the comparison result. By comparing the two thresholds, it is judged whether the middle time period within the T time period is a speech segment or a non-speech segment. The calculation method is simple and the detection accuracy is high.
[0029] In the present invention, by adaptively adjusting the reference voltage and gain in real time, the gain and the reference voltage are always adjusted to adapt to the current analog voice signal input situation, reducing energy consumption and ensuring the detection accuracy of speech segments or non-speech segments.
[0030] The device, speech data processing system and electronic device corresponding to the speech data processing method provided by the present invention also have the above advantages.
[0031] Other beneficial effects of the present invention will be described by introducing specific technical features and technical solutions in the specific implementation manner. Those skilled in the art should be able to understand the beneficial technical effects brought by the technical features and technical solutions through these introductions. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Through the following description of the embodiments of the present invention with reference to the drawings, the above and other objects, features and advantages of the present invention will become clearer. In the drawings:
[0033] Figure 1 Schematic flow chart of a speech data processing method in the present invention.
[0034] Figure 2 For Figure 1 Detailed flow chart of the specific implementation manner of step S4 therein.
[0035] Figure 3 For Figure 2 Schematic diagram for defining the T time period in step S43 therein.
[0036] Figure 4 Schematic diagram of the modules of the speech data processing device in the present invention.
[0037] Figure 5 Schematic diagram of the modules of the speech data processing system in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] The present invention will be described based on embodiments, but the present invention is not limited to these embodiments. In the following detailed description of the present invention, some specific details are described in detail. In order to avoid obscuring the essence of the present invention, well-known methods, processes, procedures, and components are not described in detail.
[0039] In addition, those of ordinary skill in the art should understand that the drawings provided herein are for illustrative purposes only, and the drawings are not necessarily drawn to scale.
[0040] Unless the context clearly requires otherwise, the words "including", "comprising", and the like in the entire specification and claims should be construed in an inclusive sense rather than an exclusive or exhaustive sense; that is, the meaning of "including but not limited to".
[0041] In the description of the present invention, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, the meaning of "a plurality of" is two or more.
[0042] First Embodiment
[0043] Please refer to Figure 1 , the first embodiment of the present invention discloses a method for processing voice data, which includes:
[0044] Step S1: Receive an analog voice signal; specifically, the source of the analog voice signal can be emitted by a user or a certain device. The device for receiving the analog voice signal can be any device capable of receiving an analog voice signal, such as a microphone.
[0045] Step S2: Perform gain amplification processing on the analog voice signal; specifically, let the gain be gain_db, and amplify the analog voice signal for further signal processing.
[0046] As a preferred embodiment, the amplification of the analog voice signal is achieved through an amplifier:
[0047]
[0048] v in is the voltage value of the original analog voice signal, v out is the voltage value of the analog voice signal amplified by the amplifier, and gain_db is the gain of the amplifier, and its unit is db. In the present invention, the gain is an adjustable parameter. As an embodiment, the gain value range is between -6 db and 42 db. When adjusting the gain, the gain interval is 1 db to 5 db, preferably 3 db.
[0049] It can be understood that the amplification process in step S2 is not limited to being implemented by an amplifier, and it can also be implemented by other software and / or hardware circuits capable of implementing the amplification function.
[0050] Step S3: Compare the voltage of the voice signal after gain amplification with a reference voltage V ref to obtain a first comparison result, where the first comparison result characterizes whether the voltage of the voice signal after gain amplification is greater than the reference voltage V ref ; As an embodiment, after amplification, the analog voice signal and the reference voltage V ref are compared to obtain a first comparison result of 0 or 1.
[0051] As a preferred embodiment, step S3 can be implemented by a comparator. As an embodiment, the comparator sets a threshold voltage, which is divided into multiple levels of voltage, and one level of voltage is selected as the reference voltage. The reference voltage is matched with the gain value, that is, there is a set correspondence between the gain and the reference voltage. Through the set correspondence, each gain matches an optimal reference voltage. In the present invention, the reference voltage is an adjustable parameter. As another embodiment, the reference voltage is adjusted in fixed or non-fixed steps. The reference voltage V ref shall not be lower than 5V and not higher than 15V.
[0052] It can be understood that the comparison operation in step S2 is not limited to being implemented by a comparator, and it can also be implemented by other software and / or hardware circuits capable of implementing the comparison operation.
[0053] Step S4: Calculate the quantity ratio of the first comparison results within t time, and compare the quantity ratio with a preset first threshold P to obtain a second comparison result; calculate the sum num T of the second comparison results within T time, where T time includes n t times; compare num T with a preset second threshold N and judge whether the analog voice signal corresponds to a voice segment or a non-voice segment according to the comparison result. If the judgment result is a voice segment, wake up the digital voice processing module to enter the voice state. If the judgment is a non-voice segment, the digital voice processing module remains in the silent state.
[0054] In step S4, calculate the noise energy of the analog voice signal according to the first comparison result, and judge whether the analog voice signal corresponds to a voice segment or a non-voice segment according to the noise energy. The judgment result is used to determine whether to wake up the digital voice processing module;
[0055] Please refer to Figure 2 , as a specific embodiment, step S4 includes:
[0056] Step S41: Calculate the quantity ratio a0 of 0 and 1 within time t. The quantity ratio ɑ0 = quantity value of 0 / quantity value of all, where the quantity value of all is the sum of the quantity values of 0 and 1 within time t, the quantity value of 0 is the number of quantity values of 0 within time t, and the quantity value of 1 is the number of quantity values of 1 within time t. For example: within time t, the comparison results include 5 zeros and 15 ones, that is, the quantity value of 0 is 5, the quantity value of 1 is 15, and ɑ0 = 5 / (5 + 15).
[0057] Step S42: Compare the quantity ratio a0 with the first threshold P. If a0 < P, the second comparison result num temp = 0; if a0 ≥ P, the second comparison result num temp = 1. For example: the first threshold P is 0.8, ɑ0 = 5 / (5 + 15) = 0.25, ɑ0 is less than 0.8, and num temp = 0.
[0058] Step S43: Calculate num T = sum(num temp1 +... num tempn), please refer to Figure 3 , where time T contains several n equal time periods t1, t1,..., tn, and num tempn is the second comparison result obtained by comparing the quantity ratio a0 with the first threshold P within time tn. For example: time T contains several 4 equal time periods t1, t2, t3, t4, and the second comparison results within time periods t1, t2, t3, t4 are: 0, 1, 1, 1. num T = 3.
[0059] Step S44: Compare num T with the second threshold N. If num T > N, then judge that tn is a non-speech segment; if num T ≤ N, then judge that tn is a speech segment. It can be understood that this judgment result can be considered as corresponding to any tn within time T, or it can be considered as corresponding to time T. For example, if the second threshold N is set to 2 and num T = 3, num T is greater than the second threshold N, then judge that tn or time T is a non-speech segment.
[0060] It can be understood that in step S4, the calculation method for calculating the noise energy of the analog voice signal according to the first comparison result is not limited, and it can also be performed by other means such as establishing a noise energy calculation model. In the above calculation method, it can be understood that in step S41, the quantity ratio a1 can also be calculated, and the quantity ratio ɑ1 = magnitude 1 / magnitude all. Correspondingly, in step S42, the quantity ratio a1 is compared with a preset first threshold P. If a0 < P, the second comparison result num temp = 0; if a0 ≥ P, the second comparison result num temp = 1. num T is compared with a preset second threshold N. If num T > N, it is determined that the time period tn is a voice segment; if num T ≤ N, it is determined that the time period tn is a non-voice segment. It can be understood that calculating the quantity ratio a1 is an equivalent way of calculating the quantity ratio a0, and both belong to the protection scope of the present invention.
[0061] As an embodiment, the second threshold values adopted in the voice state and the silent state are different. In this way, the calculation accuracy can be improved. As an embodiment, when the quantity ratio calculated is a0, the second threshold in the voice state is greater than the second threshold in the silent state.
[0062] In step S4, when it is determined that the analog voice signal corresponds to a non-voice segment, the digital voice processing module will not be awakened and will remain in the silent state. When it is determined that the analog voice signal corresponds to a voice segment, the digital voice processing module is awakened to enter the voice state, that is, the digital voice processing module is awakened to work, and the digital voice processing module performs precise digital-end voice endpoint detection, voice recognition, noise reduction, etc.
[0063] Step S5: Calculate the quantity ratio of the first comparison results within the time period S_T, and increase or decrease the gain and the reference voltage V according to the comparison result between this quantity ratio and the first threshold P ref For determining whether the analog voice signal within the next time period S_T corresponds to a voice segment or a non-voice segment.
[0064] As an embodiment, step S5 includes:
[0065] Step S51: Calculate the quantity ratio a0 of 0 and 1 within the time period S_T, and the quantity ratio ɑ0 = magnitude 0 / magnitude all, where magnitude all is the sum of the magnitudes of 0 and 1 within the time period t, magnitude 0 is the number of magnitude values of 0 within the time period t, and magnitude 1 is the number of magnitude values of 1 within the time period t; and
[0066] Step S52: Compare the quantity ratio a0 with the first threshold P. If a0 < P, increase the gain and the reference voltage; if a0 ≥ P, decrease the gain and the reference voltage.
[0067] It can be understood that the calculation method of the quantitative ratio of the first comparison result in step S5 is consistent with the calculation method mentioned in step S4. If the quantitative ratio calculated is a1, then the corresponding a1<P, reduce the gain and the reference voltage; if a1≥P, increase the gain and the reference voltage. In the next S_T time, the calculations of steps S2 and S3 are performed with the adjusted gain and reference voltage. In this way, adaptive real-time adjustment of the gain and reference voltage is achieved. It can be understood that the frequency and step size of adjusting the gain and reference voltage can be freely selected. In the present invention, the gain and the reference voltage V are adjusted with S_T time as a period. ref As an embodiment, in step S5, the gain and the reference voltage V are determined according to the difference between the quantity ratio and the first threshold value P. ref The larger the difference between the quantity ratio and the first threshold value P, the longer the adjustment step, and vice versa. It can be understood that the increase or decrease of the reference voltage and gain is not limited to the aforementioned adjustment methods, and as long as the adjustment purpose can be achieved, it belongs to the protection scope of the present invention.
[0068] As an embodiment, the number of voice signal frames included in the t time period is 4-10 frames, the number of voice signal frames included in the T time period is 70-95 frames, and the number of data frames included in the S_T time period is 100-200 frames. Under this parameter, the adaptive adjustment can be carried out stably and the calculation accuracy is also guaranteed.
[0069] As an implementation method, the adjustment of the reference voltage and the gain is linked. Since there is a set corresponding relationship between the gain and the voltage, when one of the gain or the reference voltage is adjusted, the other is adjusted accordingly according to the set corresponding relationship. For example, after adjusting the reference voltage, the corresponding gain is selected. As another implementation method, the adjustment of the reference voltage and the gain is performed separately, and the two are adjusted according to the comparison result of the quantity ratio within the S_T time and the first threshold value P. In this implementation method, the reference voltage and the gain still satisfy the set corresponding relationship. It can be understood that the aforementioned methods all belong to the protection scope of the present invention.
[0070] See also Figure 4 The present invention provides a voice data processing device 20, which is used to perform voice endpoint detection of an analog end. The voice data processing device 20 includes:
[0071] The voice receiving module 21 is used to receive analog voice signals.
[0072] The amplification module 22 is used to perform gain amplification processing on the analog voice signal.
[0073] A comparison module 23 for comparing the voltage of the speech signal after gain amplification processing with a reference voltage V ref to obtain a first comparison result, where the first comparison result characterizes whether the voltage of the speech signal after gain amplification processing is greater than the reference voltage V ref .
[0074] A processing module 24 for calculating the quantity ratio of the first comparison results within a time t, comparing the quantity ratio with a preset first threshold P to obtain a second comparison result; calculating the sum num T of the second comparison results within a time T, where the time T includes n equivalent time periods t; comparing num T with a preset second threshold N and judging whether the analog speech signal corresponds to a speech segment or a non-speech segment according to the comparison result. If the judgment result is a speech segment, the digital speech processing module is awakened to enter the speech state. If the judgment is a non-speech segment, the digital speech processing module remains in the silent state; and
[0075] An adjustment module 25 for calculating the quantity ratio of the first comparison results within a time S_T, and increasing or decreasing the gain and the reference voltage V according to the comparison result between the quantity ratio and the first threshold P ref for judging whether the analog speech signal in the next time S_T corresponds to a speech segment or a non-speech segment.
[0076] As an embodiment, the comparison module sets a threshold voltage, which is divided into multiple levels of voltages. The comparison module selects one of the levels of voltages that matches the gain value as the reference voltage.
[0077] As an embodiment, the processing module 24 includes:
[0078] A first calculation module for calculating the quantity ratio a0 of 0 and 1 within a time t, where the quantity ratio ɑ0 = quantity value 0 / quantity value all, quantity value all is the sum of the quantity values of 0 and 1 within the time t, quantity value 0 is the quantity value number of 0 within the time t, and quantity value 1 is the quantity value number of 1 within the time t.
[0079] A second calculation module for calculating num temp. The quantity ratio a0 is compared with the first threshold P. If a0 < P, the second comparison result num temp = 0. If a0 ≥ P, the second comparison result num temp = 1.
[0080] A third calculation module for calculating num T = sum(num temp1 +... num tempn). The time T includes several n equivalent time periods t1, t1,..., tn, and num tempn is the second comparison result obtained by comparing the quantity ratio a0 with the first threshold P within the time tn.
[0081] A fourth calculation module is configured to compare num T with a second threshold N. If num T > N, then tn is determined as a non-speech segment; if num T ≤ N, then tn is determined as a speech segment.
[0082] As an embodiment, the second threshold values adopted in the speech state and the silent state are different.
[0083] As an embodiment, the adjustment module determines the gain and the reference voltage V according to the difference between the quantity ratio of the first comparison results within the S_T time and the first threshold P. ref of the adjustment step size.
[0084] As an embodiment, the voice data processing device adjusts the gain and the reference voltage V in a time period of S_T. ref The duration of the S_T time is greater than the duration of the t time.
[0085] It can be understood that the processing module can be a CPU, a DSP, an MCU, etc. The adjustment module can be integrated with the processing module physically or can be separately arranged.
[0086] The voice data processing device can be regarded as the device corresponding to the voice data processing method, and the content disclosed in the voice data processing method is applicable to the voice data processing device.
[0087] Second Embodiment
[0088] Please refer to Figure 5 , the second embodiment of the present invention provides a voice data processing system 30, including a voice data processing device 31 and a digital voice processing module 32 as disclosed in the first embodiment. When the voice data processing device 31 detects a speech segment, it wakes up the digital voice processing module 32 to work. The digital voice processing module 32 performs precise digital-end voice endpoint detection, speech recognition, noise reduction, etc. on the voice signal.
[0089] Third Embodiment
[0090] The third embodiment of the present invention provides an electronic device (not shown), and the electronic device includes the voice data processing system as described in the second embodiment. The electronic device can be a mobile phone, a speaker, an earphone, a sports bracelet, a computer, a recording pen, or an electronic toy.
[0091] Those skilled in the art can understand that, on the premise of no conflict, the above-mentioned preferred solutions can be freely combined and superimposed.
[0092] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of methods, apparatuses, systems, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0093] It should be understood that the above-described embodiments are merely exemplary and not restrictive. Without departing from the basic principles of the present invention, various obvious or equivalent modifications or substitutions that those skilled in the art can make to the above details will all be included within the scope of the claims of the present invention.
Claims
1. A method for processing voice data, characterized in that, Including: Step S1: Receive an analog voice signal; Step S2: Perform gain amplification processing on the analog voice signal; Step S3: Compare the voltage of the voice signal after gain amplification processing with a reference voltage Vref to obtain a first comparison result, the first comparison result being 0 or 1, and the first comparison result indicating whether the voltage of the voice signal after gain amplification processing is greater than the reference voltage Vref; Step S4: Calculate the quantity ratio of the first comparison results within a t time period, and compare the quantity ratio with a preset first threshold P to obtain a second comparison result; Calculate the sum num_T of the second comparison results within a preset T time, where the T time includes n identical t time periods; compare num_T with a preset second threshold N and judge whether the analog voice signal corresponds to a voice segment or a non-voice segment according to the comparison result. If the judgment result is a voice segment, wake up the digital voice processing module to enter the voice state. If the judgment is a non-voice segment, the digital voice processing module remains in the silent state; and Step S5: Calculate the quantity ratio of the first comparison results within an S_T time period, and increase or decrease the gain and the reference voltage Vref according to the comparison result of the quantity ratio and the first threshold P to judge whether the analog voice signal within the next S_T time period corresponds to a voice segment or a non-voice segment; The calculation of the quantity ratio of the first comparison results within the t time period includes: Step S41: Calculate the quantity ratio a0 of 0 and 1 within the t time, the quantity ratio ɑ0 = quantity value 0 / quantity value all, where quantity value all is the sum of the quantity values of 0 and 1 within the t time, quantity value 0 is the number of quantity values of 0 within the t time, and quantity value 1 is the number of quantity values of 1 within the t time; The calculation of the quantity ratio of the first comparison results within the S_T time period includes: Step S51: Calculate the quantity ratio a0 of 0 and 1 within the S_T time, the quantity ratio ɑ0 = quantity value 0 / quantity value all, where quantity value all is the sum of the quantity values of 0 and 1 within the t time, quantity value 0 is the number of quantity values of 0 within the t time, and quantity value 1 is the number of quantity values of 1 within the t time; In the voice data processing method, the gain and the reference voltage Vref are adjusted in the S_T time period , The duration of the S_T time is longer than the duration of the T time.
2. The voice data processing method according to claim 1, characterized in that, In step S2, the gain of the gain amplification processing matches the value of the reference voltage Vref.
3. The voice data processing method according to claim 1, characterized in that, In step S5, determine the adjustment step size of the gain and the reference voltage Vref according to the difference between the quantity ratio and the first threshold P.
4. The voice data processing method according to claim 1, characterized in that The step S4 further includes: Step S42: Compare the quantity ratio a0 with the first threshold P. If a0 < P, the second comparison result num_temp = 0. If a0 ≥ P, the second comparison result num_temp = 1; Step S43: Calculate num_T = sum(num_temp1 +... num_tempn), the T time includes n identical t time periods t1, t1,..., tn, and num_tempn is the second comparison result obtained by comparing the quantity ratio a0 with the first threshold P within the tn time; and Step S44: Compare num_T with the second threshold N. If num_T > N, judge that tn is a non-voice segment. If num_T ≤ N, judge that tn is a voice segment.
5. The voice data processing method according to any one of claims 1-4, characterized in that The second threshold values adopted in the voice state and the silent state are different.
6. The voice data processing method according to claim 4, wherein The second threshold values adopted in the voice state and the silent state are different, and the second threshold value in the voice state is greater than the second threshold value in the silent state.
7. The voice data processing method according to any one of claims 1-4, characterized in that, The step S5 includes: Step S52: Compare the quantity ratio a0 with the first threshold P. If a0 < P, increase the gain and the reference voltage; if a0 ≥ P, decrease the gain and the reference voltage.
8. The voice data processing method according to any one of claims 1-4, characterized in that The number of voice signal frames included in the t time period is 4 - 10 frames, the number of voice signal frames included in the T time period is 70 - 95 frames, and the number of data frames included in the S_T time period is 100 - 200 frames.
9. A voice data processing device, characterized in that, It includes: A voice receiving module for receiving an analog voice signal; An amplifying module for performing gain amplification processing on the analog voice signal; A comparison module for comparing the voltage of the voice signal after gain amplification processing with a reference voltage Vref to obtain a first comparison result, where the first comparison result represents whether the voltage of the voice signal after gain amplification processing is greater than the reference voltage Vref; A processing module for calculating the quantity ratio of the first comparison results within the t time, and comparing the quantity ratio with a preset first threshold P to obtain a second comparison result; Calculate the sum num T of the second comparison results within the T time, where the T time includes n identical t time periods; compare num T with a preset second threshold N and judge whether the analog voice signal corresponds to a voice segment or a non-voice segment according to the comparison result. If the judgment result is a voice segment, wake up the digital voice processing module to enter the voice state. If the judgment is a non-voice segment, the digital voice processing module remains in the silent state; and An adjustment module for calculating the quantity ratio of the first comparison results within the S_T time, and increasing or decreasing the gain and the reference voltage Vref according to the comparison result of the quantity ratio and the first threshold P to be used for judging whether the analog voice signal in the next S_T time corresponds to a voice segment or a non-voice segment; The processing module includes: a first calculation module for calculating the quantity ratio a0 of 0 and 1 within the t time, and the quantity ratio ɑ0 = quantity value 0 / quantity value all, where quantity value all is the sum of the quantity values of 0 and 1 within the t time, quantity value 0 is the number of quantity values of 0 within the t time, and quantity value 1 is the number of quantity values of 1 within the t time; The adjustment module includes: a first adjustment calculation module for calculating the quantity ratio a0 of 0 and 1 within the S_T time, and the quantity ratio ɑ0 = quantity value 0 / quantity value all, where quantity value all is the sum of the quantity values of 0 and 1 within the t time, quantity value 0 is the number of quantity values of 0 within the t time, and quantity value 1 is the number of quantity values of 1 within the t time; The voice data processing device adjusts the gain and the reference voltage Vref with a period of S_T time, and the duration of the S_T time is greater than the duration of the t time.
10. The voice data processing device according to claim 9, wherein, The comparison module sets a threshold voltage, which is divided into multiple levels of voltage, and the comparison module selects one of the levels of voltage that matches the gain value as the reference voltage.
11. The voice data processing device according to claim 9, wherein, The adjustment module determines the adjustment steps of the gain and the reference voltage Vref according to the difference between the quantity ratio of the first comparison results within the S_T time and the first threshold P.
12. The voice data processing device according to claim 9, wherein The processing module includes: A second calculation module, configured to calculate num temp, compare the quantity ratio a0 with the first threshold P. If a0 < P, the second comparison result num temp = 0; if a0 ≥ P, the second comparison result num temp = 1; A third calculation module, configured to calculate num T = sum(num temp1 +... num tempn). The T time includes several n equal t time periods t1, t1,..., tn, and num tempn is the second comparison result obtained by comparing the quantity ratio a0 with the first threshold P within the tn time period; and A fourth calculation module, configured to compare num T with the second threshold N. If num T > N, it is determined that tn is a non-speech segment; if num T ≤ N, it is determined that tn is a speech segment.
13. The voice data processing device according to any one of claims 9-12, characterized in that, The second threshold values adopted in the speech state and the silent state are different.
14. The voice data processing device according to any one of claims 9-12, characterized in that, The adjustment module includes: A second adjustment calculation module; configured to compare the quantity ratio a0 with the first threshold P. If a0 < P, increase the gain and the reference voltage; if a0 ≥ P, decrease the gain and the reference voltage.
15. A voice data processing system, characterized in that, It includes the voice data processing device and the digital voice processing module according to any one of claims 9-14. When the voice data processing device detects a speech segment, it wakes up the digital voice processing module to work.
16. An electronic device, characterized in that: The electronic device includes the voice data processing system according to claim 15.
17. The electronic device according to claim 16, wherein: The electronic device is a mobile phone, a speaker, a headset, a sports bracelet, a computer, a recording pen or an electronic toy.
Citation Information
Patent Citations
Automatic gain control method and device
CN108573709A
Game voice chat volume adaptive adjustment method
CN109326298A