A method and system for collecting voice signals

By constructing an adaptive quantization function, combining the short-time energy and zero-crossing rate of the speech frame, dynamically adjusting the quantization strategy of the A-law thirteen-fold line, the problems of feature information loss and noise interference caused by traditional quantization strategies are solved, and the high fidelity and coding efficiency of speech signal acquisition are achieved.

CN120236592BActive Publication Date: 2025-08-05GUANGZHOU JIUSI INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510712168.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-08-05
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

Traditional fixed quantization strategies are prone to loss of feature information due to quantization errors in speech signal acquisition, affecting the intelligibility of speech, and may introduce additional noise interference, especially in modern intelligent voice applications with high-fidelity voice acquisition.

Method used

By analyzing the short-time energy and zero-crossing rate of the speech frame, dynamically adjusting the quantization strategy of the A-law thirteen-fold line, setting 8 positive segments and 8 negative segments, adjusting the quantization function adaptively according to the information density of the speech frame, realizing differentiated processing of different speech types.

Benefits of technology

While keeping the A-law encoding structure unchanged, the intelligent balance between speech quality and bit rate is improved, effectively avoiding the problems of loss of key speech components or noise amplification, and improving the robustness and listening quality of speech encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236592B_ABST
    Figure CN120236592B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of speech signal acquisition technology, and specifically relates to a speech signal acquisition method and system. The method comprises: sampling and framing an analog signal to obtain a plurality of speech frames; obtaining the information density of each speech frame based on the short-time energy and zero-crossing rate of each speech frame, wherein the information density is positively correlated with the short-time energy and zero-crossing rate; constructing a quantization function for each speech frame based on the information density of each speech frame; and quantizing and encoding each speech frame based on the quantization function to obtain a digital speech signal. The present invention adapts the quantization function of each speech frame to its information density, thereby improving speech clarity and intelligibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voice signal acquisition, and more particularly to a method and system for acquiring voice signals. Background Art

[0002] With the rapid development of technologies such as artificial intelligence, speech recognition, and intelligent voice interaction, the acquisition of high-quality voice signals has become a core link in improving the performance of voice processing systems.

[0003] Currently, traditional methods based on fixed quantization strategies still dominate the field of voice acquisition. For example, the A-law thirteen-fold quantization method, widely used in telephone communication systems, is used. This method simulates the characteristics of the human ear and employs a nonlinear quantization strategy: fine-grained quantization is applied to small signals, while moderate compression is applied to large signals. This approach achieves high coding efficiency while preserving subjective auditory quality.

[0004] However, in actual speech signal processing, traditional fixed quantization strategies have significant technical limitations due to the complex dynamic and time-varying characteristics of speech signals. Specifically, for speech components such as low-energy voiceless consonants and plosives with a large dynamic range, quantization errors can easily lead to feature information loss, significantly affecting speech intelligibility. Furthermore, during silent periods, ambient noise can be unduly amplified, introducing additional noise interference. These limitations are particularly prominent in modern intelligent speech applications that require high-fidelity speech acquisition.

[0005] Therefore, how to dynamically adjust the acquisition and quantization strategies according to the speech content to improve the fidelity of key speech segments and reduce data redundancy in non-key segments is an urgent problem that needs to be solved in the current field of speech signal acquisition. Summary of the Invention

[0006] To address the technical issues that the above-mentioned fixed quantization strategy is prone to loss of feature information due to quantization error, affecting speech intelligibility and potentially introducing additional noise interference, the present invention provides solutions in the following aspects.

[0007] In a first aspect, the present invention provides a method for collecting a speech signal, comprising:

[0008] An analog signal is collected by using a voice collection device, and the analog signal is sampled and framed to obtain a plurality of voice frames; the information density of each voice frame is obtained according to the short-time energy and the zero-crossing rate of each voice frame, and the information density is positively correlated with the short-time energy and the zero-crossing rate; a quantization function of each voice frame is constructed according to the information density of each voice frame, including: setting 8 positive segments, determining an adjustment index of each positive segment according to the information density of the voice frame and the sequence number of each positive segment; obtaining the adjusted size of each positive segment in the voice frame according to the adjustment index of each positive segment and the initial size of each positive segment, and constructing a quantization function of the voice frame according to the adjusted size; and quantizing and encoding each voice frame according to the quantization function of each voice frame to obtain a digital voice signal.

[0009] The present invention uses the two time-domain characteristics of speech frames, short-time energy and zero-crossing rate, to accurately evaluate the information density of speech frames. A quantization function is adaptively constructed based on the information density. By setting 8 positive segments and dynamically adjusting their ranges, non-uniform quantization of speech frames is achieved. The resulting digital speech signal retains key speech features such as low-energy voiceless consonants and plosives with a large dynamic range, while avoiding improper amplification of ambient noise in silent segments and ensuring coding efficiency.

[0010] Preferably, the obtaining of the information density of each speech frame based on the short-time energy and zero-crossing rate of each speech frame includes: performing weighted summation on the normalized result of the short-time energy of the speech frame and the normalized result of the zero-crossing rate of the speech frame to obtain the information density of the speech frame.

[0011] Short-time energy and zero-crossing rate reflect speech characteristics from the perspectives of amplitude domain and frequency domain respectively. The present invention conducts a collaborative analysis of short-time energy and zero-crossing rate to more comprehensively characterize the information content of speech frames, providing a more discriminative decision basis for the subsequent acquisition of the quantization function of the speech frames.

[0012] Preferably, the adjustment index satisfies the expression:

[0013] ;in, Indicates the The first paragraph Adjustment index under speech frames; Indicates the The information density of a speech frame; 、 is a hyperparameter; Indicates the sequence number of the positive segment; represents the maximum value function; Indicates the absolute value symbol.

[0014] The present invention adopts a segmentation strategy to generate different adjustment indices according to the information density of speech frames, thereby enhancing the representation capability of key speech components such as voiceless consonants, effectively suppressing noise amplification in silent segments, and protecting transient speech components such as plosives from being lost, thereby improving the fidelity and compression efficiency of speech coding in different semantic segments.

[0015] Preferably, the size of each positive segment after adjustment in the voice frame satisfies the expression: ;in, Indicates the The first paragraph The adjusted size under each speech frame; Indicates the The initial size of a positive segment; Indicates the The first paragraph The adjustment index under each speech frame.

[0016] Preferably, the method for obtaining the quantization function is as follows: the quantization function is a piecewise linear function, comprising 8 positive segments and 8 negative segments, wherein the 8 negative segments and the 8 positive segments are symmetrical about the origin; The quantization function of the speech frame The slope of the positive segment is , , Indicates the The first paragraph The size after adjustment under the speech frame; When, The range of the positive segment is ,when When, The range of the positive segment is , Indicates the The first paragraph The resized image is displayed for each speech frame.

[0017] The present invention adaptively adjusts the quantization strategy based on different speech content: in speech frames with extremely high information density (plosives), large signal segments are quantized more precisely to preserve key speech components. In speech frames with moderate to high information density (vowels), small signal segments are quantized more precisely to preserve key speech components. In speech frames with moderate to low information density (vowels), the quantization precision of large signal segments is appropriately relaxed to improve compression efficiency. In speech frames with extremely low information density (silence segments), the quantization precision of small signal segments is appropriately relaxed to improve compression efficiency while suppressing noise amplification. Furthermore, the quantization function in the present invention maintains consistency with the traditional A-law 13-zigzag coding structure, ensuring compatibility between the encoding and decoding process and existing PCM coding systems. It also possesses stronger content-awareness, significantly improving the overall quality and robustness of speech acquisition and encoding without adding additional complexity.

[0018] Preferably, the quantization encoding of each speech frame according to the quantization function of each speech frame includes: normalizing the signals of all sampling points in the speech frame to the range of [-1, 1], inputting the normalization result into the quantization function of the speech frame, outputting the quantization value of each sampling point, and mapping the quantization value into PCM encoding.

[0019] This invention uses different quantization functions for different speech frames, achieving differentiated quantization accuracy for different speech contents. This allows for more refined representation in speech frames with high information density (such as voiceless consonants and plosives), while improving compression efficiency in speech frames with low information density (such as silences and stationary vowels). While retaining the advantages of the A-law 13-fold structure, this invention introduces a content-aware mechanism, achieving an intelligent balance between speech quality and bit rate.

[0020] Preferably, the method further comprises: performing windowing processing on the framed speech frames.

[0021] Preferably, the short-time energy satisfies the expression: ;in, Indicates the The short-time energy of a speech frame; Indicates the first Frame speech frame The signal of the sampling point; is the sequence number of the sampling point in the speech frame, The number of sampling points contained in each speech frame.

[0022] Preferably, the zero-crossing rate satisfies the expression: ;in, Indicates the The zero-crossing rate of speech frames; Indicates the first Frame speech frame The signal of the sampling point; Indicates the first Frame speech frame The signal of the sampling point; represents a symbolic function; The number of sampling points contained in each speech frame; is the absolute value function.

[0023] In a second aspect, the present invention provides a voice signal acquisition system, comprising a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above-mentioned voice signal acquisition method is implemented.

[0024] By adopting the above technical solution, the above-mentioned voice signal acquisition method is generated into a computer program and stored in a memory so as to be loaded and executed by a processor, thereby making a terminal device based on the memory and the processor for easy use.

[0025] The beneficial effects of the present invention are:

[0026] The present invention constructs an information density function by analyzing the short-time energy and zero-crossing rate of speech frames, and dynamically adjusts the interval size of each segment in the A-law thirteen-fold curve based on the information density function, thereby generating an adaptive quantization function that matches the current speech content. While maintaining the A-law coding structure unchanged, it achieves differentiated processing of different speech types (such as plosives, voiceless consonants, vowels, and silent segments): in key speech frames with high information density (such as plosives and voiceless consonants), a finer quantization step size is used for the signal segments where the speech frames may be located to improve speech fidelity; in stationary frames with low information density (such as vowels and silent segments), the quantization precision is relaxed for the signal segments where the speech frames may be located to improve compression efficiency and suppress noise amplification, thereby achieving an intelligent balance between speech quality and bit rate overall, effectively avoiding the problems of key speech component loss or noise amplification that may occur in traditional A-law, and improving the robustness and listening quality of speech coding. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 The figure is a flow chart schematically showing a method for collecting speech signals in the present invention. DETAILED DESCRIPTION

[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work shall fall within the scope of protection of the present invention.

[0029] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0030] The embodiment of the present invention discloses a method for collecting voice signals, referring to Figure 1 , including steps S1 to S5:

[0031] S1. Use a voice acquisition device to collect an analog signal, perform sampling and framing processing on the analog signal, and obtain a number of voice frames.

[0032] A microphone or other voice collection device converts sound wave vibrations into a continuously varying electrical signal, i.e., an analog signal. This analog signal is sampled to produce a sampled signal. In this embodiment, the sampling rate is 44.1kHz. In other embodiments, the implementer can set the sampling rate based on actual implementation, such as 8kHz, 16kHz, or 48kHz.

[0033] Furthermore, the sampled signal is framed. In this embodiment, the frame length is set to 20ms and the frame shift is set to 10ms. The frame length corresponds to 882 sampling points and the frame shift corresponds to 441 sampling points. That is, each speech frame contains 882 sampling points. The first signal of adjacent speech frames is separated by 441 sampling points, so that there are 441 overlapping sampling points between adjacent speech frames. The speech frame signal can be expressed as ,in Indicates the Frame speech frame The signal of the sampling point, is the sequence number of the speech frame, is the sequence number of the sampling point in the speech frame, is the number of sampling points contained in each speech frame, is the number of sampling points corresponding to the frame shift, Indicates the sampling signal The signal of the sampling point.

[0034] It should be noted that in this embodiment, the frame shift is set to be smaller than the frame length, so that there are overlapping sampling points between adjacent speech frames. This is to reduce information loss in the speech signal and improve the stability of speech signal feature extraction. In other embodiments, implementers can set the frame length and frame shift according to actual implementation conditions.

[0035] Preferably, in order to reduce the boundary effect, each speech frame is windowed. In this embodiment, a Hamming window is used: ,in Indicates about The window function of the sampling points, is the sequence number of the sampling point in the speech frame, is the number of sampling points contained in each speech frame, is a cosine function. Then the speech frame signal after windowing can be expressed as ,in Indicates the first Frame speech frame The signal of the sampling point, Indicates the first Frame speech frame In other embodiments, the implementer may select a window function according to actual implementation conditions, such as a Hanning window, a Kaiser window, etc. It should be noted that windowing is a well-known technology and the specific principle will not be described in detail here.

[0036] S2. Obtain the information density of each speech frame according to the short-time energy and zero-crossing rate of each speech frame.

[0037] Specifically, the short-time energy of a speech frame satisfies the expression:

[0038] ;

[0039] in, Indicates the The short-time energy of a speech frame; Indicates the first Frame speech frame The signal of the sampling point; is the sequence number of the sampling point in the speech frame, The number of sampling points contained in each speech frame. Short-time energy reflects the activity level of the current speech frame. When the short-time energy is large, it means that the current speech frame contains strong speech components, such as vowels and plosives. When the short-time energy is small, it means that the current speech frame has sparse content, such as unvoiced consonants or silence.

[0040] Furthermore, the zero-crossing rate of the speech frame satisfies the expression:

[0041] ;

[0042] in, Indicates the The zero-crossing rate of speech frames; Indicates the first Frame speech frame The signal of the sampling point; Indicates the first Frame speech frame The signal of the sampling point; represents a symbolic function, when When greater than 0, ,when When it is equal to 0, ,when When it is less than 0, ; The number of sampling points contained in each speech frame; is an absolute value function. The zero-crossing rate reflects the frequency of signal changes in the speech frame. When the zero-crossing rate is large, the current speech frame contains more high-frequency components, which may be voiceless consonants or plosives. When the zero-crossing rate is small, the current speech frame is stable and may be a vowel or silence.

[0043] It should be noted that the short-time energy of a speech frame can, to a certain extent, distinguish vowels, plosives from voiceless consonants, and silent segments; the zero-crossing rate of a speech frame can, to a certain extent, distinguish voiceless consonants, plosives from vowels, and silent segments. Therefore, the present invention combines the short-time energy and zero-crossing rate of a speech frame to determine the information density of the speech frame, and uses the information density to quantify the amount of information contained in the speech frame.

[0044] Specifically, the information density of the speech frame satisfies the expression:

[0045] ;

[0046] in, Indicates the The information density of a speech frame; Indicates the The short-time energy of a speech frame; Indicates the The zero-crossing rate of speech frames; Represents the normalization function. This embodiment adopts the maximum and minimum value normalization, that is, the short-time energy of all speech frames is used to normalize the first The short-time energy of the speech frames is normalized to the maximum and minimum values, and the , according to the zero-crossing rate of all speech frames The zero-crossing rate of the speech frames is normalized to the maximum and minimum values, and we get ; represents the weight coefficient of the zero-crossing rate, represents the weight coefficient of short-time energy, The implementation personnel shall set it according to the actual implementation situation. Since light consonants and plosives are key speech components, their loss may affect the intelligibility of speech, while vowels are smooth speech and the human ear is not sensitive to slight distortion of vowels. Even if part of them is lost, it is easier to infer through the context. Therefore, light consonants and plosives contain more information than vowels, and the zero-crossing rate of light consonants and plosives is greater than that of vowels. Therefore, this embodiment sets the weight coefficient of the zero-crossing rate to If set to 0.7, the weight coefficient of short-time energy is 0.3, which makes the information density of light consonants and plosives greater than that of vowels. At the same time, since the short-term energy of plosives is greater than that of light consonants, when When the information density of a speech frame is extremely high, The speech frame is more likely to be a plosive sound. When the information density of the speech frame is moderate to large, The first speech frame is more likely to be a light consonant. The information density of each speech frame is moderate to small. The speech frame is more likely to be a vowel. When the information density of a speech frame is extremely small, The speech frames are more likely to be silence segments.

[0047] S3. Construct a quantization function for each speech frame according to the information density of each speech frame.

[0048] It should be noted that the A-law Thirteen-fold curve is a non-uniform quantization method used for voice signal compression in digital telephone communication systems. This method divides the dynamic range into 8 positive segments and 8 negative segments by constructing a continuous function, thereby realizing dynamic compression and expansion of the input signal, finely quantizing small signals (i.e., "amplifying" small signals) and roughly quantizing large signals (i.e., "compressing" large signals). The 8 positive segments of the A-law Thirteen-fold curve are symmetrical with the 8 negative segments, and the sizes of the 8 positive segments are respectively 、 、 、 、 、 、 、 The larger the size of the positive segment, the greater the degree of compression of the signal in that positive segment. The smaller the size of the positive segment, the greater the degree of amplification of the signal in that positive segment. However, because the A-law uses a fixed non-uniform quantization strategy, it cannot dynamically adjust the quantization strategy based on the content of the current speech frame. This can cause problems such as the loss of plosives and excessive amplification of silent segments. Therefore, the present invention adaptively adjusts the size of each segment in the A-law based on the information density of the speech frame, thereby constructing a quantization function for the speech frame, preventing the loss of key information in the speech frame while preventing excessive amplification of silent segments. Because the eight positive segments are symmetrical with the eight negative segments, the present invention uses the positive segment as an example for explanation.

[0049] Specifically, 8 positive segments are set, and the initial size of each positive segment is 、 、 、 、 、 、 、 For each speech frame, the adjustment index of each positive segment is determined according to the information density of the speech frame and the sequence number of each positive segment:

[0050] ;

[0051] in, Indicates the The first paragraph Adjustment index under speech frames; Indicates the The information density of a speech frame; 、 is a hyperparameter; Indicates the sequence number of the positive segment; represents the maximum function, Indicates acquisition and The maximum value in is used to prevent is 0, resulting in an adjusted index of 0; Indicates the absolute value symbol.

[0052] Among them, the hyperparameters 、 The setting method is as follows: cluster the information density of all speech frames, set the number of categories to 4, divide all speech frames into 4 categories, take the mean of the information density of all speech frames in each category as the representative information density of each category, sort all categories in descending order of representative information density, and Set to the mean of the representative information density of the second category and the representative information density of the third category. Calculate the information density of each speech frame in the second category and the third category and The absolute value of the difference between The absolute value of the difference of the maximum point is set. This embodiment does not limit the clustering algorithm, and the implementer can select a clustering algorithm according to the actual implementation situation, such as K-means clustering.

[0053] It should be noted that when the information density of the speech frame is extremely large, the speech frame is more likely to be a plosive sound, when the information density of the speech frame is moderate to large, the speech frame is more likely to be a light consonant, when the information density of the speech frame is moderate to small, the speech frame is more likely to be a vowel, and when the information density of the speech frame is extremely small, the speech frame is more likely to be a silent segment. The present invention clusters the speech frames into 4 categories according to the information density of the speech frames, and sorts all the categories in descending order according to the representative information density. The sorted categories are plosive sound, light consonant, vowel, and silent segment. The present invention takes the average of the representative information density of the second category and the representative information density of the third category as , the information density of each speech frame in the second category and the third category is The maximum absolute value of the difference between the absolute values of , so that satisfaction The speech frames of the condition are more likely to be light consonants and vowels, and meet The speech frames under these conditions are more likely to be plosives and silence segments.

[0054] Furthermore, the adjusted size of each positive segment is obtained according to the adjustment index of each positive segment:

[0055] ;

[0056] in, Indicates the The first paragraph The adjusted size under each speech frame; Indicates the The initial size of a positive segment; Indicates the The first paragraph Adjustment index under speech frames; Used for Perform normalization.

[0057] It should be noted that when the information density of the speech frame is extremely high, , the speech frame is more likely to be a plosive sound, and the plosive sound is a large signal. The positive segment corresponding to the large signal has a larger sequence number. In order to avoid the large signal being over-compressed and causing the plosive sound to be lost, the size of the positive segment with a large sequence number needs to be reduced, so that the compression degree of the speech frame is reduced. The range of the value of is [0,1], so The value range of is also [0,1]. The present invention uses gamma transformation to convert the sequence number of the positive segment As The index to get the adjusted index , then adjust the index The value range of is also [0,1]. The present invention uses the initial size of each positive segment to multiply the adjustment index to adjust the size of each positive segment. When the sequence number of the positive segment is When the value is larger, adjust the index The smaller the size, the more the initial size is multiplied by the adjustment index. The greater the degree of reduction of the positive segment, the greater the The smaller the time, the more index is adjusted The larger the size, the more the initial size is multiplied by the adjustment index. The smaller the reduction of the positive segment, the smaller the reduction of the positive segment. Used for After normalization, the size of the positive segment with a larger sequence number is reduced compared to the initial size, and the size of the positive segment with a smaller sequence number is increased compared to the initial size, thereby avoiding excessive compression of large signals and resulting in loss of plosive sounds.

[0058] When the information density of the speech frame is extremely small, , the speech frame is more likely to be a silent segment, and the silent segment is a small signal. The sequence number of the positive segment corresponding to the small signal is small. In order to avoid the small signal being over-amplified and causing the noise of the silent segment to be amplified, it is necessary to increase the size of the positive segment with a small sequence number, so that the amplification degree of the speech frame is reduced. When the value is larger, adjust the index The smaller the size, the more the initial size is multiplied by the adjustment index. The greater the degree of reduction of the positive segment, the greater the The smaller the time, the more index is adjusted The larger the size, the more the initial size is multiplied by the adjustment index. The smaller the reduction of the positive segment, the smaller the reduction of the positive segment. Used for After normalization, the size of the positive segment with a larger sequence number is reduced compared to the initial size, and the size of the positive segment with a smaller sequence number is increased compared to the initial size, thereby avoiding excessive amplification of small signals and causing amplification of noise in silent segments.

[0059] When the information density of the speech frame is moderately high, , the speech frame is more likely to be a light consonant, and the light consonant is a small signal. The sequence number of the positive segment corresponding to the small signal is small. Since the human ear is sensitive to the distortion of light consonants, it is necessary to amplify the light consonants to avoid distortion of light consonants, and then it is necessary to reduce the size of the positive segment with a small sequence number. The present invention uses the gamma transform method to amplify the sequence number of the positive segment. The reciprocal of As The index to get the adjusted index , then adjust the index The value range of is also [0,1]. The present invention uses the initial size of each positive segment to multiply the adjustment index to adjust the size of each positive segment. When the sequence number of the positive segment is The bigger it is, The smaller the value, the more the index is adjusted. The larger the size, the more the initial size is multiplied by the adjustment index. The smaller the degree of reduction of the positive segment, the smaller the number of the positive segment The smaller the time, The larger the adjustment index The smaller the size, the more the initial size is multiplied by the adjustment index. The greater the reduction of the positive segment, the greater the reduction of the positive segment. Used for After normalization, the size of the positive segment with a larger sequence number increases compared to the initial size, and the size of the positive segment with a smaller sequence number decreases compared to the initial size, thereby achieving further amplification of the light consonants.

[0060] When the information density of the speech frame is moderately small, , the speech frame is more likely to be a vowel, and the vowel is a large signal. The sequence number of the positive segment corresponding to the large signal is large. Since the vowel is highly repetitive and has high redundancy, the human ear is not sensitive to the distortion of the vowel. Therefore, the vowel can be further compressed, and the size of the positive segment with a large sequence number needs to be increased. The present invention uses the gamma transform method to convert the sequence number of the positive segment into The reciprocal of As The index to get the adjusted index , then adjust the index The value range of is also [0,1]. The present invention uses the initial size of each positive segment to multiply the adjustment index to adjust the size of each positive segment. When the sequence number of the positive segment is The bigger it is, The smaller the value, the more the index is adjusted. The larger the size, the more the initial size is multiplied by the adjustment index. The smaller the degree of reduction of the positive segment, the smaller the number of the positive segment The smaller the time, The larger the adjustment index The smaller the size, the more the initial size is multiplied by the adjustment index. The greater the reduction of the positive segment, the greater the reduction of the positive segment. Used for After normalization, the size of the positive segment with a larger sequence number increases compared to the initial size, and the size of the positive segment with a smaller sequence number decreases compared to the initial size, thereby achieving further compression of the vowels.

[0061] Furthermore, since the positive segment and the negative segment are symmetrical about the origin, for any speech frame, the present invention constructs a quantization function of the speech frame according to the adjusted size of each positive segment in the speech frame:

[0062] ;

[0063] in, Indicates the Quantization function of speech frames; 、 Respectively represent , The first paragraph The adjusted size under the speech frame, ; is the independent variable, i.e. the first The signal of each sampling point in a speech frame; Indicates the absolute value; represents a symbolic function, when When is greater than 0, ,when When it is equal to 0, ,when When it is less than 0, .

[0064] S4. Perform quantization encoding on each speech frame according to the quantization function of each speech frame to obtain a digital speech signal.

[0065] Specifically, for any speech frame, the signals of all sampling points in the speech frame are normalized to the range of [-1, 1]. The normalization result is input into the quantization function of the speech frame, and the quantized value of each sampling point is output. The range of the quantized value is [-1, 1]. The quantized value is mapped to PCM code, thereby achieving the encoding of the speech frame. PCM code is a digital speech signal.

[0066] It should be noted that the method of mapping the quantization value to the PCM code is the same as the mapping method in the A-law thirteen-fold line, and will not be described in detail here.

[0067] An embodiment of the present invention further discloses a voice signal acquisition system, including a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a voice signal acquisition method according to the present invention is implemented.

[0068] The above system also includes other components well known to those skilled in the art, such as a communication bus and a communication interface. The configuration and functions of these components are known in the art and will not be described in detail here.

Claims

1. A method for collecting speech signals, characterized in that: include: Using voice acquisition equipment to collect analog signals, sampling and framing the analog signals to obtain several voice frames; Obtaining information density of each speech frame according to the short-time energy and zero-crossing rate of each speech frame, wherein the information density is positively correlated with the short-time energy and the zero-crossing rate; The quantization function of each speech frame is constructed according to the information density of each speech frame, including: The quantization function is a piecewise linear function including 8 positive segments and 8 negative segments, wherein the 8 negative segments and the 8 positive segments are symmetrical about the origin; Set 8 positive segments, and determine the adjustment index of each positive segment according to the information density of the speech frame and the sequence number of each positive segment, satisfying the expression: , Indicates the The first paragraph The adjustment index under the speech frame, Indicates the The information density of a speech frame is 、 is a hyperparameter, Indicates the sequence number of the positive segment, represents the maximum function, Indicates the absolute value symbol; according to the adjustment index of each positive segment and the initial size of each positive segment, the adjusted size of each positive segment in the speech frame is obtained to satisfy the expression: , Indicates the The first paragraph The adjusted size under the speech frame, Indicates the The initial size of a positive segment; According to the adjusted size, a quantization function of the speech frame is constructed to satisfy the relationship: ; in, Indicates the The quantization function of speech frames, Indicates the The first paragraph The adjusted size under the speech frame, , represents a symbolic function; Each speech frame is quantized and encoded according to the quantization function of each speech frame to obtain a digital speech signal.

2. A method for collecting speech signals according to claim 1, characterized in that: The obtaining of the information density of each speech frame according to the short-time energy and the zero-crossing rate of each speech frame includes: The information density of the speech frame is obtained by performing a weighted summation on the normalized result of the short-time energy of the speech frame and the normalized result of the zero-crossing rate of the speech frame.

3. The method for collecting speech signals according to claim 1, wherein: The step of quantizing and encoding each speech frame according to the quantization function of each speech frame comprises: Normalize the signals of all sampling points in the speech frame to the range of [-1, 1], input the normalization result into the quantization function of the speech frame, output the quantization value of each sampling point, and map the quantization value to PCM code.

4. A method for collecting speech signals according to any one of claims 1 to 3, characterized in that: Also includes: Perform windowing processing on the divided speech frames.

5. A method for collecting speech signals according to claim 4, characterized in that: The short-time energy satisfies the expression: ; in, Indicates the The short-time energy of a speech frame; Indicates the first Frame speech frame The signal of the sampling point; is the sequence number of the sampling point in the speech frame, The number of sampling points contained in each speech frame.

6. A method for collecting speech signals according to claim 4, characterized in that: The zero-crossing rate satisfies the expression: ; in, Indicates the The zero-crossing rate of speech frames; Indicates the first Frame speech frame The signal of the sampling point; Indicates the first Frame speech frame The signal of the sampling point; represents a symbolic function; The number of sampling points contained in each speech frame; is the absolute value function.

7. A voice signal acquisition system, characterized in that: include: A processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a speech signal collection method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • A voice interaction analysis method and system for an intelligent voice robot

    CN119785777A

  • Pitch determiner for a speech analyzer

    US6018706A