Voice signal acquisition method and system

By constructing the information density function of the speech frame and dynamically adjusting the quantization segment size in the A-law thirteenth fold line, the problems of feature information loss and noise interference in traditional quantization strategies are solved, and higher speech signal acquisition quality and robustness are achieved.

CN120236592AActive Publication Date: 2025-07-01GUANGZHOU JIUSI INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510712168.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-01
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

Traditional fixed quantization strategies have problems of feature information loss and noise interference in speech signal processing, especially in speech components such as low energy consonants and burst sounds with large dynamic range.

Method used

By analyzing the short-time energy and zero-crossing rate of the speech frame, the interval sizes of each segment in the A-law thirteenth line are dynamically adjusted, thereby generating an adaptive quantization function. This method sets 8 positive segments and 8 negative segments, and adjusts the size of each positive segment according to the information density of the speech frame to achieve non-uniform quantization.

Benefits of technology

It effectively retains key speech features such as low energy consonants and large dynamic range blast sounds, avoids improper amplification of noise in the silent segment environment, and ensures coding efficiency and improves speech quality and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236592A_ABST
    Figure CN120236592A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of voice signal acquisition, and particularly relates to a voice signal acquisition method and system, and the method comprises the steps: carrying out the sampling and framing processing of an analog signal, obtaining a plurality of voice frames, obtaining the information density of each voice frame according to the short-time energy and zero-crossing rate of each voice frame, and obtaining the information density of each voice frame; the information density is in positive correlation with short-time energy and a zero-crossing rate, constructing a quantization function of each voice frame according to the information density of each voice frame, and performing quantization coding on each voice frame according to the quantization function of each voice frame to obtain a digital voice signal. According to the invention, the quantization function of the voice frame is self-adapted according to the information density of the voice frame, and the voice definition and intelligibility are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voice signal acquisition. More specifically, the present invention relates to a method and system for voice signal acquisition. Background Art

[0002] With the rapid development of technologies such as artificial intelligence, speech recognition, and intelligent voice interaction, the acquisition of high-quality voice signals has become a core link in improving the performance of speech processing systems.

[0003] Currently, in the field of voice acquisition, traditional methods based on fixed quantization strategies still dominate. For example, the A-law thirteen-segment quantization method widely used in telephone communication systems. This method simulates the auditory characteristics of the human ear and adopts a non-linear quantization strategy: implementing refined quantization processing for small signals and moderately compressing large signals, thereby achieving high coding efficiency while ensuring subjective auditory quality.

[0004] However, in actual voice signal processing, due to the complex dynamic characteristics and time-varying features of voice signals, traditional fixed quantization strategies have significant technical limitations. Specifically, for voice components such as voiceless consonants with low energy and plosives with a large dynamic range, it is easy to lose characteristic information due to quantization errors, significantly affecting the intelligibility of speech; at the same time, environmental noise in the silent segment may be improperly amplified, introducing additional noise interference. These limitations are particularly prominent in modern intelligent voice applications that require high-fidelity voice acquisition.

[0005] Therefore, how to dynamically adjust the acquisition and quantization strategies according to the voice content to improve the fidelity of key voice segments and reduce data redundancy in non-critical segments is an urgent problem to be solved in the current field of voice signal acquisition. Summary of the Invention

[0006] To solve the technical problems that the above-mentioned fixed quantization strategy is prone to losing characteristic information due to quantization errors, affecting the intelligibility of speech, and may introduce additional noise interference, the present invention provides solutions in the following aspects.

[0007] In a first aspect, the present invention provides a method for voice signal acquisition, including: An analog signal is collected by a voice acquisition device, and the analog signal is sampled and framed to obtain a number of voice frames; the information density of each voice frame is obtained according to the short-time energy and zero-crossing rate of each voice frame, and the information density is positively correlated with the short-time energy and zero-crossing rate; a quantization function for each voice frame is constructed according to the information density of each voice frame, including: setting 8 positive segments, and determining the adjustment index of each positive segment according to the information density of the voice frame and the serial number of each positive segment; obtaining the adjusted size of each positive segment under the voice frame according to the adjustment index of each positive segment and the initial size of each positive segment, and constructing a quantization function of the voice frame according to the adjusted size; each voice frame is quantized and encoded according to the quantization function of each voice frame to obtain a digital voice signal.

[0008] The present invention accurately evaluates the information density of a voice frame by using two time-domain features, namely, the short-time energy and zero-crossing rate of the voice frame, adaptively constructs a quantization function according to the information density, and realizes non-uniform quantization of the voice frame by setting 8 positive segments and dynamically adjusting their ranges. The finally generated digital voice signal not only retains key voice features such as voiceless consonants with low energy and plosives with a large dynamic range, but also avoids the improper amplification of environmental noise in the silent segment, and at the same time ensures the coding efficiency.

[0009] Preferably, the obtaining the information density of each voice frame according to the short-time energy and zero-crossing rate of each voice frame includes: performing weighted summation on the normalized result of the short-time energy of the voice frame and the normalized result of the zero-crossing rate of the voice frame to obtain the information density of the voice frame.

[0010] The short-time energy and zero-crossing rate respectively reflect the voice characteristics from the amplitude domain and frequency domain perspectives. The present invention conducts collaborative analysis on the short-time energy and zero-crossing rate, which can more comprehensively characterize the information content of the voice frame and provide a more discriminative decision basis for obtaining the quantization function of the voice frame subsequently.

[0011] Preferably, the adjustment index satisfies the expression: ; where represents the adjustment index of the th positive segment under the th voice frame; represents the information density of the th voice frame; , are hyperparameters; represents the serial number of the positive segment; represents the maximum value function; represents the absolute value symbol.

[0012] The present invention adopts a segmented strategy to generate different adjustment indices according to the information density of speech frames, enhancing the representation ability of key speech components such as voiceless consonants, effectively suppressing the amplification of noise in silent segments and protecting transient speech components such as plosives from being lost, and improving the fidelity and compression efficiency of speech coding under different semantic segments.

[0013] Preferably, the size of each positive segment after adjustment under a speech frame satisfies the expression: ; where represents the size of the th positive segment after adjustment under the th speech frame; represents the initial size of the th positive segment; represents the adjustment index of the th positive segment under the th speech frame.

[0014] Preferably, the method for obtaining the quantization function is as follows: the quantization function is a piecewise linear function, including 8 positive segments and 8 negative segments, and the 8 negative segments are symmetric about the origin with the 8 positive segments; the slope of the th positive segment of the quantization function of the th speech frame is , , represents the size of the th positive segment after adjustment under the th speech frame; when , the range of the th positive segment is , when , the range of the th positive segment is , represents the size of the th positive segment after adjustment under the th speech frame.

[0015] The present invention can adaptively adjust the quantization strategy according to different speech contents: in speech frames with extremely high information density (plosives), more refined quantization is performed on large signal segments to retain key speech components; in speech frames with moderately high information density (fricatives), more refined quantization is performed on small signal segments to retain key speech components; in speech frames with moderately low information density (vowels), the quantization accuracy of large signal segments is appropriately relaxed to improve the compression efficiency; in speech frames with extremely low information density (silent segments), the quantization accuracy of small signal segments is appropriately relaxed to improve the compression efficiency and suppress noise amplification. In addition, the quantization function in the present invention maintains consistency with the traditional A-law thirteen-segment coding structure, ensuring that the encoding and decoding processes are compatible with the existing PCM coding system, and at the same time having stronger content perception ability, thereby significantly improving the overall quality and robustness of speech acquisition and encoding without increasing additional complexity.

[0016] Preferably, quantizing and encoding each speech frame according to the quantization function of each speech frame includes: normalizing the signals of all sampling points in the speech frame to the range of [-1, 1], inputting the normalization result into the quantization function of the speech frame, outputting the quantization values of each sampling point, and mapping the quantization values to PCM codes.

[0017] The present invention adopts different quantization functions for different speech frames, enabling different speech contents to obtain different quantization accuracies: more refined representation is achieved in speech frames with high information density (such as fricatives and plosives), and the compression efficiency is improved in speech frames with low information density (such as silent segments and steady vowels). On the basis of retaining the advantages of the A-law thirteen-segment structure, the present invention introduces a content perception mechanism to achieve an intelligent balance between speech quality and bit rate.

[0018] Preferably, it further includes: performing windowing processing on the framed speech frames.

[0019] Preferably, the short-time energy satisfies the expression: ; where represents the short-time energy of the th speech frame; represents the signal of the nd sampling point in the th speech frame after windowing processing; is the serial number of the sampling point in the speech frame, is the number of sampling points included in each speech frame.

[0020] Preferably, the zero-crossing rate satisfies the expression: ; where represents the zero-crossing rate of the th speech frame; represents the The signal of the th sampling point in the frame of speech; represents the signal of the th sampling point in the frame of speech after windowing; represents the number of sampling points included in each frame of speech; is the absolute value function.

[0021] In a second aspect, the present invention provides a voice signal acquisition system, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above-mentioned voice signal acquisition method is implemented.

[0022] By adopting the above technical solution, a computer program is generated from the above-mentioned voice signal acquisition method and stored in the memory to be loaded and executed by the processor, so as to manufacture a terminal device according to the memory and the processor, which is convenient to use.

[0023] The beneficial effects of the present invention are as follows: The present invention constructs an information density function by analyzing the short-time energy and zero-crossing rate of speech frames, and dynamically adjusts the interval size of each segment in the A-law thirteen-segment line accordingly, so as to generate an adaptive quantization function matching the current speech content. Without changing the A-law coding structure, the present invention realizes differential processing of different speech types (such as plosives, voiceless consonants, vowels, silent segments): in key speech frames with high information density (such as plosives, voiceless consonants), a finer quantization step is adopted for the signal segment where the speech frame may be located to improve speech fidelity, and in smooth frames with low information density (such as vowels, silent segments), the quantization accuracy for the signal segment where the speech frame may be located is relaxed to improve compression efficiency and suppress noise amplification, so as to achieve an intelligent balance between speech quality and bit rate as a whole, effectively avoiding the problems of loss of key speech components or noise amplification that may occur in the traditional A-law, and improving the robustness and listening quality of speech coding. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is a flowchart schematically showing a method for acquiring a voice signal in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0026] The following will describe in detail the specific implementation manners of the present invention with reference to the accompanying drawings.

[0027] An embodiment of the present invention discloses a method for collecting voice signals. Referring to Figure 1 , it includes steps S1 to S5: S1. Use a voice collection device to collect an analog signal, and perform sampling and framing processing on the analog signal to obtain a number of voice frames.

[0028] Convert the sound wave vibration into a continuously varying electrical signal, that is, an analog signal, through a voice collection device such as a microphone. Sample the analog signal to obtain a sampled signal. In this embodiment, the sampling rate is 44.1 kHz. In other embodiments, the implementer can set the sampling rate according to the actual implementation situation, such as 8 kHz, 16 kHz, 48 kHz, etc.

[0029] Furthermore, perform framing processing on the sampled signal. In this embodiment, the frame length is set to 20 ms, and the frame shift is set to 10 ms. Then the frame length corresponds to 882 sampling points, and the frame shift corresponds to 441 sampling points. That is, each voice frame contains 882 sampling points, and there is an interval of 441 sampling points between the first signals of adjacent voice frames, so that there are 441 overlapping sampling points between adjacent voice frames. Then the voice frame signal can be expressed as , where represents the signal of the th sampling point in the th voice frame, is the serial number of the voice frame, is the serial number of the sampling point in the voice frame, is the number of sampling points included in each voice frame, is the number of sampling points corresponding to the frame shift, represents the signal of the th sampling point in the sampled signal.

[0030] It should be noted that the frame shift set in this embodiment is less than the frame length, so that there are overlapping sampling points between adjacent voice frames, in order to reduce the information loss of the voice signal and improve the stability of voice signal feature extraction. In other embodiments, the implementer can set the frame length and the frame shift according to the actual implementation situation.

[0031] Preferably, in order to reduce the boundary effect, perform windowing processing on each voice frame. In this embodiment, the Hamming window is used: , where represents the window function with respect to the th sampling point, is the serial number of the sampling point in the voice frame, is the number of sampling points included in each voice frame, is the cosine function. Then the speech frame signal after windowing can be expressed as , where represents the signal of the th sampling point in the th speech frame after windowing, represents the signal of the th sampling point in the th speech frame before windowing. In other embodiments, the implementer can select a window function according to the actual implementation situation, such as the Hanning window, the Kaiser window, etc. It should be noted that windowing is a well-known technology, and the specific principle will not be elaborated in detail here.

[0032] S2. Obtain the information density of each speech frame according to the short-time energy and zero-crossing rate of each speech frame.

[0033] Specifically, the short-time energy of the speech frame satisfies the expression: ; where represents the short-time energy of the th speech frame; represents the signal of the th sampling point in the th speech frame after windowing; is the serial number of the sampling point in the speech frame, is the number of sampling points included in each speech frame. The short-time energy reflects the activity degree of the current speech frame. When the short-time energy is large, it indicates that the current speech frame contains strong speech components, such as vowels, plosives, etc. When the short-time energy is small, it indicates that the content of the current speech frame is sparse, which may be voiceless consonants, silent segments.

[0034] Furthermore, the zero-crossing rate of the speech frame satisfies the expression: ; where represents the zero-crossing rate of the th speech frame; represents the signal of the th sampling point in the th speech frame after windowing; represents the signal of the th sampling point in the th speech frame after windowing; represents the sign function. When is greater than 0, , when is equal to 0, , when is less than 0, ; is the number of sampling points included in each speech frame; is the absolute value function. The zero-crossing rate reflects the change frequency of the signal in the speech frame. When the zero-crossing rate is relatively large, the current speech frame contains more high-frequency components, which may be voiceless consonants or plosives. When the zero-crossing rate is relatively small, the current speech frame is stable and may be a vowel or a silent segment.

[0035] It should be noted that the short-time energy of the speech frame can, to a certain extent, distinguish vowels, plosives from voiceless consonants and silent segments, and the zero-crossing rate of the speech frame can, to a certain extent, distinguish voiceless consonants, plosives from vowels and silent segments. Therefore, the present invention combines the short-time energy and the zero-crossing rate of the speech frame to confirm the information density of the speech frame, and uses the information density to quantify the amount of information contained in the speech frame.

[0036] Specifically, the information density of the speech frame satisfies the expression: ; where represents the information density of the th speech frame; represents the short-time energy of the th speech frame; represents the zero-crossing rate of the th speech frame; represents the normalization function. In this embodiment, the maximum-minimum normalization is adopted, that is, the short-time energy of the th speech frame is normalized by the maximum and minimum values of the short-time energy of all speech frames to obtain , and the zero-crossing rate of the th speech frame is normalized by the maximum and minimum values of the zero-crossing rate of all speech frames to obtain ; represents the weight coefficient of the zero-crossing rate, represents the weight coefficient of the short-time energy, which is set by the implementer according to the actual implementation situation. Since voiceless consonants and plosives are key speech components, loss may affect the intelligibility of the speech, while vowels are stable speech, and the human ear is not sensitive to slight distortion of vowels. Even if a part is lost, it is relatively easy to infer through the context. Therefore, voiceless consonants and plosives contain more information than vowels, and the zero-crossing rate of voiceless consonants and plosives is larger than that of vowels. Therefore, in this embodiment, the weight coefficient of the zero-crossing rate is set to 0.7, and the weight coefficient of the short-time energy is 0.3, so that the information density of voiceless consonants and plosives is larger than the information density of vowels. At the same time, since the short-time energy of plosives is larger than the short-time energy of voiceless consonants, therefore, when the information density of the th speech frame is extremely large, the th speech frame is more likely to be a plosive, when the When the information density of a voice frame is moderately large, the voice frame is more likely to be a voiceless consonant. When the information density of a voice frame is moderately small, the voice frame is more likely to be a vowel. When the information density of a voice frame is extremely small, the voice frame is more likely to be a silent segment.

[0037] S3. Construct a quantization function for each voice frame according to the information density of each voice frame.

[0038] It should be noted that the A-law thirteen-segment line is a non-uniform quantization method used for voice signal compression in digital telephone communication systems. This method divides the dynamic range into 8 positive segments and 8 negative segments by constructing a continuous function, realizes the dynamic companding of the input signal, finely quantizes small signals (i.e., "amplifies" small signals), and roughly quantizes large signals (i.e., "compresses" large signals). The 8 positive segments and 8 negative segments of the A-law thirteen-segment line are symmetric, and the sizes of the 8 positive segments are , , , , , , , . When the size of a positive segment is larger, the compression degree of the signal located in this positive segment is larger. When the size of a positive segment is smaller, the amplification degree of the signal located in this positive segment is larger. However, since the A-law adopts a fixed non-uniform quantization strategy and cannot dynamically adjust the quantization strategy according to the content of the current voice frame, problems such as the loss of plosive sounds and the over-amplification of silent segments may occur. Therefore, the present invention adaptively adjusts the sizes of each segment in the A-law according to the information density of the voice frame, thereby constructing a quantization function for the voice frame, preventing the loss of key information in the voice frame and preventing the over-amplification of silent segments. Since the 8 positive segments and 8 negative segments are symmetric, the present invention takes the positive segments as an example for illustration.

[0039] Specifically, set 8 positive segments, and the initial sizes of each positive segment are , , , , , , , . For each voice frame, determine the adjustment index of each positive segment according to the information density of the voice frame and the serial number of each positive segment: ; where represents the th positive segment in the The adjustment index for each speech frame; Indicates the Information density of the speech frame; 、 Are hyperparameters; Indicates the serial number of the positive segment; Indicates the maximum value function, Indicates to obtain And The maximum value in, used to prevent From being 0, resulting in the adjustment index being 0; Indicates the absolute value symbol.

[0040] Among them, the hyperparameters 、 The setting method is as follows: Cluster the information density of all speech frames, set the number of categories to 4, divide all speech frames into 4 categories, take the mean of the information density of all speech frames in each category as the representative information density of each category, sort all categories in descending order according to the representative information density, and set As the mean of the representative information density of the second category and the representative information density of the third category. Calculate the absolute value of the difference between the information density of each speech frame in the second category and the third category and And set As the absolute value of the difference at the maximum point. In this embodiment, the clustering algorithm is not limited, and the implementer can select the clustering algorithm according to the actual implementation situation, such as K-means clustering.

[0041] It should be noted that when the information density of the speech frame is extremely large, the speech frame is more likely to be a plosive sound. When the information density of the speech frame is moderately large, the speech frame is more likely to be a voiceless consonant. When the information density of the speech frame is moderately small, the speech frame is more likely to be a vowel. When the information density of the speech frame is extremely small, the speech frame is more likely to be a silent segment. The present invention clusters speech frames into 4 categories according to the information density of speech frames, and sorts all categories in descending order according to the representative information density. Then the sorted categories are plosive sounds, voiceless consonants, vowels, and silent segments in sequence. The present invention uses the mean of the representative information density of the second category and the representative information density of the third category as , and uses the maximum absolute value of the difference between the information density of each speech frame in the second category and the third category and As , so that speech frames satisfying Condition are more likely to be voiceless consonants and vowels, and speech frames satisfying Condition are more likely to be plosive sounds and silent segments.

[0042] Furthermore, obtain the adjusted size of each positive segment according to the adjustment index of each positive segment: ; Among them, represents the size of the th positive segment after adjustment under the th speech frame; represents the initial size of the th positive segment; represents the adjustment index of the th positive segment under the th speech frame; is used to perform normalization.

[0043] It should be noted that when the information density of the speech frame is extremely high, , the speech frame is more likely to be a plosive sound, and the plosive sound is a large signal. The serial number of the positive segment corresponding to the large signal is larger. To avoid the loss of the plosive sound caused by excessive compression of the large signal, it is necessary to reduce the size of the positive segment with a large serial number, so that the compression degree of the speech frame is reduced. Since has a value range of [0, 1], so that also has a value range of [0, 1]. The present invention uses the gamma transformation method to use the serial number of the positive segment as the exponent of to obtain the adjustment index . Then the value range of the adjustment index is also [0, 1]. The present invention multiplies the initial size of each positive segment by the adjustment index to realize the adjustment of the size of each positive segment. When the serial number of the positive segment is larger, the adjustment index is smaller. By multiplying the initial size by the adjustment index, the degree of reduction of the th positive segment is larger. When the serial number of the positive segment is smaller, the adjustment index is larger. By multiplying the initial size by the adjustment index, the degree of reduction of the th positive segment is smaller. After using to perform normalization on

[0044] When the information density of the speech frame is extremely low, , the speech frame is more likely to be a silent segment, and the silent segment is a small signal. The serial number of the positive segment corresponding to the small signal is smaller. To avoid the amplification of the noise of the silent segment caused by excessive amplification of the small signal, it is necessary to increase the size of the positive segment with a small serial number, so that the amplification degree of the speech frame is reduced. When the serial number When it is larger, the adjustment exponent is smaller. By multiplying the initial size by the adjustment exponent, the reduction degree of the th positive segment is greater. When the serial number of the positive segment is smaller, the adjustment exponent is larger. By multiplying the initial size by the adjustment exponent, the reduction degree of the th positive segment is smaller. After using to normalize , the size of the positive segment with a larger serial number decreases compared to the initial size, and the size of the positive segment with a smaller serial number increases compared to the initial size, thereby preventing the noise in the silent segment from being amplified due to the over-amplification of small signals.

[0045] When the information density of the speech frame is moderately large, , the speech frame is more likely to be a voiceless consonant, and the voiceless consonant is a small signal. The serial number of the positive segment corresponding to the small signal is small. Since the human ear is sensitive to the distortion of voiceless consonants, it is necessary to amplify the voiceless consonants to avoid voiceless consonant distortion. Furthermore, it is necessary to reduce the size of the positive segment with a small serial number. Through the gamma transformation, the present invention uses the reciprocal of the serial number of the positive segment as the exponent of to obtain the adjustment exponent . Then the value range of the adjustment exponent is also [0,1]. The present invention multiplies the initial size of each positive segment by the adjustment exponent to adjust the size of each positive segment. When the serial number of the positive segment is larger, is smaller, and the adjustment exponent is larger. By multiplying the initial size by the adjustment exponent, the reduction degree of the th positive segment is smaller. When the serial number of the positive segment is smaller, is larger, and the adjustment exponent is smaller. By multiplying the initial size by the adjustment exponent, the reduction degree of the th positive segment is greater. After using to normalize

[0046] , the size of the positive segment with a larger serial number increases compared to the initial size, and the size of the positive segment with a smaller serial number decreases compared to the initial size, thereby further amplifying the voiceless consonants. , the voice frame is more likely to be a vowel, and the vowel is a large signal. The serial number of the positive segment corresponding to the large signal is larger. Since the vowel has strong repeatability and high redundancy, the human ear is not sensitive to the distortion of the vowel. Therefore, the vowel can be further compressed, and then it is necessary to increase the size of the positive segment with a large serial number. The present invention uses the gamma transformation method to take the reciprocal of the serial number of the positive segment as the exponent of to obtain the adjustment exponent . Then the value range of the adjustment exponent is also [0,1]. The present invention multiplies the initial size of each positive segment by the adjustment exponent to realize the adjustment of the size of each positive segment. When the serial number of the positive segment is larger, is smaller, and the adjustment exponent is larger. By multiplying the initial size by the adjustment exponent, the degree of reduction of the th positive segment is smaller. When the serial number of the positive segment is smaller, is larger, and the adjustment exponent is smaller. By multiplying the initial size by the adjustment exponent, the degree of reduction of the th positive segment is larger. After using to normalize , the size of the positive segment with a larger serial number increases compared with the initial size, and the size of the positive segment with a smaller serial number decreases compared with the initial size, so as to realize the further compression of the vowel.

[0047] Furthermore, since the positive segment and the negative segment are symmetric about the origin, for any voice frame, the present invention constructs the quantization function of the voice frame according to the adjusted size of each positive segment under the voice frame: ; wherein, represents the quantization function of the th voice frame; , respectively represent the adjusted sizes of the th and th positive segments under the th voice frame, ; is the independent variable, that is, the signal of each sampling point in the th input voice frame; represents the absolute value size; represents the sign function. When is greater than 0, , when is equal to 0, , when is less than 0, 。

[0048] S4. Quantize and encode each speech frame according to the quantization function of each speech frame to obtain a digital speech signal.

[0049] Specifically, for any speech frame, normalize the signals of all sampling points in the speech frame to the range of [-1, 1], input the normalization result into the quantization function of the speech frame, output the quantization values of each sampling point, the value range of the quantization values is [-1, 1], map the quantization values to PCM encoding, thereby realizing the encoding of the speech frame. The PCM encoding is the digital speech signal.

[0050] It should be noted that the method of mapping the quantization values to PCM encoding is the same as the mapping method in the A-law thirteen-segment line, and will not be elaborated in detail here.

[0051] The embodiment of the present invention also discloses a speech signal acquisition system, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, a speech signal acquisition method according to the present invention is implemented.

[0052] The above system further includes other components well known to those skilled in the art such as a communication bus and a communication interface, and their settings and functions are known in the art, so they will not be elaborated here.

Claims

1. A method for collecting voice signals, characterized in that, Including: Collecting an analog signal by using a voice acquisition device, performing sampling and frame segmentation processing on the analog signal to obtain a plurality of voice frames; Obtaining the information density of each voice frame according to the short-time energy and zero-crossing rate of each voice frame, and the information density is positively correlated with the short-time energy and zero-crossing rate; Constructing a quantization function for each voice frame according to the information density of each voice frame, including: Setting 8 positive segments, and determining the adjustment index of each positive segment according to the information density of the voice frame and the serial number of each positive segment; Obtaining the adjusted size of each positive segment under the voice frame according to the adjustment index of each positive segment and the initial size of each positive segment, and constructing a quantization function of the voice frame according to the adjusted size; Performing quantization coding on each voice frame according to the quantization function of each voice frame to obtain a digital voice signal.

2. The method for collecting voice signals according to claim 1, characterized in that, The obtaining the information density of each voice frame according to the short-time energy and zero-crossing rate of each voice frame includes: Performing weighted summation on the normalization result of the short-time energy of the voice frame and the normalization result of the zero-crossing rate of the voice frame to obtain the information density of the voice frame.

3. A method for collecting voice signals according to claim 1, characterized in that, The adjustment index satisfies the expression: ; Among them, represents the adjustment index of the th positive segment under the th speech frame; represents the information density of the th speech frame; and are hyperparameters; represents the serial number of the positive segment; represents the maximum value function; represents the absolute value symbol.

4. A method for collecting voice signals according to claim 1, characterized in that, The adjusted size of each positive segment under the voice frame satisfies the expression: ; Among them, represents the resized size of the th positive segment under the th speech frame; represents the initial size of the th positive segment; represents the adjustment exponent of the th positive segment under the th speech frame.

5. A method for collecting voice signals according to claim 1, characterized in that, The method for obtaining the quantization function is: The quantization function is a piecewise linear function, including 8 positive segments and 8 negative segments, and the 8 negative segments are symmetric about the origin with the 8 positive segments; The slope of the th positive segment of the quantization function of the th voice frame is ; represents the adjusted size of the th positive segment under the th voice frame; when is the case, the range of the th positive segment is ; when is the case, the range of the th positive segment is ; represents the adjusted size of the th positive segment under the th voice frame.

6. A method for collecting voice signals according to claim 1, characterized in that, The performing quantization coding on each voice frame according to the quantization function of each voice frame includes: Normalizing the signals of all sampling points in the voice frame to the range of [-1, 1], inputting the normalization result into the quantization function of the voice frame, outputting the quantization values of each sampling point, and mapping the quantization values to PCM coding.

7. A method for collecting speech signals according to any one of claims 1-6, characterized in that, Also including: Performing windowing processing on the framed voice frames.

8. A method for collecting voice signals according to claim 7, characterized in that, The short-time energy satisfies the expression: ; Among them, represents the short-time energy of the th speech frame; represents the signal of the th sampling point in the th speech frame after windowing; is the serial number of the sampling point in the speech frame, is the number of sampling points included in each speech frame.

9. A method for collecting voice signals according to claim 7, characterized in that, The zero-crossing rate satisfies the expression: ; Among them, represents the zero-crossing rate of the th speech frame; represents the signal of the th sampling point in the th speech frame after windowing; represents the signal of the th sampling point in the th speech frame after windowing; represents the sign function; is the number of sampling points included in each speech frame; is the absolute value function.

10. A voice signal acquisition system, characterized in that, Including: A processor and a memory, and the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a voice signal acquisition method according to any one of claims 1-9 is implemented.

Citation Information

Patent Citations

  • A voice interaction analysis method and system for an intelligent voice robot

    CN119785777A

  • Computer-implemented method and system for analysing digital speech data

    EP2418643A1

  • Pitch determiner for a speech analyzer

    US6018706A

  • Method and apparatus for pitch determination of a low bit rate digital voice message

    US6418407B1

Cited By

  • Intelligent telephone customer service interaction method for power enterprise based on artificial intelligence

    CN120783769A

  • Feature extraction method based on end-side speech recognition, electronic equipment and readable medium

    CN121483237A