A cloud-based voice data storage method and system
By combining polarity and segmentation in the A-law thirteen-segment method and setting variable-length codewords for speech data encoding, the problem of low compression efficiency of PCM encoding is solved, and more efficient speech data storage is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU JIUSI INTELLIGENT TECH CO LTD
- Filing Date
- 2026-03-19
- Publication Date
- 2026-05-26
AI Technical Summary
Traditional PCM encoding methods have limited efficiency in compressing quantized speech data, and cannot effectively improve storage efficiency.
The A-law thirteen-segment method is adopted in combination with polarity and segment division. Three types of variable-length codewords of different lengths are set. Variable-length encoding is performed according to the correlation and similarity of polarity segments to enhance the correlation and similarity of adjacent speech data and improve compression efficiency.
By using variable-length coding methods, the compression efficiency of speech data is improved, the amount of data in the encoded result is reduced, and the storage efficiency is enhanced.
Smart Images

Figure CN121862129B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voice data storage technology. More specifically, this invention relates to a voice data storage method and system based on a cloud platform. Background Technology
[0002] With the rapid development of information technology, the application scenarios of voice data are becoming increasingly widespread, covering multiple fields such as intelligent voice assistants, voice recognition, voice communication, online education, telemedicine, and intelligent customer service; these applications have placed higher demands on the storage of voice data.
[0003] Traditional local storage methods face numerous challenges when dealing with massive, dynamically growing voice data, including limited storage capacity, poor scalability, complex data management, and high costs. Therefore, cloud-based voice data storage has emerged as an effective solution to these problems; simultaneously, to ensure storage efficiency, voice data typically requires encoding and compression.
[0004] Since the amplitude distribution of speech signals is usually non-uniform, with most of the energy concentrated in the lower amplitude range, speech signals are often quantized using the A-law thirteen-segment method, and the quantized speech data is encoded using PCM coding.
[0005] Among them, the A-law thirteen-segment method, as a method for non-uniform quantization, divides the amplitude range of the speech signal into multiple segments, uses a smaller quantization step size in the small amplitude region and a larger quantization step size in the large amplitude region, thereby achieving non-uniform quantization of the signal and better adapting to the dynamic range of the signal.
[0006] Speech data quantized using the A-law thirteen-segment method can be compressed and stored using PCM encoding. However, PCM encoding is a fixed-length encoding method, so its compression efficiency is limited. Summary of the Invention
[0007] To address the technical problem that PCM encoding, as a fixed-length encoding method, has limited compression efficiency for quantized speech data, this invention provides solutions in the following aspects.
[0008] In a first aspect, the present invention provides a cloud-based voice data storage method, comprising: acquiring and quantizing a voice signal, the voice signal containing multiple voice data; determining the polarity and segment corresponding to each voice data by combining the polarity and segment division results in the A-law thirteen-segment method; performing variable-length encoding on the polarity and segment corresponding to each voice data to obtain the encoding result of the polarity and segment corresponding to each voice data, including: setting three types of variable-length codewords of different lengths: the first type of codeword has a length of 1 bit and the first bit is the first digit; the second type of codeword and the third type of codeword have lengths of 3 bits and 6 bits respectively and the first bit of both is the second digit, while the second type of codeword and the third type of codeword have lengths of 3 bits and 6 bits respectively, the first bit of both being the second digit, and ... The second bit of the class codeword is a different digit; polarity and segment are combined to form 16 polarity segments; variable-length codewords are assigned to all polarity segments, and according to the assignment results and the polarity segments corresponding to each speech data, variable-length encoding is performed on each speech data, including: after encoding the previous speech data, the polarity segment corresponding to the previous speech data is recorded as the target segment, the first type codeword is assigned to the target segment, the second type codeword is assigned to the two types of polarity segments adjacent to the target segment, and the third type codeword is assigned to all remaining polarity segments; according to the assignment results, the next speech data is encoded with variable length; the encoding results of all speech data are transmitted and stored on the cloud platform.
[0009] This invention combines the polarity of the signal amplitude corresponding to speech data with the temporal correlation and similarity of segments. It combines polarity and segments to form polar segments, enhancing the correlation and similarity of polar segments corresponding to adjacent speech data. In subsequent variable-length encoding of speech data, variable-length codewords of different lengths are assigned according to the correlation and similarity between each polar segment and the polar segment corresponding to the previous speech data. This allows more speech data to be encoded with shorter codewords, thereby improving compression efficiency.
[0010] Preferably, in the A-law thirteen-segment method, the voice data is divided into positive and negative polarities based on the signal amplitude, and each polarity range is divided into 8 segments, and each segment is further divided into 16 sub-segments.
[0011] Preferably, the first digit is 0 and the second digit is 1; or the first digit is 1 and the second digit is 0.
[0012] Preferably, the method for obtaining the two polarity paragraphs adjacent to the target paragraph includes: designating the target paragraph as the first... a type of polar paragraph; when At that time, the first and the Two polarity paragraphs, as two types of polarity paragraphs adjacent to the target paragraph; when At that time, the first and the Two polarity paragraphs, as two types of polarity paragraphs adjacent to the target paragraph; when and At that time, the first and the Two polarity paragraphs, which are adjacent to the target paragraph.
[0013] Preferably, the method further includes: dividing all speech data into three categories based on the relationship between the polarity segment corresponding to each speech data and the polarity segment corresponding to the previous speech data; obtaining an estimated data volume based on the number of speech data belonging to each category and the length of the codewords in each category; obtaining a theoretical data volume based on the total number of speech data and the length of the fixed-length codewords used in PCM encoding to encode polarity and segments; if the estimated data volume is greater than or equal to the theoretical data volume, performing fixed-length encoding on the polarity and segments corresponding to each speech data using PCM encoding to obtain the encoding results of the polarity and segments corresponding to each speech data; otherwise, performing variable-length encoding on the polarity and segments corresponding to each speech data to obtain the encoding results of the polarity and segments corresponding to each speech data.
[0014] This invention calculates and compares the polarity and segment corresponding to each speech data using PCM encoding for fixed-length encoding, and selects the encoding method with the least amount of data based on the allocation results of variable-length codewords and the size of the data volume when performing variable-length encoding on each speech data corresponding to the polarity segment. This ensures that the compression efficiency is maximized when compressing and storing speech data.
[0015] Preferably, the step of dividing all speech data into three categories includes: obtaining the sequence number of the polarity segment corresponding to each speech data, and classifying the first segment into three categories. The first voice data and the first The sequence number of the polarity segment corresponding to each speech data is denoted as: and :if , No. The voice data belongs to the first category of data; if or , No. The voice data belongs to the second category of data; if or , No. The voice data belongs to the third category of data.
[0016] This invention combines the principle of allocating variable-length codewords to all polarity segments, dividing all speech data into three categories, so as to determine the amount of data for subsequent variable-length encoding of each speech data according to the category to which each speech data belongs.
[0017] Preferably, the method for obtaining the estimated data volume includes: calculating the product of the number of speech data belonging to the first type of data and the length of the first type of codeword, the product of the number of speech data belonging to the second type of data and the length of the second type of codeword, and the product of the number of speech data belonging to the third type of data and the length of the third type of codeword; the estimated data volume is equal to the sum of the three product results.
[0018] This invention combines the quantity of speech data belonging to various data types to evaluate the amount of data when performing variable-length encoding on each speech data based on the allocation results of variable-length codewords and the polarity segment corresponding to each speech data.
[0019] Preferably, the method for obtaining the theoretical data volume includes: the length of the fixed-length codeword used in PCM encoding to encode polarity. The length of the fixed-length codeword used to encode a paragraph in PCM encoding. ; Calculate the sum of the lengths of fixed-length codewords that encode polarity and segmentation in PCM encoding. The theoretical data volume is equal to the product of the obtained sum and the total number of speech data.
[0020] This invention combines the total amount of speech data with the length of the fixed-length codewords used in PCM encoding to encode polarity and segments, and evaluates the amount of data for fixed-length encoding of the polarity and segments corresponding to each speech data using PCM encoding.
[0021] Preferably, the encoding result of the speech data is composed of the encoding result of the polarity and segment corresponding to the speech data and the encoding result of the sub-segment corresponding to the speech data. The method for obtaining the encoding result of the sub-segment corresponding to the speech data is as follows: the sub-segment corresponding to the speech data is encoded with a fixed length of 4 bits, and the fixed length codewords corresponding to the 16 sub-segments are different.
[0022] Secondly, the present invention provides a cloud-based voice data storage system, including a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the aforementioned cloud-based voice data storage method is implemented.
[0023] By adopting the above technical solution, a computer program is generated from the above-mentioned cloud platform-based voice data storage method and stored in a memory so that it can be loaded and executed by a processor. In this way, a terminal device can be made based on the memory and the processor for convenient use.
[0024] The beneficial effects of this invention are as follows:
[0025] This invention combines the polarity of the signal amplitude corresponding to speech data with the temporal correlation and similarity of segments. It combines polarity and segments to form polar segments, enhancing the correlation and similarity of polar segments corresponding to adjacent speech data. In subsequent variable-length encoding of speech data, variable-length codewords of different lengths are assigned according to the correlation and similarity between each polar segment and the polar segment corresponding to the previous speech data. This allows more speech data to be encoded with shorter codewords, thereby improving compression efficiency. Attached Figure Description
[0026] Figure 1 This is a flowchart illustrating a cloud-based voice data storage method according to the present invention;
[0027] Figure 2 This is a schematic diagram illustrating the allocation result of randomly assigning all variable-length codewords to all polarity segments;
[0028] Figure 3 This is a schematic diagram illustrating the allocation results of all variable-length codewords to all polarity segments after the first speech data has been encoded;
[0029] Figure 4 This is a schematic diagram illustrating the allocation results of all variable-length codewords to all polarity segments after the third speech data has been encoded. Detailed Implementation
[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0031] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0032] This invention discloses a method for storing voice data based on a cloud platform, referring to... Figure 1 This includes steps S1-S4:
[0033] S1. Acquire and quantize the speech signal, which contains multiple speech data.
[0034] It should be noted that the A-law thirteen-segment method is a method for non-uniform quantization. It achieves non-uniform quantization of the signal by dividing the amplitude range of the speech signal into multiple segments and using different quantization step sizes in each segment.
[0035] Specifically, sound wave vibrations are converted into continuously changing electrical signals using voice acquisition devices such as microphones. The continuously changing electrical signals are sampled, and the sampling results are quantized according to the A-law thirteen-segment method to obtain a speech signal. The speech signal contains multiple speech data, and the signal amplitude of the speech data is in the range of [-1, 1].
[0036] S2. Based on the results of polarity, segment, and sub-segment division in the A-law thirteen-segment method, determine the polarity, segment, and sub-segment corresponding to each speech data.
[0037] Specifically, in the A-law thirteen-segment method, the distribution range of the signal amplitude of speech data is divided into positive and negative polarities, and each polarity range is further divided into 8 segments; each segment is further divided into 16 sub-segments; therefore, for each speech data, the polarity, segment and sub-segment corresponding to each speech data can be determined according to the signal amplitude of each speech data.
[0038] For example, when the speech signal is [-0.0567227, -0.0410239, -0.0301126, -0.0129345, 0.0051028, 0.0050209, 0.0079312, 0.0149237], the speech signal contains 8 speech data points. The polarities of these 8 speech data points are negative, negative, negative, negative, positive, positive, positive, positive, and positive, respectively. The corresponding segments are the 5th segment in the negative polarity, the 6th segment in the negative polarity, the 7th segment in the negative polarity, the 1st segment in the positive polarity, the 1st segment in the positive polarity, the 2nd segment in the positive polarity, and the 2nd segment in the positive polarity.
[0039] S3. Set up three types of variable-length codewords of different lengths, combine the polarity and segment in the A-law thirteen-segment method to form multiple polarity segments, assign all variable-length codewords to all polarity segments, and perform variable-length encoding on each speech data according to the allocation results and the polarity segment corresponding to each speech data to obtain the encoding results of the polarity and segment corresponding to each speech data.
[0040] It should be noted that conventional PCM encoding uses fixed-length encoding for speech data, resulting in limited compression efficiency. This embodiment uses variable-length encoding for speech data to improve compression efficiency. Considering the temporal correlation and similarity of the signal amplitudes of speech data, this embodiment aims for a good correlation and similarity between the signal amplitudes of subsequent speech data and those of preceding speech data when performing variable-length encoding. Therefore, shorter codewords are allocated to signal amplitudes with good correlation and similarity to the preceding speech data, while longer codewords are allocated to signal amplitudes with poor correlation and similarity to the preceding speech data. This achieves variable-length encoding of speech data to improve compression efficiency.
[0041] It should be further explained that, regarding the division of signal amplitude of speech data into polarity, segments, and sub-segments in the A-law thirteen-segment method, the polarity and segments corresponding to the signal amplitude of speech data have strong temporal correlation and similarity, while the sub-segments corresponding to the signal amplitude of speech data have weak temporal correlation and similarity. Therefore, this embodiment combines polarity and segments to form polarity segments, thereby enhancing the correlation and similarity of polarity segments corresponding to adjacent speech data, and enabling more speech data to be encoded with shorter codewords, thus improving compression efficiency.
[0042] In summary, this embodiment sets up three types of variable-length codewords of different lengths, combines the polarity and segment in the A-law thirteen-segment method to form multiple polarity segments, assigns all variable-length codewords to all polarity segments, and performs variable-length encoding on each speech data according to the allocation results and the polarity segment corresponding to each speech data, specifically as follows:
[0043] 1. Set up 3 types of variable-length codewords with different lengths.
[0044] Specifically, variable-length codewords consist of a first digit and a second digit. The first type of codeword has a length of 1 bit and the first bit is the first digit. Therefore, there is only one first type of codeword. The second type of codeword and the third type of codeword have lengths of 3 bits and 6 bits, respectively. The first bit of both is the second digit, while the second bit is a different digit. Therefore, there are 2 second type codewords and 16 third type codewords.
[0045] In one embodiment, the first digit is 0 and the second digit is 1; in another embodiment, the first digit is 1 and the second digit is 0.
[0046] For example, when the first digit is 0 and the second digit is 1, there is one and only one first type of codeword, which is 0; there are two second type of codewords, which are 100 and 101 respectively; there are 16 third type of codewords, which are 110000, 110001, 110010, 110011, 110100, 110101, 110110, 110111, 111000, 111001, 111010, 111011, 111100, 111101, 111110, 111111.
[0047] 2. Combine the polarity and paragraphs in the A-law thirteen-line method to form multiple polarity paragraphs.
[0048] Specifically, in the A-law thirteen-line method, there are 2 types of polarity and 8 types of paragraphs combined. Therefore, combining the 2 types of polarity and 8 types of paragraphs in the A-law thirteen-line method results in 16 types of polarity paragraphs.
[0049] For example, the speech signal [-0.0567227, -0.0410239, -0.0301126, -0.0129345, 0.0051028, 0.0050209, 0.0079312, 0.0149237] contains 8 speech data points, and the polarity segments corresponding to these 8 speech data points are the 5th polarity segment, the 6th polarity segment, the 7th polarity segment, the 9th polarity segment, the 9th polarity segment, the 10th polarity segment, and the 10th polarity segment, respectively.
[0050] It should be noted that this embodiment combines polarity and paragraph to form polarity paragraphs, which enhances the correlation and similarity between polarity paragraphs corresponding to adjacent speech data, enabling more speech data to be encoded with shorter codewords, thereby improving compression efficiency.
[0051] 3. Based on the relationship between the polarity segment corresponding to each speech data and the polarity segment corresponding to the previous speech data, all speech data are divided into three categories.
[0052] All voice data is divided into three categories: Category 1, Category 2, and Category 3.
[0053] Specifically, the sequence number of the polarity segment corresponding to each speech data is obtained, and the segment number is assigned to the next segment. The first voice data and the first The sequence number of the polarity segment corresponding to each speech data is denoted as: and According to the The polar segment corresponding to the first voice data and the first The relationship between the polarity segments corresponding to each speech data point is used to determine the first... The class to which each voice data belongs, and the specific methods for obtaining it include:
[0054] (1) If the first The polar segment corresponding to the first speech data and the first Each voice data point corresponds to a segment of the same polarity, meaning... At that time, the first The voice data belongs to the first category of data. This indicates taking the absolute value.
[0055] (2) If the first The polar segment corresponding to the first speech data and the first The polarity segment corresponding to each speech data point is an adjacent polarity segment, that is... or At that time, the first The voice data belongs to the second category of data.
[0056] (3) If the first The polar segment corresponding to the first speech data and the first The polarity segments corresponding to each speech data point are neither segments of the same polarity nor adjacent segments of the same polarity, that is... or At that time, the first The voice data belongs to the third category of data.
[0057] It should be noted that when At that time, for the first voice data, the first voice data is divided into the third type of data.
[0058] For example, considering the eight speech data points contained in the speech signal [-0.0567227, -0.0410239, -0.0301126, -0.0129345, 0.0051028, 0.0050209, 0.0079312, 0.0149237], the polarity segments corresponding to the 2nd, 6th, and 8th speech data points are the same as the polarity segments corresponding to the preceding speech data points; therefore, the 2nd, 6th, and 8th speech data points belong to the first category of data. The polarity segments corresponding to the 3rd, 4th, and 7th speech data points are adjacent to the polarity segments corresponding to the preceding speech data points; therefore, the 3rd, 4th, and 7th speech data points belong to the second category of data. However, the polarity segment corresponding to the 5th speech data point is neither the same as nor adjacent to the polarity segment corresponding to the preceding speech data point; therefore, the 1st and 5th speech data points belong to the third category of data.
[0059] 4. Based on the number of voice data belonging to each category and the length of the codewords in each category, obtain the estimated data volume.
[0060] It should be noted that in the subsequent process of "allocating all variable-length codewords to all polarity segments, and performing variable-length encoding on the polarity and segments corresponding to each speech data according to the allocation results and the polarity segments corresponding to each speech data": (1) assigning first-class codewords to the polarity segment corresponding to the previous speech data. Therefore, if the polarity segment corresponding to the next speech data is the same polarity segment as the polarity segment corresponding to the previous speech data, that is, if the next speech data belongs to the first-class data, the encoding length is equal to the length of the first-class codeword; (2) assigning second-class codewords to the two polarity segments adjacent to the polarity segment corresponding to the previous speech data. Therefore, if the polarity segment corresponding to the next speech data is the same polarity segment as the previous speech data, that is, if the next speech data belongs to the first-class data, the encoding length is equal to the length of the first-class codeword; The polarity segment corresponding to the speech data is an adjacent polarity segment to the polarity segment corresponding to the previous speech data. That is, if the next speech data belongs to the second type of data, the coding length is equal to the length of the second type of codeword; (3) Allocate the third type of codeword to all remaining polarity segments. That is, allocate the third type of codeword to the polarity segment corresponding to the previous speech data that is neither the same polarity segment nor an adjacent polarity segment. Therefore, if the polarity segment corresponding to the next speech data is neither the same polarity segment nor an adjacent polarity segment to the polarity segment corresponding to the previous speech data, that is, if the next speech data belongs to the third type of data, the coding length is equal to the length of the third type of codeword.
[0061] Therefore, in the subsequent process of "assigning all variable-length codewords to all polarity segments, and encoding the polarity and segment corresponding to each speech data according to the assignment results and the polarity segment corresponding to each speech data", the estimated data volume of the encoding result is equal to the sum of the products of the number of speech data belonging to the first type of data, the second type of data and the third type of data and the length of the first type of codeword, the second type of codeword and the third type of codeword.
[0062] Specifically, the lengths of the first, second, and third type codewords are 1 bit, 3 bits, and 6 bits, respectively. The number of speech data belonging to the first, second, and third types of data are respectively denoted as... , , The formula for calculating the estimated data volume is:
[0063] ;
[0064] In the formula, To estimate the amount of data, , , The numbers represent the number of voice data belonging to the first, second, and third categories, respectively.
[0065] For example, in summary, the number of voice data belonging to the first type of data, the second type of data, and the third type of data are respectively =3、 =3、 =2, therefore, the estimated data volume =3×1+3×3+2×6=24.
[0066] 5. Based on the total amount of speech data and the length of the fixed-length codewords in PCM encoding, obtain the theoretical data volume.
[0067] It should be noted that when using PCM encoding to perform fixed-length encoding on the polarity and segments corresponding to speech data, since the A-law thirteen-segment method has 2 polarities and 8 segments, the length of the fixed-length codeword for fixed-length encoding on the polarity of the speech data is... equal The length of the fixed-length codeword used for fixed-length encoding of the corresponding segment of speech data. equal Therefore, for each speech data point, the length of the encoding for its corresponding polarity and segment is equal to... Therefore, the theoretical data size when using PCM encoding to perform fixed-length encoding on the polarity and segments corresponding to all speech data is equal to the total number of speech data. The product of.
[0068] Specifically, the length of the fixed-length codeword that encodes the polarity corresponding to the speech data using PCM encoding. The length of the fixed-length codeword used for fixed-length encoding of the corresponding segment of speech data. The total number of all voice data is denoted as The formula for calculating the theoretical data volume is:
[0069] ;
[0070] In the formula, For theoretical data volume, The number of all voice data. This refers to the length of the fixed-length codeword used in PCM encoding to encode polarity. This refers to the length of the fixed-length codeword used to encode a segment in PCM encoding, and , .
[0071] For example, the speech signal [-0.0567227, -0.0410239, -0.0301126, -0.0129345, 0.0051028, 0.0050209, 0.0079312, 0.0149237] contains a total of 8 speech data points, which is the total number of speech data points. =8, then the theoretical data volume =8×(1+3)=32.
[0072] 6. When the estimated data volume is greater than or equal to the theoretical data volume, PCM coding is used to perform fixed-length coding on the polarity and segment corresponding to each speech data to obtain the coding results of the polarity and segment corresponding to each speech data. When the estimated data volume is less than the theoretical data volume, all variable-length codewords are assigned to all polarity segments. Based on the allocation results and the polarity segments corresponding to each speech data, variable-length coding is performed on the polarity and segment corresponding to each speech data to obtain the coding results of the polarity and segment corresponding to each speech data.
[0073] It should be noted that this embodiment calculates and compares the theoretical data volume when performing fixed-length encoding on the polarity and segment corresponding to each speech data using PCM encoding, and the estimated data volume when performing variable-length encoding on each speech data based on the allocation results of variable-length codewords and the polarity segment corresponding to each speech data. By selecting the encoding method with the least data volume, the compression efficiency when compressing and storing speech data can be maximized.
[0074] (1) Fixed-length encoding of the polarity and segment corresponding to each speech data is performed by PCM encoding to obtain the encoding results of the polarity and segment corresponding to the speech data, including:
[0075] Specifically, in PCM encoding, the fixed-length codeword consists of a first digit and a second digit. The polarity corresponding to the speech data is encoded by a fixed-length codeword with a length of 1 bit, and the fixed-length codewords corresponding to the two polarities are different. The segments corresponding to the speech data are encoded by a fixed-length codeword with a length of 3 bits, and the fixed-length codewords corresponding to the 8 segments are different.
[0076] For example, polarity is divided into positive polarity and negative polarity. The fixed-length codeword corresponding to positive polarity is 1 and the fixed-length codeword corresponding to negative polarity is 0. There are 8 types of paragraphs. The fixed-length codewords corresponding to the 8 types of paragraphs are 000, 001, 010, 011, 100, 101, 110, and 111, respectively.
[0077] For example, for all speech data in the speech signal [-0.0567227, -0.0410239, -0.0301126, -0.0129345, 0.0051028, 0.0050209, 0.0079312, 0.0149237], the polarity encoding results for each speech data are “0”, “0”, “0”, “0”, “1”, “1”, “1”, and the segment encoding results for each speech data are “100”, “100”, “101”, “110”, “000”, “000”, “001”, and “001”, and the total data size of the polarity and segment encoding results for all speech data is equal to 32.
[0078] (2) Allocate all variable-length codewords to all polarity segments, including: initially, when performing variable-length encoding on the first speech data, randomly assigning all variable-length codewords to all polarity segments; for the first... When encoding the first voice data into variable length, the first... The polarity segment corresponding to each speech data is denoted as the target segment. The target segment is assigned a first-class codeword, the two polarity segments adjacent to the target segment are assigned a second-class codeword, and the remaining polarity segments are assigned a third-class codeword.
[0079] The method for obtaining the two polarity paragraphs adjacent to the target paragraph includes: designating the target paragraph as the first... a type of polar paragraph; when At that time, the first and the Two polarity paragraphs, as two types of polarity paragraphs adjacent to the target paragraph; when At that time, the first and the Two polarity paragraphs, as two types of polarity paragraphs adjacent to the target paragraph; when and At that time, the first and the Two polarity paragraphs, which are adjacent to the target paragraph.
[0080] (3) Perform variable-length coding on the polarity and segment corresponding to each speech data to obtain the coding result of the polarity and segment corresponding to the speech data, including: taking the variable-length codewords assigned to the polarity segment corresponding to the speech data in the allocation result as the coding result of the speech data.
[0081] For example, for the speech signal [-0.0567227, -0.0410239, -0.0301126, -0.0129345, 0.0051028, 0.0050209, 0.0079312, 0.0149237], due to the estimated data volume... =24 and theoretical data volume =32, therefore the estimated data volume is less than the theoretical data volume. Thus, all variable-length codewords are assigned to all polarity segments. Based on the assignment results and the polarity segments corresponding to each speech data point, variable-length coding is performed on the polarity and segments corresponding to each speech data point to obtain the coding results. The specific process is as follows:
[0082] (1) Initially, when performing variable-length encoding on the first speech data, all variable-length codewords are randomly assigned to all polarity segments, and the assignment results are as follows: Figure 2As shown; the polarity segment corresponding to the first speech data is the fifth polarity segment. Therefore, the encoding result of the polarity and segment corresponding to the first speech data is "110001".
[0083] (2) When performing variable-length encoding on the second speech data, since the polarity segment corresponding to the first speech data is the fifth polarity segment, the fifth polarity segment is assigned the first type codeword "0", the sixth and seventh polarity segments adjacent to the fifth polarity segment are assigned the second type codewords "100" and "101", and all remaining polarity segments are assigned the third type codewords. The allocation result is as follows: Figure 3 As shown; the polarity segment corresponding to the second speech data is the fifth polarity segment. Therefore, the encoding result of the polarity and segment corresponding to the second speech data is "0".
[0084] (3) When performing variable-length encoding on the third speech data, since the polarity segment corresponding to the second speech data is the fifth polarity segment, the fifth polarity segment is assigned the first type codeword "0", the sixth and seventh polarity segments adjacent to the fifth polarity segment are assigned the second type codewords "100" and "101", and all remaining polarity segments are assigned the third type codewords. The allocation result is as follows: Figure 3 As shown; the polarity segment corresponding to the 3rd speech data is the 6th polarity segment. Therefore, the encoding result of the polarity and segment corresponding to the 3rd speech data is "101".
[0085] (4) When performing variable-length encoding on the fourth speech data, since the polarity segment corresponding to the third speech data is the sixth polarity segment, the sixth polarity segment is assigned the first type codeword "0", the fifth and seventh polarity segments adjacent to the sixth polarity segment are assigned the second type codewords "100" and "101", and all remaining polarity segments are assigned the third type codewords. The allocation result is as follows: Figure 4 As shown; the polarity segment corresponding to the 4th speech data is the 7th polarity segment. Therefore, the encoding result of the polarity and segment corresponding to the 4th speech data is "101".
[0086] (5) By analogy, the polarity and segment encoding results corresponding to the 5th to 8th speech data are obtained as “110101”, “0”, “101” and “0” respectively.
[0087] (6) The total data volume of the polarity and segment encoding results corresponding to all speech data is equal to 24.
[0088] (7) The polarity and segment corresponding to each speech data are encoded by PCM with fixed length, and the total data volume of the encoded results of the polarity and segment corresponding to the speech data is equal to 32. According to the allocation result and the polarity segment corresponding to each speech data, the polarity and segment corresponding to each speech data are encoded with variable length, and the total data volume of the encoded results of the polarity and segment corresponding to each speech data is equal to 24. Therefore, the method of this embodiment reduces the data volume of the encoded results of speech data and improves the compression efficiency.
[0089] It should be noted that this embodiment combines the polarity of the signal amplitude of the speech data with the temporal correlation and similarity of the segments. When performing variable-length encoding on the speech data, variable-length codewords of different lengths are assigned according to the correlation and similarity between each polarity segment and the polarity segment corresponding to the previous speech data. This allows more speech data to be encoded with shorter codewords, thereby improving compression efficiency.
[0090] S4. Combine the encoding results of the polarity and segment corresponding to the speech data with the encoding results of the sub-segments corresponding to the speech data to form the encoding result of the speech data, and transmit and store the encoding results of all speech data on the cloud platform.
[0091] 1. PCM coding is used to perform fixed-length coding on the segments corresponding to each speech data to obtain the coding results of the segments corresponding to each speech data.
[0092] Specifically, in PCM encoding, a fixed-length codeword consists of a first digit and a second digit. Since there are 16 sub-segments in the A-law 13-segment encoding method, the length of the fixed-length codeword for encoding the corresponding sub-segment of the speech data is equal to... Therefore, fixed-length codewords of 4 bits are used to encode the segments corresponding to the speech data, and the fixed-length codewords corresponding to the 16 segments are different.
[0093] For example, if there are 16 types of sub-segments, the fixed-length codewords corresponding to the 16 sub-segments are 0000, 0001, 0010, 0011, 0100, 0101, 0110, 0111, 1000, 1001, 1010, 1011, 1100, 1101, 1110, and 1111.
[0094] 2. Combine the polarity and segment encoding results of the speech data with the sub-segment encoding results of the speech data to form the speech data encoding result.
[0095] 3. Transmit and store the encoded results of all voice data on the cloud platform.
[0096] This invention also discloses a cloud-based voice data storage system, including a processor and a memory. The memory stores computer program instructions, which, when executed by the processor, implement a cloud-based voice data storage method according to the present invention.
[0097] The system also includes other components well known to those skilled in the art, such as communication buses and communication interfaces, the settings and functions of which are known in the art and will not be described in detail here.
Claims
1. A voice data storage method based on a cloud platform, characterized in that, include: The speech signal is acquired and quantized, and the speech signal contains multiple speech data. Based on the polarity and segmentation results of the A-law thirteen-segment method, the polarity and segment corresponding to each speech data are determined. Variable-length encoding is then performed on the polarity and segment corresponding to each speech data to obtain the encoding results for the polarity and segment corresponding to each speech data, including: Three types of variable-length codewords with different lengths are set: the first type of codeword has a length of 1 bit and the first bit is the first digit; the second type of codeword and the third type of codeword have lengths of 3 bits and 6 bits respectively and the first bit of the second digit of the second type of codeword and the third type of codeword are different digits; Polarity and segment are combined to form 16 polarity segments; variable-length codewords are assigned to all polarity segments. Based on the assignment results and the polarity segments corresponding to each speech data, variable-length encoding is performed on each speech data, including: after encoding the previous speech data, the polarity segment corresponding to the previous speech data is recorded as the target segment, a first-class codeword is assigned to the target segment, a second-class codeword is assigned to the two polarity segments adjacent to the target segment, and a third-class codeword is assigned to all remaining polarity segments; based on the assignment results, variable-length encoding is performed on the next speech data. All voice data encoding results are transmitted and stored on the cloud platform.
2. The voice data storage method based on a cloud platform according to claim 1, characterized in that, In the aforementioned A-law thirteen-segment method, the distribution range of the signal amplitude of the speech data is divided into positive and negative polarities. Each polarity range is further divided into 8 segments, and each segment is further divided into 16 sub-segments.
3. The voice data storage method based on a cloud platform according to claim 1, characterized in that, The first digit is 0 and the second digit is 1; or the first digit is 1 and the second digit is 0.
4. The voice data storage method based on a cloud platform according to claim 1, characterized in that, The method for obtaining the two polarity paragraphs adjacent to the target paragraph includes: Label the target paragraph as paragraph number 1. Polarity paragraphs; when At that time, the first and the Two types of polarity paragraphs, which are adjacent to the target paragraph; when At that time, the first and the Two types of polarity paragraphs, which are adjacent to the target paragraph; when and At that time, the first and the Two polarity paragraphs, which are adjacent to the target paragraph.
5. The voice data storage method based on a cloud platform according to claim 1, characterized in that, The method of combining the polarity and segment division results of the A-law thirteen-segment method to determine the polarity and segment corresponding to each speech data, performing variable-length encoding on the polarity and segment corresponding to each speech data to obtain the encoding results of the polarity and segment corresponding to each speech data, further includes: Based on the relationship between the polarity segment corresponding to each speech data and the polarity segment corresponding to the previous speech data, all speech data are divided into three categories; based on the number of speech data belonging to each category and the length of the codewords in each category, the estimated data volume is obtained; based on the total number of speech data and the length of the fixed-length codewords used in PCM encoding to encode polarity and segments, the theoretical data volume is obtained. If the estimated data volume is greater than or equal to the theoretical data volume, PCM coding is used to perform fixed-length coding on the polarity and segment corresponding to each speech data to obtain the coding results of the polarity and segment corresponding to each speech data; otherwise, variable-length coding is used on the polarity and segment corresponding to each speech data to obtain the coding results of the polarity and segment corresponding to each speech data.
6. The voice data storage method based on a cloud platform according to claim 5, characterized in that, The document states that all voice data is divided into three categories, including: Obtain the sequence number of the polarity segment corresponding to each speech data point, and then assign the first segment to the next polarity segment. The first voice data and the first The sequence number of the polarity segment corresponding to each speech data is denoted as: and : if , No. The voice data belongs to the first category of data; if or , No. The voice data belongs to the second category of data; if or , No. The voice data belongs to the third category of data.
7. A voice data storage method based on a cloud platform according to claim 6, characterized in that, The method for obtaining the estimated data volume includes: Calculate the product of the number of speech data belonging to the first type of data and the length of the first type of codeword, the product of the number of speech data belonging to the second type of data and the length of the second type of codeword, and the product of the number of speech data belonging to the third type of data and the length of the third type of codeword; the estimated data volume is equal to the sum of the three product results.
8. The voice data storage method based on a cloud platform according to claim 5, characterized in that, The method for obtaining the theoretical data volume includes: The length of the fixed-length codeword that encodes polarity in PCM encoding. The length of the fixed-length codeword used to encode a paragraph in PCM encoding. ; Calculate the sum of the lengths of fixed-length codewords that encode polarity and segmentation in PCM encoding. The theoretical data volume is equal to the product of the obtained sum and the total number of speech data.
9. A voice data storage method based on a cloud platform according to claim 2, characterized in that, The encoding result of the speech data is composed of the encoding result of the polarity and segment corresponding to the speech data and the encoding result of the sub-segment corresponding to the speech data. The method for obtaining the encoding result of the sub-segment corresponding to the speech data is as follows: the sub-segment corresponding to the speech data is encoded using a fixed-length codeword with a length of 4 bits, and the fixed-length codewords corresponding to the 16 sub-segments are different.
10. A voice data storage system based on a cloud platform, characterized in that, include: A processor and a memory, wherein the memory stores computer program instructions that, when executed by the processor, implement a cloud-based voice data storage method according to any one of claims 1-9.