Voice coding and decoding device, method and computer program product

By decompressing and replacing the first compressed data in the voice encoder, the problem of restoring one frame of content during decoding in the prior art is solved, and the sound effect quality is improved and the frame number consistency is maintained.

CN119993170APending Publication Date: 2025-05-13NUVOTON
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411164033.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-13
Filing Date
2024-08-23
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing voice encoding and decompression technology will generate unnecessary frame data during compression and decompression, resulting in one more frame of content being restored during decoding, affecting the quality of sound effects.

Method used

By decompressing the first compressed data in the voice encoder, preset feature data is obtained and saved, and replaced with the original first compressed data, so that the voice decoder can directly use the preset feature data during decoding, avoiding decompression of the unnecessary first compressed data.

Benefits of technology

It effectively avoids the generation of unnecessary frame data during decoding, improves the sound quality, and ensures frame number consistency during the speech restoration process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993170A_ABST
    Figure CN119993170A_ABST
Patent Text Reader

Abstract

The invention provides a voice coding and decoding device, method and computer program product, comprising the following steps of: compressing multiple frames of original voice into multiple compressed data, and decompressing first compressed data in the multiple compressed data to obtain and store preset characteristic data, and replacing the first compressed data with the preset feature data for voice restoration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to audio codec, and more particularly to an audio codec device, method, and computer program product. Background Art

[0002] A voice frame is usually 320 points. One voice compression method is to use interleave compression technology, where adjacent original frames each provide part of the content, which are combined and compressed to generate compressed data.

[0003] Taking the encoding of N frames of original speech as an example, the second half 160 points of the first original frame FrameO_1 are combined with the first half 160 points of the second original frame FrameO_2 and compressed into compressed data (corresponding to the second original frame FrameO_2, numbered as the second compressed data Data_2). In particular, the first compressed data Data_1 corresponding to the first original frame FrameO_1 is obtained by combining the 160 points of the initial dummy content (which can all be 0) with the first 160 points of the first original frame FrameO_1. As for the second half of the last original frame FrameO_N, it is combined with the ending dummy content and compressed to form the (N+1)th compressed data Data_N+1.

[0004] In this way, N frames of original speech FrameO_1...FrameO_N are compressed into (N+1) pieces of compressed data Data_1...Data_N+1. When decoding, one more frame of content will be restored. Summary of the invention

[0005] A speech encoding and decoding method implemented according to an embodiment of the present application includes: compressing multiple frames of original speech into multiple compressed data, wherein a first compressed data among the multiple compressed data is also decompressed to obtain and save a preset feature data, and the first compressed data is replaced by the preset feature data for speech restoration, wherein the first compressed data is generated by compressing a first original frame among the multiple frames of original speech, and the first original frame is the starting frame among the multiple frames of original speech.

[0006] In one embodiment, the speech encoding and decoding method further includes: decompressing the multiple frames to restore the speech, wherein a second compressed data in the multiple compressed data is decompressed using the preset characteristic data. The second compressed data is generated by compressing the first original frame and a second original frame. The second original frame is inferior to the first original frame in the multiple frames of original speech.

[0007] In one embodiment, the number of frames of the multiple frames of original speech is N, which are numbered from the first original frame to the Nth original frame in time sequence. The multiple compressed data are N+1, which are numbered from the first compressed data to the (N+1)th compressed data in time sequence. The (N+1)th compressed data is obtained by compressing an Nth original frame. The number of frames of the multiple frames of restored speech is N, which are numbered from the first restored frame to the Nth restored frame in time sequence. The first restored frame is obtained by decompressing the second compressed data with the preset feature data. The Nth restored frame is obtained by decompressing the (N+1)th compressed data with an Nth feature data, wherein the Nth feature data is generated when decompressing an Nth compressed data.

[0008] The above concept can be implemented into a computer program product, including a computer program, which is executed by a computer to implement the above-mentioned speech coding and decoding method.

[0009] The above concept is also used to implement a speech encoding and decoding device, including a speech encoder and a speech decoder. The speech encoder is used to compress multiple frames of original speech into multiple compressed data. The speech decoder is used to decompress the multiple compressed data into multiple frames of restored speech. The initial frame in the multiple frames of original speech is a first original frame. The speech encoder compresses the first original frame to generate a first compressed data in the multiple compressed data. The speech encoder also decompresses the first compressed data to obtain and save a preset feature data, and replaces the first compressed data with the preset feature data for the speech decoder to perform speech restoration.

[0010] In one embodiment, the speech decoder decompresses a second compressed data among the plurality of compressed data using the preset characteristic data. The second compressed data is generated by the speech encoder compressing the first original frame and a second original frame. The second original frame is next to the first original frame in the plurality of frames of original speech.

[0011] The following is a detailed description of the present invention with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 A speech encoding and decoding device 100 implemented according to an embodiment of the present application includes a speech encoder 102 and a speech decoder 104;

[0013] Figure 2 The operation of the speech encoder 102 is illustrated in a timing diagram;

[0014] Figure 3 The operation of the speech decoder 104 is illustrated in a timing diagram;

[0015] Figure 4 A flowchart illustrating a speech coding method according to an embodiment of the present application;

[0016] Figure 5 A flowchart illustrating a speech decoding method according to an embodiment of the present application; and

[0017] Figure 6 The structure of a computer program product 602 is illustrated in FIG. 6 .

[0018]

Explanation of symbols

[0019] 100: speech coding and decoding device

[0020] 102: Voice Codec

[0021] 104: Voice Decoder

[0022] 602: Computer program product

[0023] 604: Computer

[0024] 606: Computer Program

[0025] 608: Processor

[0026] 610: Memory

[0027] Data_1…Data_N+1: compressed data

[0028] Feature_1: Preset feature data

[0029] Feature_i: Feature data

[0030] FrameO_1…FrameO_N: original frame

[0031] FrameR_1…FrameR_N: Restore frames

[0032] S402…S420, S502…S510: Steps DETAILED DESCRIPTION

[0033] The following description lists various embodiments of the present invention to introduce the basic concepts of the present invention and is not intended to limit the content of the present invention. The actual scope of the invention should be defined in accordance with the claims. The various functional structures mentioned below can be implemented by a combination of hardware, software, and firmware, and can also include special circuits. The various functional structures are not limited to being implemented separately, but can also be combined together to share certain functions.

[0034] Figure 1A speech coding and decoding device 100 implemented according to an embodiment of the present application includes a speech encoder 102 and a speech decoder 104. The speech encoder 102 is used to compress multiple frames of original speech (hereinafter referred to as original frames, numbered as the first original frame FrameO_1 to the Nth original frame FrameO_N in time sequence, N is a number) into multiple compressed data (numbered as the first compressed data Data_1 to the (N+1)th compressed data Data_N+1 in time sequence). The speech decoder 104 is used to decompress multiple frames of restored speech (hereinafter referred to as restored frames, numbered as the first restored frame FrameR_1 to the Nth restored frame FrameR_N in time sequence).

[0035] In particular, the speech encoder 102 does not simply perform compression, but also decompresses specific compression results.

[0036] The speech encoder 102 first combines the initial dummy content (e.g., fill 0) with the first original frame FrameO_1 and compresses it to generate the first compressed data Data_1. After obtaining the first compressed data Data_1, the speech encoder 102 also decompresses the first compressed data Data_1 to obtain and save a preset feature data Feature_1. As for the subsequent other original frames, the speech encoder 102 will not decompress them after compression. For example, the second compressed data Data_2 is generated by the speech encoder 102 combining the first original frame FrameO_1 with a second original frame FrameO_2 and compressing them, but the speech encoder 102 will not decompress the second compressed data Data_2.

[0037] The preset feature data Feature_1 is reserved for use by the speech decoder 104. The speech decoder 104 decompresses the second compressed data Data_2 based on the preset feature data Feature_1 to obtain a first restored frame FrameO_1.

[0038] With such a design, the voice decoder 104 can omit the decompression of the first compressed data Data_1. Since the original content of the first compressed data Data_1 includes the initial virtual content, the decompression result (hereinafter referred to as FrameR_0) is meaningless. The voice decoder 104 omits the decompression of the first compressed data Data_1, that is, omits the meaningless restoration frame FrameR_0. If it is a sinusoidal wave signal, the meaningless restoration frame FrameR_0 will cause the restored sound effect to have a complex sound (such as, bop...bop...bop...). The restored sound effect of this application will not have a meaningless restoration frame FrameR_0, and the decompression restoration effect is good.

[0039] In one embodiment, the speech encoder 102 deletes the first compressed data Data_1 after obtaining and saving the preset feature data Feature_1, so that the speech decoder 104 does not decompress the first compressed data Data_1. The preset feature data Feature_1 replaces the first compressed data Data_1 and is supplied to the speech decoder 104 for speech restoration.

[0040] like Figure 1 As shown, starting from the second compressed data Data_2, the speech decoder 104 generates feature data Feature_i by decompressing each compressed data Data_i (i is an integer greater than 1) for decompression of the next compressed data Data_i+1. The restored frames FrameR_2...FrameR_N are generated accordingly.

[0041] Under this architecture, the number of frames of the original speech is N (FrameO_1~FrameO_N), and the number of restored frames is also N (FrameR_1~FrameR_N), which is consistent with the number of original speech frames.

[0042] Figure 2 The operation of the speech encoder 102 is illustrated in a timing diagram.

[0043] The speech encoder 102 combines the initial dummy content (e.g., 160 points of "0" value) with the first half of the first restored frame FrameO_1, and compresses it into the first compressed data Data_1. The speech encoder 102 also decompresses the first compressed data Data_1, generates and stores the preset feature data Feature_1. Next, the speech encoder 102 allows two adjacent original frames to each provide partial content to be combined and compressed into compressed data. For example, the speech encoder 102 combines the second half of the first original frame FrameO_1 with the first half of the second original frame FrameO_2. 160 points, and compresses them into the second compressed data Data_2, and the same processing is performed all the way to generate the Nth compressed data Data_N. Finally, the speech encoder 102 combines the second half of the Nth original frame FrameO_N with the ending dummy content (e.g., 160 points of "0" value), and compresses them into the (N+1)th compressed data Data_N+1. The diagram clearly shows that a decompression occurs during the encoding stage of the present application (decoding the first compressed data Data_1 to generate the preset feature data Feature_1).

[0044] Figure 3 The operation of the speech decoder 104 is illustrated in a timing diagram.

[0045] The speech decoder 104 first decompresses the second compressed data Data_2 according to the preset feature data Feature_1 generated in the encoding stage to generate the first restored frame FrameR_1 and the second feature data Feature_2. The second feature data Feature_2 will be used to decompress the next compressed data (Data_3, not shown in the figure). According to this rule, after the Nth compressed data Data_N is decompressed, in addition to generating the (N-1)th restored frame FrameR_N-1, the Nth feature data Feature_N is also generated. The speech decoder 104 decompresses the (N+1)th compressed data Data_N+1 based on the Nth feature data Feature_N to generate the final Nth restored frame FrameR_N.

[0046] Figure 4 The present invention is a flowchart illustrating a speech coding method according to an embodiment of the present application.

[0047] Step S402 combines the initial dummy content with the first original frame FrameO_1, and compresses it in step S404 to obtain the first compressed data Data_1. Step S406 decompresses the first compressed data Data_1, obtains the preset feature data Feature_1, and stores it. Step S408 sets the variable i to 2. Step S410 combines the two adjacent original frames FrameO_i-1 and FrameO_i, and compresses it in step S412 to obtain the i-th compressed data Data_i. Step S414 determines whether the original frame FrameO_i is the last frame. If not, step S416 increments the variable i. If so, step S418 combines the original frame FrameO_i with the ending dummy content, and compresses it in step S420 to obtain the (N+1)-th compressed data Data_N+1, and the process ends.

[0048] Figure 5 The present invention is a flowchart illustrating a speech decoding method according to an embodiment of the present application.

[0049] Step S502 decompresses the second compressed data Data_2 based on the preset feature data Feature_1, generates the first restored frame FrameR_1, and collects the second feature data Feature_2. Step S504 sets the variable i to 2. Step S506 decompresses the (i+1)th compressed data Data_i+1 based on the i-th feature data Feature_i, generates the i-th restored frame FrameR_i, and collects the (i+1)th feature data Feature_i+1. Step S508 determines whether the i+1th compressed data Data_i+1 is the last compressed data. If so, the process ends. If not, step S510 increments the variable i.

[0050] In addition to being implemented as hardware, the speech codec of this application can also be implemented as software.

[0051] Figure 6 To illustrate the structure of a computer program product 602 implemented according to an embodiment of the present application, a computer 604 is loaded to execute a computer program 606, which is run by a processor 608 to implement the aforementioned speech coding and decoding technology. The memory 610 can temporarily store the data used for the operation, including original speech frames (FrameO_1 to FrameO_N), preset feature data (Feature_1), compressed data (Data_1 to Data_N+1), feature data (Feature_i, i>1), and speech restoration frames (FrameR_1 to FrameR_N).

[0052] The above various embodiments may be implemented in various ways. Any technology that performs decompression at the beginning of encoding and uses preset feature data for decoding may fall within the technical scope of this application.

[0053] In other implementations, the encoding details may be slightly different from the above. For example, the first original frame FrameO_1 may not be compressed in combination with the virtual content at the beginning of half a frame, and may have other deformations; the Nth original frame FrameO_N may not be compressed in combination with the virtual content at the end of half a frame, and may have other deformations. Two adjacent frames may not be compressed in combination with equal data volume, and may have other deformations.

[0054] Although the present invention has been disclosed as above with preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art may make some changes and modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be determined by the claims.

Claims

1. A speech encoding and decoding method, characterized in that: include: Compress multiple frames of original speech into multiple compressed data, wherein a first compressed data among the multiple compressed data is also decompressed to obtain and save a preset feature data, and the preset feature data is used to replace the first compressed data for speech restoration, wherein the first compressed data is generated by compressing a first original frame among the multiple frames of original speech, and the first original frame is the initial frame among the multiple frames of original speech.

2. The speech encoding and decoding method according to claim 1, characterized in that: Also includes: Decompressing the multiple frames of restored speech, wherein a second compressed data among the multiple compressed data is decompressed using the preset characteristic data; in: The second compressed data is generated by compressing the first original frame and a second original frame; and The second original frame is subsequent to the first original frame in the multiple frames of original speech.

3. The speech encoding and decoding method according to claim 1, characterized in that: Also includes: The first compressed data is replaced by the preset characteristic data so that decompression of the first compressed data no longer occurs.

4. The speech encoding and decoding method according to claim 2, characterized in that: The number of frames of the multiple frames of original speech is N, and they are numbered in time sequence as the first original frame to the Nth original frame; The plurality of compressed data are N+1 pieces, which are numbered from the first compressed data to an (N+1)th compressed data in time sequence, wherein the (N+1)th compressed data is obtained by compressing an Nth original frame; The number of frames of the multi-frame restored speech is N, which are numbered from a first restored frame to an Nth restored frame in time sequence; The first restored frame is obtained by decompressing the second compressed data with the preset characteristic data; and The Nth restored frame is obtained by decompressing the (N+1)th compressed data using an Nth characteristic data, wherein the Nth characteristic data is generated when decompressing an Nth compressed data.

5. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the speech coding and decoding method according to any one of claims 1 to 4 is implemented.

6. A speech encoding and decoding device, characterized in that: include: A speech encoder, used for compressing multiple frames of original speech into multiple compressed data; as well as A voice decoder, used for decompressing the plurality of compressed data into a plurality of frames to restore the voice; in: The initial frame in the multiple frames of original speech is a first original frame; The speech encoder compresses the first original frame to generate a first compressed data among the plurality of compressed data; and The speech encoder also decompresses the first compressed data to obtain and save a preset feature data, and replaces the first compressed data with the preset feature data for the speech decoder to perform speech restoration.

7. The speech encoding and decoding device according to claim 6, characterized in that: The speech decoder decompresses a second compressed data among the plurality of compressed data using the preset characteristic data; The second compressed data is generated by the speech encoder compressing the first original frame and a second original frame; and The second original frame is subsequent to the first original frame in the multiple frames of original speech.

8. The speech encoding and decoding device according to claim 6, characterized in that: Also includes: The speech encoder replaces the first compressed data with the preset characteristic data, so that the speech decoder does not decompress the first compressed data.

9. The speech encoding and decoding device according to claim 7, characterized in that: The number of frames of the multiple frames of original speech is N, and they are numbered in time sequence as the first original frame to the Nth original frame; The plurality of compressed data are N+1 pieces, which are numbered from the first compressed data to an (N+1)th compressed data in time sequence, wherein the (N+1)th compressed data is obtained by compressing an Nth original frame; The number of frames of the multi-frame restored speech is N, which are numbered from a first restored frame to an Nth restored frame in time sequence; The first restored frame is obtained by decompressing the second compressed data with the preset characteristic data; and The Nth restored frame is obtained by the speech decoder decompressing the (N+1)th compressed data with an Nth characteristic data, wherein the Nth characteristic data is generated when the speech decoder decompresses the Nth compressed data.