Method for improving voice quality of receiving end of voice communication system

By using a combination technology of frame synchronization and RS encoding in the voice communication system, the problems of delay and disordered order caused by heterogeneous network transmission are solved, the voice quality of the receiver is improved, and the integrity and continuity of the voice frame are ensured.

CN120050344AInactive Publication Date: 2025-05-27NO 30 INST OF CHINA ELECTRONIC TECH GRP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510195179.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the voice communication system, after the voice signal is transmitted to the receiving end through a heterogeneous network, problems such as delay, disordered order, error code, packet loss, packet interval jitter, unbalanced, and instability are prone to problems such as degradation of the voice quality at the receiving end.

Method used

The encoding post-processing module is built on the sending end, and the frame synchronization head and RS code are attached; the UDP voice packet processing module and the decoding driver module are built on the receiving end, and pre-processing such as cache, sorting, sliding average, rate adjustment and frame synchronization are carried out to ensure the integrity and continuity of the voice frame.

Benefits of technology

It effectively alleviates the influence of unstable factors on the transmission path on the voice, ensures the integrity and continuity of the voice frame, improves the voice quality at the receiver, and ensures the balance and stability of the voice code rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050344A_ABST
    Figure CN120050344A_ABST
Patent Text Reader

Abstract

The invention discloses a method for improving voice quality of a receiving end of a voice communication system, and relates to the technical field of voice communication. According to the invention, in the aspect of protocol design, a mode of combining frame synchronization with RS coding is adopted, so that the influence of unstable factors such as disorder, packet loss and burst error codes introduced on a transmission path on voice can be reduced, and the integrity and continuity of voice frames are ensured; a coding post-processing module is designed at a sending end, a frame synchronization head and an RS code can be added to a voice frame, a UDP voice packet processing module and a decoding driving module are designed at a receiving end, and a UDP voice packet is fully preprocessed before being sent to a decoding chip; the processing method comprises the steps of caching, sorting, UDP voice packet number moving average in a cache pool, rate adjustment state circulation, a rate dynamic adjustment mechanism, frame synchronization, RS decoding and the like, the problems of delay, packet interval jitter and the like can be effectively relieved after preprocessing, the overall balance and stability of the voice code rate are guaranteed, and the voice quality of a receiving end is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voice communication, and particularly to a method for improving the voice quality at the receiving end of a voice communication system. Background Art

[0002] The statements in this section only provide background information related to the present disclosure and may not constitute prior art.

[0003] In a typical voice communication system, its transmission path usually includes heterogeneous networks, such as Figure 1 As shown, after the audio at the sending end undergoes analog-to-digital conversion, audio encoding, and modulation, it is sent to the receiving system through a wireless channel.

[0004] The receiving system demodulates the received data, then encapsulates the baseband signal into UDP voice packets (the packet interval time is determined by the packet length and the encoding rate at the sending end), and forwards them to the receiving end through a computer network and a gateway. After processing the voice packets at the receiving end, they are sent to a decoding chip, and after decoding, the audio is output through digital-to-analog conversion.

[0005] During the voice communication process, the voice signal is forwarded through a wireless channel, a computer network, and a gateway. Due to the diversity and complexity of the transmission path, as well as the unreliable characteristics of UDP being connectionless, the voice packets received at the receiving end may have many problems such as delay, out-of-order, error code, packet loss, packet interval jitter, imbalance, and instability. Summary of the Invention

[0006] The purpose of the present invention is: aiming at the problems that after the voice signal in the current voice communication system is transmitted through a heterogeneous network and reaches the receiving end, there are often many problems such as delay, out-of-order, error code, packet loss, packet interval jitter, imbalance, and instability, a method for improving the voice quality at the receiving end of a voice communication system is provided. In terms of protocol design, the method of frame synchronization combined with RS coding is adopted, which can reduce the impact of unstable factors such as out-of-order, packet loss, and burst error codes introduced in the transmission path on the voice, and ensure the integrity and continuity of the voice frames. In terms of implementation, as Figure 1 shown, the sending end designs an after-coding processing module, which can attach a frame synchronization header (hereinafter referred to as "frame header") and RS code to the voice frames. The receiving end designs a UDP voice packet processing module and a decoding driver module, and performs sufficient preprocessing before sending the UDP voice packets to the decoding chip. The processing methods include caching, sorting, moving average of the number of UDP voice packets in the cache pool (hereinafter referred to as "cache number"), rate adjustment state transition (including smooth processing of state commutation), rate dynamic adjustment mechanism, frame synchronization, RS decoding, etc. After preprocessing, problems such as delay and packet interval jitter can be effectively alleviated, the overall balance and stability of the voice code rate are ensured, and the voice quality at the receiving end is improved.

[0007] The technical solution of the present invention is as follows:

[0008] A method for improving the voice quality at the receiving end of a voice communication system, comprising:

[0009] Construct an encoded post-processing module at the sending end to append a frame synchronization header and RS code to the voice frame;

[0010] Construct a UDP voice packet processing module and a decoding driver module at the receiving end, and perform sufficient preprocessing before the UDP voice packet is sent to the decoding chip.

[0011] Further, the processing flow of the encoded post-processing module is as follows:

[0012] Step A: Obtain the voice data output by the encoding chip, with each V bytes forming a frame. Append an H-byte frame header before the voice frame, and reserve a T-byte RS coding result area at the end of the voice frame;

[0013] Step B: Adopt an RS(2 m-1 , k) coding method based on the GF(2m) domain to perform forward error correction coding on the voice frame, and fill the result into the reserved RS coding result area; where m represents the number of bits contained in the coding information unit, 2 m-1 represents the number of all information units in the RS coding group, k represents the number of valid information units in the group, and (2 m-1 -k) / 2 represents the number of information units that can be corrected in the group.

[0014] Further, the UDP voice packet processing module includes:

[0015] 3 inputs and 2 outputs; used to receive the UDP voice packet from the previous stage, the enable signal and frame synchronization indication from the subsequent stage, and output a rate adjustment signal and a buffered and sorted voice packet to the subsequent stage; the UDP voice packet processing module is implemented by a CPU, and the CPU processing software includes three threads, including: Thread One, Thread Two, and Thread Three, and the three threads work simultaneously.

[0016] Further, the processing flow of the Thread One is as follows:

[0017] Step S1: Receive the UDP voice packet on the network. Each UDP voice packet contains its own packet sequence number; the number of voice frames in the packet is any value; the voice frame data is a bit stream;

[0018] Step S2: Insert the received UDP voice packet into the buffer pool and sort it according to its own packet sequence number;

[0019] Step S3: Increment the buffer count by 1.

[0020] Further, the processing flow of the Thread Two is as follows:

[0021] Step a: When processing data for the first time, the cache pool is filled with a set number of UDP voice packets; if the "first processing flag" is 1, that is, processing UDP voice packets for the first time, then enter Step b, otherwise enter Step c;

[0022] Step b: If the cache count is not greater than the set value, it means that the set number of UDP voice packets has not been filled, go back to Step a, otherwise, set the "first processing flag" to 0 and enter Step c;

[0023] Step c: If the cache count is greater than 0, record the current time T0 and enter Step d, otherwise enter Step f;

[0024] Step d: If the enable signal of the decoding driver module is not 1, enter Step e, otherwise send data to the decoding driver module, decrement the cache count by 1, and go back to Step a;

[0025] Step e: Delay for x microseconds, record the current time T1. If T1 - T0 is greater than the set value, it is determined that a timeout has occurred and go back to Step a, otherwise go back to Step d;

[0026] Step f: Read the frame synchronization indication from the decoding driver module. If synchronized, go back to Step a, otherwise it means that the cache pool has entered the "read empty" state, set the "first processing flag" to 0, and go back to Step a. At this time, the cache pool will be filled with a set number of UDP voice packets again before starting the subsequent processing flow.

[0027] Further, the processing flow of Thread 3 is as follows:

[0028] Step Ⅰ: Set the rate initial state to "normal state", and design the low value and high value of rate adjustment according to the mapping relationship between the set value of the cache count and the high and low values of rate adjustment;

[0029] Step Ⅱ: Delay for x seconds and obtain the current cache count in real time;

[0030] Step Ⅲ: Insert the obtained cache count into the tail of the queue, and increment the number of data in the queue by 1;

[0031] Step Ⅳ: If the number of data in the queue is not equal to N, go back to Step Ⅱ, otherwise calculate the mean MEAN of all data in the queue and enter Step Ⅴ;

[0032] Step Ⅴ: Remove the data at the head of the queue, and decrement the number of data in the queue by 1;

[0033] Step Ⅵ: Start to adjust the rate. The rate adjustment adopts a smooth state transition processing mechanism, and a "transition zone" is added during the transition. When the MEAN value has repeated positive or negative deviations from the target value due to disturbances, the rate adjustment state machine will not frequently switch between the corresponding two states.

[0034] Further, step VI includes:

[0035] Step VI1: Obtain the current rate status. If it is in the "normal state", go to step VI2; if it is in the "slow state", go to step VI3; if it is in the "fast state", go to step VI4.

[0036] The processing steps for the "normal state" in step VI2 are as follows:

[0037] Step VI21: If low ≤ MEAN ≤ high, maintain the status and return to step II.

[0038] Step VI22: If MEAN > high, send a "fast" instruction to the decoding and driving module through the "rate adjustment signal" interface, enter the "fast state", and return to step II.

[0039] Step VI23: If MEAN < low, send a "slow" instruction to the decoding and driving module through the "rate adjustment signal" interface, the rate status enters the "slow state", and return to step II.

[0040] The processing steps for the "slow state" in step VI3 are as follows:

[0041] Step VI31: If MEAN < target, maintain the status and return to step II.

[0042] Step VI32: If MEAN > high, send a "fast" instruction to the decoding and driving module through the "rate adjustment signal" interface, the rate status enters the "fast state", and return to step II.

[0043] Step VI33: If target ≤ MEAN ≤ high, send a "normal" instruction to the decoding and driving module through the "rate adjustment signal" interface, the rate status enters the "normal state", and return to step II.

[0044] The processing steps for the "fast state" in step VI4 are as follows:

[0045] Step VI41: If MEAN ≥ target, maintain the status and return to step II.

[0046] Step VI42: If MEAN < low, send a "slow" instruction to the decoding and driving module through the "rate adjustment signal" interface, the rate status enters the "slow state", and return to step II.

[0047] Step VI43: If target > MEAN ≥ low, send a "normal" instruction to the decoding and driving module through the "rate adjustment signal" interface, the rate status enters the "normal state", and return to step II.

[0048] Further, the decoding and driving module includes: 2 inputs and 3 outputs; it is used to receive the rate adjustment signal and voice packet from the previous stage, output an enable signal and a frame synchronization indication to the previous stage, and output triple frame synchronization and the voice frame after RS decoding to the next stage; the decoding and driving module is implemented by an FPGA with pure parallel processing.

[0049] Further, the parallel processing includes:

[0050] Step ①: Initialize the power-on parameters, with the enable signal output being 1 and the frame synchronization indication output being 0;

[0051] Step ②: Receive the voice packet sent by the UDP voice packet processing module and store it in a FIFO with 8-bit input and 1-bit output;

[0052] Step ③: If the quantity in the FIFO is 0, the frame synchronization indication output is 0; otherwise, when the quantity in the FIFO is greater than 2 / 3 of the capacity, the enable signal output is 0, and when it is less than 1 / 3 of the capacity, the enable signal output is 1;

[0053] Step ④: When there is data in the FIFO, read the data from the FIFO and insert it into the shift register, detect the voice data bit by bit. After receiving three complete frames starting with the frame header and having the frame length continuously, the synchronization is successful and the frame synchronization indication output is 1; after successful synchronization, the frame header search is no longer performed bit by bit, but directly jump to the position of the next frame header according to the frame length for comparison; if the frame header is not found at the corresponding frame header position three times or no data is received within the timeout, it is considered out of sync and the frame synchronization indication output is 0;

[0054] Step ⑤: Perform RS(2 m-1 , k) decoding on the synchronized voice frame;

[0055] Step ⑥: Read the input of the rate adjustment signal and select the processing rate of the voice frame. The rate selection is jointly determined by the coding rate and the dynamic range of the voice code rate input of the decoding chip;

[0056] Step ⑦: Send the decoded voice frame to the decoding chip.

[0057] Further, Step ⑥ includes:

[0058] Step ⑥1: When the input of the rate adjustment signal is in the normal state, select the rate median, and the rate of taking data from the FIFO is moderate;

[0059] Step ⑥2: When the input of the rate adjustment signal is in the fast state, select the rate upper limit, and the rate of taking data from the FIFO is slightly faster, which makes the data volume in the FIFO decrease slowly, shortens the interval time of the enable signal from output 0 to output 1, and can make the cache number in the UDP voice packet processing module decrease slowly;

[0060] Step ⑥3: When the rate adjustment signal input is in the slow state, select the rate lower limit. The rate of fetching data from the FIFO is slightly slower, causing the data volume in the FIFO to increase slowly. This shortens the interval time of the enable signal from output 1 to output 0, and enables the buffer count in the UDP voice packet processing module to increase slowly.

[0061] Compared with the existing technologies, the beneficial effects of the present invention are as follows:

[0062] 1. The present invention uses RS coding for forward error correction. Within the allowable error correction range, it can effectively avoid the burst interference introduced on the heterogeneous network transmission path. A frame header is added to the voice frame header, and a three-frame synchronization method is introduced. This method retrieves the frame header bit by bit, bringing great convenience to the baseband processing of the upper-level receiving end. It is not necessary to consider bit synchronization (the frame header is byte-aligned), and only the demodulated bit stream needs to be directly assembled into UDP voice packets of any byte length. The three-frame synchronization can tolerate no more than 2 consecutive frame header bit errors in the voice frame. Even if there are packet losses in the UDP transmission network, invalid voice data can be eliminated, ensuring the integrity and continuity of the overall voice frame.

[0063] 2. Inside the UDP voice packet processing module of the present invention, a cache pool is used to cache and sort voice packets. When the cache pool is in the "empty" state, it will be filled with a set number of voice packets, which can isolate the packet interval jitter introduced on the transmission path and ensure the smoothness of the subsequent voice frame processing. A transmission rate adjustment mechanism is designed to perform a moving average on the cache count. According to the relationship between the mean value and the set cache count, the processing rate of the lower-level module (decoding driver module) is adjusted through the three-state transition of "normal state", "fast state" and "slow state" (an additional "transition zone" is added during the state change to achieve smooth state switching), so that the cache count is maintained near the set value, further ensuring the overall balance and stability of the voice frame code rate and improving the voice quality of the receiving end. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 It is a schematic diagram of a voice communication system including a heterogeneous network;

[0065] Figure 2 It is a flow chart of post-processing after coding;

[0066] Figure 3 It is a flow chart of voice frame transmission;

[0067] Figure 4 It is a block diagram of the UDP voice packet processing and decoding driver module;

[0068] Figure 5 It is a flow chart of the first thread processing of the UDP voice packet processing module;

[0069] Figure 6It is the processing flow chart of Thread 2 of the UDP voice packet processing module;

[0070] Figure 7 It is the processing flow chart of Thread 3 of the UDP voice packet processing module;

[0071] Figure 8 It is the schematic diagram of the decoding driver module processing. Specific implementation mode

[0072] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0073] The features and performance of the present invention will be further described in detail below in conjunction with the embodiments.

[0074] Embodiment 1

[0075] Please refer to Figure 1 , a method for improving the voice quality of the receiving end of a voice communication system, specifically including:

[0076] Construct an encoding post-processing module at the sending end, and append a frame synchronization header and RS code to the voice frame;

[0077] Construct a UDP voice packet processing module and a decoding driver module at the receiving end, and perform sufficient preprocessing before the UDP voice packet is sent to the decoding chip.

[0078] In this embodiment, specifically, as Figure 2 shown, the processing flow of the encoding post-processing module is as follows:

[0079] Step A: Obtain the voice data output by the encoding chip, form a frame with every V bytes, append an H-byte frame header before the voice frame, and reserve a T-byte RS coding result area at the end of the voice frame; as Figure 3 shown, particularly: take the frame header as 0xEB90, H = 2;

[0080] Step B: Adopt RS(2 based on the GF(2m) domain m-1, k) Encoding method, performing forward error correction encoding on voice frames and filling the results into the reserved RS encoding result area; where m represents the number of bits contained in the encoding information unit, 2 m-1 represents the total number of information units in the RS encoding group, k represents the number of valid information units in the group, and (2 m-1 -k) / 2 represents the number of correctable information units in the group; specifically: take m = 8, k = 223, an information unit contains 8-bit data, that is, encoding is processed in bytes, the group length is 255 bytes, of which the valid data is 223 bytes and the additional redundancy is 32 bytes. Any 16 bytes in the 255-byte group length can be corrected by RS decoding. Correspondingly, V = 223 and T = 32 in step A.

[0081] In this embodiment, specifically, as Figure 4 shown, the UDP voice packet processing module includes:

[0082] 3 inputs and 2 outputs; used to receive the UDP voice packets from the previous stage, the enable signal and frame synchronization indication from the subsequent stage, and output the rate adjustment signal and the buffered and sorted voice packets to the subsequent stage; the UDP voice packet processing module is implemented by the CPU, and the CPU processing software includes three threads, including: thread one, thread two, and thread three, and the three threads work simultaneously.

[0083] In this embodiment, specifically, as Figure 5 shown, the processing flow of the thread one is as follows:

[0084] Step S1: Receive the UDP voice packets on the network, as Figure 3 shown, each UDP voice packet contains its own packet sequence number (added after demodulation to the baseband); the number of voice frames in the packet can be any value, even a decimal (voice frames are transmitted in UDP segments); the voice frame data is a bit stream and does not need to be byte-aligned (the frame header will only appear after bit shifting of the data).

[0085] Step S2: Insert the received UDP voice packets into the buffer pool and sort them according to their own packet sequence numbers;

[0086] Step S3: Increment the buffer count by 1.

[0087] In this embodiment, specifically, as Figure 6 shown, the processing flow of the thread two is as follows:

[0088] Step a: When processing data for the first time, the buffer pool will be filled with a set number of UDP voice packets, which can avoid the problem of uneven packet interval jitter on the transmission path and ensure the smoothness of subsequent voice frame processing; if the "first processing flag" is 1, that is, processing UDP voice packets for the first time, then enter step b, otherwise enter step c;

[0089] Step b: If the cache count is not greater than the set value, it indicates that the set number of UDP voice packets has not been filled. Return to step a. Otherwise, set the "first processing flag" to 0 and enter step c;

[0090] Step c: If the cache count is greater than 0, record the current time T0 and enter step d. Otherwise, enter step f;

[0091] Step d: If the enable signal of the lower module (decoding and driving module) is not 1, enter step e. Otherwise, send data to the decoding and driving module, decrement the cache count by 1, and return to step a;

[0092] Step e: Delay for x microseconds, record the current time T1. If T1 - T0 is greater than the set value, it is determined that a timeout has occurred and return to step a. Otherwise, return to step d; Specifically, x = 1000 microseconds and the set value is 1 second;

[0093] Step f: Read the frame synchronization indication from the lower module (decoding and driving module). If synchronized, return to step a. Otherwise, it indicates that the cache pool has entered the "read empty" state. Set the "first processing flag" to 0 and return to step a. At this time, the cache pool will start the subsequent processing flow only after being filled with the set number of UDP voice packets again.

[0094] In this embodiment, specifically, as Figure 7 shown, the processing flow of the third thread is as follows:

[0095] Step Ⅰ: Set the rate initial state to "normal state", and design the low value and high value of rate adjustment according to the mapping relationship between the cache count set value and the high and low values of rate adjustment (specifically shown in Table 1);

[0096] Table 1 Mapping relationship between cache count set value and high and low values of rate adjustment

[0097]

[0098] Step Ⅱ: Delay for x seconds and obtain the current cache count in real time; Specifically, x = 1 second;

[0099] Step Ⅲ: Insert the obtained cache count into the tail of the queue, and increment the number of data in the queue by 1;

[0100] Step Ⅳ: If the number of data in the queue is not equal to N, return to step Ⅱ. Otherwise, calculate the mean MEAN of all data in the queue and enter step Ⅴ; Specifically, N = 10;

[0101] Step Ⅴ: Remove the data at the head of the queue, and decrement the number of data in the queue by 1;

[0102] Step VI: Start to adjust the rate. The rate adjustment adopts a stable state commutation processing mechanism, adding a "transition zone" during commutation. When the MEAN value has repeated positive or negative deviations from the target value due to disturbances, the rate adjustment state machine will not frequently switch between the corresponding two states.

[0103] In this embodiment, specifically, the said Step VI includes:

[0104] Step VI1: Obtain the current rate state. If it is the "normal state", enter Step VI2; if it is the "slow state", enter Step VI3; if it is the "fast state", enter Step VI4.

[0105] The processing steps for the "normal state" in Step VI2 are as follows:

[0106] Step VI21: If low ≤ MEAN ≤ high, maintain the state and return to Step II.

[0107] Step VI22: If MEAN > high, send a "fast" instruction to the lower-level module (decoding and driving module) through the "rate adjustment signal" interface, enter the "fast state", and return to Step II. Example of the current commutation processing mechanism: If there is a minimal disturbance ε in the next-round MEAN value, MEAN = high - ε, due to the buffering of the transition zone [target, high), the state switching condition of Step VI43 is not met, and the rate state will not return from the "fast state" to the "normal state" again. Similarly, transition zones are designed for state switching in subsequent steps.

[0108] Step VI23: If MEAN < low, send a "slow" instruction to the lower-level module (decoding and driving module) through the "rate adjustment signal" interface, and the rate state enters the "slow state" and returns to Step II.

[0109] The processing steps for the "slow state" in Step VI3 are as follows:

[0110] Step VI31: If MEAN < target, maintain the state and return to Step II.

[0111] Step VI32: If MEAN > high, send a "fast" instruction to the decoding and driving module through the "rate adjustment signal" interface, and the rate state enters the "fast state" and returns to Step II.

[0112] Step VI33: If target ≤ MEAN ≤ high, send a "normal" instruction to the decoding and driving module through the "rate adjustment signal" interface, and the rate state enters the "normal state" and returns to Step II.

[0113] The processing steps for the "fast state" in Step VI4 are as follows:

[0114] Step Ⅵ41: If MEAN≥target, maintain the status and return to Step Ⅱ;

[0115] Step Ⅵ42: If MEAN<low, send a "slow" instruction to the decoding and driving module through the "rate adjustment signal" interface, the rate status enters the "slow state", and return to Step Ⅱ;

[0116] Step Ⅵ43: If target>MEAN≥low, send a "normal" instruction to the decoding and driving module through the "rate adjustment signal" interface, the rate status enters the "normal state", and return to Step Ⅱ.

[0117] In this embodiment, specifically, as Figure 4 shown, the decoding and driving module includes: 2 inputs and 3 outputs; it is used to receive the rate adjustment signal and voice packet from the previous stage, output an enable signal and a frame synchronization indication to the previous stage, and output three-frame synchronization and the voice frame after RS decoding to the next stage; the decoding and driving module is implemented by FPGA and performs pure parallel processing.

[0118] In this embodiment, specifically, as Figure 8 shown, the parallel processing includes:

[0119] Step ①: Initialize the power-on parameters, the enable signal output is 1, and the frame synchronization indication output is 0;

[0120] Step ②: Receive the voice packet sent by the upper module (UDP voice packet processing module) and store it in the FIFO with 8-bit input and 1-bit output;

[0121] Step ③: If the number in the FIFO is 0, the frame synchronization indication output is 0; otherwise, when the number in the FIFO is greater than 2 / 3 of the capacity, the enable signal output is 0, and when it is less than 1 / 3 of the capacity, the enable signal output is 1;

[0122] Step ④: When there is data in the FIFO, read the data from the FIFO and insert it into the shift register, detect the voice data bit by bit. After receiving three complete frames starting with the frame header and having the frame length continuously, the synchronization is successful and the frame synchronization indication output is 1; after successful synchronization, the frame header search is no longer performed bit by bit, but directly jump to the position of the next frame header according to the frame length for comparison; if the frame header is not found at the corresponding frame header position three times or no data is received within the timeout, it is considered out of sync and the frame synchronization indication output is 0. For the detailed state transition, see Figure 8 ;

[0123] Step ⑤: Perform RS(2 m-1 , k) decoding on the synchronized voice frame; specifically, corresponding to the encoding end, m = 8;

[0124] Step ⑥: Read the input of the rate adjustment signal, and select the processing rate of the voice frame. The rate selection is jointly determined by the coding rate and the dynamic range of the voice code rate input of the decoding chip;

[0125] Step ⑦: Send the decoded voice frame to the decoding chip.

[0126] In this embodiment, specifically, Step ⑥ includes:

[0127] Step ⑥1: When the input of the rate adjustment signal is in the normal state, select the median rate, and the rate of fetching data from the FIFO is moderate;

[0128] Step ⑥2: When the input of the rate adjustment signal is in the fast state, select the upper limit of the rate. The rate of fetching data from the FIFO is slightly faster, which makes the data volume in the FIFO decrease slowly, shortens the interval time of the enable signal from output 0 to output 1, and can make the cache number in the UDP voice packet processing module decrease slowly;

[0129] Step ⑥3: When the input of the rate adjustment signal is in the slow state, select the lower limit of the rate. The rate of fetching data from the FIFO is slightly slower, which makes the data volume in the FIFO increase slowly, shortens the interval time of the enable signal from output 1 to output 0, and can make the cache number in the UDP voice packet processing module increase slowly;

[0130] By processing in this way and combining with the upper-level module (UDP voice packet processing module), the joint control of the overall voice code rate at the receiving end is realized, which can keep the code rate balanced and stable as a whole, and improve the voice quality at the receiving end.

[0131] The above-described embodiments only represent the specific implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the protection scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the technical solution of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application.

[0132] This background technology section is provided to generally present the context of the present invention. The work of the currently named inventors, to the extent described in this background technology section, and aspects of the work that are not prior art at the time of filing this application are neither expressly nor impliedly admitted to be prior art of the present invention.

Claims

1. A method for improving the voice quality of a receiving end of a voice communication system, characterized in that: include: Construct a post-coding processing module at the transmitting end to add a frame synchronization header and RS code to the voice frame; A UDP voice packet processing module and a decoding driver module are constructed at the receiving end to perform sufficient preprocessing before the UDP voice packet is sent to the decoding chip.

2. A method for improving the voice quality of a receiving end of a voice communication system according to claim 1, characterized in that: The processing flow of the post-encoding processing module is as follows: Step A: Acquire the voice data output by the encoding chip, each V bytes constitute a frame, add an H-byte frame header before the voice frame, and reserve a T-byte RS encoding result area at the end of the voice frame; Step B: Using RS(2 m-1 , k) Coding method, forward error correction coding is performed on the voice frame, and the result is filled into the reserved RS coding result area; where m represents the number of bits contained in the coding information unit, 2 m-1 represents the number of all information units in the RS coding group, k represents the number of valid information units in the group, (2 m-1 -k) / 2 represents the number of correctable information units in the group.

3. A method for improving the voice quality of a receiving end of a voice communication system according to claim 1, characterized in that: The UDP voice packet processing module comprises: 3 inputs and 2 outputs; used to receive UDP voice packets from the previous stage, enable signals and frame synchronization indications from the next stage, and output rate adjustment signals and cached sorted voice packets to the next stage; the UDP voice packet processing module is implemented by the CPU, and the CPU processing software contains three threads, including: thread one, thread two, and thread three, and the three threads work simultaneously.

4. A method for improving the voice quality of a receiving end of a voice communication system according to claim 3, characterized in that: The processing flow of thread one is as follows: Step S1: receiving a UDP voice packet on the network, each UDP voice packet contains its own packet sequence number; the number of voice frames in the packet is an arbitrary value; the voice frame data is a bit stream; Step S2: insert the received UDP voice packets into the buffer pool and sort them according to their own packet sequence numbers; Step S3: The cache number is increased by 1.

5. A method for improving the voice quality of a receiving end of a voice communication system according to claim 4, characterized in that: The processing flow of thread 2 is as follows: Step a: When processing data for the first time, the buffer pool will be filled with a set number of UDP voice packets; if the "first processing flag" is 1, that is, the UDP voice packet is processed for the first time, then go to step b, otherwise go to step c; Step b: If the buffer number is not greater than the set value, it means that the set number of UDP voice packets is not fully stored, and the process returns to step a. Otherwise, the "first processing flag" is set to 0, and the process goes to step c. Step c: If the cache number is greater than 0, record the current time T0 and proceed to step d, otherwise proceed to step f; Step d: If the enable signal of the decoding driver module is not 1, go to step e, otherwise send data to the decoding driver module, reduce the cache number by 1, and return to step a; Step e: Delay x microseconds, record the current time T1, if T1-T0 is greater than the set value, determine the timeout, return to step a, otherwise return to step d; Step f: Read the frame synchronization indication from the decoding driver module. If it is synchronized, return to step a. Otherwise, it means that the buffer pool has entered the "read empty" state. Set the "first processing flag" to 0 and return to step a. At this time, the buffer pool will be filled with the set number of UDP voice packets again before starting the subsequent processing flow.

6. A method for improving the voice quality of a receiving end of a voice communication system according to claim 5, characterized in that: The processing flow of thread three is as follows: Step I: Set the initial state of the rate to "normal state", and design the low and high values ​​of the rate adjustment according to the mapping relationship between the cache number setting value and the high and low values ​​of the rate adjustment; Step II: Delay x seconds and obtain the current cache count in real time; Step III: Insert the acquired cache number into the tail of the queue, and increase the number of data in the queue by 1; Step IV: If the number of data in the queue is not equal to N, return to step II, otherwise calculate the mean MEAN of all data in the queue and go to step V; Step V: Remove the queue header data and reduce the number of data in the queue by 1; Step VI: Start to adjust the rate. The rate adjustment adopts a stable state commutation processing mechanism, adding a "transition zone" during commutation. When the MEAN value has repeated positive or negative deviations from the target value due to disturbances, the rate adjustment state machine will not frequently switch between the corresponding two states.

7. A method for improving the voice quality of a receiving end of a voice communication system according to claim 6, characterized in that: The said Step VI includes: Step VI1: Obtain the current rate state. If it is in the "normal state", enter Step VI2; if it is in the "slow state", enter Step VI3; if it is in the "fast state", enter Step VI4. The processing steps in the "normal state" of Step VI2 are as follows: Step VI21: If low ≤ MEAN ≤ high, maintain the state and return to Step II. Step VI22: If MEAN > high, send a "fast" instruction to the decoding and driving module through the "rate adjustment signal" interface, enter the "fast state", and return to Step II. Step VI23: If MEAN < low, send a "slow" instruction to the decoding and driving module through the "rate adjustment signal" interface, and the rate state enters the "slow state", then return to Step II. The processing steps in the "slow state" of Step VI3 are as follows: Step VI31: If MEAN < target, maintain the state and return to Step II. Step VI32: If MEAN > high, send a "fast" instruction to the decoding and driving module through the "rate adjustment signal" interface, and the rate state enters the "fast state", then return to Step II. Step VI33: If target ≤ MEAN ≤ high, send a "normal" instruction to the decoding and driving module through the "rate adjustment signal" interface, and the rate state enters the "normal state", then return to Step II. The processing steps in the "fast state" of Step VI4 are as follows: Step VI41: If MEAN ≥ target, maintain the state and return to Step II. Step VI42: If MEAN < low, send a "slow" instruction to the decoding and driving module through the "rate adjustment signal" interface, and the rate state enters the "slow state", then return to Step II. Step VI43: If target > MEAN ≥ low, send a "normal" instruction to the decoding and driving module through the "rate adjustment signal" interface, and the rate state enters the "normal state", then return to Step II.

8. A method for improving the voice quality of a receiving end of a voice communication system according to claim 1, characterized in that: The said decoding and driving module includes: 2 inputs and 3 outputs; it is used to receive the rate adjustment signal and voice packet from the previous stage, output the enable signal and frame synchronization indication to the previous stage, and output three-frame synchronization and the voice frame after RS decoding to the next stage; the decoding and driving module is implemented by FPGA with pure parallel processing.

9. A method for improving the voice quality of a receiving end of a voice communication system according to claim 8, characterized in that: The said parallel processing includes: Step ①: Initialize the power-on parameters, with the enable signal output being 1 and the frame synchronization indication output being 0. Step ②: Receive the voice packet sent by the UDP voice packet processing module and store it in the FIFO with 8-bit input and 1-bit output. Step ③: If the quantity in the FIFO is 0, the frame synchronization indication output is 0; otherwise, when the quantity in the FIFO is greater than 2 / 3 of the capacity, the enable signal output is 0, and when it is less than 1 / 3 of the capacity, the enable signal output is 1. Step ④: When there is data in FIFO, read the data from FIFO and insert it into the shift register, detect the voice data bit by bit, and only after receiving three consecutive complete frames starting with the frame header and with a length equal to the frame length, the synchronization is successful, and the frame synchronization indication output is 1; after the synchronization is successful, the frame header is no longer searched by bit, but directly jumps to the position where the next frame header is located for comparison according to the frame length; if the frame header is not found at the corresponding frame header position three times or the data is not received within the timeout period, it is considered to be out of step, and the frame synchronization indication output is 0; Step ⑤: Perform RS(2) on the synchronized voice frames m-1 , k) decoding; Step ⑥: Read the rate adjustment signal input and select the processing rate of the voice frame. The rate selection is determined by the encoding rate and the dynamic range of the voice bit rate input of the decoding chip; Step 7: Send the decoded voice frame to the decoding chip.

10. A method for improving the voice quality of a receiving end of a voice communication system according to claim 9, characterized in that: The step ⑥ comprises: Step ⑥1: When the rate adjustment signal input is in a normal state, the median rate is selected, and the rate of taking data from the FIFO is moderate; Step ⑥2: When the rate adjustment signal input is in the fast state, select the rate upper limit, and the rate of taking data from the FIFO is slightly faster, so that the amount of data in the FIFO decreases slowly, which will shorten the interval time from output 0 to output 1 of the enable signal, and the number of buffers in the UDP voice packet processing module can be slowly reduced; Step ⑥3: When the rate adjustment signal input is in a slow state, select the lower limit of the rate, and the rate of taking data from the FIFO is slightly slower, so that the amount of data in the FIFO increases slowly, which will shorten the interval time from output 1 to output 0 of the enable signal, and the number of buffers in the UDP voice packet processing module can increase slowly.