Method and system for processing out-of-sync noise in digitally encoded speech for pcm
Patent Information
- Application Number
- CN202410036679.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-10
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2044-01-10
AI Technical Summary
[0004]本发明旨在至少解决现有技术中存在的失步噪音完全消除困难、不能及时判断出话音失步、无法提供失步噪音反馈的技术问题之一
[0034] This invention can quickly detect and eliminate out-of-step noise within 20ms, while also providing an accurate out-of-step noise feedback mechanism. It can be applied to various types of digital voice communication systems. It solves the problems of current noise processing methods, such as the difficulty in completely eliminating out-of-step noise, the inability to promptly detect voice out-of-step, and the lack of feedback on out-of-step noise.
Smart Images

Figure CN117789733B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech out-of-sync noise processing technology, and more specifically, to a method and system for processing digital speech out-of-sync noise for PCM encoding. Background Technology
[0002] Currently, the general method for handling out-of-synchronization noise in voice communication is to convert the analog voice signal into a digital signal and use digital filtering algorithms on platforms such as processors and FPGAs to eliminate the out-of-synchronization noise.
[0003] like Figure 1 As shown, current methods for handling out-of-synchronization noise directly process the noise signal. While this can attenuate and eliminate some of the noise, completely eliminating it is extremely difficult. Furthermore, these methods only have the function of eliminating out-of-synchronization noise and lack a mechanism for detecting it, thus failing to provide feedback. When a large amount of noise is generated in voice communication due to out-of-synchronization, the aforementioned noise processing methods are unable to completely eliminate the noise and cannot determine whether out-of-synchronization has occurred, thus failing to provide feedback on the noise level. Summary of the Invention
[0004] The present invention aims to at least solve one of the technical problems existing in the prior art, namely, the difficulty in completely eliminating out-of-step noise, the inability to promptly detect out-of-step speech, and the inability to provide feedback on out-of-step noise.
[0005] Therefore, the first aspect of the present invention provides a method for processing out-of-sync noise in digital speech encoded in PCM.
[0006] A second aspect of the present invention provides a digital speech out-of-sync noise processing system for PCM encoding.
[0007] This invention provides a method for processing out-of-sync noise in digital speech encoded with PCM, comprising:
[0008] Convert speech signals into linear PCM data;
[0009] The audio amplitude of each sampling point is monitored in each detection cycle of the speech signal, and the sampling point whose audio amplitude exceeds a first set threshold is defined as a high audio amplitude point. There are multiple sampling points in each detection cycle.
[0010] Based on the number of high-frequency amplitude points within the detection period, it is determined whether there is speech sync noise within the detection period, and a silence flag is used to mark whether there is speech sync noise within the detection period.
[0011] The marked speech signal is delayed, and the speech signal after delay is switched according to the silence flag. When the silence flag indicates that there is speech out-of-step noise in the detection period, the speech signal in the detection period is muted. When the silence flag indicates that there is no speech out-of-step noise in the detection period, the speech signal in the detection period is output.
[0012] The digital speech out-of-sync noise processing method for PCM encoding according to the above-described technical solution of the present invention may further have the following additional technical features:
[0013] In the above technical solution, the conversion of the speech signal into linear PCM data includes:
[0014] The speech signal at each sampling point is converted into 16-bit linear PCM data;
[0015] The sign bit of the 16-bit linear PCM data is removed, and after two's complement conversion, 15-bit data is generated. The data volume PCM_VAL of this data is defined as the audio amplitude of the corresponding sampling point.
[0016] In the above technical solution, defining the sampling point whose audio amplitude exceeds the first set threshold as a high-frequency amplitude point includes:
[0017] Sampling points with PCM_VAL greater than or equal to 0x2800 are defined as high-frequency amplitude points.
[0018] In the above technical solution, the step of determining whether there is speech sync noise in the detection period based on the number of high-frequency amplitude points within the detection period, and marking the presence of speech sync noise in the detection period using a silence flag, includes:
[0019] When the number of high-frequency amplitude points in the detection period exceeds the second set threshold, the mute flag is set to 1 to indicate that there is speech out-of-sync noise in the detection period.
[0020] When speech out-of-sync noise is detected, the conditions for determining speech recovery include: there are no high-frequency amplitude points in N consecutive detection cycles, and the mute flag is set to 0.
[0021] In the above technical solution, the delay time for the marked voice signal is not less than the detection period.
[0022] In the above technical solution, the delay time for the marked voice signal is twice the length of the detection period.
[0023] The above technical solution also includes:
[0024] If the number of high-frequency amplitude points exceeds the third set threshold within M consecutive detection cycles, a step-out noise notification will be generated, and the voice call process will be restarted.
[0025] This invention provides a digital speech out-of-sync noise processing system for PCM encoding, comprising:
[0026] The PCM codec chip receives voice data and converts it into linear PCM data.
[0027] The out-of-step noise detection module is connected to the PCM codec chip. It monitors the audio amplitude of each sampling point in each detection cycle and defines the sampling point whose audio amplitude exceeds the first set threshold as a high-frequency amplitude point. Based on the number of high-frequency amplitude points in the detection cycle, it determines whether there is speech out-of-step noise in the detection cycle. It also uses a silence flag to mark whether there is speech out-of-step noise in the detection cycle.
[0028] The delay module, connected to the step loss noise detection module, delays the marked speech signal.
[0029] The voice switching module, connected to the delay module, switches the call line to mute or outputs voice within the detection period that is free of voice synchronization noise, based on the mute flag of the detection period.
[0030] In the above technical solution, in the step loss noise detection module, when the number of high-frequency amplitude points in the detection period exceeds the second set threshold, the silence flag is set to 1 to indicate that there is speech step loss noise in the detection period; when speech step loss noise is determined to occur, the conditions for determining speech recovery include: there are no high-frequency amplitude points in N consecutive detection periods, and the silence flag is set to 0.
[0031] In the above technical solution, the digital speech out-of-sync noise processing system for PCM encoding also includes:
[0032] The out-of-step feedback algorithm decision module is connected to the out-of-step noise detection module. When the number of high-frequency amplitude points exceeds a third set threshold within M consecutive detection cycles, it generates an out-of-step noise notification and sends the out-of-step noise notification to the system processor to restart the voice call process.
[0033] In summary, due to the adoption of the above-mentioned technical features, the beneficial effects of the present invention are:
[0034] This invention can quickly detect and eliminate out-of-step noise within 20ms, while also providing an accurate out-of-step noise feedback mechanism. It can be applied to various types of digital voice communication systems. It solves the problems of current noise processing methods, such as the difficulty in completely eliminating out-of-step noise, the inability to promptly detect voice out-of-step, and the lack of feedback on out-of-step noise.
[0035] Specifically, this invention can not only quickly detect and eliminate out-of-step noise, preventing users from hearing harsh noise and improving the user experience, but also has an accurate out-of-step noise feedback mechanism that can promptly notify the voice terminal processing unit to restart the call process and restore normal voice in a timely manner.
[0036] Additional aspects and advantages of the invention will become apparent in the following description or may be learned by practice of the invention. Attached Figure Description
[0037] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:
[0038] Figure 1 This is a schematic diagram of traditional step loss noise processing technology;
[0039] Figure 2 This is a flowchart of a method for processing out-of-sync noise in digital speech encoded with PCM according to an embodiment of the present invention;
[0040] Figure 3 This is a schematic diagram of a module of a digital speech out-of-sync noise processing system for PCM encoding according to an embodiment of the present invention;
[0041] Figure 4 This is a schematic diagram illustrating an application scenario of a digital speech out-of-sync noise processing method for PCM encoding according to an embodiment of the present invention. Detailed Implementation
[0042] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0043] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.
[0044] The following reference Figures 2 to 4 This describes a method for processing out-of-sync noise in PCM-encoded digital speech, provided by some embodiments of the present invention.
[0045] Some embodiments of this application provide a method for processing out-of-sync noise in digital speech encoded with PCM.
[0046] Figure 2The first embodiment of the present invention illustrates a method for processing out-of-sync noise in digital speech for PCM encoding, including steps S1-S4.
[0047] S1. Convert the speech signal into linear PCM data.
[0048] In step S1, the speech signal of each sampling point is converted into 16-bit linear PCM data; the sign bit is removed from the 16-bit linear PCM data, and after two's complement conversion, 15-bit data is generated. The data size PCM_VAL is defined as the audio amplitude of the corresponding sampling point. The larger the PCM_VAL, the larger the audio amplitude of the sampling point.
[0049] S2. Monitor the audio amplitude of each sampling point in each detection cycle of the speech signal, and define the sampling point whose audio amplitude exceeds the first set threshold as the high audio amplitude point, wherein there are multiple sampling points in each detection cycle.
[0050] Understandably, the duration of the detection period and the number of sampling points within the detection period can be set as needed. In this application, a detection period of 20ms is used, with 160 sampling points within the 20ms detection period. During normal calls, the audio amplitude of these sampling points is not high, but when the voice loses synchronization, the audio signal undergoes large-scale, irregular jumps, resulting in some sampling points with very high audio amplitude.
[0051] In one specific embodiment, sampling points with PCM_VAL greater than or equal to 0x2800 are defined as high-frequency amplitude points. It can be understood that 0x2800 represents the first set threshold in hexadecimal, and when using decimal, the first set threshold is 10240.
[0052] S3. Based on the number of high-frequency amplitude points within the detection period, determine whether there is speech sync noise within the detection period, and use a silence flag to mark whether there is speech sync noise within the detection period.
[0053] Specifically, through principle analysis and a large amount of experimental data statistics, speech out-of-sync noise can be detected by judging the number of high-amplitude audio signal points within a 20ms period.
[0054] In some embodiments, when the number of high-frequency amplitude points in a detection period exceeds a second preset threshold, the mute flag is set to 1 to indicate that there is speech out-of-sync noise in the detection period; when speech out-of-sync noise is determined to occur, the conditions for determining speech recovery include: there are no high-frequency amplitude points in N consecutive detection periods, and the mute flag is set to 0.
[0055] In one specific embodiment, the condition for determining speech out-of-sync noise is: when the number of high-frequency amplitude points within 20ms is >= 4, the mute flag is set to 1. The condition for determining speech recovery is: when there are no high-frequency amplitude points within 8 consecutive 20ms intervals, the mute flag is set to 0.
[0056] When speech synchronization noise is detected, the mute flag changes. The speech data to be played can be switched to mute data based on the change in the mute flag's state; this method is simple and easy to implement. However, the user may have already heard some noise before the synchronization noise is detected. Therefore, this disclosure has designed step S4.
[0057] S4. Perform delay processing on the completed marked speech signal, and switch the speech signal after delay processing according to the silence flag. When the silence flag indicates that there is speech out-of-step noise within the detection period, the speech signal within the detection period is silenced. When the silence flag indicates that there is no speech out-of-step noise within the detection period, the speech signal within the detection period is output.
[0058] In step S4, a voice data buffering method is used. The voice signal is first stored in a data buffer. If the mute flag is 0 (low level), normal voice data from the data buffer is selected for playback by the user. If the mute flag is 1 (high level), the voice data to be played is switched to mute data. This method ensures that the user cannot hear noise during the out-of-step noise detection process.
[0059] The delay time for the completed marked speech signal shall not be less than the detection period.
[0060] In one specific embodiment, when the detection period is 20ms, the delay duration of 20ms is sufficient to meet the requirements. To further improve the fault tolerance, the delay is doubled, that is, the delay duration for the marked speech signal is 40ms, which better achieves the effect of eliminating out-of-step noise.
[0061] Using the method described above, when out-of-sync noise is detected, the mute flag is set to 1, the out-of-sync noise data is delayed and buffered, and the call line is switched to mute, at which point the user hears silence. When voice out-of-sync is restored, the mute flag is set to 0, the call line switches back to voice signal, and the user hears normal speech.
[0062] In some embodiments, the method for processing out-of-sync noise in digital speech encoded with PCM further includes step S5.
[0063] S5. If the number of high-frequency amplitude points exceeds the third set threshold within M consecutive detection cycles, a step loss noise notification will be generated, and the voice call process will be restarted.
[0064] Specifically, steps S1-S4 are sufficient to identify out-of-synchronization noise within a 20ms detection period, but false alarms may occur. In some embodiments, the results of M detection periods are algorithmically evaluated to accurately determine out-of-synchronization noise within a fixed time, and then the results are fed back to the system processor. This causes the system processor to restart the voice call process, achieving voice resynchronization and completely eliminating out-of-synchronization noise.
[0065] In one specific embodiment, if the number of high-frequency amplitude points in each detection cycle is greater than 12 within 50 consecutive detection cycles (i.e., 1 second), a step loss noise notification is generated and fed back to the system processor, which then restarts the voice call process, fundamentally solving the step loss noise problem.
[0066] Figure 3 This invention illustrates a digital speech out-of-sync noise processing system for PCM encoding, which includes at least a PCM codec chip, an out-of-sync noise detection module, a delay module, and a voice switching module.
[0067] The system includes a PCM codec chip that receives voice data and converts it into linear PCM data; a step-out noise detection module connected to the PCM codec chip that monitors the audio amplitude of each sampling point in each detection cycle and defines sampling points whose audio amplitude exceeds a first set threshold as high-frequency amplitude points; a module that determines whether there is voice step-out noise in the detection cycle based on the number of high-frequency amplitude points in the detection cycle; and a mute flag that marks the presence of voice step-out noise in the detection cycle; a delay module connected to the step-out noise detection module that performs a delay buffer on the marked voice signal; and a voice switching module connected to the delay module that switches the call line to mute or outputs voice in the detection cycle where there is no voice step-out noise based on the mute flag of the detection cycle.
[0068] In some embodiments, within the out-of-step noise detection module, when the number of high-frequency amplitude points within a detection period exceeds a second preset threshold, a silence flag is set to 1 to indicate that speech out-of-step noise exists in that detection period. When speech out-of-step noise is determined to exist, the conditions for determining speech recovery include: if no high-frequency amplitude points exist within N consecutive detection periods, the silence flag is set to 0. The second preset threshold can be set according to actual conditions and accuracy requirements.
[0069] In some embodiments, the digital speech out-of-sync noise processing system for PCM encoding further includes:
[0070] The out-of-step feedback algorithm decision module, connected to the out-of-step noise detection module, generates an out-of-step noise notification when the number of high-frequency amplitude points in each of the 50 consecutive detection cycles exceeds a third set threshold. The out-of-step noise notification is then sent to the system processor to restart the voice call process.
[0071] Figure 4 This paper illustrates a specific application scenario of a digital speech out-of-sync noise processing method for PCM encoding according to an embodiment of the present invention.
[0072] like Figure 4 As shown, a method for handling out-of-synchronization noise in digital voice for PCM encoding is embedded in digital phones A and B. During a voice call between phones A and B, when the voice sent from phone A to phone B is normal voice, phone B outputs the voice after a 40ms delay following processing by this method. When the voice sent from phone A to phone B is out-of-synchronization noise, phone B can detect the out-of-synchronization noise within 20ms and outputs mute after processing by this method. When phone B receives more than 50 consecutive packets of out-of-synchronization noise data, phone B restarts the call process after out-of-synchronization feedback using this method, restoring the out-of-synchronization noise between phones A and B to normal voice.
[0073] In this specification, the illustrative expressions of the terms used do not necessarily refer to the same embodiments or examples. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0074] Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this invention shall be included within the scope of protection of this invention.
Claims
1. A method for processing out-of-sync noise in digital speech encoded with PCM, characterized in that, include: Convert speech signals into linear PCM data; The audio amplitude of each sampling point is monitored in each detection cycle of the speech signal, and the sampling point whose audio amplitude exceeds a first set threshold is defined as a high audio amplitude point. There are multiple sampling points in each detection cycle. Based on the number of high-frequency amplitude points within the detection period, it is determined whether there is speech sync noise within the detection period, and a silence flag is used to mark whether there is speech sync noise within the detection period. The marked speech signal is delayed, and the speech signal after delay is switched according to the silence flag. When the silence flag indicates that there is speech out-of-step noise in the detection period, the speech signal in the detection period is muted. When the silence flag indicates that there is no speech out-of-step noise in the detection period, the speech signal in the detection period is output.
2. The method for processing out-of-sync noise in digital speech encoding according to claim 1, characterized in that, The process of converting the speech signal into linear PCM data includes: The speech signal at each sampling point is converted into 16-bit linear PCM data; The sign bit of the 16-bit linear PCM data is removed, and after two's complement conversion, 15-bit data is generated. The data volume PCM_VAL of this data is defined as the audio amplitude of the corresponding sampling point.
3. The method for processing out-of-sync noise in digital speech encoding according to claim 2, characterized in that, The step of defining sampling points whose audio amplitude exceeds a first preset threshold as high-frequency amplitude points includes: Sampling points with PCM_VAL greater than or equal to 0x2800 are defined as high-frequency amplitude points.
4. The method for processing digital speech out-of-sync noise for PCM encoding according to claim 3, characterized in that, The step of determining whether speech synchronization noise exists within a detection period based on the number of high-frequency amplitude points within that detection period, and marking the presence of speech synchronization noise within the detection period using a silence flag, includes: When the number of high-frequency amplitude points in the detection period exceeds the second set threshold, the mute flag is set to 1 to indicate that there is speech out-of-sync noise in the detection period. When speech out-of-sync noise is detected, the conditions for determining speech recovery include: there are no high-frequency amplitude points in N consecutive detection cycles, and the mute flag is set to 0.
5. The method for processing out-of-sync noise in digital speech encoding according to claim 1, characterized in that, The delay time for the completed marked speech signal shall not be less than the detection period.
6. The method for processing out-of-sync noise in digital speech encoding according to claim 5, characterized in that, The delay time for the completed marked speech signal is twice the length of the detection period.
7. The method for processing out-of-sync noise in digital speech encoding according to claim 1, characterized in that, Also includes: If the number of high-frequency amplitude points in each of the M consecutive detection cycles is greater than the third set threshold, a step loss noise notification will be generated, and the voice call process will be restarted.
8. A digital speech out-of-sync noise processing system for PCM encoding, characterized in that, include: The PCM codec chip receives voice data and converts it into linear PCM data. The out-of-step noise detection module is connected to the PCM codec chip. It monitors the audio amplitude of each sampling point in each detection cycle and defines the sampling point whose audio amplitude exceeds the first set threshold as a high-frequency amplitude point. Based on the number of high-frequency amplitude points in the detection cycle, it determines whether there is speech out-of-step noise in the detection cycle. It also uses a silence flag to mark whether there is speech out-of-step noise in the detection cycle. The delay module, connected to the step loss noise detection module, delays the marked speech signal. The voice switching module, connected to the delay module, switches the call line to mute or outputs voice within the detection period that is free of voice synchronization noise, based on the mute flag of the detection period.
9. The digital speech out-of-sync noise processing system for PCM encoding according to claim 8, characterized in that, In the out-of-step noise detection module, when the number of high-frequency amplitude points in the detection period exceeds the second set threshold, the mute flag is set to 1 to indicate that there is speech out-of-step noise in the detection period. When speech out-of-sync noise is detected, the conditions for determining speech recovery include: there are no high-frequency amplitude points in N consecutive detection cycles, and the mute flag is set to 0.
10. The digital speech out-of-sync noise processing system for PCM encoding according to claim 8, characterized in that, Also includes: The out-of-step feedback algorithm decision module is connected to the out-of-step noise detection module. When the number of high-frequency amplitude points exceeds a third set threshold within M consecutive detection cycles, it generates an out-of-step noise notification and sends the out-of-step noise notification to the system processor to restart the voice call process.
Citation Information
Patent Citations
PCM code flow voice detection method
CN101046955A
Voice out-of-synchronism detection method based on PCM coding characteristics
CN106612168A