A method and system for regulating voice data payload information enhancement
By using technologies for separating, reducing noise, and enhancing control voice signals in the air traffic control system, the problems of multiple voices intertwining and noise interference have been solved, achieving efficient and accurate speech recognition and improving the overall performance of the air traffic control system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SICHUAN JIUZHOU AIR TRAFFIC CONTROL TECHNOLOGY CO LTD
- Filing Date
- 2023-10-19
- Publication Date
- 2026-06-23
Smart Images

Figure CN117636889B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech signal processing technology, specifically to a method and system for enhancing the effective information of controlled speech data. Background Technology
[0002] Traditional air traffic control methods rely on verbal communication between controllers and pilots. This method is subject to many limitations, such as human error, intermingling of multiple voices, and communication interference, all of which can affect aviation safety. Furthermore, with the growth of air traffic, the number of flights that manual control can handle has reached its limit.
[0003] In recent years, with the rapid development of artificial intelligence and machine learning technologies, these technologies have achieved significant breakthroughs in many fields, one of which is speech recognition. Modern speech recognition technology can achieve high-accuracy speech-to-text conversion, which makes its application in air traffic control systems possible. Using this technology, controllers can handle communication tasks more efficiently, reduce errors, and improve the overall efficiency of the system.
[0004] However, the challenges of speech recognition in air traffic control systems differ from those in other application areas. Against this backdrop, the demand for intelligent air traffic control systems is increasingly urgent, and a key technology is efficient and accurate control voice signal processing. This requires high-quality preprocessing of the input speech signal to ensure the smooth progress of subsequent speech recognition and ultimately achieve intelligent control operations. However, current speech recognition technology still faces challenges when applied to air traffic control scenarios, such as multiple speakers simultaneously, difficulty in recognizing weak and effective signals, and background noise interference, which limits the recognition rate. Summary of the Invention
[0005] The technical problem to be solved by this invention is how to separate multi-person voice signals and reduce noise and enhance effective information of the signals. The purpose is to provide a method and system for enhancing effective information of controlled voice data, which realizes the separation, noise reduction and enhancement of voice signals.
[0006] This invention is achieved through the following technical solution:
[0007] A method for enhancing the effective information of regulatory voice data includes the following steps in sequence: separation of regulatory voice signals, noise reduction of regulatory voice signals, and enhancement of regulatory voice signals;
[0008] Specific methods for controlling voice signal separation include:
[0009] The first step is to acquire the regulated voice signal;
[0010] The second step is to identify the interleaved signal segments;
[0011] The third step is to establish a prediction model for preceding signals.
[0012] The fourth step is to predict the possible states of the preceding signal in the interleaving section;
[0013] The fifth step is to separate the signal that is consistent with the preceding signal from the interleaved segment and output the complete signal;
[0014] The sixth step is to splice the remaining signal of the interleaved segment with the subsequent signal and determine whether there is an interleaved segment in the spliced signal. If there is no interleaved segment, the spliced signal is output as a valid signal; if there is an interleaved segment, the process jumps to the second step.
[0015] Specifically, the method for identifying interleaved signal segments in the second step includes:
[0016] A1. Construct a mathematical model of the phase change of the input control voice signal relative to time t: S(t)=m(t)cos(wt+a), where m is the communication baseband signal, w is the standard call frequency, and a is the frequency jitter relative to the baseband signal;
[0017] A2. Set the sampling interval and sample the controlled voice signal according to the sampling interval;
[0018] A3. Select sampling point t k and its next sampling point t k+1 And obtain the corresponding signal model S(t) k ) and S(t k+1 );
[0019] A4. Set the phase transition threshold T, and use |S(t) k )-S(t k+1 )|>T determines whether a phase change has occurred. If no phase change has occurred, return to step A3; if a phase change has occurred, record the change position and increment the number of change positions by 1.
[0020] A5. Determine if the number of mutation positions is 2. If not, return to step A3; if yes, mark the control voice signal segment between the two mutation positions as an interleaving segment.
[0021] A6. Reset the number of mutation locations to 0, and determine whether all sampling points have been detected. If not, proceed to step A3; if so, end the process.
[0022] Specifically, the methods for establishing the preceding signal prediction model in the third step include:
[0023] B1. Obtain the preceding non-interleaved signal of the interleaved signal segment;
[0024] B2. Calculate the standard frequency difference Δw at each sampling location based on the standard communication frequency w. iand phase difference Δφ i ;
[0025] B3. Through mathematical modeling analysis, establish the fitting function w(i) for the standard frequency difference and the fitting function φ(i) for the phase difference;
[0026] B4. Establish a prediction model for preceding signals: Where i is the number of samples taken relative to the start time of model establishment at the predicted location time t;
[0027] B5. Establish a prediction model for the preceding signal in the interleaving segment: Wherein, φ0 is the phase of the controlled voice signal relative to the standard communication signal;
[0028] Specifically, the method for separating the signal consistent with the preceding signal from the interleaved segment in the fifth step includes:
[0029] C1. Calculate the number of samples i1 at time t1 and the number of samples i2 at time t2, where t1 is the time when interleaving starts, t2 is the time when interleaving ends, and i is the number of samples taken between time points t1 and t2 relative to the start time t0 of the preceding model.
[0030] C2. Take sampling positions k in sequence and obtain the predicted signal Phase(k) at k.
[0031] C3. Separate the prediction signal from the sampled value V(k) of the effective signal at point k. in, Let γ(k) be the remaining signal after separation, and let γ(k) be the separation calculation function.
[0032] C4. Judgment If the signal is valid, proceed to C5; otherwise, proceed to C6.
[0033] C5. Incorporate Phase(k) into the effective separated audio signal. Incorporate the signal to be processed and proceed to step C7;
[0034] C6. The preceding signal and the effective audio signal are spliced together and output. The signal to be processed is spliced together with the remaining unseparated interlaced signal and the subsequent un-interlaced signal and output to complete the separation.
[0035] C7. Determine whether the interleaving segment has been processed. If not, proceed to step C2. If it has been processed, splice the preceding signal and the effective separated audio signal together and output them. Splice the processed signal together with the subsequent uninterleaved signal and output them to complete the separation.
[0036] Among them, the pre-sequence signal and the effective separation signal of the spliced output are the effective signals.
[0037] Alternatively, methods for controlling voice signal noise reduction include:
[0038] Step 1: Acquire valid signals and perform frame-by-frame sampling on the valid signals;
[0039] Step 2: Obtain the sampling point set, and obtain the abnormal point set and the normal point set;
[0040] Step 3: Perform Fourier transform on the normal point set and the abnormal point set;
[0041] Step four: Represent the noise spectrum by calculating the variance of the point set;
[0042] Step 5: Perform inverse Fourier transform to obtain the denoised speech signal.
[0043] Specifically, the specific methods for step one include:
[0044] S1. Obtain the valid signal S, with a time length of t. s ;
[0045] S2. Perform frame sampling on the effective signal data frequency to obtain the sampling point set F = {f(t), t|0<t<t} consisting of time t and frequency f(t). s};
[0046] The specific methods for step two include:
[0047] S3. Use the local outlier factor algorithm to identify the set of outliers E, where E∈F, from set F;
[0048] S4. Obtain the normal point set G = FE;
[0049] The specific methods for step three include:
[0050] S5. For each point in the outlier set, the corresponding time t e Calculate the Fourier transform γ(t) of S in the corresponding frame. e );
[0051] S6. Time t corresponding to each point in the normal point set. g Calculate the Fourier transform γ(t) of S in the corresponding frame. g );
[0052] The specific methods for step four include:
[0053] S7. Calculate the upper frequency limit γ of the Fourier transform result corresponding to the normal point set. u and frequency lower limit γ d ;
[0054] S8. Calculate and obtain the noise signal:
[0055] S9. Calculate the variance of the noise signal: Where n is the number of outlier sets E, δ j The difference corresponding to the j-th point;
[0056] The specific methods for step five include:
[0057] S10. Based on the spectral subtraction theory, obtain the corrected signal spectrum. Among them, Y ω Let w be the Fourier transform of the effective signal S, w be the sampling frame of the effective signal, α be the over-subtraction factor, and β be the lower limit parameter of the spectrum.
[0058] S11. Obtain the speech phase value from the effective signal S, and perform inverse Fourier transform on the corrected signal spectrum to obtain the denoised speech signal.
[0059] Alternatively, methods for controlling voice signal enhancement include:
[0060] s1. Obtain the denoised control voice data;
[0061] s2. Discrete the noisy control voice data through discrete wavelet transform to form a multi-scale decomposition;
[0062] s3. Based on the discrete wavelet transform results, calculate the modulus maxima;
[0063] s4. Determine the maximum scale e;
[0064] s5. Obtain the modulus maxima at scale e through thresholding.
[0065] s6. Take the propagation point of the maximum value of the modulus at scale e from scale e-1;
[0066] s7. Let e = e-1, and repeat step s4 until e = 2;
[0067] s8. At the position where there is a modulus maximum when e=2, retain the corresponding modulus maximum point when e=1, and set the remaining coefficients to zero;
[0068] s9. Construct wavelet coefficients using the retained modulus maxima, and then perform inverse discrete wavelet transform using the constructed wavelet coefficients to obtain the enhanced voice signal.
[0069] Optionally, the formula for performing the discrete wavelet transform is: Where u is a discrete scaling index, v is a discrete translation index, c0 is the scaling factor, d0 is the translation step size, and the scaling parameter is... Translation parameters
[0070] The formula for calculating wavelet coefficients is: Where f(t) is the signal for wavelet transform analysis;
[0071] The formula for performing the inverse discrete wavelet transform is: Where g is a constant.
[0072] A system for enhancing the effectiveness of regulated voice data includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of a method for enhancing the effectiveness of regulated voice data as described above.
[0073] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0074] This invention reduces interference during speech recognition by separating the intertwined voice signals of multiple people into effective single-person voice signals. It also improves the clarity of the voice signals by denoising the voice signals, filtering out background noise and process noise caused by deinterleaving. Finally, the voice signal enhancement technology makes the effective signals clearer, further improving the voice quality.
[0075] This invention maintains a stable high recognition rate in complex air traffic control environments where multiple people are speaking simultaneously, weak effective signals exist, or background noise interference exists. This avoids controllers having to process and judge a large amount of voice information, reduces errors, and improves the overall efficiency of the system. Attached Figure Description
[0076] The accompanying drawings illustrate exemplary embodiments of the present invention and, together with the description thereof, serve to explain the principles of the invention. These drawings are included to provide a further understanding of the invention and are incorporated in and constitute a part of this specification, but do not constitute a limitation on the embodiments of the present invention.
[0077] Figure 1 This is a flowchart illustrating a method for enhancing the effective information of controlled voice data according to the present invention.
[0078] Figure 2 This is a flowchart of the control voice signal separation according to the present invention.
[0079] Figure 3 This is a schematic diagram of the process for identifying interlaced segments according to the present invention.
[0080] Figure 4 This is a schematic diagram of the process for establishing a preceding signal prediction model according to the present invention.
[0081] Figure 5 This is a schematic diagram of the effective signal separation process according to the present invention.
[0082] Figure 6This is a schematic diagram of the process for noise reduction of controlled voice signals according to the present invention. Detailed Implementation
[0083] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0084] It should also be noted that, for ease of description, only the parts relevant to the present invention are shown in the accompanying drawings.
[0085] Where there is no conflict, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0086] Example 1
[0087] A method for enhancing the effective information of controlled voice data, such as Figure 1 As shown, the steps include, in sequence: separation of the control voice signal, noise reduction of the control voice signal, and enhancement of the control voice signal.
[0088] Controlled speech signal separation technology is based on deinterleaving technology of controlled speech signal phase frequency characteristics. It utilizes the differences in center frequency and phase of different speech signals to separate controlled speech interleaved by multiple people into effective single-person speech signals.
[0089] Controlled voice signal noise reduction extracts the clean controlled voice signal from the voice data of noisy controlled voice calls, filters out background noise and process noise caused by deinterleaving, and improves the clarity of the voice signal.
[0090] Controlled voice signal enhancement technology performs lossless enhancement of the phase frequency characteristics of voice signals, solving the problem of weak effective signals that cannot be effectively identified and applied due to related reasons, and improving the quality of voice signals.
[0091] The following examples illustrate in detail the separation, noise reduction, and enhancement of the controlled voice signal.
[0092] Example 2
[0093] like Figure 3 As shown, the specific methods for separating control voice signals include:
[0094] The first step is to acquire the regulated voice signal;
[0095] The second step is to identify the interleaved signal segment. Since the voice signal has a fixed baseband signal output, it is modulated onto a fixed transmission frequency using a mixer. Because the mixer's local oscillator has a slight frequency jitter, the center frequency of the voice signal will have a certain difference. The frequencies at which this difference begins and ends are identified as the interleaved signal segment. For example... Figure 4 The general steps for identifying interleaved signal segments are shown below:
[0096] A1. Construct a mathematical model of the phase change of the input control voice signal relative to time t: S(t)=m(t)cos(wt+a), where m is the communication baseband signal, w is the standard call frequency, and a is the frequency jitter relative to the baseband signal;
[0097] A2. Set the sampling interval and sample the controlled voice signal according to the sampling interval;
[0098] A3. Select sampling point t k and its next sampling point t k+1 And obtain the corresponding signal model S(t) k ) and S(t k+1 );
[0099] A4. Set the phase transition threshold T, and use |S(t) k )-S(t k+1 )|>T determines whether a phase change has occurred. If no phase change has occurred, return to step A3; if a phase change has occurred, record the change position and increment the number of change positions by 1.
[0100] A5. Determine if the number of mutation positions is 2. If not, return to step A3; if yes, mark the control voice signal segment between the two mutation positions as an interleaving segment.
[0101] A6. Reset the number of mutation locations to 0, and determine whether all sampling points have been detected. If not, proceed to step A3; if so, end the process.
[0102] The third step is to establish a prediction model for the preceding signal. This prediction model is based on the characteristics of the adjacent non-interleaved signals before the interleaving segment, and generates a model of the preceding signal in the interleaving section. For example... Figure 4 As shown, the methods for establishing a preceding signal prediction model include:
[0103] B1. Obtain the preceding non-interleaved signal of the interleaved signal segment;
[0104] B2. Calculate the standard frequency difference Δw at each sampling location based on the standard communication frequency w. i and phase difference Δφ i ;
[0105] B3. Through mathematical modeling and analysis, establish the fitting function w(i) for the standard frequency difference and the fitting function φ(i) for the phase difference. Under different characteristic audio frequencies, the fitting function may be a piecewise curve function or a continuous curve function related to trigonometric functions. The fitting can be performed according to the actual situation.
[0106] B4. Establish a prediction model for preceding signals: Where i is the number of samples taken relative to the start time of model establishment at the predicted location time t;
[0107] Taking the interleaving of the control voice signals S1(t) and S2(t) of two aircraft as an example, the mathematical model of the control voice signal is: S1(t)=m1(t)cos(w1t+φ1), S2(t)=m2(t)cos(w2t+φ2), where m1 and m2 are the standard voice signals, w1 and w2 are the frequencies of the control voice signals, and φ1 and φ2 are the phases of the control voice signals relative to the standard communication signals.
[0108] The interleaved signal can be represented as: After solving the Costas ring problem, the signal amplitude after interleaving can be expressed as:
[0109] As can be seen from the above formula, in the case of interleaving, the signal amplitude after interleaving is not only related to the signal power, but also to the phase and frequency of the two interleaved signals. When two voice signals are interleaved, it can be seen from the V(t) signal that the demodulation phase of the interleaved symbols is related to Δw1, Δw2, φ1, and φ2; the uninterleaved symbols (m1, m2, or 0), The demodulation phase is determined only by its own frequency difference and phase difference.
[0110] B5. Since the phase of the preceding speech signal remains continuous and its frequency characteristics remain unchanged in the interleaving section, based on the above analysis, the predicted phase model of the preceding signal in the interleaving section can be expressed as: Wherein, φ0 is the phase of the controlled voice signal relative to the standard communication signal.
[0111] The fourth step is to predict the possible states of the preceding signal in the interleaving segment. Based on the prediction model of the preceding control voice signal, the signal of each sampling point in the interleaving segment is predicted. Let t0 represent the start time of the preceding model establishment, t1 represent the start time of the interleaving segment, t2 represent the end time of the interleaving segment, and the number of samplings of the time point t between t1 and t2 relative to time t0 is i.
[0112] The fifth step is to separate the signal consistent with the preceding signal from the interleaved segment, output the complete signal, and determine the portion of the preceding signal belonging to sampling position i in the interleaved segment by comparing the predicted signal with the interleaved signal. Figure 5 The general steps for separating the effective signal shown are as follows:
[0113] C1. Calculate the number of samples i1 at time t1 and the number of samples i2 at time t2, where t1 is the time when interleaving starts, t2 is the time when interleaving ends, and i is the number of samples taken between time points t1 and t2 relative to the start time t0 of the preceding model.
[0114] C2. Take sampling positions k in sequence and obtain the predicted signal Phase(k) at k.
[0115] C3. Separate the prediction signal from the sampled value V(k) of the effective signal at point k. in, Let γ(k) be the remaining signal after separation, and let γ(k) be the separation calculation function.
[0116] C4. Judgment If the signal is valid, proceed to C5; otherwise, proceed to C6.
[0117] C5. Incorporate Phase(k) into the effective separated audio signal. Incorporate the signal to be processed and proceed to step C7;
[0118] C6 indicates that the preceding signal features are no longer present at this point. The preceding signal and the valid audio signal are spliced together and output. The signal to be processed is spliced together with the remaining unseparated interlaced signal and the subsequent uninterlaced signal and output, thus completing the separation.
[0119] C7. Determine whether the interleaving segment has been processed. If not, proceed to step C2. If it has been processed, splice the preceding signal and the effective separated audio signal together and output them. Splice the processed signal together with the subsequent uninterleaved signal and output them to complete the separation.
[0120] Among them, the pre-sequence signal and the effective separation signal of the spliced output are the effective signals.
[0121] In steps C6 and C7, the preceding signal and the effectively separated audio signal are spliced together and output, which means that a complete signal consistent with the preceding signal is separated from the interleaved segment.
[0122] The output of the processed signal in step C6, which is spliced together with the remaining unseparated interleaved signal and the subsequent uninterleaved signal, and the output of the signal to be processed in step C7, which is spliced together with the subsequent uninterleaved signal, serve as inputs for further deinterleaving judgment and deinterleaving calculation.
[0123] The sixth step is to splice the remaining signal of the interleaved segment with the subsequent signal and determine whether there is an interleaved segment in the spliced signal. If there is no interleaved segment, the spliced signal is output as a valid signal; if there is an interleaved segment, the process jumps to the second step.
[0124] Example 3
[0125] Improved spectral subtraction for denoising control voice signals extracts clean control voice signals from noisy control voice call data, filtering out background noise and deinterleaving noise to improve voice signal clarity. Based on the fact that the separated voice signal only includes one person's voice and does not necessarily include background silence, sampling is based on the spectral average variance of frequency outliers, replacing the statistical noise variance method in standard spectral subtraction, thus making control voice denoising applicable to spectral subtraction. Figure 6 The general steps for noise reduction of controlled speech data based on the improved spectral subtraction method of this invention are as follows:
[0126] Step 1: Acquire valid signals and perform frame sampling on the valid signals.
[0127] S1. Obtain the valid signal S, with a time length of t. s ;
[0128] S2. Perform frame sampling on the effective signal data frequency to obtain the sampling point set F = {f(t), t|0<t<t} consisting of time t and frequency f(t). s};
[0129] Step 2: Obtain the sampling point set, and obtain the abnormal point set and the normal point set.
[0130] S3. Use the local outlier factor algorithm to identify the set of outliers E, where E∈F, from set F;
[0131] S4. Obtain the normal point set G = FE;
[0132] Step 3: Perform Fourier transform on the normal point set and the abnormal point set.
[0133] S5. For each point in the outlier set, the corresponding time t e Calculate the Fourier transform γ(t) of S in the corresponding frame. e );
[0134] S6. Time t corresponding to each point in the normal point set. g Calculate the Fourier transform γ(t) of S in the corresponding frame. g );
[0135] Step four: Calculate the variance of the point set to represent the noise spectrum.
[0136] S7. Calculate the upper frequency limit γ of the Fourier transform result corresponding to the normal point set. u and frequency lower limit γ d ;
[0137] S8. Calculate the difference between the upper and lower limits of the Fourier transform of each point in the outlier set and the normal set, which is used to represent the noise signal. Calculate and obtain the noise signal:
[0138] S9. Calculate the variance of the noise signal: Where n is the number of outlier sets E, δ j The difference corresponding to the j-th point;
[0139] Step 5: Perform inverse Fourier transform to obtain the denoised speech signal.
[0140] S10. Based on the spectral subtraction theory, obtain the corrected signal spectrum. Among them, Y ω Let be the Fourier transform of the effective signal S, w be the sampling frame of the effective signal, α be the over-subtraction factor, β be the lower spectral limit parameter, and D be the noise spectrum. a > 1; the larger the value of a, the more severe the speech distortion, affecting intelligibility. 0 < β < 1; increasing the value of β may result in residual noise, while a value that is too small may introduce other noise.
[0141] S11. Obtain the speech phase value from the effective signal S, and perform inverse Fourier transform on the corrected signal spectrum to obtain the denoised speech signal.
[0142] Example 4
[0143] Controlled speech signal enhancement technology losslessly enhances the phase frequency characteristics of the speech signal, solving the problem of weak effective signals and ineffective identification and application due to related factors, thus improving the quality of the speech signal. Based on the characteristic that controlled speech content is not necessarily continuous, sampling is based on discrete wavelet transform to enhance the controlled speech signal. The general steps of controlled speech signal enhancement are as follows:
[0144] s1. Obtain the denoised control voice data;
[0145] s2. The noisy control voice data is discretized using discrete wavelet transform to form a multi-scale decomposition; the formula for discrete wavelet transform is: Where u is a discrete scaling index, v is a discrete translation index, c0 is the scaling factor, d0 is the translation step size, and the scaling parameter is... Translation parameters
[0146] s3. Based on the discrete wavelet transform results, calculate the modulus maxima;
[0147] s4. First, thresholding is performed on the largest scale j to determine the largest scale e;
[0148] s5. Obtain the modulus maxima at scale e through thresholding.
[0149] s6. Take the propagation point of the maximum value of the modulus at scale e from scale e-1;
[0150] s7. Suppose that a cone-shaped region of propagation is simulated by the maximum point on the scale of e, and then processed (keeping the maximum points within the region and discarding the maximum points outside the region), let e = e-1, and repeat step s4 until e = 2;
[0151] s8. At the position where there is a modulus maximum when e=2, retain the corresponding modulus maximum point when e=1, and set the remaining coefficients to zero;
[0152] s9. Construct wavelet coefficients using the retained modulus maxima, and then perform inverse discrete wavelet transform using the constructed wavelet coefficients to obtain the enhanced voice signal.
[0153] The formula for calculating wavelet coefficients is: Where f(t) is the signal for wavelet transform analysis;
[0154] The formula for performing the inverse discrete wavelet transform is: Where g is a constant that is uncorrelated with the signal.
[0155] Example 5
[0156] A system for enhancing the effectiveness of regulated voice data includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method for enhancing the effectiveness of regulated voice data as described above.
[0157] Memory is used to store software programs and modules. The processor executes various terminal functions and data processing by running the software programs and modules stored in memory. Memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, at least one executable program required for a given function, etc.
[0158] The storage data area can store data created based on the use of the terminal. Furthermore, the memory can include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0159] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of a method for enhancing the validity of controlled voice data as described above.
[0160] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instruction data structures, program modules, or other data. Computer storage media includes RAM, ROM, EPROM, EEPROM, flash memory or other solid-state storage technologies, CD-ROM, DVD or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that computer storage media are not limited to the above-mentioned types. The aforementioned system memories and mass storage devices can be collectively referred to as memory.
[0161] In the description of this specification, the references to terms such as "one embodiment / mode," "some embodiments / modes," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment / mode or example is included in at least one embodiment / mode or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment / mode or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments / modes or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments / modes or examples described in this specification, as well as the features of different embodiments / modes or examples.
[0162] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0163] Those skilled in the art should understand that the above embodiments are merely for illustrating the present invention and are not intended to limit the scope of the invention. Those skilled in the art can make other changes or modifications based on the above invention, and these changes or modifications still fall within the scope of the present invention.
Claims
1. A method for enhancing the effective information of controlled voice data, characterized in that, The steps include, in sequence: separation of the control voice signal, noise reduction of the control voice signal, and enhancement of the control voice signal; Specific methods for controlling voice signal separation include: The first step is to acquire the regulated voice signal; The second step is to identify the interleaved signal segments; The third step is to establish a prediction model for preceding signals. The fourth step is to predict the possible states of the preceding signal in the interleaved signal segment; The fifth step is to separate the signal that is consistent with the preceding signal from the interleaved signal segment and output the complete signal; Step 6: Concatenate the remaining signal of the interleaved signal segment with the subsequent signal, and determine whether the concatenated signal has an interleaved signal segment. If there is no interleaved signal segment, output the concatenated signal as a valid signal; if there is an interleaved signal segment, proceed to step 2. The method for identifying interleaved signal segments in the second step includes: A1. Construct the input control voice signal relative to time. Mathematical model of phase change: ,in, For communication baseband signals, For standard call frequencies, This refers to frequency jitter relative to the baseband signal; A2. Set the sampling interval and sample the controlled voice signal according to the sampling interval; A3. Select sampling points and its next sampling point And obtain the corresponding signal model. and ; A4. Set the phase change threshold and through Determine whether a phase change has occurred. If no phase change has occurred, return to step A3. If a phase change has occurred, record the change position and increment the number of change positions by 1. A5. Determine if the number of mutation positions is 2. If not, return to step A3; if yes, mark the control voice signal segment between the two mutation positions as an interleaved signal segment. A6. Reset the number of mutation locations to 0, and determine whether all sampling points have been detected. If not, proceed to step A3; if so, end the process. The methods for establishing the preceding signal prediction model in the third step include: B1. Obtain the preceding non-interleaved signal of the interleaved signal segment; B2, Based on standard call frequencies Calculate the standard frequency difference at each sampling location and difference ; B3. Establish a fitting function for the standard frequency deviation through mathematical modeling analysis. Fitting function of sum and difference ; B4. Establish a prediction model for preceding signals: , in, To predict location time The number of samples relative to the start time of model building; B5. Establish a prediction model for the preceding signal in the interleaved signal segment: , ,in, This is to control the phase of the voice signal relative to the standard communication signal.
2. The method for enhancing effective information in controlled voice data according to claim 1, characterized in that, The fifth step involves separating the signal that is consistent with the preceding signal from the interleaved signal segment, including: C1. Calculation Number of samples at time 1 and Number of samples at time 1 ,in, The time at which the interlacing begins, The time when the interlacing ends, for arrive The time points between Relative to the start time of the preceding model establishment The number of sampling times between; C2. Take sampling positions in sequence. , obtain Predicted signal at location ; C3, From the effective signal in Sample value at Separate the predicted signal in the middle. ,in, The remaining signal after separation To separate the computation functions; C4. Judgment If the signal is valid, proceed to C5; otherwise, proceed to C6. C5, will Incorporating effectively separated audio signals, Incorporate the signal to be processed and jump to C7; C6. The preceding signal and the effective separated audio signal are spliced together and output. The signal to be processed is spliced together with the remaining unseparated interleaved signal and the subsequent non-interleaved signal and output to complete the separation. C7. Determine whether the interleaved signal segment has been processed. If not, jump to C2. If it has been processed, splice the preceding signal and the effective separated audio signal and output them. Splice the signal to be processed with the subsequent non-interleaved signal and output it to complete the separation. Among them, the pre-sequence signal and the effective separated audio signal output by splicing are the effective signals.
3. The method for enhancing effective information in controlled voice data according to claim 1, characterized in that, Methods for controlling voice signal noise reduction include: Step 1: Acquire valid signals and perform frame-by-frame sampling on the valid signals; Step 2: Obtain the sampling point set, and obtain the abnormal point set and the normal point set; Step 3: Perform Fourier transform on the normal point set and the abnormal point set; Step four: Represent the noise spectrum by calculating the variance of the point set; Step 5: Perform inverse Fourier transform to obtain the denoised speech signal.
4. The method for enhancing effective information in controlled voice data according to claim 3, characterized in that, The specific methods for step one include: S1. Obtain a valid signal The duration is ; S2. Perform frame-by-frame sampling of the effective signal data frequency to obtain the time. With frequency The set of sampling points ; The specific methods for step two include: S3. Employ the local outlier factor algorithm from the set Anomaly detection set , ; S4. Obtain the normal point set ; The specific methods for step three include: S5. The time corresponding to each point in the anomaly set. calculate Fourier transform of the corresponding frame ; S6. Time corresponding to each point in the normal point set calculate Fourier transform of the corresponding frame ; The specific methods for step four include: S7. Calculate the upper frequency limit of the Fourier transform result corresponding to the normal point set. and frequency lower limit ; S8. Calculate and obtain the noise signal: ; S9. Calculate the variance of the noise signal: ,in, For the set of outliers Quantity, For the first The noise signal difference corresponding to each point; The specific methods for step five include: S10. Based on the spectral subtraction theory, obtain the corrected signal spectrum. ,in, Valid signal Fourier transform, For sampling frames of valid signals, For over-subtraction factor, This is the lower limit parameter of the spectrum; S11, From valid signal The speech phase value is obtained, and the denoised speech signal is obtained by performing an inverse Fourier transform on the corrected signal spectrum.
5. A system for enhancing the effectiveness of controlled voice data, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method for enhancing effective information of controlled voice data as described in any one of claims 1-4.
Citation Information
Patent Citations
Method for suppressing transient noise in voice
CN103440871A
A method for processing speech
WO2004064040A1