A method, medium and device for eliminating FSK noise in cockpit voice recording

Through FSK audio template matching and multi-channel recording similarity replacement method, the problem of removing FSK noise in the cockpit voice recorder is solved, and a natural and comfortable speech denoising effect is achieved, which is suitable for automatic analysis and processing.

CN120496561BActive Publication Date: 2025-10-03HANGKE TECH DEV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510970489.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-10-03
Estimated Expiration
2045-07-15

AI Technical Summary

Technical Problem

In cockpit voice recorders, FSK noise aliasing is difficult to remove from the speech waveform, affecting the inspection and speech recognition accuracy. Existing low-pass filtering methods will damage valuable signals.

Method used

FSK audio template matching is used to identify FSK noise segments. The audio segments containing FSK noise are replaced by the similarity between the driver's microphone channels recorded in multiple channels, and the reference audio segments are used for denoising.

Benefits of technology

It realizes the automatic removal of FSK noise, maintains the naturalness and comfort of speech, is suitable for automatic analysis and processing scenarios, and is applicable to single or batch processing of records.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496561B_ABST
    Figure CN120496561B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of noise elimination, and in particular to a method, medium, and device for eliminating FSK noise in cockpit voice recordings. The method comprises: determining whether the FSK noise segment and the reference audio segment on the recorded audio to be processed are background noise; if both the FSK noise segment and the reference audio segment are non-background noise, and the frequency domain similarity between the first discrimination audio and the first reference audio is greater than a first similarity threshold, and the time domain similarity between the first discrimination audio and the first reference audio is greater than a second similarity threshold, then replacing the audio in the FSK noise segment with the audio in the reference audio segment. This method utilizes the characteristics of CVR multi-channel recording, compares the audio similarity of two microphone channels, and automatically eliminates noise. The method is suitable for rapidly processing single or multiple records and improving the efficiency of cockpit voice analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of noise elimination, and in particular to a method, a medium and a device for eliminating FSK noise in cockpit voice recording. Background Art

[0002] Frequency Shift Keying (FSK) signals are used in cockpit voice recorders (CVRs) to encode and record time information. Specifically, an FSK signal, known as a "beep," is recorded every four seconds (with an error range of ±2 milliseconds). This technology is used on the CVRs of some aircraft, including the ATR, Airbus series (except the A380 and A350), and Fokker 50 aircraft. This mechanism is primarily used to provide a precise timestamp function, ensuring that the data recorded by the CVR can be associated with accurate time information. This is crucial for post-event analysis, such as accident investigations, as knowing the exact time at which each event occurred helps reconstruct the event process.

[0003] However, improper recorder downloading, or direct recording without FSK separation when using a Quick Access Cockpit Voice Recorder (QCVR) to record cockpit voice, can destroy the original FSK structure, making the UTC information contained in FSK unrecognizable. Furthermore, FSK sound can become aliased within the cockpit voice waveform and cannot be removed, resulting in FSK noise. This harsh FSK noise frequently occurs during cockpit voice inspections and can also overlap with speech sounds, disrupting the cockpit voice inspector's work and affecting the accuracy of subsequent speech recognition applications. Therefore, it is essential to treat FSK sound as noise and remove it.

[0004] A common approach to FSK removal is to low-pass filter the entire recording. Because the carrier frequency used by FSK overlaps with the frequency range (8-16kHz) recorded by CVR, using a filter will inevitably damage the voice and other valuable signal sounds in the cabin sound, adversely affecting the recognition and realism of other signals. Therefore, it is not suitable for FSK removal in this scenario. Summary of the Invention

[0005] In order to solve one of the above technical problems, the present invention adopts the following technical solution:

[0006] According to one aspect of the present invention, a method for eliminating FSK noise in a cockpit voice recording is provided, the method comprising the following steps:

[0007] Use the FSK audio template to match the recorded audio to be processed to determine the FSK noise segment on the recorded audio to be processed;

[0008] The audio of a preset length adjacent to and preceding the FSK noise segment on the recorded audio to be processed is used as the first discrimination audio corresponding to the FSK noise segment;

[0009] On the reference recorded audio, obtaining the audio of the same segment corresponding to the first discrimination audio as the first reference audio; the reference recorded audio and the to-be-processed recorded audio are audio recorded by two driver microphone channels respectively;

[0010] Determining, based on the first discrimination audio, whether the FSK noise segment on the recorded audio to be processed is background noise;

[0011] determining, based on the first reference audio, whether a reference audio segment corresponding to the FSK noise segment on the reference recorded audio is background noise;

[0012] If both the FSK noise segment and the reference audio segment are non-background noise, and the frequency domain similarity between the first discrimination audio and the first reference audio is greater than the first similarity threshold, and / or the time domain similarity between the first discrimination audio and the first reference audio is greater than the second similarity threshold, then the audio in the reference audio segment is used to replace the audio in the FSK noise segment.

[0013] According to a second aspect of the present invention, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned method for eliminating FSK noise in a cockpit voice recording.

[0014] According to a third aspect of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for eliminating FSK noise in a cockpit voice recording is implemented.

[0015] The present invention has at least one of the following beneficial effects:

[0016] Therefore, this paper proposes an FSK noise removal method based on the characteristics of CVR multi-channel recording. For example, CVRs typically contain four channels. Channels 2 and 3 typically record the air-ground conversation and the voices recorded by the pilots' respective microphones. The audio captured by microphone channel 2 typically contains FSK noise, representing the processed recording, while the audio captured by microphone channel 3 is FSK-free and represents the reference recording. This method leverages the similarity of the sound content within these two channels to identify the most similar audio segments as interchangeable segments. The FSK-noise-free segments are then used to replace the FSK-noise segments, achieving FSK noise removal.

[0017] While this method doesn't guarantee the complete elimination of FSK noise, it does, overall, produce a natural and comfortable sound after denoising, achieving a good overall denoising effect. Furthermore, this method can be automated, allowing for rapid completion of single recordings and batch processing of multiple recordings, making it suitable for automated analysis and processing of cabin sound. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0019] Figure 1 A flowchart of a method for eliminating FSK noise in cockpit voice recordings provided by an embodiment of the present invention;

[0020] Figure 2 Schematic diagrams of FSK noise provided in embodiments of the present invention; (a) is a schematic diagram of audio containing FSK noise acquired by the second microphone channel; (b) is a schematic diagram of audio containing the entire FSK noise; (c) is an enlarged schematic diagram of (b); in the figures of the present invention, the horizontal axis represents the sampling points of the data, which can be understood as the time axis, and the vertical axis represents the amplitude of the original audio;

[0021] Figure 3 Schematic diagram of audio correlation comparison between the second and third channels provided in an embodiment of the present invention; (a) is a schematic diagram of audio obtained by the second microphone channel containing FSK noise; (b) is a schematic diagram of audio obtained by the third microphone channel without FSK noise;

[0022] Figure 4Schematic diagram of a comparison of a CVR recording of a flight before and after processing provided by an embodiment of the present invention; (a) is a schematic diagram of the audio containing FSK noise obtained by the second microphone channel of the flight; (b) is the audio obtained by the third microphone channel of the flight without FSK noise; (c) is the audio information obtained by the second microphone channel of the flight after being processed by the FSK elimination method of the present invention;

[0023] Figure 5 Schematic diagram of the self-filling effect of x(n) when x(n) has no signal (i.e., background noise) and y(n) has no signal, provided by an embodiment of the present invention;

[0024] Figure 6 Schematic diagram of the self-filling effect of x(n) when x(n) has no signal and y(n) has a signal (i.e., non-background noise) according to an embodiment of the present invention;

[0025] Figure 7 Schematic diagram of the effect of using the substitution method when x(n) has a signal, y(n) has a signal, and x(n) is correlated with y(n) according to an embodiment of the present invention;

[0026] Figure 8 This is a schematic diagram of not processing when x(n) has a signal, y(n) has a signal but x(n) and y(n) are not correlated, or y(n) has no signal, provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.

[0028] As a possible embodiment of the present invention, Figure 1 As shown, a method for eliminating FSK noise in cockpit voice recording is provided, the method comprising the following steps:

[0029] S100: Using the FSK audio template to perform matching on the recorded audio to be processed, and determining the FSK noise segment on the recorded audio to be processed.

[0030] The recorded audio to be processed in this embodiment is as follows: Figure 2The figure below shows audio typically captured from microphone channel 2 with FSK noise. FSK occurs every 4 seconds (±2 milliseconds), and each FSK signal is fixed at 32 bits. Furthermore, FSK transmits information by varying the frequency of the carrier signal. In FSK modulation, digital signals (such as binary "0" and "1") are mapped to different frequency values. For example, a higher frequency is typically defined to represent a logical "1," while a lower frequency represents a logical "0." Because information is encoded in frequency rather than amplitude, this modulation scheme makes FSK signals highly robust to amplitude variations and noise during transmission.

[0031] Based on the above characteristics, it can be seen that when FSK is used to encode time information and record it in the cockpit voice recorder (CVR), although the content of the time information transmitted each time varies (for example, the time value changes), the encoding format, structure, and transmission method remain consistent, resulting in the FSK noise signal having a basically consistent morphology. Furthermore, in the acquired recorded audio information to be processed, the FSK noise signal appears every four seconds.

[0032] Based on this, in this embodiment, an FSK audio template is set and matched against the recorded audio to be processed to identify the FSK noise segment in the recorded audio to be processed. For example, in the FSK audio template, an FSK encoded signal corresponding to a timestamp can be set at intervals of 4 seconds. A total of 5-8 groups of FSK encoded signals corresponding to timestamps can be set to form the FSK audio template.

[0033] The FSK audio template is then compared against the recorded audio to be processed, with the locations where the similarity exceeds the threshold being considered successful matches. The corresponding segments of the FSK segments in the FSK audio template on the recorded audio to be processed are then treated as FSK noise segments, and subsequent audio segments after every four seconds are treated as FSK noise segments.

[0034] S200: The audio of a preset length adjacent to and preceding the FSK noise segment on the recorded audio to be processed is used as the first discrimination audio corresponding to the FSK noise segment.

[0035] In this embodiment, both channels 2 and 3 record the air-ground call audio, and when the intercom switch is turned on, the microphone signals of the left and right seats will also be sent to the other party's headphones. Figure 3 As shown, it can be seen that sometimes there are speech / non-speech signals between the two channels and the content is relatively consistent, but the volume is different; sometimes one channel has a signal, while the other channel has only background noise (generally considered to be a stable signal); sometimes both channels have only background noise, of course, their respective noise levels may be different.

[0036] Therefore, while the content of these two channels isn't exactly the same, there's still a high degree of overlap. Furthermore, on models that record FSK, one and only one of these two channels often contains FSK noise. This forms the basis for losslessly replacing the audio from one channel with the audio from the other.

[0037] Since subsequent steps require determining whether an audio segment in the reference recorded audio can be used to replace the FSK noise segment in the processed recorded audio based on the similarity between the reference recorded audio and the recorded audio to be processed, it is necessary to determine the audio segment to be used for similarity comparison. However, since the FSK noise segment in the processed recorded audio is mixed with the FSK signal, it significantly changes the original sound signal, making similarity comparison impossible. Therefore, in this step, audio of a preset length immediately preceding and / or following the FSK noise segment is selected as the audio segment for similarity comparison. Specifically, the preset length can be determined based on actual needs and needs to be less than 4 seconds, such as 200 milliseconds.

[0038] S300: From the reference recorded audio, obtain the audio of the same segment as the first determination audio as the first reference audio. The reference recorded audio and the to-be-processed recorded audio are audio recorded by two driver microphone channels respectively.

[0039] In this embodiment, the reference recorded audio may be the audio obtained by the third microphone channel and is audio without FSK noise. Before performing S300 , the reference recorded audio is time-aligned with the recorded audio to be processed.

[0040] S400: Determine, based on the first discrimination audio, whether the FSK noise section on the recorded audio to be processed is background noise.

[0041] S500: Determine, based on the first reference audio, whether a reference audio segment corresponding to the FSK noise segment on the reference recorded audio is background noise.

[0042] The purpose of signal detection at S400 and S500 is to distinguish between background noise (stationary signals) and signal sounds (non-stationary signals such as speech, warning sounds, and mechanical sounds). Considering that the left and right cabin audio channels in this embodiment are both headset or handheld microphone channels, their noise is mostly air conditioning noise and has a high signal-to-noise ratio. Therefore, this embodiment uses one or a combination of short-term energy detection, short-term zero-crossing rate, or a posteriori signal-to-noise ratio to determine whether the audio in the current segment is background noise.

[0043] For example, let the recorded audio data to be processed be x(n), where n = 0, 1, 2, ..., N-1, where N is the signal length, and let the reference recorded audio data be y(n), where n = 0, 1, 2, ..., N-1. FSK noise occurs within a time interval [n1, n2] of x(n). The time interval corresponding to the first discrimination audio and the first reference audio is [n1-Δ, n1)∪(n2, n2+Δ].

[0044] Short-time energy detection: Calculate the short-time energy E1 of signals x(n) and y(n) in the interval [n1-△, n1)∪(n2, n2+△] x and E1 y , if E1 x or E1 y Exceeding a certain threshold T E , then the interval is considered to contain signal sound; otherwise, it is considered to be mainly background noise.

[0045] Short-time zero-crossing rate: Calculate the short-time zero-crossing rate ZCR of signals x(n) and y(n) in the interval [n1-△, n1)∪(n2, n2+△] x and ZCR y , if ZCR x or ZCR y Exceeding a certain threshold T ZCR , then the interval is considered to contain signal sound.

[0046] Posterior signal-to-noise ratio: Calculate the corresponding posterior signal-to-noise ratio (SNR) of signals x(n) and y(n) in the interval [n1-△, n1)∪(n2, n2+△] postx and SNR posty In the specific calculation, the signal is divided into frames with a length of 25 milliseconds and a frame shift of 10 milliseconds. The posterior signal-to-noise ratio is calculated as e(m) and e v Logarithmic ratio of (m):

[0047] ;

[0048] Where m is the frame index, e(m) is the energy of the mth frame of x(n), and e v (m) is the noise energy of the mth frame of x(n). v The calculation method of (m) is as follows:

[0049] First, every 200 frames in x(n) are divided into a supersegment (approximately 2 seconds): x(p)=s(p)+v(p), p=1,…,P, where P is the total number of supersegments in x(n). s(p) and v(p) are the signal and noise in each supersegment x(p), respectively. For each supersegment x(p), the noise energy e v(p) may be the energy corresponding to the lowest energy frame in the top 10% after the signals in the super segment are sorted by energy.

[0050] Afterwards, in order to make the noise energy e v The value of (p) is smoother. In this embodiment, when calculating the noise energy of the current super segment, the influence of the noise energy of the previous super segment is introduced, and this influence is taken into account by setting the forgetting factor to 0.9. The specific calculation is as follows:

[0051] ;

[0052] Among them, the noise energy e of the mth frame v (m) is taken from the e of the pth super segment to which the mth frame belongs v The energy value of (p).

[0053] If the SNR in the interval [n1-△, n1)∪(n2, n2+△] postx or SNR posty Exceeding a certain threshold T SNR , then the interval is considered to contain signal sound.

[0054] S600: If both the FSK noise segment and the reference audio segment are non-background noise, and the frequency domain similarity between the first discrimination audio and the first reference audio is greater than a first similarity threshold, and / or the time domain similarity between the first discrimination audio and the first reference audio is greater than a second similarity threshold, then the audio in the reference audio segment is used to replace the audio in the FSK noise segment. Figure 7 shown.

[0055] Specifically, in this step, the similarity between the first discrimination audio and the first reference audio can be determined by using an existing similarity calculation method. In this embodiment, similarity is evaluated based on the time domain similarity and frequency domain similarity between the two audio segments. The specific similarity calculation is as follows:

[0056] The frequency domain similarity S between the first determination audio and the first reference audio xy Meet the following requirements:

[0057] ;

[0058] Where X(f) and Y(f) are the spectra of x(n) and y(n) on [n1-△, n1)∪(n2, n2+△]. If S xy Exceeding a certain threshold T S , then x(n) and y(n) are considered to be highly similar in this interval.

[0059] Specifically, the time domain similarity R between the first discrimination audio and the first reference audio xy(k) The following conditions are met:

[0060] ;

[0061] Where k is the time delay, N=[n1-△, n1)∪(n2, n2+△]. In this algorithm, the similarity between two audios is determined by calculating the correlation between the two audios in the time domain. If R xy The maximum value of (k) exceeds a certain threshold T R , then x(n) and y(n) are considered to be highly similar in this interval.

[0062] Considering that the volume levels of x(n) and y(n) are likely to be different, the volume level of the replacement signal should be adjusted to be comparable to the volume level of the replaced signal before signal substitution. This can be achieved by performing volume normalization on y(n). Specifically, the energy of x(n) and y(n) within the interval [n1, n2] can be calculated. However, since x(n) has been contaminated by FSK within the interval [n1, n2], the volume normalization should refer to the surrounding signals of [n1, n2] rather than [n1, n2] itself. Assume that we choose the interval [n1-△, n1)∪(n2, n2+△] as the reference interval, where △ is the preset audio length, such as the audio length corresponding to the preset audio length of 200 milliseconds.

[0063] Based on the above method, specifically replacing the audio of the FSK noise segment with the audio in the reference audio segment includes:

[0064] S601: Generate a volume normalization factor α based on the audio in the initial reference audio segment and the first discrimination audio. α satisfies the following conditions:

[0065] ;

[0066] ;

[0067] ;

[0068] Where x(n) and y(n) are the functions of the recorded audio to be processed and the reference recorded audio, respectively. n1 and n2 are the two boundary points of the FSK noise segment on the recorded audio to be processed. △ is the preset audio length. E x is the energy of x(n) on [n1-△, n1)∪(n2, n2+△]. y is the energy of y(n) on [n1-△, n1)∪(n2, n2+△].

[0069] S602: Generate target replacement audio y based on α and the audio in the initial reference audio segment scaled(n). y scaled (n) The following conditions are met:

[0070] y scaled (n)=α×y(n), n∈[n1, n2].

[0071] S603: Use y scaled (n) Replace the audio of the FSK noise segment.

[0072] After volume normalization, the volume transition between the replaced sound and the original sound can be made smoother and more natural.

[0073] After the FSK noise segment and the reference audio segment are both non-noise floor segments, the method further includes:

[0074] S610: If the frequency domain similarity between the first discrimination audio and the first reference audio is less than the first similarity threshold, and the time domain similarity between the first discrimination audio and the first reference audio is less than the second similarity threshold, retain the audio of the FSK noise segment, or perform low-pass filtering on the audio of the FSK noise segment. Figure 8 shown.

[0075] For audio segments where FSK noise cannot be eliminated through replacement, consider retaining them as is or using other denoising methods. Specifically, if the audio is used for investigation-related scenarios, it is best to keep the original signal unchanged, that is, retain the audio of the FSK noise segment. For batch routine flight inspections, denoising methods such as low-pass filtering can be considered to deal with the noise.

[0076] After determining whether the FSK noise segment and the reference audio segment are background noise, the method further includes:

[0077] S620: If the FSK noise segment is non-background noise and the reference audio segment is background noise, retain the audio of the FSK noise segment or perform low-pass filtering on the audio of the FSK noise segment.

[0078] S630: If the FSK noise segment is background noise, the background noise audio on the recorded audio to be processed is used to replace the audio of the FSK noise segment. Figure 5 and Figure 6 shown.

[0079] The method from S600 to S630 in this embodiment mainly makes the following types of judgments based on whether x(n) and y(n) around the FSK noise segment are background signals and the similarity between the two:

[0080] x(n) has no signal, and y(n) has no signal or has a signal: x(n) should be kept consistent with the surrounding noise floor characteristics to ensure auditory coherence and comfort. For example, the simplest method is to use the surrounding noise floor to fill in the noise. To avoid discontinuities or sudden changes at signal splicing points, thereby reducing auditory unnaturalness, windowing can be performed around [n1, n2] during filling.

[0081] x(n) has a signal, y(n) has a signal, and x(n) and y(n) are correlated: Use the substitution method.

[0082] x(n) has a signal, y(n) has a signal but x(n) and y(n) are uncorrelated, or y(n) has no signal: In this case, the signal of x(n) cannot be replaced by y(n) and should be retained as is or other denoising methods should be used.

[0083] Take the CVR recording of a certain flight as an example. The audio is 3 hours and 36 minutes long and the sampling rate is 16kHz. Of the 4 recording channels, 2 channels are contaminated by FSK noise, while 3 channels are free of FSK noise (see Figure 4 The total length of FSK is 180ms, with the middle 42ms representing the main FSK signal and flanked by side lobes. Using the FSK elimination method in this embodiment, each FSK signal is located based on its regular intervals, and the conditions of channels 2 and 3 are calculated. Different processing methods are then applied based on the aforementioned method to eliminate FSK noise. Since this audio segment is for investigation purposes, the FSK signal was left unchanged when deciding on the processing solution.

[0084] Specifically, the final FSK elimination overall effect is as follows Figure 4 As shown in (c) in Figure 4 It can be seen that the original x(n) contains a lot of FSK noise, so when it is reduced, the waveform is completely covered by the FSK noise with large amplitude. removeFSK As can be seen from the waveform, since almost all FSK is eliminated, the true appearance of x(n) is revealed, which is very similar to y(n) of the reference channel in most cases.

[0085] As another possible embodiment of the present invention, after obtaining the first discrimination audio and the first reference audio corresponding to all FSK noise segments, the method further includes:

[0086] S700: Obtain a first identification audio sequence corresponding to the recorded audio to be processed, which is a time sequence of the first identification audio.

[0087] S701: Acquire a first reference audio sequence corresponding to a reference recorded audio. The sequence is a time sequence of the first reference audio.

[0088] S702: Using the same sliding window, select corresponding discrimination sub-audio sequences and reference sub-audio sequences from the first discrimination audio sequence and the first reference audio sequence, respectively. With each sliding window, the length of the sliding window decreases. Each time the sliding window changes length, a corresponding set of discrimination sub-audio sequences and reference sub-audio sequences is generated. The minimum sliding window length is greater than 4 seconds, and the maximum length can be 12 seconds. This allows for a large-scale audio similarity comparison to be performed, and the similarity is used to determine whether direct replacement is possible, thus reducing the computational complexity of steps S400 through S600 in the above embodiment.

[0089] S703: Obtaining the similarity between the determined sub-audio sequence and the reference sub-audio sequence at each sequence length.

[0090] S704: If, before the sequence length reaches the minimum length, the similarity between any set of discrimination sub-audio sequences and the reference sub-audio sequence is greater than a preset threshold, the audio segment corresponding to the reference sub-audio sequence on the reference recorded audio is used to replace the audio segment corresponding to the discrimination sub-audio sequence on the recorded audio to be processed. At the same time, the sliding window is stopped, restored to the maximum length, and slides to the next position.

[0091] S705: If the similarity between all group identification sub-audio sequences and the reference sub-audio sequence is less than a preset threshold when the sequence length reaches the minimum length, the steps from S400 to S600 are used to eliminate the FSK noise.

[0092] Furthermore, although the steps of the method of the present disclosure are described in a particular order in the accompanying drawings, this does not require or imply that the steps must be performed in this particular order, or that all steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0093] Through the description of the above embodiments, it will be readily understood by those skilled in the art that the example embodiments described herein can be implemented via software or via a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, or mobile hard drive) or on a network and includes several instructions for enabling a computing device (such as a personal computer, server, mobile terminal, or network device) to execute the methods according to the embodiments of the present disclosure.

[0094] In an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided.

[0095] Those skilled in the art will appreciate that various aspects of the present invention may be implemented as systems, methods, or program products. Therefore, various aspects of the present invention may be implemented in the following forms: entirely in hardware, entirely in software (including firmware, microcode, etc.), or in a combination of hardware and software, collectively referred to herein as "circuits," "modules," or "systems."

[0096] The electronic device according to this embodiment of the present invention is merely an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0097] The electronic device is implemented as a general-purpose computing device. Components of the electronic device may include, but are not limited to, the at least one processor, the at least one memory, and a bus connecting different system components (including the memory and the processor).

[0098] The storage stores program codes, which can be executed by the processor, so that the processor executes the steps according to various exemplary embodiments of the present invention described in the above “Exemplary Method” section of this specification.

[0099] The memory may include readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory, and may further include read only memory (ROM).

[0100] The storage may also include a program / utility having a set (at least one) of program modules, such program modules including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0101] The bus may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures.

[0102] The electronic device may also communicate with one or more external devices (e.g., a keyboard, pointing device, Bluetooth device, etc.), one or more devices that enable a user to interact with the electronic device, and / or any device that enables the electronic device to communicate with one or more other computing devices (e.g., a router, modem, etc.). This communication may occur via an input / output (I / O) interface. Furthermore, the electronic device may communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter. The network adapter communicates with other modules of the electronic device via a bus. It should be understood that, although not shown in the figures, other hardware and / or software modules may be used in conjunction with the electronic device, including but not limited to microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0103] Through the description of the above embodiments, it will be readily understood by those skilled in the art that the example embodiments described herein can be implemented via software or via a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, or mobile hard drive) or on a network and includes several instructions for enabling a computing device (such as a personal computer, server, terminal device, or network device) to execute the methods according to the embodiments of the present disclosure.

[0104] In exemplary embodiments of the present disclosure, a computer-readable storage medium is also provided, on which is stored a program product capable of implementing the methods described above. In some possible implementations, various aspects of the present invention may also be implemented in the form of a program product comprising program code that, when executed on a terminal device, causes the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the "Exemplary Methods" section above.

[0105] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0106] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0107] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0108] Program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and conventional procedural programming languages ​​such as C or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0109] Furthermore, the above-described figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes illustrated in the above-described figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0110] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.

[0111] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A method for eliminating FSK noise in cockpit voice recording, characterized in that: The method comprises the following steps: Use the FSK audio template to match the recorded audio to be processed to determine the FSK noise segment on the recorded audio to be processed; The audio of a preset length adjacent to and preceding the FSK noise segment on the recorded audio to be processed is used as the first discrimination audio corresponding to the FSK noise segment; On the reference recorded audio, obtaining the audio of the same segment corresponding to the first discrimination audio as the first reference audio; the reference recorded audio and the recorded audio to be processed are audio recorded by two driver microphone channels respectively; Determining, based on the first discrimination audio, whether the FSK noise segment on the recorded audio to be processed is background noise; determining, based on the first reference audio, whether a reference audio segment corresponding to the FSK noise segment on the reference recorded audio is background noise; If both the FSK noise segment and the reference audio segment are non-background noise, and the frequency domain similarity between the first discrimination audio and the first reference audio is greater than the first similarity threshold, and / or the time domain similarity between the first discrimination audio and the first reference audio is greater than the second similarity threshold, then the audio in the reference audio segment is used to replace the audio in the FSK noise segment.

2. The method according to claim 1, characterized in that After the FSK noise segment and the reference audio segment are both non-noise floor segments, the method further includes: If the frequency domain similarity between the first discrimination audio and the first reference audio is less than the first similarity threshold, and the time domain similarity between the first discrimination audio and the first reference audio is less than the second similarity threshold, the audio of the FSK noise segment is retained, or the audio of the FSK noise segment is low-pass filtered.

3. The method according to claim 1, characterized in that After determining whether the FSK noise segment and the reference audio segment are background noise, the method further includes: If the FSK noise segment is non-background noise and the reference audio segment is background noise, the audio of the FSK noise segment is retained or low-pass filtered.

4. The method according to claim 1, wherein After determining whether the FSK noise segment and the reference audio segment are background noise, the method further includes: If the FSK noise section is the background noise, the background noise audio on the recorded audio to be processed is used to replace the audio of the FSK noise section.

5. The method according to claim 1, wherein The audio that replaces the FSK noise segment with the audio from the reference audio segment includes: A volume normalization factor α is generated based on the audio in the initial reference audio segment and the first discrimination audio; α satisfies the following conditions: ; ; ; Where x(n) and y(n) are the functions of the recorded audio to be processed and the reference recorded audio, respectively; n1 and n2 are the two boundary points of the FSK noise segment on the recorded audio to be processed; △ is the preset audio length; E x is the energy of x(n) on [n1-△, n1)∪(n2, n2+△]; E y is the energy of y(n) on [n1-△, n1)∪(n2, n2+△); Generate target replacement audio y based on α and the audio in the initial reference audio segment scaled (n); y scaled (n) The following conditions are met: the scaled (n)=α×y(n),n∈[n1,n2]; Use y scaled (n) Replace the audio of the FSK noise segment.

6. The method according to claim 5, characterized in that The frequency domain similarity S between the first determination audio and the first reference audio xy Meet the following requirements: ; Among them, X(f) and Y(f) are the frequency spectra of x(n) and y(n) on [n1-△, n1)∪(n2, n2+△] respectively.

7. The method according to claim 5, characterized in that The time domain similarity R between the first determination audio and the first reference audio xy (k) The following conditions are met: ; Where k is the delay, N=[n1-△,n1)∪(n2,n2+△].

8. The method according to claim 1, characterized in that Use one or a combination of short-time energy detection, short-time zero-crossing rate, or posterior signal-to-noise ratio to determine whether the audio in the current segment is background noise.

9. A non-transitory computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for eliminating FSK noise in a cockpit voice recording according to any one of claims 1 to 8 is implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for eliminating FSK noise in a cockpit voice recording according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Method and device for conducting self-adaption spectrum reduction and wavelet packet noise elimination processing on voice signals

    CN104269178A

  • Voice changing processing method and device, equipment and storage medium

    CN117975981A