A Speech Segmentation Method and System for an Air Traffic Control Voice Recorder

By extracting the characteristics of voice data in the air-baric voice recorder system, accurately segmenting and controlling voice from unit voice, the problem of inaccurate voice segmentation in the existing system is solved, and the recording quality and accuracy of signal recognition are improved.

CN118968970BActive Publication Date: 2025-06-20GUANGZHOU ZHONGNANMIN AVIATION GUAN COMM NETWORK TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410943143.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-15
Publication Date
2025-06-20
Estimated Expiration
2044-07-15

AI Technical Summary

Technical Problem

In the existing air-tube voice recorder system, voice segmentation is inaccurate, resulting in the inability to effectively distinguish between controlled voice and unit voice, affecting recording quality and signal recognition.

Method used

By obtaining the voice data of the mixed channel and the controlled channel frame by frame, extracting the speech data characteristics, including signal strength and Mel frequency cepspectral coefficient, to identify the unit voice data in the mixed channel.

Benefits of technology

It realizes the accurate segmentation of controlled voice and unit voice, improves recording quality and signal recognition accuracy, and reduces the transformation and hardware investment in the tower's intra-tower voice system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118968970B_ABST
    Figure CN118968970B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for voice segmentation of an air traffic control voice recorder, including: obtaining mixed-channel voice data and control-channel voice data output by the voice recorder frame by frame, each frame of voice data being accompanied by a timestamp and a channel number, where only control voice data exists in the control channel, and the mixed channel includes crew voice data and control voice data in the control channel; extracting voice data features frame by frame, and identifying whether the voice data in the mixed channel with the same timestamp is control voice data according to the control-channel voice data features frame by frame, and removing the voice data corresponding to the frame in the mixed channel to obtain the crew voice data in the mixed channel. The present invention can remove the control voice in the mixed channel in the air traffic control system according to the voice data features, and finally obtain separate control voice data and separate crew data, and output them through different channels, ensuring the accuracy of other applications such as subsequent voice recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of air traffic control communication, and in particular to a voice segmentation method and system for an air traffic control voice recorder. Background Art

[0002] The air traffic control voice recorder system mainly records air traffic control ground-to-air and plane conversations, and provides corresponding data output capabilities for external systems to use. Usually, the recorder system uses one channel to record the voice of a control seat (frequency point). Since the air traffic control unit has tower control, approach control, and regional control departments, each control department has dozens of control seats. Each control seat is divided into a transmitting signal and a receiving signal. In order to meet business needs and save the number of channels, the receiving and transmitting signals of each seat (frequency point) are usually merged in a half-duplex form and stored through a channel number, that is, the control voice and crew voice of the same control seat are merged and stored to form a mixed voice channel where the control voice and crew voice exist crosswise, which will make it impossible to prepare and judge the voice role. If the merger is not performed and the daily business needs cannot be met, at least three voice channels are required to achieve the separation of the receiving and transmitting signals and the daily business needs, namely the control voice channel, the crew voice channel, and the mixed voice channel. When there are too many voice channels, it is usually necessary to configure the intercom system in a complex way to obtain the crew voice channel, which has certain safety risks and is not convenient for later maintenance. Moreover, in complex environments, such as when the channel noise is large or there is interference from other wireless signals, the accurate recognition of the signal may be affected.

[0003] However, due to the limitation of the design principle of the voice recorder system, when the recorder is recording and monitoring, it usually stops recording after 1 second if it does not detect any high-intensity voice signal. Since the recorder will record and forward the voice signal when it detects a voice signal, if the timestamp of the control channel is directly used to remove the data packets of the mixed channel, the problem of false removal may occur. Therefore, a method that can accurately separate the control voice and the crew voice is currently needed. Summary of the invention

[0004] To solve the above problems, the present invention provides a voice segmentation method and system for an air traffic control voice recorder, which performs voice segmentation based on voice data features, thereby solving the problem of inaccurate voice segmentation in existing air traffic control voice systems.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A voice segmentation method for an air traffic control voice recorder comprises the following steps:

[0007] S1. Acquire the mixed channel voice data and the control channel voice data output by the voice recorder frame by frame. Each frame of voice data carries a timestamp and a channel number. The control channel only contains the control voice data. The mixed channel includes the crew voice data and the control voice data in the control channel. At the same time, the control voice data and the crew voice data in the mixed channel cannot exist at the same time.

[0008] S2. Extract the voice data features of each frame of the mixed channel voice data and each frame of the control channel voice data in units of frames to obtain the mixed channel voice data features and the control channel voice data features;

[0009] S3. According to the voice data characteristics of the control channel, identify whether the voice data in the mixed channel with the same timestamp is the control voice data frame by frame, remove the voice data of the corresponding frame in the mixed channel, and obtain the crew voice data in the mixed channel.

[0010] Furthermore, in step S1, the mixed channel voice data and the controlled channel voice data output by the voice recorder are obtained frame by frame. The specific implementation method is: receiving ADPCM data packets frame by frame, the ADPCM data packets including the channel number, the voice data block and the timestamp, judging whether the voice data block is mixed channel voice data or controlled channel voice data according to the channel number, and converting the ADPCM voice data block into PCM format voice data.

[0011] Furthermore, in step S2, the voice data features of each frame of mixed channel voice data and each frame of controlled channel voice data are extracted respectively on a frame basis, and the specific implementation thereof includes: calculating the signal strength of each frame of mixed channel voice data and each frame of controlled channel voice data, performing spectral analysis on the voice data, and obtaining Mel-frequency cepstral coefficients.

[0012] Furthermore, the calculation formula of the signal strength is:

[0013]

[0014] Wherein, I represents the signal strength, p represents the decimal number represented by the converted PCM format voice data item, and n / 2 represents the number of PCM format voice data items.

[0015] Furthermore, the Mel-frequency cepstral coefficients are calculated by performing a fast Fourier transform on each frame of speech data to obtain a spectrum of each frame of speech data, converting the spectrum of each frame of speech data to a Mel-frequency domain to obtain a Mel-frequency spectrum of the speech data, and performing a discrete cosine transform on the Mel-frequency spectrum of each frame of speech data to obtain a Mel-frequency cepstral coefficient of each frame of speech data.

[0016] Further, before step S3, it further includes: determining whether the average signal intensity of each frame of voice data exceeds a preset threshold. If it does not exceed the preset threshold, then this frame of voice data is a silent frame.

[0017] Further, before identifying whether the voice data in the mixed channel at the same timestamp is regulated voice data, it further includes: at the same timestamp, when the voice data in the regulated channel is a silent frame while the voice data in the mixed channel is not a silent frame, then this frame of voice data in the mixed channel is crew voice data. When subsequently identifying whether the voice data in the mixed channel is regulated voice data, this frame of voice data is skipped.

[0018] Further, in step S3, frame by frame, according to the characteristics of the voice data in the regulated channel, it is identified whether the voice data in the mixed channel at the same timestamp is regulated voice data. The specific implementation method is: comparing the Mel-frequency cepstral coefficients of the voice data in the regulated channel and the mixed channel at the same timestamp. When the Mel-frequency cepstral coefficients of the two are equal or similar, then this frame of voice data in the mixed channel is regulated voice data.

[0019] Further, ADPCM data packets are received through the UDP protocol.

[0020] Through the above technical solution, the characteristics of the voice data in the regulated channel and the mixed channel are extracted frame by frame, and it is identified whether the current frame of voice is regulated voice according to the characteristics of the voice data. Thus, without the need to transform the existing tower interphone system and increase or decrease the hardware investment, the regulated voice in the mixed channel is accurately removed, ensuring the accuracy of voice segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a schematic diagram of the overall process of the voice segmentation method of an air traffic control voice recorder according to the present invention.

[0022] Figure 2 It is a schematic diagram of the structure of a voice segmentation system of an air traffic control voice recorder in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0024] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0025] Example 1

[0026] See also Figure 1 , a voice segmentation method for an air traffic control voice recorder, comprising the following steps:

[0027] S1. Acquire the mixed channel voice data and the control channel voice data output by the voice recorder frame by frame. Each frame of voice data carries a timestamp and a channel number. The control channel only contains the control voice data. The mixed channel includes the crew voice data and the control voice data in the control channel. At the same time, the control voice data and the crew voice data in the mixed channel cannot exist at the same time.

[0028] S2. Extract the voice data features of each frame of the mixed channel voice data and each frame of the control channel voice data in units of frames to obtain the mixed channel voice data features and the control channel voice data features;

[0029] S3. According to the voice data characteristics of the control channel, identify whether the voice data in the mixed channel with the same timestamp is the control voice data frame by frame, remove the voice data of the corresponding frame in the mixed channel, and obtain the crew voice data in the mixed channel.

[0030] The full name of ATC is air traffic control, which is one of the pillars of aviation safety. To ensure flight safety, air traffic controllers communicate with crew members through wireless telecommunications to command and coordinate the entire flight process of each aircraft from takeoff to landing.

[0031] In an optional embodiment, in step S1, the mixed channel voice data and the controlled channel voice data output by the voice recorder are obtained frame by frame. The specific implementation method is: receiving ADPCM data packets frame by frame, the ADPCM data packets including the channel number, the voice data block and the timestamp, judging whether the voice data block is mixed channel voice data or controlled channel voice data according to the channel number, and converting the ADPCM voice data block into PCM format voice data.

[0032] The full name of ADPCM in Chinese is Adaptive Differential Pulse Code Modulation, which is a differential coding modulation with adaptive prediction and quantization functions. The quantization level of its quantizer can be adaptively adjusted with the change of signal value to improve the quantization signal-to-noise ratio. At the same time, the prediction coefficient of its predictor can be adaptively adjusted with the change of signal statistical characteristics to reduce the prediction error and reduce the transmission bit error rate. ADPCM combines the differential characteristics of differential pulse code modulation and the adaptive characteristics of adaptive pulse code modulation. It has the advantages of simple algorithm and low delay, so it is widely used in the field of digital speech processing.

[0033] ADPCM can be decoded through an inverse quantizer and a predictor to be restored to PCM format.

[0034] In an alternative embodiment, in step S2, the speech data features of each frame of mixed-channel speech data and each frame of control-channel speech data are extracted respectively in units of frames. The specific implementation includes: calculating the signal strength of each frame of mixed-channel speech data and each frame of control-channel speech data, and performing spectral analysis on the speech data to obtain Mel-frequency cepstral coefficients.

[0035] The features of the speech data are obtained by calculating the signal strength and Mel-frequency cepstral coefficients. The signal strength is used to determine whether the user is speaking in this frame, while the Mel-frequency cepstral coefficients are used to extract the user's voiceprint features, so as to determine the identity of the current speaking user according to the voiceprint features.

[0036] In an alternative embodiment, the calculation formula of the signal strength is:

[0037]

[0038] where I represents the signal strength, p represents the decimal number represented by the converted PCM data item. Since the recorder uses 16-bit sampling and occupies 2 bytes, the maximum value of p is 2^16, n / 2 represents the number of PCM data items. Usually, after an ADPCM voice data block is converted into PCM format, n = 1024, and usually n = 3072 or 4096 for an ADPCM voice data block.

[0039] In an alternative embodiment, the calculation method of the Mel-frequency cepstral coefficients is: performing a fast Fourier transform on each frame of speech data to obtain the spectrum of each frame of speech data, converting the spectrum of each frame of speech data to the Mel-frequency domain to obtain the Mel-spectrum of the speech data, and performing a discrete cosine transform on the Mel-spectrum of each frame of speech data to obtain the Mel-frequency cepstral coefficients of each frame of speech data.

[0040] Mel-frequency cepstral coefficients are features widely used in automatic speech and speaker recognition. The sound signal itself is a one-dimensional time-domain signal, and it is intuitively difficult to analyze the spectrum change law, which is not convenient for extracting and analyzing its features. Therefore, the short-time Fourier transform is used to process the sound signal. The long signal is framed and windowed, and then the Fourier transform is performed on each frame of the signal, so as to extract the frequency change of the sound signal in the time dimension, and thus extract the features of the sound to better identify the sound.

[0041] The Mel Frequency Cepstral Coefficient (MFCC) is a transformation of the spectrogram based on the Short-Time Fourier Transform (STFT), taking into account the characteristics of the human ear. Since the human ear can hear frequencies in the range of 20 - 20,000 Hz, and the human ear's perception of Hz is not linear, we use the Mel frequency scale to transform the original spectrum onto the cepstrum. This makes the resulting spectrum more in line with the human ear's perception of frequency changes, simulating the characteristics of human auditory perception processing. Further feature extraction is then performed on the transformed spectrum to improve the accuracy of subsequent speech recognition.

[0042] The calculation formula for the Mel Frequency Cepstral Coefficient is as follows:

[0043]

[0044] Where MFCCi is the i-th MFCC coefficient, X[k] is the spectral energy in the Mel frequency domain, and Hk[j] is the response of the k-th Mel filter at the j-th frequency point.

[0045] In an optional embodiment, before step S3, it further includes: determining whether the average signal strength of each frame of speech data exceeds a preset threshold. If it does not exceed the preset threshold, then this frame of speech data is a silent frame.

[0046] In the embodiment, the average signal strength when there is no communication by personnel in the channel is less than 45 decibels. Therefore, 45 decibels is used as the threshold for silent frames, and speech data frames with an average signal strength less than 45 decibels are silent frames.

[0047] In an optional embodiment, before identifying whether the speech data in the mixed channel at the same timestamp is regulated speech data, it further includes: at the same timestamp, when the speech data in the regulated channel is a silent frame and the speech data in the mixed channel is not a silent frame, then this frame of speech data in the mixed channel is crew speech data. When subsequently identifying whether the speech data in the mixed channel is regulated speech data, this frame of speech data is skipped.

[0048] Using silent frames to pre-screen the speech data in the channel reduces the amount of subsequent comparison operations and saves the time required for speech recognition.

[0049] In an optional embodiment, in step S3, frame by frame, based on the characteristics of the speech data in the regulated channel, it is identified whether the speech data in the mixed channel at the same timestamp is regulated speech data. The specific implementation method is: comparing the Mel Frequency Cepstral Coefficients of the speech data in the regulated channel and the mixed channel at the same timestamp. When the Mel Frequency Cepstral Coefficients of the two are equal or similar, then this frame of speech data in the mixed channel is regulated speech data.

[0050] In this embodiment, Mel-frequency cepstral coefficients are used as the dataset to train a convolutional neural network model, and then the trained convolutional neural network model is used to identify the Mel-frequency cepstral coefficients of each frame, so as to identify the channel to which the frame of voice data belongs.

[0051] In an alternative embodiment, ADPCM data packets are received through the UDP protocol.

[0052] The full name of the UDP protocol is the User Datagram Protocol. This is a transport layer protocol that can send encapsulated data packets without establishing a connection. The UDP protocol does not need to maintain a connection, does not check whether data reception and transmission are in error, and does not need to perform functions such as sorting and flow control on data transmission. Therefore, it has a short transmission time and high efficiency, and is a good choice for audio data.

[0053] Example 2

[0054] See Figure 2 , a voice segmentation system for an air traffic control voice recorder, comprising:

[0055] A voice reception module, configured to obtain mixed-channel voice data and control-channel voice data output by the voice recorder frame by frame. Each frame of voice data is provided with a timestamp and a channel number. Only control voice data exists in the control channel. The mixed channel includes crew voice data and control voice data in the control channel. At the same moment, the control voice data and crew voice data in the mixed channel cannot exist simultaneously;

[0056] A feature extraction module, configured to extract voice data features of each frame of mixed-channel voice data and each frame of control-channel voice data respectively in units of frames, to obtain mixed-channel voice data features and control-channel voice data features;

[0057] A voice elimination module, configured to identify whether the voice data in the mixed channel with the same timestamp is control voice data frame by frame according to the control-channel voice data features, and eliminate the voice data of the corresponding frame in the mixed channel to obtain the crew voice data in the mixed channel.

[0058] The embodiments disclosed in this specification are only an illustration of the unilateral features of the present invention. The protection scope of the present invention is not limited to this embodiment, and any other functionally equivalent embodiments fall within the protection scope of the present invention. For those skilled in the art, various corresponding changes and deformations can be made according to the technical solutions and concepts described above, and all these changes and deformations should fall within the protection scope of the claims of the present invention.

Claims

1. A voice segmentation method for an air traffic control voice recorder, characterized in that: The following steps are involved: S1. Acquire the mixed channel voice data and the control channel voice data frame by frame, each frame of voice data has a timestamp and a channel number, the control channel only has control voice data, the mixed channel includes the crew voice data and the control voice data in the control channel, and the control voice data and the crew voice data in the mixed channel are prohibited from existing at the same time; S2. Extract the voice data features of each frame of the mixed channel voice data and each frame of the control channel voice data in units of frames to obtain the mixed channel voice data features and the control channel voice data features; S3. Identify whether the mixed channel voice data with the same timestamp is controlled voice data frame by frame based on the characteristics of the controlled channel voice data. If the mixed channel voice data of the frame is controlled channel voice data, remove the mixed channel voice data of the corresponding frame in the mixed channel. If not, do not remove it and obtain the crew voice data in the mixed channel.

2. The voice segmentation method of an air traffic control voice recorder according to claim 1, characterized in that: In step S1, the mixed channel voice data and the controlled channel voice data output by the voice recorder are obtained frame by frame. The specific implementation method is: receiving ADPCM data packets frame by frame, the ADPCM data packets including the channel number, the voice data block and the timestamp, judging whether the voice data block is mixed channel voice data or controlled channel voice data according to the channel number, and converting the ADPCM voice data block into PCM format voice data.

3. The voice segmentation method of an air traffic control voice recorder according to claim 2, characterized in that: In step S2, the voice data features of each frame of mixed channel voice data and each frame of controlled channel voice data are extracted respectively in units of frames. The specific implementation thereof includes: calculating the signal strength of each frame of mixed channel voice data and each frame of controlled channel voice data, performing spectrum analysis on the voice data, and obtaining Mel-frequency cepstral coefficients.

4. The voice segmentation method of an air traffic control voice recorder according to claim 3, characterized in that: The calculation formula of the signal strength is: Wherein, I represents the signal strength, p represents the decimal number represented by the converted PCM format voice data item, and n / 2 represents the number of PCM format voice data items.

5. The voice segmentation method of an air traffic control voice recorder according to claim 3, characterized in that: The Mel-frequency cepstral coefficients are calculated by performing a fast Fourier transform on each frame of speech data to obtain a spectrum of each frame of speech data, converting the spectrum of each frame of speech data to a Mel-frequency domain to obtain a Mel-frequency spectrum of the speech data, and performing a discrete cosine transform on the Mel-frequency spectrum of each frame of speech data to obtain a Mel-frequency cepstral coefficient of each frame of speech data.

6. The voice segmentation method of an air traffic control voice recorder according to claim 3, characterized in that: Before step S3, the method further includes: determining whether the average signal strength of each frame of voice data exceeds a preset threshold; if the average signal strength of each frame of voice data does not exceed the preset threshold, the frame of voice data is a silent frame.

7. The method for voice segmentation of an air traffic control voice recorder according to claim 6, characterized in that: Before identifying whether the voice data in the mixed channel at the same timestamp is the controlled voice data, it also includes: at the same timestamp, when the controlled channel voice data is a silent frame, and the mixed channel voice data is not a silent frame, then the mixed channel voice data of this frame is the crew voice data, and when subsequently identifying whether the voice data in the mixed channel is the controlled voice data, this frame of voice data is skipped.

8. The method for voice segmentation of an air traffic control voice recorder according to claim 7, characterized in that: In step S3, whether the voice data in the mixed channel at the same timestamp is controlled voice data is identified frame by frame based on the characteristics of the controlled channel voice data. The specific implementation method is: compare the Mel-frequency cepstral coefficients of the controlled channel voice data and the mixed channel voice data at the same timestamp. When the Mel-frequency cepstral coefficients of the two are equal or similar, the mixed channel voice data of this frame is controlled voice data.

9. The method for voice segmentation of an air traffic control voice recorder according to claim 2, characterized in that: Receive ADPCM packets via UDP protocol.

10. A voice segmentation system for an air traffic control voice recorder, characterized in that: include: A voice receiving module is used to obtain the mixed channel voice data and the control channel voice data output by the voice recorder frame by frame. Each frame of voice data carries a timestamp and a channel number. The control channel only contains the control voice data. The mixed channel contains the crew voice data and the control voice data in the control channel. At the same time, the control voice data and the crew voice data in the mixed channel cannot exist at the same time. A feature extraction module is used to extract the voice data features of each frame of mixed channel voice data and each frame of controlled channel voice data in units of frames, so as to obtain the mixed channel voice data features and the controlled channel voice data features; The voice rejection module is used to identify whether the voice data in the mixed channel with the same timestamp is controlled voice data according to the voice data characteristics of the controlled channel frame by frame, and to reject the voice data of the corresponding frame in the mixed channel to obtain the crew voice data in the mixed channel.

Citation Information

Patent Citations

  • Voice recognition method and voice recognition device in air traffic control system

    CN101916565A

  • Training assessment method based on embedded automatic data acquisition technology

    CN102163380A