Air-ground communication sound source identity identification method, air traffic control system and readable storage medium

By acquiring the RSSI value of the VHF receiver, using the changes in the RSSI value to determine the start and end times of the call, and calculating the arithmetic mean, the problems of low accuracy and insufficient real-time performance in air-to-ground communication command acquisition and speaker differentiation are solved, achieving efficient, accurate, and real-time air-to-ground communication command acquisition and speaker differentiation.

CN121459825BActive Publication Date: 2026-07-24中国民用航空珠海进近管制中心
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
中国民用航空珠海进近管制中心
Filing Date
2025-09-19
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

In the field of civil aviation air traffic control, the accuracy of collecting, segmenting and distinguishing the caller in air-to-ground communication commands is low. In particular, it is difficult to achieve effective collection and segmentation when commands overlap and are continuous during busy periods. Moreover, existing methods are costly and lack real-time performance, making it difficult to meet the needs of air traffic control operations.

Method used

By acquiring the RSSI value of the VHF receiver at the VHF station, the start and end times of the call are determined by the changes in the RSSI value, and the arithmetic mean is calculated. Combined with a threshold, the speaker's identity is determined, thus realizing the real-time acquisition, segmentation, and speaker differentiation of air-to-ground communication commands. The existing civil aviation VHF receiver data interface does not need to be modified.

Benefits of technology

It achieves efficient, accurate, and real-time collection and speaker differentiation of air-to-ground communication commands, and has the advantages of high security, high accuracy, and good real-time performance. Moreover, it does not require complex neural networks and large-scale labeled data, resulting in low deployment costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121459825B_ABST
    Figure CN121459825B_ABST
Patent Text Reader

Abstract

The application provides a land-air communication sound source identity identification method, an air traffic control system and a readable storage medium. The land-air communication sound source identity identification method comprises the following steps: acquiring an RSSI value corresponding to a very high frequency receiver of a very high frequency station according to a preset acquisition frequency; judging whether the RSSI value is greater than a first threshold value, and if yes, recording a current time stamp as a start time stamp; after recording the start time stamp, judging whether the RSSI value is less than or equal to the first threshold value, and if yes, recording a current time stamp as an end time stamp; calculating an arithmetic mean value of all RSSI values collected between the start time stamp and the end time stamp; and if the arithmetic mean value is greater than or equal to a second threshold value, determining that the sound source identity is a controller, and if the arithmetic mean value is greater than the first threshold value and less than the second threshold value, determining that the sound source identity is a pilot. The application further provides an air traffic control system and a readable storage medium applying the method. The method can efficiently, accurately, reliably and conveniently realize land-air communication instruction acquisition, segmentation and speaker differentiation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of civil aviation air traffic control technology, specifically to a method for identifying the source of a ground-to-air communication voice source, an air traffic control system that applies the method, and a readable storage medium that applies the method. Background Technology

[0002] In the field of civil aviation air traffic control, air traffic control recorders can output streaming air-to-ground communication audio data. The traditional approach to extract each air-to-ground communication instruction from this streaming audio data relies on the silent intervals between instructions for segmentation. However, in actual air traffic control production scenarios, this method has a low recognition accuracy because air-to-ground communication is fast-paced, and instructions are often transmitted continuously. The silent intervals between different instructions are either short or long and unstable. In addition, when the communication channel is busy, multiple instructions often overlap and occur consecutively. Relying solely on silent intervals is no longer sufficient for effective acquisition and segmentation. Moreover, this type of streaming air-to-ground communication audio data itself cannot directly distinguish the speaker's identity in the air-to-ground communication instruction.

[0003] To address the shortcomings in the aforementioned air-to-ground communication command collection, segmentation, and speaker identification, the following additional methods are currently available:

[0004] The first approach relies on artificial intelligence models to collect, segment, and distinguish the speaker in air-to-ground communication commands. This type of technology introduces AI models such as speech recognition and semantic understanding to achieve automatic command collection and intelligent segmentation through deep analysis of audio content. While this technology theoretically can accurately capture command boundaries by modeling the semantic logic of speech signals using algorithms, it requires building complex neural network architectures and depends heavily on large-scale, high-quality labeled data for training, making it difficult and costly, thus limiting the efficiency of its practical application.

[0005] The second approach involves collecting, segmenting, and differentiating speakers for air-to-ground communication commands based on speech features and speaker recognition. This technology learns and extracts voice features to build a speaker recognition model, enabling the differentiation of speakers for air-to-ground communication commands. Its core logic is to extract unique voiceprint features, intonation patterns, and other biometric features from the audio to establish an individual voice model for identity matching. However, in actual air traffic control environments, this technology still has significant shortcomings: First, accuracy is insufficient. Especially when there is noise interference, differences in accents among different speakers, and fluctuations in speech rate in the air-to-ground communication audio, the model's extraction of voice features may deviate, leading to misjudgment of the speaker for air-to-ground communication commands and a decrease in speaker differentiation accuracy, failing to meet the high requirements of air traffic control operations. Second, real-time performance is lagging. Due to the complexity of feature extraction algorithms or the large computational load of the model, this technology struggles to respond quickly during real-time air-to-ground communication, resulting in a significant delay in the speaker differentiation results and failing to meet real-time operational requirements. Summary of the Invention

[0006] To address the aforementioned problems, the primary objective of this invention is to provide a method for identifying the source of a ground-to-air communication voice source that is efficient, accurate, reliable, and easy to implement for collecting, segmenting, and distinguishing the speaker in ground-to-air communication commands.

[0007] The second objective of this invention is to provide an air traffic control system that implements the above-described method for identifying the source of a ground-to-air communication voice.

[0008] A third objective of this invention is to provide a readable storage medium that applies the above-described method for identifying the source of a land-to-air communication voice.

[0009] To achieve the main objective of this invention, the present invention provides a method for identifying the source of a ground-to-air communication voice source, comprising: acquiring the RSSI value corresponding to the VHF receiver of a VHF station according to a preset acquisition frequency; determining whether the RSSI value is greater than a first threshold, and if so, recording the current timestamp as the start timestamp; after recording the start timestamp, determining whether the RSSI value is less than or equal to the first threshold, and if so, recording the current timestamp as the end timestamp; calculating the arithmetic mean of all RSSI values ​​acquired between the start timestamp and the end timestamp; if the arithmetic mean is greater than or equal to a second threshold, determining the voice source as an air traffic controller, and if the arithmetic mean is greater than the first threshold and less than the second threshold, determining the voice source as a pilot.

[0010] As can be seen from the above, by collecting the RSSI value corresponding to the VHF receiver, the functions of collecting, segmenting, and distinguishing the speaker of civil aviation air traffic control ground-to-air communication instructions can be realized based on the changes in the RSSI value without modifying or affecting other air traffic control equipment. This not only ensures the real-time and accurate collection and segmentation of ground-to-air communication instructions, but also improves the accuracy of speaker distinction, realizing the integrated function of collecting, segmenting, and distinguishing the speaker of ground-to-air communication instructions. It also has the advantages of high security, high accuracy, and good real-time performance.

[0011] A further approach is to determine whether the arithmetic mean of the VHF receivers of at least one VHF station is greater than or equal to the second threshold if there is at least one VHF station. If so, the source of the sound is identified as an air traffic controller. If the arithmetic mean of the VHF receivers of any VHF station is greater than the first threshold and less than the second threshold, the source of the sound is identified as a pilot.

[0012] As can be seen above, an air traffic control unit typically has two or more VHF stations. When an air traffic controller makes a call, the arithmetic mean of the RSSI of at least one VHF receiver will always be greater than or equal to the second threshold (depending on which VHF transmitter the controller chooses to use to issue the instruction). When a pilot makes a call, the arithmetic mean of the RSSI of all VHF receivers at all VHF stations will be greater than the first threshold and less than the second threshold. Utilizing this characteristic, the RSSI value data can be used to accurately and quickly distinguish the speaker of civil aviation air traffic control air-to-ground communication instructions.

[0013] A preferred embodiment is that the air-to-ground communication voice source identification method further includes: saving the audio data between the start timestamp and the end timestamp as a first audio file; saving the entire air-to-ground communication audio data as a second audio file; saving all RSSI values ​​collected between the start timestamp and the end timestamp; and saving the start timestamp and the end timestamp.

[0014] As can be seen from the above, this design makes the process and results of data collection, segmentation and speaker differentiation traceable, and facilitates the review and viewing of relevant data information. It also enables the archiving of air-to-ground communication content.

[0015] A further proposed solution is that the start timestamp includes a first time node corresponding to the second audio file, and / or a second time node corresponding to the time zone; the end timestamp includes a third time node corresponding to the second audio file, and / or a fourth time node corresponding to the time zone.

[0016] As can be seen from the above, this design facilitates the precise retrieval of a specific air-to-ground communication command within the second audio file; and / or helps users understand the time of occurrence of the air-to-ground communication command in its respective time zone.

[0017] Another preferred approach is that the step of obtaining the RSSI value corresponding to the VHF receiver of the VHF station according to the preset acquisition frequency includes: obtaining the voltage value of a specific pin of the VHF receiver; and converting the voltage value into an RSSI value according to a preset relationship table.

[0018] As can be seen from the above, by collecting the voltage value of the received signal strength of a specific pin of the VHF receiver and converting it into the corresponding RSSI value, the acquisition of RSSI value can be made more convenient and faster. At the same time, it does not require modification of VHF receivers, VHF stations and other air traffic control equipment, and will not affect the normal operation of air traffic control equipment. It can achieve efficient, accurate, timely, reliable and fast segmentation of air-to-ground communication commands and differentiation of the speaker.

[0019] Another preferred option is to preset the sampling frequency to between 150 milliseconds and 250 milliseconds.

[0020] As can be seen from the above, by designing the preset collection frequency, it is possible to ensure the accurate collection, segmentation, and speaker differentiation of air-to-ground communication commands, avoiding omissions, miscollections, incorrect segmentation of command statements, and incorrect speaker identification.

[0021] Another preferred option is to have a first threshold of -100dBm and a second threshold of -30dBm.

[0022] As can be seen from the above, by designing the first threshold and the second threshold, it is possible to more accurately determine whether the current state is idle, whether the controller is the current speaker, or whether the pilot is the current speaker.

[0023] A further proposed solution is that the air-to-ground communication voice source identification method also includes: displaying a curve of the changes in all RSSI values ​​collected between the start and end timestamps; if the arithmetic mean is greater than a first threshold and less than a second threshold, controlling the change curve corresponding to the RSSI value to be displayed in a first color; if the arithmetic mean is greater than or equal to the second threshold, controlling the change curve corresponding to the RSSI value to be displayed in a second color.

[0024] As can be seen from the above, the design allows users to quickly know whether the curve corresponding to the current RSSI value is in an idle state, or whether the controller or the pilot is the current speaker when checking.

[0025] To achieve the second objective of the present invention, the present invention provides an air traffic control system including a VHF station, a control center, and an aircraft. The control center and the aircraft establish communication through the VHF station. The control center includes a processor and a memory, wherein the memory stores a computer program. When the computer program is executed by the processor, it implements the various steps of the above-mentioned air-to-ground communication voice source identification method.

[0026] To achieve the third objective of the present invention, the present invention provides a readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the various steps of the above-described method for identifying the source of a land-to-air communication voice. Attached Figure Description

[0027] Figure 1 This is a schematic diagram of the communication interface of a VHF receiver in an embodiment of the air-to-ground communication voice source identification method of the present invention.

[0028] Figure 2 This is a flowchart of an embodiment of the air-to-ground communication voice source identification method of the present invention.

[0029] The present invention will be further described below with reference to the accompanying drawings and embodiments. Detailed Implementation

[0030] The method for identifying the source of air-to-ground communication provided by this invention aims to provide an efficient, accurate, reliable and easy-to-implement solution. By collecting RSSI (Received Signal Strength Indicator) data through the interface of existing civil aviation VHF receivers, it is possible to achieve real-time and accurate collection, segmentation and speaker differentiation of air-to-ground communication commands in a non-intrusive manner.

[0031] Implementation Example of a Method for Identifying Voice Sources in Air-to-Ground Communication

[0032] In the civil aviation VHF communication system of the civil air traffic control system, communication with aircraft is carried out through VHF stations, which include VHF transmitters and VHF receivers.

[0033] The air-to-ground communication voice source identification method of this embodiment is applied to the control center of the air traffic control system. The following is a combination of... Figure 1 and Figure 2 This paper introduces a method for identifying the source of voice communication in air-to-ground communication.

[0034] First, step S1 is executed to obtain the RSSI value corresponding to the VHF receiver of the VHF station according to the preset acquisition frequency. Specifically, the step of obtaining the RSSI value corresponding to the VHF receiver of the VHF station according to the preset acquisition frequency includes: obtaining the voltage value of a specific pin of the VHF receiver; and converting the voltage value into an RSSI value according to a preset relationship table.

[0035] like Figure 1As shown, the VHF receiver has a communication interface 100. This embodiment uses the DB15 interface [D-Sub (D-subminiature)] as an example. The DB15 interface has 15 pins distributed in two rows, and its shape is "D". Among them, pin 1 is the ground pin; pin 12 can be used to receive the voltage value of the signal strength, with a voltage range of 0V to 10V. Therefore, the RSSI value can be calculated through this pin.

[0036] When a VHF receiver receives a radio signal, it calculates the RSSI value based on the voltage value of the received signal strength at pin 12. As mentioned earlier, a pre-set correspondence table between RSSI values ​​(dBm, dB milliwatt) and the voltage value (V) of pin 12 on the VHF receiver's DB15 interface can be used to convert the real-time voltage value of pin 12 into an RSSI value. For example, when the voltage value is 1V, the corresponding RSSI value is -110dBm; when the voltage value is 1.75V, the corresponding RSSI value is -100dBm; when the voltage value is 2.5V, the corresponding RSSI value is -90dBm; ...; when the voltage value is 9.25V, the corresponding RSSI value is 0dBm; and when the voltage value is 10V, the corresponding RSSI value is 10dBm. It is understood that the correspondence between voltage values ​​and RSSI values ​​may be adjusted adaptively due to factors such as the VHF receiver model and the operating environment. The above correspondence is only provided as an example for illustrative purposes.

[0037] Furthermore, the aforementioned preset sampling frequency is preferably between 150 milliseconds and 250 milliseconds; more preferably, the preset sampling frequency is 200 milliseconds, that is, the voltage value of pin 12 of the DB15 interface is sampled every 200 milliseconds and converted into the corresponding RSSI value. This design can ensure accurate sampling, segmentation, and speaker differentiation of air-to-ground communication commands, avoiding omissions, missampling, incorrect segmentation of command statements, and incorrect speaker identification.

[0038] Next, step S2 is executed to determine whether the RSSI value is greater than the first threshold. If the collected RSSI value is greater than the first threshold, it indicates that someone is speaking in the current frequency band. Preferably, the first threshold is -100dBm. It is understood that in other embodiments, the specific value of the first threshold may be adaptively adjusted according to factors such as the VHF receiver model and the usage environment. The specific value of the first threshold mentioned above is only used as an example for auxiliary illustration.

[0039] Since the RSSI value of a VHF receiver generally ranges from -120dBm to 10dBm, by comparing the RSSI values ​​of the signals, the strength of the signals can be determined, thereby identifying and distinguishing the speaker of different control instructions. Typically, when the real-time RSSI value of a certain VHF frequency for air-to-ground communication is greater than -100dBm, it can be determined that someone is speaking on that VHF frequency channel.

[0040] When the collected RSSI value is less than the first threshold, it indicates that no one is speaking in the VHF channel at the current time point, so return to step S1.

[0041] When the collected RSSI value is greater than the first threshold, it indicates that someone is speaking in the VHF frequency channel at the current time point. Therefore, step S3 is executed to record the current timestamp as the start timestamp.

[0042] After recording the start timestamp, step S4 is executed to continue acquiring the RSSI value corresponding to the VHF receiver according to the preset acquisition frequency; the RSSI value acquired again is used to determine whether the speech has ended in the VHF frequency channel at the current time node.

[0043] Next, step S5 is executed to determine whether the RSSI value is less than or equal to the first threshold. If the RSSI value is greater than the first threshold, it means that someone is still speaking in the VHF channel at the current time point, and step S4 is executed. If the RSSI value is less than or equal to the first threshold, it means that speaking has stopped in the VHF channel at the current time point. At this time, step S6 is executed to record the current timestamp as the end timestamp.

[0044] After recording the end timestamp, step S7 is executed to collect and segment the streaming air-to-ground communication audio data using the start and end timestamps, resulting in a first audio file for an air-to-ground communication command.

[0045] The audio data between the start timestamp and the end timestamp is saved as the first audio file (i.e., the aforementioned first audio file for air-to-ground communication instructions). By designing the preset acquisition frequency, one air-to-ground communication instruction can be implemented as one first audio file. Of course, by designing the preset acquisition frequency, one first audio file can also contain multiple consecutive air-to-ground communication instructions.

[0046] In addition, the entire air-to-ground communication audio data is saved as a second audio file, which means that all air-to-ground communication audio data within a certain time period is saved as a second audio file.

[0047] Furthermore, it also saves all RSSI values ​​collected between the start and end timestamps for later review and calculation of the arithmetic mean of the RSSI values.

[0048] In addition, start timestamps and end timestamps are also saved. The start timestamp includes the first time node corresponding to the second audio file, and the end timestamp includes the third time node corresponding to the second audio file, facilitating precise retrieval of a specific air-to-ground communication command within the second audio file; and / or the start timestamp also includes the second time node corresponding to its time zone, and the end timestamp also includes the fourth time node corresponding to its time zone, thus allowing users to understand when the air-to-ground communication command appeared in their respective time zone.

[0049] The above design makes the process and results of data collection, segmentation and speaker identification traceable, and facilitates the review and viewing of relevant data information. It also enables the archiving of air-to-ground communication content.

[0050] After obtaining the first audio file based on the start and end timestamps, step S8 is executed to calculate the arithmetic mean of all RSSI values ​​collected between the start and end timestamps.

[0051] After obtaining the arithmetic mean, step S9 is executed to determine whether the arithmetic mean is greater than or equal to the second threshold. Preferably, the second threshold is -30dBm. It is understood that in other embodiments, the specific value of the second threshold may be adaptively adjusted according to factors such as the VHF receiver model and the usage environment. The specific value of the second threshold mentioned above is only used as an example for auxiliary illustration.

[0052] If the arithmetic mean is greater than or equal to the second threshold, proceed to step S10 to determine that the sound source is an air traffic controller. If the arithmetic mean is less than the second threshold, proceed to step S11 to determine that the sound source is a pilot. It can be understood that, since the end timestamp is not recorded after the start timestamp is recorded, all RSSI values ​​collected before the end timestamp are greater than the first threshold, the arithmetic mean of all RSSI values ​​collected between the start and end timestamps is also greater than the first threshold. That is, if the arithmetic mean is greater than the first threshold and less than the second threshold, proceed to step S11 to determine that the sound source is a pilot.

[0053] Here, examples are given in the following cases:

[0054] First, within the same VHF station, because the transmitting antenna of the VHF transmitter and the receiving antenna of the receiver are generally close together, the average RSSI value of the VHF receiver will be greater than or equal to the second threshold during the duration of the controller's communication. However, aircraft are generally farther from the receiving antennas of the VHF receivers at each VHF station. Therefore, during the duration of the pilot's communication, the arithmetic mean of all RSSI values ​​collected by the VHF receiver will fall within the range of being greater than the first threshold and less than the second threshold.

[0055] The first scenario occurs when there is only one VHF receiver on a VHF frequency for air-to-ground communication, meaning that there is only one VHF station or only one VHF station in use on that VHF frequency:

[0056] Assuming the current RSSI value collected on the VHF frequency for the air-to-ground communication is -12 dBm, it indicates that someone is speaking on the current VHF channel. Therefore, the current timestamp (i.e., the current time point) is recorded as the start timestamp. When the communication ends, the collected RSSI value is less than -100 dBm; therefore, the current timestamp is recorded as the end timestamp. Subsequently, using the saved first audio file, the start timestamp, the end timestamp, and all RSSI values ​​collected from the start timestamp to the end timestamp, the arithmetic mean of all RSSI values ​​collected from the start timestamp to the end timestamp is calculated. Based on the above, if this arithmetic mean is greater than or equal to the second threshold (e.g., the arithmetic mean ≥ -30 dBm), it can be determined that this air-to-ground communication was initiated by air traffic controllers.

[0057] Assuming the current RSSI value collected on the VHF frequency for the air-to-ground communication is -72 dBm, it indicates that someone is speaking on the current VHF channel. Therefore, the current timestamp (i.e., the current time point) is recorded as the start timestamp. When the communication ends, the collected RSSI value is less than -100 dBm; therefore, the current timestamp is recorded as the end timestamp. Subsequently, using the saved first audio file, the start timestamp, the end timestamp, and all RSSI values ​​collected from the start timestamp to the end timestamp, the arithmetic mean of all RSSI values ​​collected from the start timestamp to the end timestamp is calculated. Based on the above, if this arithmetic mean is greater than a first threshold and less than a second threshold (e.g., -100 dBm < arithmetic mean < -30 dBm), it can be determined that this air-to-ground communication was made by a pilot.

[0058] The second scenario occurs when a VHF frequency for air-to-ground communication includes two or more VHF receivers, meaning that two or more VHF stations are in use on that VHF frequency.

[0059] Controllers can choose VHF transmitters from different stations to issue control instructions. Assuming an air traffic control unit has two VHF stations, A and B, when a controller uses the VHF transmitter from station A to issue a control instruction, based on the above, the arithmetic mean of the RSSI values ​​of the VHF receivers at station A will be greater than or equal to the second threshold during the duration of the controller's instruction. However, the receiving antenna of the VHF receiver at station B is farther from the transmitting antenna of the VHF transmitter at station A. Therefore, during the duration of the controller's instruction, the arithmetic mean of the RSSI values ​​of the VHF receivers at station B will fall within the range of being greater than the first threshold and less than the second threshold.

[0060] Based on this, when the pilot speaks, during the duration of the pilot's speech, the arithmetic mean of the RSSI values ​​of the VHF receivers at stations A and B will be within the range of being greater than the first threshold and less than the second threshold.

[0061] Specifically, an air traffic control unit typically has two or more VHF stations. When an air traffic controller makes a call, the arithmetic mean of the RSSI values ​​from at least one VHF receiver will always be greater than or equal to a second threshold (depending on which station's VHF transmitter the controller chooses to use to issue the instruction). When a pilot makes a call, the arithmetic mean of the RSSI values ​​from all VHF receivers will fall within the range of being greater than a first threshold and less than a second threshold. Therefore, utilizing this characteristic, RSSI value data can be used to accurately and quickly distinguish the sender of civil aviation air traffic control air-to-ground communication instructions.

[0062] Understandably, when an air traffic control unit includes two or more VHF stations, the execution of step S9 changes accordingly: determine whether the arithmetic mean of the VHF receivers of at least one VHF station is greater than or equal to the second threshold. If so, proceed to step S10 to determine that the sound source is an air traffic controller. If the arithmetic mean of the VHF receivers of any VHF station is within the range of being greater than the first threshold and less than the second threshold, proceed to step S11 to determine that the sound source is a pilot.

[0063] Furthermore, if the RSSI value of the VHF receiver at the VHF station, obtained according to the preset acquisition frequency, is less than the first threshold, it proves that the current frequency band is idle.

[0064] Furthermore, the air-to-ground communication voice source identification method also includes: displaying a curve showing the changes in all RSSI values ​​collected between the start and end timestamps. For example, the curve showing the changes in all RSSI values ​​collected between the start and end timestamps can be displayed on an external display screen. When displayed:

[0065] If the arithmetic mean is greater than a first threshold and less than a second threshold, the curve corresponding to the RSSI value is displayed in the first color; if the arithmetic mean is greater than or equal to the second threshold, the curve corresponding to the RSSI value is displayed in the second color. This design allows users to quickly determine whether the curve corresponding to the current RSSI value is in an idle state, or whether the controller or pilot is the current speaker.

[0066] For example, when the arithmetic mean is greater than the first threshold and less than the second threshold, the change curve corresponding to the RSSI value can be displayed in blue; when the arithmetic mean is greater than or equal to the second threshold, the change curve corresponding to the RSSI value can be displayed in red.

[0067] Furthermore, when the RSSI value of the VHF receiver of the VHF station, obtained according to the preset acquisition frequency, is continuously less than the first threshold, the change curve of the RSSI value within this duration can be displayed in a third color, such as black.

[0068] In summary, compared to existing technologies in the civil aviation air traffic control industry that rely on artificial intelligence models or speech features and speaker recognition for air-to-ground communication data acquisition, segmentation, and speaker differentiation, the air-to-ground communication source identification method provided by this invention acquires RSSI data without the need for complex neural networks or large-scale speech annotation data. It can be implemented directly using the data interface provided by existing civil aviation VHF receivers, offering advantages such as zero hardware intrusion, low deployment cost, and high security. Furthermore, this invention utilizes RSSI data, employing a VHF receiver to receive air-to-ground communication signals. It calculates the arithmetic mean RSSI value of a communication segment based on its start and end times, and then distinguishes between the air traffic controller and the pilot based on the magnitude of the arithmetic mean RSSI value. This allows for rapid and accurate acquisition and segmentation of air-to-ground communication instructions with simple logic and minimal computational load, significantly outperforming existing solutions that require high computing power. It also avoids interference from accents, noise, and speech rate, resulting in higher speaker differentiation accuracy. Furthermore, by acquiring RSSI values ​​from civil aviation VHF receivers at preset acquisition frequencies and combining them with streaming air-to-ground communication audio data output by air traffic control recorders, the system automatically acquires air-to-ground communication data based on changes in RSSI values. This allows the streaming air-to-ground communication audio data to be segmented into audio files containing only one air-to-ground communication instruction. Then, it determines whether the speaker of each air-to-ground communication instruction is an air traffic controller or a pilot. This achieves non-intrusive, accurate acquisition, segmentation, and speaker differentiation of air-to-ground communication instructions. It boasts strong real-time performance, high accuracy, and high security, making it suitable for scenarios such as civil aviation air traffic control instruction monitoring and voice recognition preprocessing.

[0069] Air Traffic Control System Implementation Examples

[0070] The air traffic control system includes VHF stations, a control center, and aircraft. The control center and aircraft establish communication through VHF stations. The control center includes a processor and a memory. The memory stores a computer program, which is executed by the processor to implement the various steps of the above-mentioned air-to-ground communication voice source identification method.

[0071] The control center may include, but is not limited to, processors and memory. Those skilled in the art will understand that the control center may include more or fewer components, or combinations of certain components, or different components; for example, the control center may also include input / output devices, network access devices, buses, etc.

[0072] The controller can be a Central Processing Unit (CPU), or other general-purpose controllers, Digital Signal Processors (DSPs), Application-Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose controller can be a microcontroller or any conventional controller. The controller is the control center of the control center, connecting all parts of the control center through various interfaces and lines.

[0073] The memory can be used to store computer programs and / or modules. The controller implements various functions of the control center by running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory. For example, the memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound receiving function, sound to text conversion function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, text data, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0074] Readable storage medium embodiments

[0075] The air traffic control system ground-to-air communication source identification method described in the above embodiments can be stored as a computer program in a computer-readable storage medium. When the computer program is executed, it can complete the steps of the air traffic control system ground-to-air communication source identification method described above. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to: electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0076] Finally, it should be emphasized that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for identifying the source of a voice source in air-to-ground communication, characterized in that, include: The method involves acquiring the RSSI value corresponding to the VHF receiver of the VHF station according to a preset acquisition frequency, including acquiring the voltage value of a specific pin of the VHF receiver and converting the voltage value into the RSSI value according to a preset relationship table. Determine whether the RSSI value is greater than the first threshold. If so, record the current timestamp as the start timestamp. After recording the start timestamp, determine whether the RSSI value is less than or equal to the first threshold. If so, record the current timestamp as the end timestamp. Calculate the arithmetic mean of all RSSI values ​​collected between the start timestamp and the end timestamp; If the arithmetic mean is greater than or equal to the second threshold, the sound source is identified as an air traffic controller; if the arithmetic mean is greater than the first threshold and less than the second threshold, the sound source is identified as a pilot.

2. The method for identifying the source of a ground-to-air communication voice source according to claim 1, characterized in that: If there are two or more VHF stations, determine whether the arithmetic mean of the VHF receivers of at least one VHF station is greater than or equal to the second threshold. If so, determine that the sound source is an air traffic controller. If the arithmetic mean of the VHF receivers of any of the VHF stations is greater than the first threshold and less than the second threshold, then the sound source is identified as a pilot.

3. The method for identifying the source of a ground-to-air communication voice source according to claim 2, characterized in that: The method for identifying the source of a ground-to-air communication voice source also includes: Save the audio data between the start timestamp and the end timestamp as a first audio file; Save the entire air-to-ground communication audio data as a second audio file; Save all RSSI values ​​collected between the start timestamp and the end timestamp; Save the start timestamp and the end timestamp.

4. The method for identifying the source of a ground-to-air communication voice source according to claim 3, characterized in that: The start timestamp includes its first time node corresponding to the second audio file, and / or its second time node corresponding to the time zone to which it belongs; The end timestamp includes its corresponding third time node to the second audio file, and / or its corresponding fourth time node to the time zone.

5. The method for identifying the source of a ground-to-air communication voice source according to claim 2, characterized in that: The preset acquisition frequency is between 150 milliseconds and 250 milliseconds.

6. The method for identifying the source of a ground-to-air communication voice source according to claim 2, characterized in that: The first threshold is -100dBm; The second threshold is -30dBm.

7. The method for identifying the source of a ground-to-air communication voice source according to any one of claims 1 to 6, characterized in that: The method for identifying the source of a ground-to-air communication voice source also includes: Display a graph showing the change in all RSSI values ​​collected between the start timestamp and the end timestamp; If the arithmetic mean is greater than the first threshold and less than the second threshold, the change curve corresponding to the RSSI value is displayed in the first color. If the arithmetic mean is greater than or equal to the second threshold, the change curve corresponding to the RSSI value is displayed in a second color.

8. An air traffic control system, comprising a VHF station, a control center, and an aircraft, wherein the control center and the aircraft establish communication through the VHF station, and the control center includes a processor and a memory, characterized in that: The memory stores a computer program that, when executed by the processor, implements the steps of the air-to-ground communication voice source identification method as described in any one of claims 1 to 7.

9. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the various steps of the air-to-ground communication voice source identification method as described in any one of claims 1 to 7.