A method for identifying waypoints in air traffic control speech based on spatial position information

By building a content library and using a waypoint voice extraction model, the air-controlled voice data is sliced ​​and matched, which solves the problem of inaccurate recognition of a waypoint in the existing technology of hollow-controlled voice, and achieves a higher recognition accuracy.

CN116704821BActive Publication Date: 2025-05-13THE 28TH RES INST OF CHINA ELECTRONICS TECH GROUP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310556245.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-17
Publication Date
2025-05-13
Estimated Expiration
2043-05-17

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify waypoints in air traffic voice, resulting in inaccurate recognition results.

Method used

By building a content library, waypoint information is associated with sector data, and the target waypoint collection is determined using spatial position information. The waypoint speech extraction model is used to slice the air-bucket speech data, and combine the classifier and attention mechanism to achieve accurate identification of waypoints.

Benefits of technology

This method reduces the linear pattern of speech-text-correction, directly matches waypoints from speech, reduces the possibility of error transmission, and significantly improves the accuracy of waypoint recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116704821B_ABST
    Figure CN116704821B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for identifying waypoints in air traffic control speech based on spatial position information. The method uses the source of speech data as an auxiliary judgment condition to achieve accurate identification of waypoints in air traffic control speech data. First, a supervised learning method is used to train a waypoint speech extraction model to achieve partial extraction of waypoint speech in speech data, and then a waypoint recognition method based on spatial position information is designed to obtain accurate waypoint recognition results. Finally, the waypoint recognition results are spliced ​​into the original speech recognition results to achieve accurate recognition of the entire section of air traffic control speech containing waypoints. This method optimizes the recognition accuracy of the current air traffic control speech recognition algorithm, which is helpful for the development of subsequent air traffic control services, simulated flights, and other work.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a waypoint recognition method, in particular to a waypoint recognition method in air traffic control speech based on spatial position information. Background Art

[0002] In the current field of air traffic control speech recognition, the recognition of waypoints is one of the main difficulties. Nowadays, the requirements for speech recognition in simulation flight systems and air traffic control automation systems are constantly increasing, and the demand for automatic recognition of flight intentions is constantly increasing. The accurate recognition of waypoints has become a difficult problem that needs to be solved urgently in this field.

[0003] Currently, some researchers have proposed to achieve waypoint error correction in speech recognition results by means of text error correction. However, the recognition results of waypoints in speech recognition results may be greatly different from the correct results in text, so this method is often difficult to work.

[0004] Therefore, a relatively accurate end-to-end waypoint identification method is urgently needed. Summary of the invention

[0005] Purpose of the invention: The technical problem to be solved by the present invention is to provide a method for identifying waypoints in air traffic control speech based on spatial position information in view of the deficiencies in the prior art.

[0006] In order to solve the above technical problems, the present invention discloses a method for identifying waypoints in air traffic control speech based on spatial position information, comprising the following steps:

[0007] Step 1, build a content library, associate waypoint information and sector data, and store them in the content library;

[0008] In the content library, the associated data includes at least: sector ID, adjacent sector list, waypoint list within the sector and waypoint voice data.

[0009] Step 2, determine the location according to the source of the air traffic control voice data, and use the content library to obtain the target waypoint set, specifically including:

[0010] Step 2-1, determining the sector ID corresponding to the air traffic control voice data, the specific method includes:

[0011] Step 2-1-1: For air traffic control voice data transmitted by analog signals, directly use the signal frequency to determine its corresponding sector ID;

[0012] Step 2-1-2, for air traffic control voice data that uses data link to transmit signals, use the IP address to obtain the sector ID corresponding to the signal.

[0013] Step 2-2, constructing a target waypoint set: searching the content library for waypoints in the sector according to the sector ID, and using the waypoints obtained from the search as the target waypoint set.

[0014] The construction of the target waypoint set is to expand the scope of the target waypoint set to the adjacent sectors of the current sector according to the task requirements, that is, to search for an adjacent sector list in the content library according to the sector ID, and then search for the waypoints in the adjacent sector according to the information in the adjacent sector list, and also add the waypoints in the adjacent sector to the target waypoint set.

[0015] Step 3, construct a waypoint speech extraction model, slice the waypoint part of the air traffic control speech data, and obtain the sliced ​​waypoint speech;

[0016] The slicing of the waypoint part in the air traffic control voice data specifically includes:

[0017] Step 3-1, constructing a waypoint speech extraction model, wherein the waypoint speech extraction model includes: a waypoint speech encoder, an air traffic control speech recognition encoder, and a positioning module;

[0018] Wherein, the waypoint voice encoder is used to obtain the encoding of the waypoint voice data in the content library;

[0019] The air traffic control speech recognition encoder is used to obtain the phoneme posterior matrix of the waypoint speech data in the content library;

[0020] The positioning module uses the attention mechanism to slice the entire air traffic control speech, that is, according to the encoding of the waypoint speech data, the position and duration of the waypoint speech in the entire air traffic control speech are obtained in the phoneme posterior matrix of the waypoint speech data, that is, the waypoint speech after slicing is obtained;

[0021] The waypoint speech encoder is a pre-trained encoder.

[0022] The pre-trained data source includes waypoints and their corresponding audios.

[0023] Step 3-2, constructing a training data set for constructing a waypoint speech extraction model, the data set comprising: speech data, text data corresponding to the speech data, and annotations of the waypoint part included in the text data; and using the training data set to train the waypoint speech extraction model;

[0024] Step 3-3, using the trained waypoint speech extraction model, slicing the waypoint part of the air traffic control speech data to obtain the sliced ​​waypoint speech;

[0025] Step 3-4, extract the phoneme posterior matrix of the waypoint part from the phoneme posterior matrix of the entire air traffic control voice data according to the position and duration information obtained during slicing.

[0026] Step 4, matching the sliced ​​waypoint voice with the waypoints in the target waypoint set to obtain a matching result, specifically including:

[0027] Step 4-1, build a classifier, match the extracted phoneme posterior matrix to all waypoints in the content library, and obtain the matching probability of each waypoint;

[0028] Step 4-2, select the waypoint with the highest matching probability in the target waypoint set as the matching result, and set a threshold Δ. When the classification probability of all target waypoints is less than Δ, it is considered that the slicing result in step 3 is incorrect, and there is no waypoint voice in the current air traffic control voice data, that is, there is no matching result.

[0029] Step 5, matching result splicing: splice the matching result into the air traffic control speech recognition result given by the existing speech recognition model to realize the recognition information of the air traffic control speech including the waypoints.

[0030] The matching result splicing specifically includes:

[0031] Step 5-1, using the existing speech recognition model to perform speech recognition on the air traffic control speech data to obtain a preliminary recognition text;

[0032] Step 5-2, when there is a matching result in step 4-3, according to the position and duration obtained when slicing in step 3, in the preliminary recognition text, replace the text information of the corresponding position and length with the matching result in step 4-3 to obtain the final air traffic control speech recognition text, and complete the waypoint recognition in the air traffic control speech based on spatial position information.

[0033] Beneficial effects:

[0034] The present invention changes the linear mode of "speech-text-error correction" to a mode of directly matching waypoints from speech, thereby reducing an intermediate step, thereby reducing the possibility of error transmission and significantly improving the accuracy of waypoint recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.

[0036] Figure 1 A schematic diagram of the content library.

[0037] Figure 2The present invention is a flow chart of a method for identifying waypoints in air traffic control speech based on spatial position information.

[0038] Figure 3 Replace the flowchart with text. DETAILED DESCRIPTION

[0039] A method for identifying waypoints in air traffic control speech based on spatial position information comprises the following steps:

[0040] Step 1: Build a content library to associate waypoint information with sector data;

[0041] Step 2, determine the location according to the source of the voice data, and use the content library information to obtain the target waypoint set;

[0042] Step 3: construct a waypoint speech extraction model to accurately slice the waypoint part of the speech data;

[0043] Step 4, matching the sliced ​​waypoint voice with the waypoints in the target waypoint set to obtain a matching result;

[0044] Step 5, splicing the matching results into the air traffic control speech recognition results given by the speech recognition model to achieve accurate recognition of the air traffic control speech including waypoints.

[0045] Furthermore, step 2 includes:

[0046] Step 2-1: The current air traffic control voice data is transmitted using analog signals, and there is a fixed correspondence between its frequency and sector, so the corresponding sector can be directly determined using the signal frequency;

[0047] Step 2-2: In the future, air traffic control voice may use data link to transmit signals. At that time, the sector ID corresponding to the signal can also be accurately obtained by using the IP address;

[0048] Step 2-3, in the content library, according to the sector ID, obtain which waypoints are included in the sector, and these waypoints together constitute the target waypoint set. Depending on the mission requirements, the scope of the target waypoint set can be expanded to the surrounding sectors of the current sector;

[0049] Furthermore, step 3 includes:

[0050] Step 3-1, constructing a training data set for the speech extraction model, including speech data, its corresponding text data, and annotations of waypoints in the text data;

[0051] Step 3-2, construct a waypoint speech encoder, obtain the waypoint speech encoding in the content library by pre-training, and the pre-trained data source includes the waypoint and its corresponding audio;

[0052] Step 3-3, constructing an air traffic control speech recognition encoder to obtain a phoneme posterior matrix of the speech data;

[0053] Step 3-4, construct a positioning model, use the attention mechanism, and obtain the position and duration of the waypoint part in the entire speech in the phoneme posterior matrix through the waypoint speech encoding;

[0054] Furthermore, step 4 includes:

[0055] Step 4-1, extracting the phoneme posterior matrix of the waypoint part from the phoneme posterior matrix of the entire air traffic control speech data according to the position and duration information obtained in step 3;

[0056] Step 4-2, construct a classifier, classify the extracted phoneme posterior matrix into all waypoints in the content library, and obtain the classification probability of each waypoint;

[0057] Step 4-3, select the waypoint with the highest classification probability in the target waypoint set as the classification result, and set a threshold Δ. When the classification probability of all target waypoints is less than Δ, it is considered that the extraction result in step 3 is wrong, and there is no waypoint speech in the current segment, so there is no need to select the waypoint corresponding to the speech;

[0058] Furthermore, step 5 includes:

[0059] Step 5-1, using a speech recognition algorithm including a speech and text alignment method to perform speech recognition on the air traffic control speech data;

[0060] Step 5-2, when the classification probability of a waypoint in step 4-3 is greater than Δ, the text information at the corresponding position in the recognized text is deleted according to the position and duration of the speech part extracted in step 3, and the classification result in step 4-3 is filled in the deleted information part to obtain a complete speech recognition text.

[0061] The principle of the present invention is to use a waypoint speech extraction model to obtain waypoint speech slices, then use the speech data source to obtain the target waypoint set information, then match waypoints for the waypoint speech slices in the set, and finally splice the matching results into the output results of the speech recognition model to realize air traffic control speech recognition including waypoints.

[0062] Example:

[0063] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention is further described below in conjunction with the accompanying drawings and embodiments. Since the description of the embodiments is specific and cannot cover all embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making innovative work belong to the protection scope of the present invention.

[0064] Figure 1 It is a schematic diagram of the content library. The sector ID is the number of the recorded sector, such as ZSSSAP03. The list of adjacent control sectors is the set of sectors that share the same line segment with the sector on the spatial boundary. The list of waypoints in the sector represents all numbered waypoints within the spatial range of the sector ID, regardless of altitude. The waypoint voice data contains the pronunciation of all waypoints in the sector. It can be recorded by a single person, or multiple people can record the same waypoint and record all of them to increase the robustness of the model. The content library data is manually filled in after being exported from the Aeronautical Information Publication (AIP) data. When determining adjacent sectors, they must have common edges to be considered adjacent. Only common points are not considered adjacent. The granularity is the control sector.

[0065] Figure 2 This is a flow chart of a method for identifying waypoints in air traffic control speech based on spatial position information. The figure shows the process of speech recognition when receiving new air traffic control speech after the model is trained. It is mainly divided into four main parts: general speech recognition, waypoint segmentation, waypoint classification, and text replacement. During the training process, the general speech recognition model is trained separately and needs to be fed with additional labeled air traffic control speech and text data; the waypoint speech encoding and the waypoint speech slicing part based on the attention mechanism are jointly trained so that the waypoint speech encoding method can better adapt to the needs of the waypoint segmentation algorithm; the waypoint classification algorithm is re-trained separately, without using the waypoint speech encoding method in the previous step, and the classification model is rebuilt.

[0066] Figure 3 This is a flowchart of text replacement. The purpose of text replacement is to extract the part that needs waypoint recognition from the text generated by the recognition result of the initial speech recognition model, and replace it with the classification result of the subsequent waypoint classifier, so as to obtain better waypoint recognition accuracy, and then obtain better air traffic control speech recognition accuracy. As shown in the figure, the phoneme posterior matrix of the waypoint speech and its position information in the original speech are the inputs of the flowchart. Figure 2The speech recognition model based on the Transformer structure is input, and the probability distribution of the waypoint speech classification to each waypoint is obtained through the classification model, and the waypoints whose classification probability exceeds the threshold Δ are tried to be matched in the target waypoint set. If there is a waypoint whose classification probability given by the classification model is greater than Δ and exists in the target waypoint set, then the waypoint is considered to be the classification result of the waypoint speech; if no waypoint meets the conditions, it is considered that there is a problem with the previous waypoint speech segmentation, and the original speech recognition result will be directly output without replacing the waypoint part; if there are multiple waypoints that meet the conditions, the waypoints with high classification probability are given priority.

[0067] Taking a real ATC voice as an example, the conversation in the original voice is "Eastern 3984, turn left ADBAS". First, the frequency band of this voice broadcast will be recorded and found to be the broadcast frequency band corresponding to ZSSSAR15. So it is considered that this broadcast is a voice conversation for the aircraft in Sector 15 of Shanghai Control Sector, which may be said by the controller or the pilot. Secondly, this voice will be recognized into a text by a speech recognition model based on Transformer, such as "Eastern 3984, turn left AB", and the time axis of each character in the original voice will be marked. For example, the time axis corresponding to "Eastern" is from 0 seconds and 120 milliseconds to 0 seconds and 398 milliseconds, the time axis corresponding to "Fang" is from 0 seconds and 398 milliseconds to 0 seconds and 661 milliseconds, and so on. The time axes corresponding to A and B will also be marked. All these works can be obtained by the existing speech recognition models. Subsequently, the whole voice will be compared with the encoded results of all waypoint voices through the attention mechanism to obtain the possible position and duration of the waypoint voice segment in the whole voice. This step finally gets that there is a waypoint voice at 3 seconds and 317 milliseconds to 3 seconds and 890 milliseconds of this voice, and the phoneme posterior matrix of this part of the voice is extracted and given to the subsequent steps. Then, the trained classifier classifies this phoneme posterior matrix into the set composed of all waypoints, and obtains the classification probability value and its sorting result. ADBAS obtains a classification probability of 94%. Subsequently, the set of waypoints included in Sector 15 of Shanghai Control Sector is found in the content library, and it is found that ADBAS is within Sector 15 of Shanghai Control Sector. So ADBAS is the only waypoint that meets the conditions and needs to be replaced into the text given by the initial speech recognition model. Finally, by checking the text content at 3 seconds and 317 milliseconds to 3 seconds and 890 milliseconds where the waypoint voice is located, it is found that the time axis of "A" in the text "Eastern 3984, turn left AB" input to the model is from 3 seconds and 301 milliseconds to 3 seconds and 342 milliseconds, and the time axis of "B" is from 3 seconds and 342 milliseconds to 3 seconds and 864 milliseconds, and the coincidence degree of the time axes with the detected waypoint voice part exceeds 70%. So "AB" in the text is deleted and replaced with "ADBAS", and thus the whole voice recognition result is adjusted to "Eastern 3984, turn left ADBAS".

[0068] In specific implementation, the present application provides a computer storage medium and a corresponding data processing unit. Among them, the computer storage medium can store a computer program, and when the computer program is executed by the data processing unit, it can run the invention content of a method for identifying waypoints in ATC voice based on spatial position information provided by the present invention and some or all of the steps in each embodiment. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), etc.

[0069] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present invention can be implemented by means of computer programs and their corresponding general hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention are essentially or partly contributed to the prior art can be embodied in the form of a computer program, i.e., a software product, which can be stored in a storage medium and includes several instructions for enabling a device including a data processing unit (which can be a personal computer, a server, a single-chip microcomputer, a MUU or a network device, etc.) to execute the methods described in various embodiments of the present invention or certain parts of the embodiments.

[0070] The present invention provides a method and idea for identifying waypoints in air traffic control speech based on spatial position information. There are many methods and approaches to implement the technical solution. The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be considered as the protection scope of the present invention. All components not specified in this embodiment can be implemented using existing technologies.

Claims

1. A method for identifying waypoints in air traffic control speech based on spatial position information, characterized in that: The following steps are involved: Step 1, build a content library, associate waypoint information and sector data, and store them in the content library; Step 2, determine the location according to the source of the air traffic control voice data, and use the content library to obtain the target waypoint set; Step 3, construct a waypoint speech extraction model, slice the waypoint part of the air traffic control speech data, and obtain the sliced ​​waypoint speech; Step 4, matching the sliced ​​waypoint voice with the waypoints in the target waypoint set to obtain a matching result; Step 5, matching result splicing: splicing the matching result into the air traffic control speech recognition result given by the existing speech recognition model to realize the recognition information of the air traffic control speech including the waypoints; The matching result splicing described in step 5 specifically includes: Step 5-1, using the existing speech recognition model to perform speech recognition on the air traffic control speech data to obtain a preliminary recognition text; Step 5-2, when there is a matching result in step 4, according to the position and duration obtained when slicing in step 3, in the preliminary recognition text, replace the text information of the corresponding position and length with the matching result in step 4, obtain the final air traffic control speech recognition text, and complete the waypoint recognition in the air traffic control speech based on spatial position information.

2. The method for identifying waypoints in air traffic control speech based on spatial position information according to claim 1, characterized in that: In the content library described in step 1, the associated data includes at least: sector ID, adjacent sector list, waypoint list within the sector, and waypoint voice data.

3. The method for identifying waypoints in air traffic control speech based on spatial position information according to claim 2, characterized in that: Step 2 specifically includes: Step 2-1, determining the sector ID corresponding to the air traffic control voice data; Step 2-2, constructing a target waypoint set: searching the content library for waypoints in the sector according to the sector ID, and using the waypoints obtained from the search as the target waypoint set.

4. The method for identifying waypoints in air traffic control speech based on spatial position information according to claim 3, characterized in that: The specific method for determining the sector ID corresponding to the air traffic control voice data in step 2-1 includes: Step 2-1-1: For air traffic control voice data transmitted by analog signals, directly use the signal frequency to determine its corresponding sector ID; Step 2-1-2, for air traffic control voice data that uses data link to transmit signals, use the IP address to obtain the sector ID corresponding to the signal.

5. The method for identifying waypoints in air traffic control speech based on spatial position information according to claim 4, characterized in that: The construction of the target waypoint set described in step 2-2, that is, according to the task requirements, the scope of the target waypoint set is expanded to the adjacent sectors of the current sector, that is, the adjacent sector list is searched in the content library according to the sector ID, and then the waypoints in the adjacent sector are searched according to the information in the adjacent sector list, and the waypoints in the adjacent sector are also added to the target waypoint set.

6. The method for identifying waypoints in air traffic control speech based on spatial position information according to claim 5, characterized in that: The slicing of the waypoint part in the air traffic control voice data described in step 3 specifically includes: Step 3-1, constructing a waypoint speech extraction model, wherein the waypoint speech extraction model includes: a waypoint speech encoder, an air traffic control speech recognition encoder, and a positioning module; Wherein, the waypoint voice encoder is used to obtain the encoding of the waypoint voice data in the content library; The air traffic control speech recognition encoder is used to obtain the phoneme posterior matrix of the waypoint speech data in the content library; The positioning module uses the attention mechanism to slice the entire air traffic control speech, that is, according to the encoding of the waypoint speech data, the position and duration of the waypoint speech in the entire air traffic control speech are obtained in the phoneme posterior matrix of the waypoint speech data, that is, the waypoint speech after slicing is obtained; Step 3-2, constructing a training data set for constructing a waypoint speech extraction model, the data set comprising: speech data, text data corresponding to the speech data, and annotations of the waypoint part included in the text data; and using the training data set to train the waypoint speech extraction model; Step 3-3, using the trained waypoint speech extraction model, slicing the waypoint part of the air traffic control speech data to obtain the sliced ​​waypoint speech; Step 3-4, extract the phoneme posterior matrix of the waypoint part from the phoneme posterior matrix of the entire air traffic control voice data according to the position and duration information obtained during slicing.

7. The method for identifying waypoints in air traffic control speech based on spatial position information according to claim 6, characterized in that: The step 4 of matching the sliced ​​waypoint voice with the waypoints in the target waypoint set specifically includes: Step 4-1, build a classifier, match the extracted phoneme posterior matrix to all waypoints in the content library, and obtain the matching probability of each waypoint; Step 4-2, select the waypoint with the highest matching probability in the target waypoint set as the matching result, and set a threshold Δ. When the classification probability of all target waypoints is less than Δ, it is considered that the slicing result in step 3 is incorrect, and there is no waypoint voice in the current air traffic control voice data, that is, there is no matching result.

8. The method for identifying waypoints in air traffic control speech based on spatial position information according to claim 7, characterized in that: The waypoint speech encoder described in step 3-1 is a pre-trained encoder.

9. The method for identifying waypoints in air traffic control speech based on spatial position information according to claim 8, characterized in that: The pre-trained data source described in step 3-1 includes waypoints and their corresponding audios.

Citation Information

Patent Citations

  • Aircraft and instrumentation system for voice transcription of radio communications

    CN106716523A

  • Chinese civil aviation air traffic control speech recognition method and system

    CN113160798A