A method and device for identifying speaker roles in land-air conversations based on feature fusion
By adopting feature fusion method in land-air call scenarios and combining the multimodal speaker role recognition model of deep neural networks, the problem of speaker role recognition in land-air call is solved, and the recognition accuracy and communication efficiency of air traffic management are improved.
Patent Information
- Application Number
- CN202210841849.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-18
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-07-18
AI Technical Summary
The prior art is difficult to efficiently and accurately identify the speaker's role information in land-air call scenarios, resulting in inefficient voice communication in air traffic management.
A multimodal speaker role recognition model is constructed by using feature fusion-based methods, using real-time reception and noise reduction to process speech signals, extracting voice fragments and transcribing them into text information, and combining deep neural networks to build a multimodal speaker role recognition model, extracting representations from speech and text features and performing fusion recognition.
It improves the accuracy of speaker role recognition in voice of land-air and air conversations, solves the problem of speaker role recognition, and improves the performance of air traffic control oral command understanding system and the communication efficiency between air traffic controllers and pilots.
Smart Images

Figure CN115240651B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of civil aviation air traffic control voice communication, and in particular to a method and device for identifying the role of a speaker in air-to-ground conversation based on feature fusion. Background Art
[0002] Voice communication between air traffic controllers and pilots is one of the most important ways of interaction in the field of air traffic management and is also a basic means to ensure the effective implementation of air traffic management. Since air traffic controllers communicate with several pilots on a single frequency via radio, it is difficult to distinguish the speaker role from the communication data link. According to the communication rules recommended by the International Civil Aviation Organization (ICAO), air traffic controllers will call the aircraft call sign before sending instructions to the target aircraft, and the pilots will repeat the instructions and then report their call signs. Generally speaking, most controller-pilot voice communications follow these rules, which makes the method of classifying speaker roles based on text information effective. Voice can be considered as another representation of the speaker: on the one hand, the equipment used by air traffic controllers and pilots in voice communication is different, and their equipment signal characteristics and background environmental noise are different; on the other hand, the controller-pilot voice communication signal implies other representation information, which means that it will further provide more identification knowledge for the speaker role recognition task.
[0003] In recent years, applications based on speech semantic understanding have received extensive attention and research in the field of air traffic management, such as air traffic control safety detection systems based on speech semantic understanding, air traffic controller workload analysis systems, etc. Among them, the speaker role is an indispensable key information in air traffic control speech semantic understanding, but the speaker role information cannot be directly distinguished from the communication data link, which has brought challenges to air traffic control speech-related applications to a certain extent.
[0004] According to the characteristics of air-to-ground call speech, the methods currently commonly used to complete the air-to-ground call speaker role recognition task mainly include single-modal speaker role recognition methods such as text-based speaker role recognition and speech-based speaker role recognition. However, the performance of text-based methods usually depends on air traffic management grammar, while speech-based methods are closely related to the communication environment (equipment and background noise, etc.). If the voice instructions deviate from the predefined air traffic management grammar, the performance of the text-based method will be significantly reduced. Similarly, when the speech-based method works on an unseen dataset (i.e., the new dataset is not covered by the training set), its accuracy will be poor.
[0005] Therefore, the present invention proposes a method and device for identifying the speaker role in air-to-ground calls based on speech-text feature fusion, aiming to solve the problem that the speaker role information cannot be efficiently and accurately identified in air-to-ground call scenarios, further improve the performance of the air traffic control oral command understanding system, and improve the efficiency of communication between air traffic controllers and pilots. Summary of the invention
[0006] The purpose of the present invention is to overcome the above-mentioned deficiencies in the prior art and to provide a method and device for identifying the speaker role in land-to-air conversations based on feature fusion.
[0007] In order to achieve the above-mentioned object of the invention, the present invention provides the following technical solutions:
[0008] A method for identifying the role of a speaker in a land-air call based on feature fusion, comprising the following steps:
[0009] S1: receiving the voice signal of the land-air call in real time and performing noise reduction processing on the voice signal;
[0010] S2: continuously monitoring and extracting a single-sentence speech segment containing a human voice from the speech signal after noise reduction processing;
[0011] S3: transcribing the single sentence speech segment into text information;
[0012] S4: inputting the single-sentence speech segment and the text information into a pre-built multimodal speaker role recognition model for recognition, wherein the multimodal speaker role recognition model extracts speech feature representation and text feature representation from the single-sentence speech segment and the text information, respectively, and outputs speaker role information corresponding to the speech signal according to the speech feature representation and the text feature representation;
[0013] The speaker role information includes the controller and the pilot; the construction of the multimodal speaker role recognition model includes the following steps:
[0014] A: Constructing a preliminary model of multimodal speaker role recognition based on a deep neural network; the preliminary model of multimodal speaker role recognition includes a text pre-training module, a speech pre-training module, and a classification module based on a modal attention mechanism;
[0015] B: Setting initial values of hyperparameters and training parameters of the preliminary multimodal speaker role recognition model;
[0016] C: The multimodal speaker role recognition preliminary model is trained by the corpus labeled with speaker role information, and the multimodal speaker role recognition model is output after the network converges. The present invention proposes a method for speaker role recognition in land-to-air calls based on feature fusion, which improves the accuracy of speaker role recognition in land-to-air call speech by comprehensively considering the feature representation of land-to-air call speech signals and the feature representation of text information, solves the problem of speaker role recognition in land-to-air calls, and provides corresponding speaker role information for applications such as air traffic control safety protection, air traffic control data analysis, and air traffic control business training.
[0017] As a preferred solution of the present invention, the noise reduction process in step S1 adopts a Kalman filter algorithm.
[0018] As a preferred embodiment of the present invention, step S2 comprises the following steps:
[0019] S21: preprocessing the speech signal after the noise reduction process; the preprocessing includes normalization, pre-emphasis, windowing and framing operations;
[0020] S22: dividing the preprocessed speech signal into a plurality of single-sentence speech segments according to a preset time value, and acquiring and outputting the single-sentence speech segments with human voice in the single-sentence speech segments.
[0021] As a preferred solution of the present invention, the extraction operation in step S2 is implemented by a pre-built voice activity detection model; the voice activity detection model includes a feature extraction module and a classification decision module;
[0022] The feature extraction module includes several groups of interconnected convolutional layers and pooling layers; the convolutional layers are used to extract feature vectors; the pooling layers are used to prevent overfitting;
[0023] The classification decision module includes several fully connected layers and an output layer; the fully connected layer is used to globally integrate the feature vector; the output layer is used to extract and output a single sentence speech segment containing human voice based on the globally integrated feature vector.
[0024] As a preferred embodiment of the present invention, the transcription operation in step S3 is implemented by a pre-trained speech recognition model;
[0025] The speech recognition model is a CNN-RNN-CTC architecture, including a convolutional neural network layer, a recurrent neural network layer and a fully connected layer.
[0026] As a preferred embodiment of the present invention, step S3 comprises the following steps:
[0027] S31: extracting spectrogram features after performing frame division and windowing processing on the single-sentence speech segment;
[0028] S32: Input the spectrogram features into the speech recognition model, and transcribe text information corresponding to the spectrogram features.
[0029] As a preferred solution of the present invention, the text pre-training module uses a MASK task to perform model training, which is used to extract text feature representation from the text information; the text pre-training module includes an Embedding unit, a Transformer unit and a prediction layer;
[0030] The Embedding unit includes a word embedding layer and a position embedding layer; the word embedding layer is used to convert each text word into a vector of fixed dimension; the position embedding layer is used to give different vector representations to the same text word at different positions;
[0031] The Transformer unit is composed of several encoder units;
[0032] The prediction layer is used to predict the text units that are masked to drive the neural network to learn the high-dimensional representation of text features. As a preferred embodiment of the present invention, the speech pre-training module adopts self-supervised learning to perform model training, which is used to extract speech feature representation from the single-sentence speech segment; the speech pre-training module includes a dynamic convolution audio feature extractor, an encoder, a Transformer unit and a quantizer;
[0033] The dynamic convolution audio feature extractor includes three dynamic convolution units connected in series, and the step of extracting high-dimensional preliminary speech features by the dynamic convolution unit includes: adding the output of the first dynamic convolution unit and the output of the second dynamic convolution unit through residual linking, and then inputting the output into the third dynamic convolution unit for processing and outputting preliminary speech features;
[0034] The encoder includes a plurality of convolutional neural network layers for extracting potential speech representation information of the preliminary speech features;
[0035] The Transformer unit is used to obtain context representation information;
[0036] The quantizer is used to construct a self-supervised training objective;
[0037] The speech pre-training module includes the following operating steps:
[0038] The dynamic convolution audio feature extractor is used to extract preliminary speech features from the single-sentence speech segment; the encoder is used to extract the potential speech representation information of the preliminary speech features; the deep representation information and the quantitative representation in the potential speech representation information are obtained through the Transformer unit and the quantizer respectively; and the speech feature representation corresponding to the single-sentence speech segment is output. The present invention proposes a text pre-training module and a speech pre-training module, which are used to extract features from text information and speech signals respectively. Compared with a single ordinary CNN network, the pre-training module has a stronger feature extraction capability and can give full play to the advantages of data-driven to discover potential features in a large amount of data, thereby effectively improving the accuracy of character recognition.
[0039] As a preferred solution of the present invention, the classification module is used to output the speaker role information corresponding to the speech signal according to the speech feature representation and the text feature representation; the classification module includes a modal attention mechanism unit, a pooling layer and a classifier:
[0040] The modal attention mechanism unit is used to fuse the high-dimensional feature representation of the speech feature representation and the text feature representation. ; Its operation formula is:
[0041] ;
[0042] in, is the preset trainable parameter, tanh is the activation function, The time step is When the correlation vector between the speech feature representation and the text feature representation is , the vector representing the text feature is , The sequence lengths of the vectors representing the output features of the speech pre-training module and the text pre-training module respectively; To generate modal attention weights through the Softmax function; The time step is The speech feature representation and time step at The correlation score between the text feature representations at ; is the time step variable, ;
[0043] The pooling layer is used to represent the high-dimensional features Pooling into a one-dimensional feature vector;
[0044] The classifier is used to perform speaker role recognition and classification according to the one-dimensional feature vector and output corresponding speaker role information; the classifier includes a fully connected layer with two output nodes and a Softmax activation function unit.
[0045] A device for identifying speaker roles in land-to-air conversations based on feature fusion, comprising at least one processor, at least one signal receiver communicatively connected to the at least one processor, and a memory communicatively connected to the at least one processor; the signal receiver is used to receive a voice signal and send the voice signal to the processor for processing, the memory stores instructions executable by the at least one processor, the instructions are executed by the at least one processor, so that the at least one processor can execute any of the methods described above.
[0046] Compared with the prior art, the present invention has the following beneficial effects:
[0047] 1. The present invention proposes a method for identifying the speaker role in land-to-air calls based on feature fusion. By comprehensively considering the feature representation of land-to-air call voice signals and the feature representation of text information, the accuracy of speaker role identification in land-to-air call voice is improved, the problem of speaker role identification in land-to-air calls is solved, and corresponding speaker role information is provided for applications such as air traffic control safety protection, air traffic control data analysis, and air traffic control business training.
[0048] 2. The present invention proposes a text pre-training module and a voice pre-training module, which are used to extract features from text information and voice signals respectively. Compared with a single ordinary CNN network, the pre-training module has a stronger feature extraction capability and can give full play to the advantages of data-driven, discovering potential features in a large amount of data, thereby effectively improving the accuracy of character recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 This is a flow chart of a method for identifying the speaker role in land-to-air calls based on feature fusion as described in Example 1 of the present invention.
[0050] Figure 2 This is a schematic diagram of the composition of a voice signal access module in a method for identifying speaker roles in land-to-air conversations based on feature fusion as described in Example 2 of the present invention.
[0051] Figure 3 This is a structural schematic diagram of a voice activity detection model based on a convolutional neural network in a method for identifying speaker roles in land-to-air conversations based on feature fusion as described in Example 2 of the present invention.
[0052] Figure 4This is a structural schematic diagram of a speech recognition model in a method for identifying speaker roles in land-to-air conversations based on feature fusion as described in Example 2 of the present invention.
[0053] Figure 5 This is a schematic diagram of the composition structure of a multimodal speaker role recognition model in a method for speaker role recognition in land-to-air communication based on feature fusion as described in Example 2 of the present invention.
[0054] Figure 6 This is a spectrogram of the speech of pilots and controllers in the field of air traffic management in a method for identifying speaker roles in air-to-ground conversations based on feature fusion as described in Example 2 of the present invention.
[0055] Figure 7 This is a schematic diagram of the structure of a dynamic convolution audio feature extractor in a method for identifying speaker roles in land-to-air conversations based on feature fusion as described in Example 2 of the present invention.
[0056] Figure 8 This is a schematic structural diagram of a device for identifying a speaker role for land-to-air conversation based on feature fusion described in Example 4 of the present invention, which utilizes the method for identifying a speaker role for land-to-air conversation based on feature fusion described in Example 1. DETAILED DESCRIPTION
[0057] The present invention is further described in detail below in conjunction with test examples and specific implementation methods. However, this should not be understood as the scope of the above subject matter of the present invention being limited to the following embodiments, and all technologies realized based on the content of the present invention belong to the scope of the present invention.
[0058] Example 1
[0059] like Figure 1 As shown, a method for identifying the speaker role in land-air calls based on feature fusion includes the following steps:
[0060] S1: receiving the voice signal of the land-air call in real time, and performing noise reduction processing on the voice signal; wherein the noise reduction processing adopts the Kalman filter algorithm.
[0061] S2: continuously monitoring and extracting a single-sentence speech segment containing a human voice from the speech signal after noise reduction processing;
[0062] S3: transcribing the single sentence speech segment into text information;
[0063] S4: inputting the single-sentence speech segment and the text information into a pre-built multimodal speaker role recognition model for recognition, wherein the multimodal speaker role recognition model extracts speech feature representation and text feature representation from the single-sentence speech segment and the text information, respectively, and outputs speaker role information corresponding to the speech signal according to the speech feature representation and the text feature representation;
[0064] The speaker role information includes the controller and the pilot; the construction of the multimodal speaker role recognition model includes the following steps:
[0065] A: A preliminary model for multimodal speaker role recognition is constructed based on a deep neural network; the preliminary model for multimodal speaker role recognition includes a text pre-training module, a speech pre-training module, and a classification module based on a modal attention mechanism.
[0066] B: Setting the initial values of the hyperparameters and the training parameters of the preliminary multimodal speaker role recognition model.
[0067] C: The multimodal speaker role recognition preliminary model is trained by using corpus with speaker role information annotated, and the output after network convergence is a multimodal speaker role recognition model.
[0068] According to the above scheme, the present invention can be combined with other means such as big data and cloud computing to assist air traffic control agencies in quickly classifying and counting the voice information of air traffic controllers and pilots for in-depth analysis and research (such as traffic controller workload statistics, air traffic communication terminology standard training, air traffic communication link optimization, etc.), thereby improving the daily work efficiency of air traffic control practitioners. During the review and analysis of flight accidents, the present invention can assist air traffic control agencies in analyzing the voice communications between air traffic controllers and pilots, and assist in determining the corresponding speaker role of each voice, which is conducive to quickly distinguishing the responsibilities of all parties, analyzing the causes of the accident, and formulating improvement measures, thereby improving the control safety factor and air traffic control command efficiency.
[0069] Example 2
[0070] This embodiment is a specific implementation of the method described in Example 1, comprising the following steps:
[0071] S1: receiving the voice signal of the land-air call in real time, and performing noise reduction processing on the voice signal; wherein the noise reduction processing adopts the Kalman filter algorithm.
[0072] The purpose of this step is to collect voice signals of air-ground conversations. The corpus A used in this embodiment is collected from a real air traffic control environment and then manually annotated. There are about 26.52 hours (25,765 items) of controller speech and 31.29 hours (35,895 items) of pilot speech in corpus A. During the annotation process, a small number of speech without speaker role information was marked as unknown and excluded from this example. The sampling rate of all samples in corpus A is 8000 Hz.
[0073] In order to further evaluate the performance and robustness of the method used in the present invention, in addition to using the test set a in the corpus A in the evaluation stage, a real-time air traffic control voice stream was introduced as a supplementary test set b. This real-time air traffic control voice stream was collected from another different air traffic management center and was not covered by the test set a. The total duration of the real-time air traffic control voice stream is about 2 hours (1930 voices), and the labels of the voices are also manually marked. The main purpose of introducing test set b is to evaluate the robustness of the model for an unknown air traffic management environment. In the training stage, the models used in this embodiment are all trained with test set a, and the parameters are adjusted on test set b.
[0074] This embodiment receives the voice signal using a voice signal access module, and its structure is as follows: Figure 2 As shown, specifically, the voice signal access module includes the following functions:
[0075] 1) The voice signal access module includes two voice signal access modes: linear access and radio reception, and supports two output modes: linear output and voice playback monitoring.
[0076] Linear access refers to accessing analog voice signals through the linear input interface (such as a 3.5mm audio port) of the linear access module in the input module; radio reception is to receive voice signals through the built-in radio module in the input module, by selecting the corresponding ground-to-air call frequency band through the FM knob.
[0077] The output module consists of a sound card and an amplifier. Whether it is a linear access or a radio access, the voice signal is first connected to the sound card to pre-process the input voice, and then output from the output interface of the voice signal access module. When in the voice playback monitoring mode, the output signal is connected to the amplifier device and plays the voice.
[0078] 2) The voice signal access module includes a voice signal noise reduction function, which specifically uses a Kalman filter to perform noise reduction processing on the input voice signal.
[0079] The Kalman filter is a minimum mean square error estimate for discrete linear system states. It uses the statistical information of noise and system states to minimize the mean square error as the optimization goal to give the optimal estimate of the original input signal. The Kalman filter can be used for both stationary processes and complex non-stationary processes. The general form of the Kalman filter equation is expressed as follows:
[0080] (1) One-step prediction equation:
[0081] ;
[0082] (2) Prediction equation of mean square error at time k:
[0083] ;
[0084] (3) Calculate the filter gain:
[0085] ;
[0086] (4) Calculation of variance matrix for one-step prediction:
[0087] ;
[0088] (5) Estimated variance matrix calculation:
[0089] ;
[0090] In the above Kalman filter equations, means System status at the moment; means Observation value of system state at each moment; It means using Observation at all times Estimate the state at the moment; It means from Time has come The state transfer matrix at time; It means obtaining After The minimum variance estimate of ; It refers to the linear minimum variance estimation; means The measurement matrix of the time system; means The variance matrix of ; The observation noise The variance matrix of ; means The variance matrix of ; means The noise driving matrix at the moment; The noise in the state equation The variance matrix of ; means The variance matrix of ; refers to the identity matrix; the superscript is the transposed matrix.
[0091] S2: continuously monitoring and extracting a single-sentence speech segment containing a human voice from the speech signal after the noise reduction process.
[0092] S21: preprocessing the speech signal after the noise reduction processing; the preprocessing includes normalization, pre-emphasis, windowing and framing operations.
[0093] S22: dividing the preprocessed speech signal into a plurality of single-sentence speech segments according to a preset time value, and acquiring and outputting the single-sentence speech segments with human voice in the single-sentence speech segments.
[0094] The extraction operation in step S2 is implemented by a pre-built voice activity detection model. Figure 3 As shown, the voice activity detection model is built based on a convolutional neural network architecture, specifically consisting of 1 input layer, several convolutional layers, several pooling layers, several fully connected layers and 1 output layer, and each convolutional layer is followed by 1 pooling layer.
[0095] Among them, several groups of interconnected convolutional layers and pooling layers constitute the feature extraction module. The convolutional layer is used to extract feature vectors. Its main feature is that it contains local receptive fields and weight sharing mechanisms. The local receptive fields and weight sharing mechanisms are improved methods proposed by convolutional neural networks to address the problems of too many parameters, large resource usage, and long calculation time in processing word vector matrices in fully connected neural networks. The ultimate goal is to reduce network parameters and calculations. The pooling layer mainly performs pooling operations, which are essentially downsampling operations. Its main function is to extract the most representative features in the region, reduce the output dimension of the features, and thus reduce the number of parameters.
[0096] Several fully connected layers and an output layer form a classification decision module. The feature extraction module obtains the feature vector representation of the text, and then the fully connected layer is responsible for globally integrating the high-dimensional feature information extracted and mapped by the upper layer. The output layer is used to extract and output a single sentence speech segment containing human voice based on the globally integrated feature vector.
[0097] S3: transcribe the single-sentence speech segment into text information; the transcription operation is implemented through a pre-trained speech recognition model.
[0098] S31: After framing and windowing the single-sentence speech segment, extracting spectrogram features; wherein the frame length of the spectrogram is 25 ms, the step length is 15 ms, and the dimension is 80 dimensions.
[0099] S32: Input the spectrogram features into the speech recognition model, and transcribe text information corresponding to the spectrogram features.
[0100] like Figure 4 As shown, the speech recognition model is a CNN-RNN-CTC architecture, including a convolutional neural network layer, a recurrent neural network (RNN) layer and a fully connected layer, and the model is optimized using a CTC (Connectionist Temporal Classification) loss function. The input of the model is a 25ms, 15ms step, 80-dimensional spectrogram feature; the model adopts an end-to-end modeling paradigm, with Chinese characters and English letters as the basic modeling units, that is, given the input speech features, the model directly outputs the corresponding transcribed text after decoding. The model is trained using training set a.
[0101] S4: Input the single-sentence speech segment and the text information into a pre-built multimodal speaker role recognition model for recognition, wherein the multimodal speaker role recognition model extracts speech feature representation and text feature representation from the single-sentence speech segment and the text information, respectively, and outputs speaker role information corresponding to the speech signal according to the speech feature representation and the text feature representation. The construction process of the multimodal speaker role recognition model is as follows:
[0102] A: A preliminary model for multimodal speaker role recognition is constructed based on a deep neural network. The preliminary model for multimodal speaker role recognition includes a text pre-training module, a speech pre-training module, and a classification module based on a modal attention mechanism, such as Figure 5 As shown, the specific composition structure is as follows:
[0103] 1) Text pre-training module:
[0104] The core idea of the text-based speaker role identification method is based on the air traffic management speaking rules issued by the International Civil Aviation Organization. Air traffic controllers and pilots should speak in strictly formulaic sentences. Air traffic controllers must indicate the call sign of the target flight before speaking the details of the instructions. In contrast, pilots report their call signs after repeating the instructions in the procedure. However, in practice, some air traffic management instructions violate the air traffic management speaking rules (for example, pilot instructions start with a call sign), which brings additional burden to the text-based method. In short, air traffic management grammatical rules are the theoretical basis of the text-based speaker role identification method, which can achieve good results in speaker role identification tasks. For voice instructions that violate air traffic management rules, it is expected that the model will learn to distinguish feature representations from a large number of data sets. Therefore, in the multimodal speaker role identification network of this example, a training module of the BERT model is designed to learn text-based high-level representations.
[0105] The text pre-training module uses the MASK task to perform model training, which is used to extract text feature representation from the text information; the text pre-training module includes an Embedding unit, a Transformer unit and a prediction layer;
[0106] The Embedding unit includes a word embedding layer and a position embedding layer; the word embedding layer converts each word into a vector of fixed dimension, and the head of each sequence is always the CLS mark, which is expressed as (1, n, 768); the role of the position embedding layer is to make the text pre-training module understand that the same word in different positions should have different vector representations, overcoming the disadvantage that Transformers cannot encode the sequentiality of the input sequence. The position embedding layer is not a fixed position code, but is obtained through learning, which is expressed as (1, n, 768).
[0107] The Transformer unit is composed of several encoder units; this embodiment only uses the encoder part in the classic Transformer architecture, completely discards the decoder part, has 12 hidden layers, outputs a 768-dimensional vector, and has a total of 12 self-attention heads.
[0108] The prediction layer: After being processed by the middle layer Transformer, the prediction layer makes corresponding adjustments according to different task requirements.
[0109] The text pre-training module is trained in a MASK manner, that is, for each sentence of input text, some words are randomly selected as the objects to be predicted, and then a special symbol (such as [MASK]) is used to replace them. After that, the text pre-training module will learn the words to fill in the replaced places based on the original correct labels. The process of the MASK task is:
[0110] (1) In the original training text, 15% of the words are randomly selected as the objects participating in the MASK task.
[0111] (2) Among these selected words, the data generator does not convert them all into [MASK] tags, but divides them into three situations: first, with a probability of 80%, the selected word is replaced with the [MASK] tag; second, with a probability of 10%, the selected word is replaced with a random word; third, with a probability of 10%, the selected word is kept unchanged.
[0112] (3) The text pre-training module tries its best to learn the semantics of the word in the context under highly uncertain conditions. At the same time, because the number of words involved in the MASK operation is relatively small, accounting for only 15% of the original text, the operation has little impact on the expressiveness and language rules of the original language.
[0113] 2) Voice pre-training module:
[0114] In the field of air traffic management, the voice of controller-pilot communication is transmitted via VHF radiotelephone, where the voice of the air traffic controller is ground-to-air, and the voice of the pilot is air-to-ground. Therefore, the radio transceiver equipment, microphone equipment and background noise of the environment (control room and aircraft cockpit) used by both parties are different, and these different characteristics will appear in the voice signal. Figure 6 As shown in FIG. 1 , in the spectrogram of the air traffic management speech, the characteristic intensities of the pilot and controller speech are distributed at different frequencies. For example, above 3000 Hz, the frequency energy distribution of the air traffic controller speech is stronger than that of the pilot. In this embodiment, different background noise models are given for different speech. Specifically, the background noise distribution of the pilot speech is uniform, while the background noise distribution of the air traffic controller is unstable.
[0115] In order to make full use of the key features in the above-mentioned speech information, the present invention designs a speech pre-training module, which is characterized in that it uses a self-supervised learning method to learn the representation information of the audio. The speech pre-training module is used to extract speech feature representation from the single-sentence speech segment; the speech pre-training module includes a dynamic convolution audio feature extractor, an encoder, a Transformer unit and a quantizer.
[0116] The dynamic convolution audio feature extractor includes three dynamic convolution units connected in series, and its design principle is as follows: Figure 7As shown in the figure, it is mainly used to solve the problem of extracting speech signal features in air-ground calls, where the speech speed is often fast and often accompanied by unstable noise. The convolution layer used in this extractor is different from the convolution layer commonly used in deep learning. Dynamic convolution uses a set of parallel convolution kernels instead of a single convolution kernel per layer. For each individual speech signal input, these parallel convolution kernels are dynamically aggregated through the input-dependent attention mechanism. Parallel convolution kernels share output channels by aggregation, which does not increase the depth or width of the network. The perceptron of common static convolution can be expressed as: ,in and are the weight matrix and the bias matrix respectively, is the transposed row and column symbol, is the activation function (such as ReLU function, Sigmoid function, tanh function, etc.), and are the input and output of the dynamic convolution respectively. According to the working principle of the dynamic convolution perceptron, its linear equation can be defined as follows:
[0117] ;
[0118] in, and It is obtained by the dynamic aggregation of K convolution kernels, where K is a hyperparameter:
[0119] ;
[0120] and, The following constraints are met:
[0121] ;
[0122] is the kth linear equation Attention weight, aggregation weight and deviation is a function of the input and has the same attention weights. It is not fixed, but varies according to the input.
[0123] The attention mechanism in dynamic convolution applies Squeeze-and-Excitation (SE) to calculate kernel attention. The global spatial information is first compressed by global average pooling, and then two fully connected layers (with a ReLU function between them) and Softmax function are used to generate the normalized convolution kernel attention weights. The number of neurons in the fully connected layer is the same as the vocabulary size of speech recognition.
[0124] The dynamic convolution audio feature extractor is composed of three dynamic convolution units, and the steps of extracting high-dimensional audio features by each dynamic convolution unit include:
[0125] First, the K convolution kernels in the dynamic convolution layer perform convolution operations on the input audio features to obtain a series of feature representation vectors. The attention mechanism in the dynamic convolution aggregates the features of the K convolution kernels and outputs them to the next layer of the neural network. K is a hyperparameter of the neural network model.
[0126] Secondly, the output of the dynamic convolution is fed into the Batch Normalization (BN) layer to reduce the need for regularization and speed up network convergence. It can also better prevent the problem of gradient explosion or gradient disappearance during training and prevent model overfitting.
[0127] Finally, after the batch normalization layer, the ReLU nonlinear activation function is used to perform a nonlinear transformation on the output of the dynamic convolution layer and output it to the next neural network module.
[0128] The output of the first dynamic convolution unit and the output of the second dynamic convolution unit of the dynamic convolution audio feature extractor are added through residual linking and then input into the third dynamic convolution unit to enhance the feature extraction capability of the model.
[0129] The encoder includes several convolutional neural network layers, each of which includes layer normalization (LN) and an activation function for extracting potential speech representation information of the preliminary speech features.
[0130] The Transformer unit is used to obtain context representation information; using the self-attention mechanism, the speech pre-training module generates output after fully considering the global information, rather than being limited to only seeing historical information.
[0131] The quantizer is used to construct a self-supervised training target. The quantizer splits the original d-dimensional continuous space into G subspaces, each of which has a dimension of d / G. Clustering is performed in each subspace to obtain V centers and their central features, and the central features are used to replace the features of each category. Finally, the infinite feature expression space is collapsed into a finite discrete space, making the features more robust and unaffected by a small amount of disturbance.
[0132] The speech pre-training module includes the following operating steps:
[0133] The dynamic convolution audio feature extractor is used to extract preliminary speech features from the single-sentence speech segment; the encoder is used to extract potential speech representation information of the preliminary speech features; the deep representation information and the quantized representation in the potential speech representation information are obtained through the Transformer unit and the quantizer respectively; and the speech feature representation corresponding to the single-sentence speech segment is output.
[0134] 3) Classification module based on modal attention mechanism:
[0135] The classification module is used to generate a final probability of a speaker role according to the speech feature representation and the text feature representation, and output speaker role information corresponding to the speech signal. The classification module includes a modal attention mechanism unit, a pooling layer, and a classifier.
[0136] Given that the vector output by the speech pre-training module is , the output vector of the text pre-training module is ,in The sequence lengths of the vectors representing the output features of the speech pre-training module and the text pre-training module are respectively, and the high-dimensional feature representation obtained based on the modal attention mechanism The recursive formula is as follows:
[0137] First, use the Score function to calculate the time step The speech features and time steps The correlation score between the text features :
[0138] ;
[0139] is a trainable parameter, is the time step variable, ;
[0140] Secondly, the modal attention weight is generated through the Softmax function :
[0141] ;
[0142] Then, the time step is calculated by a weighted sum operation The correlation vector between the speech features and text features :
[0143] ;
[0144] Finally, the features are fused according to the formula to generate a high-dimensional feature representation :
[0145] ;
[0146] are trainable parameters, and tanh is an activation function.
[0147] The purpose of the modality attention mechanism is to calculate the correlation representation between the text representation and the acoustic representation, so as to enhance the implicit text information and the features related to the speaker role recognition task in the auditory representation, especially the position and semantic features of the call sign.
[0148] The pooling layer is used to pool the high-dimensional feature representation into a one-dimensional feature vector;
[0149] The classifier is used to perform speaker role recognition classification according to the one-dimensional feature vector and output the corresponding speaker role information; the classifier includes a fully connected layer with two output nodes and a Softmax activation function unit.
[0150] B: Set the initial values of the hyperparameters and training parameters (such as loss function, learning rate, etc.) of the multi-modal speaker role recognition preliminary model.
[0151] C: Train the multi-modal speaker role recognition preliminary model through the corpus (i.e., training set a) with labeled speaker role information. After the network converges, the output is the multi-modal speaker role recognition model, and test the performance of the model on training set b.
[0152] Embodiment 3
[0153] Set the dimension of the word embedding to 512, and set the number of neurons in the two fully connected layers in the classifier to 256 and 2 respectively. The output dimensions of the feature encoder and the modality fusion module are both set to 512.
[0154] In the experiment of the speech-based method, spectrograms are generated by 80 log filter banks with a 25-ms window and a 15-ms step. Due to the multilingual nature of corpus A, in the text-based method, Chinese characters and English words are used as the basic vocabulary. There are 1284 tokens in the vocabulary, including 698 Chinese characters, 584 English words, and two special tokens, namely <PAD>, <UNK>.
[0155] In this example, the open-source deep learning framework PyTorch 1.4.0 is used to build and train all models. The server configuration is as follows: Install the Ubuntu 16.04 operating system, equipped with two NVIDIA GeForce RTX 2080Ti GPUs, an Intel Xeon E5-2630 CPU, and 128 GB of memory. Use an initial learning rate of 10 -4The Adam optimizer is used to complete the training task and the cross entropy is used as the loss function.
[0156] To verify the experimental results, in the embodiment, a text-based speaker role recognition method and a speech-based speaker role recognition method are designed for comparison with the method proposed in the present invention:
[0157] The text-based method is described as follows:
[0158] (1) TextCNN: Three CNN blocks with different convolution kernels are used to extract position feature information from the original text. Then, the feature map is generated by cascading the outputs of the CNN blocks. The CNN block contains a Conv2D layer, a ReLU activation function, and a pooling layer. The size of the convolution filter is set to (3, 4, 5), corresponding to different receptive fields.
[0159] (2) Transformer: The backbone network consists of 4 Transformer blocks, each of which contains a multi-head attention mechanism module, layer normalization, and a position feedforward direction network. In this example, we use 4 attention heads in the attention module and the dimension of the feedforward layer is set to 512.
[0160] The speech-based method is described as follows:
[0161] (1) X-vector: The X-vector model is used to extract DNN embedding features for speaker recognition. In this example, the front-end X-vector model is used as a feature encoder, which includes four time-delay deep neural network layers and a statistical pooling layer. The time-delay deep neural network layer mainly learns frame-level representations, and the statistical aggregation layer aggregates them into sentence-level embedding feature representations.
[0162] (2) SincNet: SincNet is a novel and effective CNN architecture that uses raw waveforms as input to complete speaker speech recognition tasks. In the SincNet module, bandpass filters are applied to convolve the waveform instead of standard CNN filters. Compared with the CNN module, it has better performance and is easier to converge.
[0163] The samples of air traffic controllers and pilots are unbalanced in corpus A, as shown in Table 1. The performance of all models in this example is measured by the accuracy (ACC) and F1-score of the model on the test set. Specifically, the accuracy index is used to consider the ratio of correctly predicted samples, and the F1 score is an indicator used in statistics to measure the accuracy of a binary classification model. When calculating these metrics in this embodiment, the pilot's voice is regarded as the positive class.
[0164] Table 1 Performance experimental results of various models on test set a and test set b
[0165]
[0166] As shown in Table 1, among the text-based methods, for the proposed unimodal speaker role recognition framework, the two text-based models achieved comparable performance, with an accuracy of 96%-97% on test set a. However, the F1 score dropped significantly on test set b, i.e., by about 2%. The performance of speech-based methods generally outperformed that of text-based methods, and they all achieved an accuracy of more than 97%.
[0167] For the proposed multimodal speaker role recognition framework, experimental results show that using both speech and text modal features is a feasible approach for speaker role recognition tasks. The proposed multimodal speaker role recognition network achieved the best accuracy and F1-score on both test set a and test set b, achieving 98.56% and 98.08% accuracy and 98.87% and 98.39% F1-scores, respectively. This result can be understood as the acoustic and text features of ATM speech complement each other in the speaker role recognition task. In addition, through the modal attention mechanism, the multimodal speaker role recognition network can properly consider the acoustic characteristics of speech and ATM grammar.
[0168] Example 4
[0169] like Figure 8 As shown, a device for identifying the role of a speaker in a land-to-air conversation based on feature fusion includes at least one processor, at least one signal receiver in communication connection with the at least one processor, and a memory in communication connection with the at least one processor; the signal receiver is used to receive a voice signal and send the voice signal to the processor for processing; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for identifying the role of a speaker in a land-to-air conversation based on feature fusion described in the above embodiment. The input and output interface may include a display, a keyboard, a mouse, and a USB interface for inputting and outputting data; the power supply is used to provide power to the device for identifying the role of a speaker in a land-to-air conversation based on feature fusion.
[0170] Those skilled in the art can understand that: all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiments; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), disks or optical disks, etc. Various media that can store program codes.
[0171] When the above-mentioned integrated unit of the present invention is implemented in the form of a software functional unit and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention can be essentially or partly reflected in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
Claims
1. A method for identifying the speaker role in land-air calls based on feature fusion, characterized in that: The following steps are involved: S1: receiving the voice signal of the land-air call in real time and performing noise reduction processing on the voice signal; S2: continuously monitoring and extracting a single-sentence speech segment containing a human voice from the speech signal after noise reduction processing; S3: transcribing the single sentence speech segment into text information; S4: inputting the single-sentence speech segment and the text information into a pre-built multimodal speaker role recognition model for recognition, wherein the multimodal speaker role recognition model extracts speech feature representation and text feature representation from the single-sentence speech segment and the text information, respectively, and outputs speaker role information corresponding to the speech signal according to the speech feature representation and the text feature representation; The speaker role information includes the controller and the pilot; the construction of the multimodal speaker role recognition model includes the following steps: A: Constructing a preliminary model of multimodal speaker role recognition based on a deep neural network; the preliminary model of multimodal speaker role recognition includes a text pre-training module, a speech pre-training module, and a classification module based on a modal attention mechanism; B: Setting initial values of hyperparameters and training parameters of the preliminary multimodal speaker role recognition model; C: training the preliminary multimodal speaker role recognition model by using corpus labeled with speaker role information, and outputting the multimodal speaker role recognition model after network convergence; The speech pre-training module adopts self-supervised learning to perform model training, and is used to extract speech feature representation from the single-sentence speech segment; the speech pre-training module includes a dynamic convolution audio feature extractor, an encoder, a Transformer unit and a quantizer; The dynamic convolution audio feature extractor includes three dynamic convolution units connected in series, and the step of extracting high-dimensional preliminary speech features by the dynamic convolution unit includes: adding the output of the first dynamic convolution unit and the output of the second dynamic convolution unit through residual linking, and then inputting the output into the third dynamic convolution unit for processing and outputting preliminary speech features; The encoder includes a plurality of convolutional neural network layers for extracting potential speech representation information of the preliminary speech features; The Transformer unit is used to obtain context representation information; The quantizer is used to construct a self-supervised training objective; The speech pre-training module includes the following operating steps: The dynamic convolution audio feature extractor is used to extract preliminary speech features from the single-sentence speech segment; the encoder is used to extract potential speech representation information of the preliminary speech features; the deep representation information and the quantized representation in the potential speech representation information are obtained through the Transformer unit and the quantizer respectively; and the speech feature representation corresponding to the single-sentence speech segment is output.
2. The method for identifying the speaker role in land-air conversation based on feature fusion according to claim 1, characterized in that: The noise reduction process in step S1 adopts Kalman filtering algorithm.
3. The method for identifying the speaker role in land-air conversation based on feature fusion according to claim 1, characterized in that: The step S2 comprises the following steps: S21: preprocessing the speech signal after the noise reduction process; the preprocessing includes normalization, pre-emphasis, windowing and framing operations; S22: dividing the preprocessed speech signal into a plurality of single-sentence speech segments according to a preset time value, and acquiring and outputting the single-sentence speech segments with human voice in the single-sentence speech segments.
4. The method for identifying the speaker role in land-air conversation based on feature fusion according to claim 3 is characterized in that: The extraction operation in step S2 is implemented by a pre-built voice activity detection model; the voice activity detection model includes a feature extraction module and a classification decision module; The feature extraction module includes several groups of interconnected convolutional layers and pooling layers; the convolutional layers are used to extract feature vectors; the pooling layers are used to prevent overfitting; The classification decision module includes several fully connected layers and an output layer; the fully connected layer is used to globally integrate the feature vector; the output layer is used to extract and output a single sentence speech segment containing human voice based on the globally integrated feature vector.
5. The method for identifying the speaker role in land-air conversation based on feature fusion according to claim 1, characterized in that: The transcription operation in step S3 is implemented by a pre-trained speech recognition model; The speech recognition model is a CNN-RNN-CTC architecture, including a convolutional neural network layer, a recurrent neural network layer and a fully connected layer.
6. The method for identifying the speaker role in land-air conversations based on feature fusion according to claim 5, characterized in that: The step S3 comprises the following steps: S31: extracting spectrogram features after performing frame division and windowing processing on the single-sentence speech segment; S32: Input the spectrogram features into the speech recognition model, and transcribe text information corresponding to the spectrogram features.
7. The method for identifying the speaker role in land-air conversations based on feature fusion according to claim 1, characterized in that: The text pre-training module uses the MASK task to perform model training, which is used to extract text feature representation from the text information; the text pre-training module includes an Embedding unit, a Transformer unit and a prediction layer; The Embedding unit includes a word embedding layer and a position embedding layer; the word embedding layer is used to convert each text word into a vector of fixed dimension; the position embedding layer is used to give different vector representations to the same text word at different positions; The Transformer unit is composed of several encoder units; The prediction layer is used to predict the text units that are masked to drive the neural network to learn the high-dimensional representation of text features.
8. The method for identifying the speaker role in land-air conversations based on feature fusion according to claim 1, characterized in that: The classification module is used to output the speaker role information corresponding to the speech signal according to the speech feature representation and the text feature representation; the classification module includes a modal attention mechanism unit, a pooling layer and a classifier: The modal attention mechanism unit is used to fuse the high-dimensional feature representation f of the speech feature representation and the text feature representation i ; Its operation formula is: Among them, β and α are preset trainable parameters, tanh is the activation function, and r i is the correlation vector between the speech feature representation and the text feature representation when the time step is i, and the vector of the speech feature representation is The vector representing the text feature is m and n are the sequence lengths of the vectors representing the output features of the speech pre-training module and the text pre-training module respectively; w ij To generate the modal attention weight through the Softmax function; e ij is the correlation score between the speech feature representation at time step i and the text feature representation at time step j; i and j are time step variables, 1≤i≤m, 1≤j≤n; The pooling layer is used to represent the high-dimensional feature f i Pooling into a one-dimensional feature vector; The classifier is used to perform speaker role recognition and classification according to the one-dimensional feature vector and output corresponding speaker role information; the classifier includes a fully connected layer with two output nodes and a Softmax activation function unit.
9. A device for identifying the speaker role in land-air conversations based on feature fusion, characterized in that: It includes at least one processor, at least one signal receiver communicatively connected to the at least one processor, and a memory communicatively connected to the at least one processor; the signal receiver is used to receive a voice signal and send the voice signal to the processor for processing, and the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in any one of claims 1 to 8.