Remote control tower scene-based controller voice fatigue prediction method
By collecting and processing audio samples in remote tower scenarios, extracting multi-dimensional features and using improved Bi-LSTM-GRU models, the problem of low accuracy of controller speech fatigue prediction in the prior art is solved, and high accuracy and real-time fatigue state prediction in complex environments are achieved.
Patent Information
- Application Number
- CN202510602935.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The existing controller voice fatigue prediction methods have low prediction accuracy, poor robustness and poor real-time performance in complex working environments, making it difficult to adapt to changes in fatigue states under high pressure and complex environments.
The controller's voice fatigue prediction method based on remote tower scene is adopted, and audio samples are collected and denoised, and the pronunciation, sound quality and spectral characteristics are extracted, and the improved Bi-LSTM-GRU model is used for training to predict the fatigue state of the controller in real time.
Improves the accuracy and robustness of fatigue prediction, and can accurately capture fatigue state changes in high noise and changeable environments, provide timely and accurate fatigue state prediction, ensuring flight safety and controller health.
Smart Images

Figure CN120108386A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voice data processing, and in particular to a method for predicting voice fatigue of a controller based on a remote tower scenario. Background Art
[0002] Voice fatigue detection technology determines whether the speaker is in a fatigued state by analyzing the feature changes in the voice signal. With the increasing demand for fatigue detection, researchers have proposed a variety of fatigue detection methods based on machine learning and deep learning. At present, the prediction method of controller voice fatigue mainly relies on the feature extraction of the controller's voice signal and the analysis of traditional machine learning models. Traditional methods usually use classic algorithms such as support vector machines (SVM) and decision trees. Although these methods can achieve good results in some simple scenarios, when faced with complex voice signals, the feature selection and model training process are easily affected by noise and data imbalance, resulting in low prediction accuracy and difficulty in adapting to the changes in the fatigue state of controllers in high-pressure and high-intensity working environments.
[0003] Therefore, the existing prediction methods for controller voice fatigue have the defects of low prediction accuracy, poor robustness, and poor real-time performance when facing complex working environments, dynamic fatigue states, and data imbalance. Therefore, a new prediction method is urgently needed to improve the accuracy, robustness, and real-time performance of controller fatigue prediction, especially in high-pressure and complex environments. Summary of the invention
[0004] The purpose of the present invention is to improve the prediction accuracy of the controller's voice fatigue state and to provide a controller's voice fatigue prediction method based on a remote tower scenario.
[0005] In order to achieve the above-mentioned purpose of the invention, the embodiment of the present invention provides the following technical solution, a method for predicting controller voice fatigue based on a remote tower scenario, comprising the following steps:
[0006] Step 1: collect audio samples containing controller voices in remote tower call data, perform denoising, and obtain denoised voice signals; and mark the fatigue status of the voice signals;
[0007] Step 2: extracting rhythm features, sound quality features and spectral features from a number of annotated speech signals and fusing them into comprehensive features;
[0008] Step 3, training the Bi-LSTM-GRU model based on the extracted comprehensive features to output fatigue prediction results; the Bi-LSTM-GRU model includes an input layer, a backward layer, a forward layer, and an output layer, wherein the input layer transmits the comprehensive features of the time step to the backward layer and the forward layer respectively, the backward layer processes the reverse time series on the time step, the forward layer processes the forward time series on the time step, and the output layer splices the outputs in the forward and backward directions;
[0009] Step 4: Collect the real-time voice data of the controller, input the real-time voice data into the trained Bi-LSTM-GRU model, and predict the fatigue status of the controller.
[0010] Compared with the prior art, the present invention aims at solving the technical problems in the existing method for predicting the voice fatigue of controllers by introducing innovative technical means such as the improved Bi-LSTM-GRU model (bidirectional long short-term memory network), fusion of subjective and objective features, and dynamic threshold adjustment mechanism, and effectively overcomes the above technical problems, producing significant technical effects, which are specifically reflected in the following aspects.
[0011] (1) Overcoming the feature dependency problem and improving the accuracy of fatigue prediction, the present invention combines multi-dimensional features such as rhythm, sound quality and spectrum, so that the prediction of fatigue status no longer depends on a single static feature. The comprehensive use of such multi-dimensional features greatly improves the performance of the model in complex environments, especially in the high-noise and variable remote tower control working environment, effectively avoiding the interference of external factors on feature data and improving the accuracy of fatigue prediction. At the same time, the introduction of subjective fatigue scales (such as the Samn-Perelli 7-level fatigue scale and the SART scale) combined with objective reaction time (such as the PVT test) can effectively and accurately mark the fatigue and non-fatigue of controllers after standardization. Through standardized subjective and objective scoring, the present invention can capture subtle changes in fatigue earlier and more accurately, especially when the controller is under high-pressure and high-intensity working conditions. This innovative method effectively solves the problem of misjudgment of fatigue status by traditional methods, especially avoiding missing subtle changes in the early stage of fatigue, providing timely and accurate fatigue status prediction, and ensuring flight safety and the work health of controllers.
[0012] (2) The present invention overcomes the limitations of traditional machine learning methods and improves the adaptability and efficiency of the system. The Bi-LSTM-GRU model based on deep learning is adopted. Compared with traditional methods such as support vector machine (SVM) and decision tree, the Bi-LSTM-GRU model can better process time series data and capture the dynamic changes of the controller's fatigue state. Through the structure of forward LSTM and reverse GRU, the model not only considers historical information, but also makes full use of future information, greatly improving the model's sensitivity to changes in fatigue state. In addition, the Bi-LSTM-GRU model has a strong feature learning ability, avoiding the problem of relying on artificial feature selection in traditional methods, and improving the system's adaptability to fatigue changes under different working conditions. Especially in a high-pressure environment, it can accurately capture the slight changes in fatigue, avoiding the failure of early warning caused by the excessive reliance on static features of traditional methods.
[0013] (3) The problem of static characteristics and progressive fatigue is solved to achieve fatigue prediction. The present invention can effectively identify the gradual changes in the fatigue state of the controller, especially the progressive characteristics of fatigue after long-term work, by using the Bi-LSTM-GRU model for deep learning of time series features. Unlike traditional methods that rely on static characteristics, this model can capture subtle changes in fatigue during long-term, high-intensity work and identify fatigue signals in advance, thereby effectively avoiding the traditional method from missing early fatigue, ensuring the timely response of the early warning system, and providing more accurate predictions for flight safety.
[0014] (4) Overcoming the high cost and data imbalance problems and improving the robustness of the model. By adopting the self-learning ability of the deep learning model, the present invention obtains the controller's voice data in real time for training and prediction without relying on a large amount of manually labeled data, thus reducing the dependence on manually labeled data during the system training process. In addition, the model can cope with the data imbalance problem through adaptive adjustment during the training process, improving the prediction accuracy and robustness when fatigue samples are scarce. This enables the present invention to adapt to the working conditions and environments of different controllers in practical applications and maintain a high fatigue prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0016] Figure 1 is a flow chart of the method of the present invention;
[0017] Figure 2 This is a schematic diagram of the Bi-LSTM-GRU network structure of the present invention. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present invention.
[0019] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used to distinguish the description, and cannot be understood as indicating or implying relative importance, or implying any such actual relationship or order between these entities or operations. In addition, the terms "connected", "connected", etc. can be directly connected between elements, or indirectly connected via other elements.
[0020] The present invention is achieved through the following technical solutions: Figure 1 As shown, a method for predicting controller voice fatigue based on a remote tower scenario includes the following steps:
[0021] Step 1: collect audio samples containing controller voices from remote tower call data, perform denoising, and obtain denoised voice signals; and label the voice signals for fatigue status.
[0022] Audio samples with controller voices are collected from real remote tower call data. It is ensured that these audio samples cover different working conditions, such as different working time periods (morning shift, night shift, etc.), and different environmental noise backgrounds, to ensure the wide applicability of the analysis results. The audio samples should also include the voices of controllers in a non-fatigued state and in a fatigued state. According to the actual situation, voices from different airports and different controllers are selected to ensure the diversity of audio samples.
[0023] Since the controllers in the remote tower are not only facing a single airport, but also facing the scenario of multiple airports integrated into one, the control and command situation is complicated, and the noise in the collected call data is relatively large, which is not conducive to the subsequent fatigue prediction. Therefore, the collected audio samples need to be denoised using a noise suppression algorithm, such as spectral subtraction denoising, to remove background noise and interference sounds to ensure the clarity of the voice signal. Adjust the volume and sound quality, unify the volume level of the audio, to reduce the feature extraction error caused by volume differences. Then remove the silent segment of the voice signal, remove the segments that do not contain valid voice content, and ensure that only the voice of the controller is included in the subsequent analysis.
[0024] In detail, the time-frequency domain representation of the original speech signal in the collected audio samples is,
[0025] Where f is frequency; t is time; x(t) is the original speech signal; X(f,t) is the complex value of the original speech signal at time t and frequency f; w(tt 0 ) is the window function; t 0 is the center time of the current window; F is the Fourier transform operation.
[0026] In spectral subtraction, the power spectrum is used to reduce the noise of the signal. Assume that the noise spectrum estimate is , denoising is performed by spectral subtraction to obtain the denoised speech signal,
[0027] where |X(f,t)| 2 is the square of the modulus of X(f,t), that is, the power spectrum of the original speech signal; is the power spectrum of the noise estimate; S(f,t) is the power spectrum of the speech signal after denoising. The max operation is used to ensure that the spectrum value does not become negative to avoid unreasonable results.
[0028] By performing an inverse short-time Fourier transform on the power spectrum S(f,t) of the denoised speech signal, we can get the corresponding time domain signal.
[0029] in, represents the time domain signal corresponding to S(f,t); F -1 Represents an inverse short-time Fourier transform operation.
[0030] The de-noised speech signal is labeled, and the speech signal is divided into several speech samples, and the fatigue state (fatigue, non-fatigue) of each speech sample is labeled. The labeling of fatigue state comprehensively considers the subjective fatigue perception and objective fatigue test of the controller.
[0031] Subjective fatigue perception assessment,Each speech sample was annotated using a subjective fatigue scale (such as the Samn-Perelli 7-level fatigue scale shown in Table 1 and the SART scale shown in Table 2) to assess the controller's subjective fatigue score.
[0032] Table 1 Samn-Perelli 7-level fatigue scale
[0033] Table 2 SART scale
[0034] Objective fatigue test evaluation, PVT (Psychomotor Vigilance Test) is a psychological test commonly used to evaluate individual reaction time. It is usually used to measure an individual's alertness, attention and responsiveness over a period of time. The PVT test tests an individual's reaction time (RT) by presenting visual or auditory stimuli. It refers to the time interval from the presentation of the stimulus (the appearance of a target on the screen) to the subject's response (such as pressing a button). The subject is scored according to the reaction time scoring table shown in Table 3. It is usually measured in milliseconds (ms), and the calculation formula is,
[0035] Wherein, RT represents the reaction time of the subject obtained through the PVT test; t response Indicates the time point when the subject responds; t stimulus Indicates the time point of stimulus presentation.
[0036] Table 3 Reaction time scoring table
[0037] Calculate the comprehensive fatigue score,
[0038] Among them, Fatigue score represents the comprehensive fatigue score; SP score SART stands for Samn-Perelli 7-level fatigue rating score. score represents the SART scale score; RT score Indicates the PVT test score.
[0039] When the comprehensive fatigue score is Fatigue score When the score exceeds 6, the subject is judged to be in a fatigue state. score The speech signal is marked. If it exceeds 6 points, it is marked as fatigue, otherwise it is marked as non-fatigue.
[0040] Step 2: extract rhythm features, sound quality features and spectral features from several annotated speech signals and fuse them into comprehensive features.
[0041] Several annotated speech signals form a corpus, which is the basis for subsequent feature extraction and analysis. It needs to contain enough samples to cover various possible situations, including but not limited to individual differences among different controllers (such as gender, age, experience level), different working hours (such as day shift, night shift, continuous duty), different control tasks (such as approach, departure, tower, regional control), different workloads (such as low, medium, and high traffic control), different scenario complexity (such as single airport, multi-airport joint command, emergency response), etc. Each speech sample in the corpus should contain speech signals and corresponding annotations.
[0042] A variety of features are extracted from the speech signals in the corpus, including rhythm features, sound quality features, and spectrum features. These features can reflect various aspects of the controller's voice and are crucial for identifying fatigue status.
[0043] 1. Prosodic features include speaking speed, pause frequency, and pause duration.
[0044] A. Speech Rate (SR), which is usually a unit of pronunciation per second (such as the number of syllables or words), is calculated as follows:
[0045] Among them, SR represents the speaking speed; N s Indicates the total number of syllables pronounced in the speech sample; T d Indicates the duration of the speech sample segment.
[0046] B. Pause Frequency (PF), Pause frequency refers to the number of pauses in a speech sample. It can be calculated by detecting the silent segments in the speech. The calculation formula is:
[0047] Where PF represents the pause frequency; N p Indicates the number of pauses in the speech sample; T d Indicates the duration of the speech sample segment.
[0048] C. Pause Duration (PD), the pause duration is expressed as the average duration of each pause, and the calculation formula is,
[0049] Among them, PD represents the pause duration; Indicates the duration of each pause; i indicates the index of the i-th pause, i=1,2,...,N p .
[0050] The calculation formula of rhythmic features is:
[0051] Among them, [SR, PF, PD] means splicing the speech rate, pause frequency and pause duration according to the time dimension; F prosody A vector representing rhythmic features. Since SR, PF, and PD are all one-dimensional values, concatenating them forms a 1*3 row vector.
[0052] 2. Sound quality characteristics reflect the clarity, stability and other factors related to the pronunciation quality of the sound. Sound quality characteristics include sound pressure level and fundamental frequency.
[0053] A. Sound Pressure Level (SPL), which measures the loudness of speech, is calculated as follows:
[0054] Where SPL represents the sound pressure level; p represents the current sound pressure; p 0 Indicates the reference sound pressure, usually 20×10 -6 Pa.
[0055] B. Fundamental Frequency (F 0 ), the basic frequency reflects the pitch of the sound, and the calculation formula is,
[0056] Among them, F 0 Indicates the basic frequency; N cycles Indicates the number of vibrations in a cycle; N period Indicates the cycle duration.
[0057] The calculation formula for sound quality features is:
[0058] Among them, [SPL,F 0 ] means to concatenate the sound pressure level and fundamental frequency according to the time dimension; F qualityA vector representing the sound quality features. Since SPL and F 0 Each is a one-dimensional value, and concatenating them forms a 1*2 row vector.
[0059] 3. Spectral features include formant and spectral entropy.
[0060] A. Formant Frequencies, F 1 , F 2 , F 3 ), the formant is the peak in the spectrum, which usually reflects the pronunciation of vowels. The formant can be estimated by linear prediction cepstrum analysis (LPCC), and the calculation formula is,
[0061] Among them, F n Indicates the frequency of the nth resonance peak, n=1,2,3; LPCC n It indicates that the frequency of the nth resonance peak is estimated by cepstrum analysis.
[0062] B. Spectral Entropy (SE): Spectral entropy measures the degree of disorder of the spectrum and is used to describe the complexity of speech. The calculation formula is:
[0063] Wherein, SE represents spectral entropy; P(f) represents the probability density function corresponding to frequency f.
[0064] The calculation formula of the spectral feature is:
[0065] Among them, [F 1 ,F 2 ,F 3 ,SE] means that the frequency of the first resonance peak, the frequency of the second resonance peak, the frequency of the third resonance peak and the spectrum entropy are spliced according to the time dimension; F spectral A vector representing the spectral features. 1 、F 2 、F 3 Both and SE are one-dimensional values, and concatenating them forms a 1*4 row vector.
[0066] The above rhythmic features, sound quality features and spectral features are integrated to form a comprehensive feature.
[0067] Among them, F fusion Represents the comprehensive features after fusion; F prosody A vector representing rhythmic features; Fquality A vector representing the sound quality features; F spectral A vector representing the spectral features; w 1 Indicates F prosody The weight of 2 Indicates F quality The weight of 3 Indicates F spectral When calculating the comprehensive features, it is necessary to take the vector F of the rhythmic features into account. prosody Add an element 0 to the end of F to make its dimension the same as the vector F of the spectral class feature spectral The dimension is the same as that of quality Add two elements of 0 at the end to make its dimension the same as the vector F of the spectral class feature spectral The dimensions are the same.
[0068] Step 3, training the Bi-LSTM-GRU model based on the extracted comprehensive features to output fatigue prediction results; the Bi-LSTM-GRU model includes an input layer, a backward layer, a forward layer, and an output layer, wherein the input layer transmits the comprehensive features of the time step to the backward layer and the forward layer respectively, the backward layer processes the reverse time series on the time step, the forward layer processes the forward time series on the time step, and the output layer splices the outputs in the forward and backward directions.
[0069] The labeled speech signals are divided into training set and validation set in a certain ratio. The common division ratio is 7:3 to ensure balanced data distribution and cover different fatigue states (fatigue and non-fatigue). The divided data has three types of features (i.e. rhythmic features, sound quality features and spectral features). The data after feature extraction is normalized to ensure consistent feature scales, reduce the impact of numerical range differences on subsequent model training, and improve the stability and convergence speed of the model.
[0070] Among traditional technologies, the deep learning method Bi-LSTM has achieved remarkable results in speech timing prediction tasks. It can effectively capture the temporal dependency of speech signals and is widely used in tasks such as speech sentiment analysis and speech recognition. However, Bi-LSTM still has some limitations, mainly including the following points.
[0071] (1) High computational complexity. Due to its bidirectional structure and complex gating mechanism, Bi-LSTM has high computational overhead when modeling long time series, which affects real-time performance.
[0072] (2) Gradient vanishing problem. Although LSTM has better long-term memory capability than RNN, it may still cause gradient vanishing when processing ultra-long sequence data, making it difficult to effectively transmit long-range dependency information.
[0073] (3) The model converges slowly. LSTM relies on multiple gated units to store and update states. It is prone to overfitting on small-scale data sets and takes a long time to train.
[0074] Therefore, the present invention improves the Bi-LSTM structure and proposes a Bi-LSTM-GRU model, which combines GRU (Gated Recurrent Unit) to optimize the shortcomings of LSTM. Based on Bi-LSTM, it extracts long-term time dependencies and can maintain the modeling ability of complex speech patterns; and uses GRU to replace part of the LSTM structure to reduce computational complexity and improve training and reasoning efficiency. This solution enhances the memory capacity of the model, alleviates the gradient vanishing problem, and improves the model's detection accuracy of fatigue status through the improved Bi-LSTM-GRU model.
[0075] like Figure 2 As shown, the Bi-LSTM-GRU model includes an input layer, a backward layer, a forward layer, and an output layer.
[0076] 1. Input layer, input comprehensive features of T time steps, comprehensive features x of each time step t They all contain rhythmic features, sound quality features and spectral features.
[0077] 2. The backward layer includes T GRU units. The comprehensive features of T time steps are input into T GRU units one by one. Each GRU unit includes an update gate and a reset gate. The calculation formula is as follows:
[0078] A. Update gate, used to determine the comprehensive feature x of the current time step input t and the hidden state h at the next time step t+1 The hidden state h at the current time step t The influence of, and thus control the degree of retention of previous information at the current time step,
[0079] Among them, Z t is the activation value of the update gate; represents the sigmoid activation function; W z is the update gate weight matrix; b z is the bias term of the update gate; h t+1 Represents the hidden state of the GRU unit from the next time step during back propagation; t∈T, T is the total number of time steps.
[0080] B. Reset gate, used to control the comprehensive feature x input at the current time step t and the hidden state h at the next time step t+1 The degree of integration,
[0081] Among them, r t is the activation value of the reset gate; W r is to reset the gate weight matrix; b r is the bias term for the reset gate.
[0082] C. Candidate hidden state, calculate the new information of the current time step,
[0083] in, is the candidate hidden state; tanh represents the hyperbolic tangent activation function; W h is the candidate hidden state pair x t The weight matrix of h is the candidate hidden state pair h t+1 The weight matrix of h is the candidate hidden state bias.
[0084] D. The final hidden state of the GRU unit,
[0085] in, is the output (final hidden state) of the GRU unit at time step t, i.e. h t .
[0086] 3. The forward layer includes T LTSM units. The comprehensive features of T time steps are input into T LTSM units one by one. Each LSTM unit includes an input gate, a forget gate, an output gate (cell state) and a cell state. The calculation formula is as follows:
[0087] A. Input gate, used to determine the comprehensive feature x of the current time step input t and the hidden state h at the previous time step t-1 The impact on the current cell state, thereby controlling the degree of retention of subsequent information at the current time step,
[0088] Among them, i t is the activation value of the input gate; represents the sigmoid activation function; W i is the input gate weight matrix; b i represents the bias term of the input gate; h t-1 Represents the hidden state of the LSTM unit from the previous time step during forward propagation; t∈T.
[0089] B. Forget gate, used to control the cell state c of the previous time step t-1 What information needs to be forgotten?
[0090] Among them, f t is the activation value of the forget gate; W f is the forget gate weight matrix; b f is the bias term of the forget gate.
[0091] C. Cell state, which determines the retention and updating of long-term memory,
[0092] in, is the candidate cell state at the current time step; tanh represents the hyperbolic tangent activation function; W c is the cell state weight matrix; b c is the bias term of the cell state; c t is the cell state at the current time step.
[0093] D. Output gate, used to control the hidden state h output at the current time step t , which depends on the cell state c at the current time step t ,
[0094] Among them, t is the activation value of the output gate; W o is the output gate weight matrix; b o is the bias term of the output gate; is the output (final hidden state) of the LTSM unit at time step t, i.e. h t .
[0095] 4. The output layer consists of the output of the GRU unit and the output of the LSTM unit at each time step, where the GRU unit processes the reverse time series from the Tth time step to the 1st time step, and the LSTM unit processes the forward time series from the 1st time step to the Tth time step; finally, the outputs of the two directions are concatenated or summed at the same time step to obtain the bidirectional hidden state.
[0096] Among them, H t Represents the bidirectional hidden state output at the current time step.
[0097] 5. Prediction output, after T time steps, the output layer outputs T bidirectional hidden states H t , t∈T, T bidirectional hidden states H t After the fully connected layer, it is fused into H t `, the fatigue prediction result output by the Bi-LSTM-GRU model is:
[0098] in, is the fatigue prediction result (fatigue, non-fatigue) output by the Bi-LSTM-GRU model; Softmax represents the Softmax activation function, which converts the output into a probability form; W out is the output layer weight matrix; b out is the output layer bias term.
[0099] The threshold can be set to 0.5. When it is greater than 0.5, the fatigue prediction result is considered to be fatigue, otherwise it is considered to be non-fatigue. .
[0100] The Bi-LSTM-GRU model performs splicing on time steps through the back propagation of GRU units and the forward propagation of LSTM units, and finally outputs the fatigue prediction result. The GRU unit selects the weight matrix W of the candidate hidden state on multiple time steps. h , the LSTM unit selects the output gate weight matrix W at multiple time steps o , but because the GRU unit and LSTM unit are relatively reverse transmission, so at the same time step for W h and W o The choice of is extremely important, which is related to the continuity and bidirectional compatibility of the output data in the time series. To this end, this solution designs its loss function based on the structural function of the Bi-LSTM-GRU model to constrain the accuracy of the selection of weights at any time step, so that the output results can maintain continuity and compatibility in the case of bidirectional propagation fusion:
[0101] Where L is the loss function; x t is the real comprehensive feature at the tth time step; is the comprehensive feature predicted at the tth time step; is a hyperparameter, , in the initial state ; is the dynamic optimization weight of the t-th time step, and we have,
[0102] in, is the dynamic optimization weight for the t-1th time step; Dynamic optimization weights for the t+1th time step; Optimize weights for the initialization; is the learning rate, and ; exp is the natural exponential function; is the adjustment rate, ; t is the index of the time step.
[0103] Step 4: Collect the real-time voice data of the controller, input the real-time voice data into the trained Bi-LSTM-GRU model, and predict the fatigue status of the controller.
[0104] The real-time voice data of the controller is collected, the corresponding comprehensive features are extracted, and the comprehensive signs are input into the trained Bi-LSTM-GRU model. The model outputs the predicted probability of fatigue state. If the probability value is greater than 0.5, the controller is judged to be "fatigued"; if the probability value is less than or equal to 0.5, the controller is judged to be "non-fatigued".
[0105] More specifically, according to the fatigue status prediction results, the system triggers the corresponding alarm mechanism. When the controller is predicted to be "fatigued", the system issues fatigue warnings of different levels according to the predicted probability value. Specifically, if the probability value is less than or equal to 0.5, the system determines that the controller is in a "non-fatigue" state and does not trigger any alarm; if the probability value is in the (0.5,0.6] interval, the system determines it as "mild fatigue", "mild fatigue" is a type of "fatigue", and prompts the controller to take a proper rest; if the probability value is in the (0.6,0.7] interval, the system determines it as "moderate fatigue", "moderate fatigue" is a type of "fatigue", triggering a medium-priority fatigue warning, and recommending rotation or short repair; if the probability value is greater than 0.7, the system determines it as "severe fatigue", "severe fatigue" is a type of "fatigue", and a high-priority warning is issued, requiring the controller to rest immediately, change shifts, or perform other necessary interventions. All fatigue status predictions and alarm results will be recorded in the system log to support subsequent monitoring, evaluation, and safety management analysis.
[0106] In actual application, the system will continuously monitor the fatigue status of the controller and adjust the fatigue prediction threshold in real time based on feedback to ensure accuracy and timeliness. If the system detects changes in fatigue status, the relevant department will immediately receive warning information and take necessary measures, such as handovers, increased rest time, etc., to ensure the health of the controller and flight safety.
[0107] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A method for predicting controller voice fatigue based on a remote tower scenario, characterized in that: The following steps are included: Step 1: collect audio samples containing controller voices in remote tower call data, perform denoising, and obtain denoised voice signals; and mark the fatigue status of the voice signals; Step 2: extracting rhythm features, sound quality features and spectral features from a number of annotated speech signals and fusing them into comprehensive features; Step 3, training the Bi-LSTM-GRU model based on the extracted comprehensive features to output fatigue prediction results; the Bi-LSTM-GRU model includes an input layer, a backward layer, a forward layer, and an output layer, wherein the input layer transmits the comprehensive features of the time step to the backward layer and the forward layer respectively, the backward layer processes the reverse time series on the time step, the forward layer processes the forward time series on the time step, and the output layer splices the outputs in the forward and backward directions; Step 4: Collect the real-time voice data of the controller, input the real-time voice data into the trained Bi-LSTM-GRU model, and predict the fatigue status of the controller.
2. The method for predicting controller voice fatigue based on remote tower scenario according to claim 1 is characterized in that: In step 1, the audio samples containing the controller's voice in the remote tower call data are collected and denoised to obtain the denoised voice signal, including: The time-frequency domain representation of the original speech signal in the collected audio samples is: Where f is frequency; t is time; x(t) is the original speech signal; X(f,t) is the complex value of the original speech signal at time t and frequency f; w(t-t0) is the window function; t0 is the center time of the current window; F is the Fourier transform operation; Assume that the noise spectrum is estimated to be , denoising is performed by spectral subtraction to obtain the denoised speech signal, where |X(f,t)| 2 is the power spectrum of the original speech signal; is the power spectrum of the noise estimation; S(f,t) is the power spectrum of the speech signal after denoising; Perform inverse short-time Fourier transform on the power spectrum S(f,t) of the denoised speech signal to obtain the corresponding time domain signal. in, represents the time domain signal corresponding to S(f,t); F -1 Represents an inverse short-time Fourier transform operation.
3. The method for predicting controller voice fatigue based on remote tower scenario according to claim 1, characterized in that: In step 1, the step of marking the fatigue state of the speech signal includes: Subjective fatigue perception and objective fatigue test are conducted on the voice signals of the controllers. The subjective fatigue perception is evaluated based on the Samn-Perelli 7-level fatigue scale and the SART scale; the objective fatigue test is evaluated based on the psychomotor alertness test; The score SP of the controller is evaluated based on the Samn-Perelli 7-level fatigue scale score ; SART score for evaluating controllers based on the SART scale score ; Scored RTs for evaluating controllers based on the psychomotor vigilance test score ; Calculate the comprehensive fatigue score, Among them, Fatigue score Indicates the comprehensive fatigue score; when the comprehensive fatigue score Fatigue score When the set threshold is exceeded, the subject is marked as being in a fatigued state.
4. The method for predicting controller voice fatigue based on remote tower scenario according to claim 1, characterized in that: In step 2, the step of extracting prosodic features from a number of annotated speech signals includes: Prosodic features include speech rate, pause frequency, and pause duration; The formula for calculating speech speed is: Among them, SR represents the speaking speed; N s Indicates the total number of syllables pronounced in the speech sample; T d Indicates the duration of the speech sample segment; The calculation formula for the pause frequency is: Where PF represents the pause frequency; N p Indicates the number of pauses in the speech sample; T d Indicates the duration of the speech sample segment; The calculation formula for the pause duration is: Among them, PD represents the pause duration; Indicates the duration of each pause; i indicates the index of the i-th pause, i=1,2,...,N p ; The calculation formula of rhythmic features is: Among them, [SR, PF, PD] means splicing the speech rate, pause frequency and pause duration according to the time dimension; F prosody A vector representing prosodic features.
5. The method for predicting controller voice fatigue based on remote tower scenario according to claim 4 is characterized in that: In step 2, the step of extracting sound quality features from a plurality of annotated speech signals includes: Sound quality characteristics include sound pressure level and fundamental frequency; The formula for calculating the sound pressure level is, Among them, SPL represents sound pressure level; p represents current sound pressure; p0 represents reference sound pressure; The fundamental frequency is calculated as, Where F0 represents the basic frequency; N cycles Indicates the number of vibrations in a cycle; N period Indicates the cycle duration; The calculation formula for sound quality features is: Where [SPL, F0] represents the concatenation of the sound pressure level and fundamental frequency in the time dimension; F quality A vector representing the sound quality features.
6. The method for predicting controller voice fatigue based on remote tower scenario according to claim 5 is characterized in that: In step 2, the step of extracting spectral features from a plurality of annotated speech signals includes: Spectral features include formant and spectral entropy; The formula for calculating the resonance peak is: Among them, F n Indicates the frequency of the nth resonance peak, n=1,2,3; LPCC n Indicates the frequency of the nth resonance peak estimated by cepstrum analysis; The calculation formula of spectrum entropy is: Where SE represents spectral entropy; P(f) represents the probability density function corresponding to frequency f; The calculation formula of the spectral feature is: Among them, [F1, F2, F3, SE] means that the frequency of the first formant, the frequency of the second formant, the frequency of the third formant and the spectrum entropy are spliced according to the time dimension; F spectral A vector representing the spectrogram-like features.
7. The method for predicting controller voice fatigue based on remote tower scenario according to claim 6, characterized in that: In step 2, the step of fusing into comprehensive features includes: The rhythm features, sound quality features and spectral features are integrated to form a comprehensive feature. Among them, F fusion represents the comprehensive features after fusion; w1 represents F prosody The weight of F quality The weight of F spectral The weight of .
8. The method for predicting controller voice fatigue based on remote tower scenario according to claim 1, characterized in that: In step 3, the backward layer of the Bi-LSTM-GRU model includes T GRU units, and the comprehensive features of T time steps are input into T GRU units one by one; each GRU unit includes an update gate and a reset gate, and the calculation formula is as follows: A. Update gate, used to determine the comprehensive feature x of the current time step input t and the hidden state h at the next time step t+1 The hidden state h at the current time step t The influence of, and thus control the degree of retention of previous information at the current time step, Among them, Z t is the activation value of the update gate; represents the sigmoid activation function; W z is the update gate weight matrix; b z is the bias term of the update gate; h t+1 Represents the hidden state of the GRU unit from the next time step during back propagation; t∈T, T is the total number of time steps; B. Reset gate, used to control the comprehensive feature x input at the current time step t and the hidden state h at the next time step t+1 The degree of integration, Among them, r t is the activation value of the reset gate; W r is to reset the gate weight matrix; b r is the bias term for resetting the gate; C. Candidate hidden state, calculate the new information of the current time step, in, is the candidate hidden state; tanh represents the hyperbolic tangent activation function; W h is the candidate hidden state pair x t The weight matrix of h is the candidate hidden state pair h t+1 The weight matrix of h is the candidate hidden state bias; D. The final hidden state of the GRU unit, in, is the final hidden state output by the GRU unit at time step t.
9. The method for predicting controller voice fatigue based on remote tower scenario according to claim 8, characterized in that: In step 3, the forward layer of the Bi-LSTM-GRU model includes T LTSM units, and the comprehensive features of T time steps are input into the T LTSM units in a one-to-one correspondence; each LSTM unit includes an input gate, a forget gate, an output gate and a cell state, and the calculation formula is as follows: A. Input gate, used to determine the comprehensive feature x of the current time step input t and the hidden state h of the previous time step t-1 The impact on the current cell state, thereby controlling the degree of retention of subsequent information at the current time step, Among them, i t is the activation value of the input gate; represents the sigmoid activation function; W i is the input gate weight matrix; b i represents the bias term of the input gate; h t-1 Represents the hidden state of the LSTM unit from the previous time step during forward propagation; t∈T; B. Forget gate, used to control the cell state c of the previous time step t-1 The information that needs to be forgotten, Among them, f t is the activation value of the forget gate; W f is the forget gate weight matrix; b f is the bias term of the forget gate; C. Cell state, which determines the retention and updating of long-term memory, in, is the candidate cell state at the current time step; tanh represents the hyperbolic tangent activation function; W c is the cell state weight matrix; b c is the bias term of the cell state; c t is the cell state at the current time step; D. Output gate, used to control the hidden state h output at the current time step t , which depends on the cell state c at the current time step t , Among them, t is the activation value of the output gate; W o is the output gate weight matrix; b o is the bias term of the output gate; is the final hidden state output by the LTSM unit at time step t.
10. The method for predicting controller voice fatigue based on remote tower scenario according to claim 9, characterized in that: In step 3, the output layer of the Bi-LSTM-GRU model is composed of the output of the GRU unit and the output of the LSTM unit at each time step, wherein the GRU unit processes the reverse time series from the Tth time step to the first time step in time steps, and the LSTM unit processes the forward time series from the first time step to the Tth time step in time steps; finally, the outputs of the two directions are concatenated or summed at the same time step to obtain a bidirectional hidden state, Among them, H t Represents the bidirectional hidden state output at the current time step; After T time steps, the output layer outputs T bidirectional hidden states H t , t∈T, T bidirectional hidden states H t After the fully connected layer, it is fused into H t `, the fatigue prediction result output by the Bi-LSTM-GRU model is: in, is the fatigue prediction result output by the Bi-LSTM-GRU model; Softmax represents the Softmax activation function; W out is the output layer weight matrix; b out is the output layer bias term.
Citation Information
Patent Citations
Fatigue monitoring method based on land-air communication voice of air-traffic controller
CN110164471A
Deep learning-based voice fatigue detection method
CN114403878A
Handwritten mathematical formula identification method based on ResNet and attention mechanism
CN116721429A
Generative adversarial network architecture search method and system based on hybrid convolution operation
CN119150925A
Self-adaptive multi-modal feature fusion well logging interpretation method based on multi-task joint learning
CN119760633A