Flight manual intelligent query method and device based on voice recognition
By optimizing speech recognition technology to handle complex pronunciation situations and combining it with deep learning and dynamic time warping, the problems of time-consuming flight manual queries and the shortcomings of traditional technologies have been resolved, enabling fast and accurate queries and responses in emergency situations, thereby improving flight safety and operational efficiency.
Patent Information
- Application Number
- CN202510815364.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-26
AI Technical Summary
The existing flight manual query method takes a long time in emergency situations and cannot meet the needs of rapid response. In addition, traditional voice recognition technology is sensitive to environmental noise and lacks support for non-standard accents, making it unable to provide immediate guidance.
An intelligent flight manual query method based on speech recognition is adopted. By optimizing the recognition technology to handle pronunciation conditions such as unclear stress, rapid enunciation, local accent, low volume and large changes in speaking speed, combined with deep learning and dynamic time warping technology, it can identify commonly used flight terms and emergency instructions, generate text instructions and query the flight manual plan.
It achieves real-time and accurate recognition and response of voice input under complex pronunciation conditions, significantly improving the efficiency and safety of flight manual queries, and ensuring the rapid acquisition and execution of disposal plans in emergency situations.
Smart Images

Figure CN120708610A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of avionics systems, and in particular to a method and device for intelligently querying a flight manual based on speech recognition. Background Art
[0002] With the advancement of aviation technology, flight safety has become a shared focus for airlines and manufacturers. When faced with an emergency during flight, the flight crew must quickly and accurately retrieve and execute the procedures outlined in the flight manual. This is crucial for ensuring flight safety. Traditional manual review has demonstrated numerous shortcomings. It is not only time-consuming but also prone to errors in high-pressure environments. Internationally, the aviation industry has begun exploring and applying voice recognition technology to optimize the flight manual review process. For example, Boeing and Airbus both have proprietary voice recognition systems. While these systems have improved query efficiency to a certain extent, they still have some limitations in practical application, such as high sensitivity to ambient noise and insufficient support for non-standard accents.
[0003] The following is an explanation of the flight manual query method in the prior art:
[0004] 1. Manually consult the flight manual
[0005] Current situation: Flight crews must manually flip through thick and heavy flight manuals to find emergency response plans. Actual operational data shows that searching through paper manuals is excessively time-consuming, with emergency manuals often taking more than 30 seconds to complete. For example, in the event of a cabin fire, the average search time is 36.5 seconds, far exceeding the required emergency response time of no more than 30 seconds.
[0006] Disadvantages: Flight manual search takes a long time, typically exceeding 5 seconds, and in emergency situations, it can even take over 30 seconds. This fails to meet the critical need for quick emergency response times, potentially delaying optimal action and increasing flight risks. Manual search distracts the flight crew, potentially affecting their ability to monitor other flight operations, further increasing flight safety risks.
[0007] 2. Document US9719799B2 discloses the next generation of electronic flight bag (EFB) applications. Although this technology has improved data display capabilities, it lacks support for rapid response to emergencies, especially in terms of limited voice interaction capabilities. In an emergency, the flight crew may not be able to quickly find the required disposal plan. This technology does not integrate voice recognition functions and cannot provide effective assistance when the flight crew is in a hurry. This defect is particularly obvious in situations where quick decision-making is required. The use of EFB still requires manual operation by the flight crew, and it is impossible to completely free their hands, which affects the operational efficiency in emergency situations.
[0008] 3. Document US20240201696A1 discloses flight obstacle detection based on voice recognition. This technology is mainly aimed at autonomous flight of unmanned aerial vehicles and is not suitable for emergency response of commercial aircraft. Commercial aircraft require detailed response plans in emergency situations, not just obstacle avoidance. This technology lacks the function of providing instant guidance to the flight crew in emergency situations and cannot meet the needs of commercial aircraft for rapid response to emergencies. Emergency handling of commercial aircraft requires more complex and comprehensive guidance. The application scenarios of this technology are relatively single and cannot cover the needs of commercial aircraft in various emergency situations.
[0009] 4. Document CN118072563A discloses a method for detecting mid-air conflicts between aircraft based on semantic analysis of control speech. This method focuses on flight safety monitoring rather than providing specific emergency response plans. In an emergency, the flight crew requires specific response steps, not just conflict detection. This method cannot directly provide the flight crew with immediate guidance in emergency situations, especially in situations where rapid decision-making is required. This shortcoming may result in missing the optimal time to handle the situation. This method mainly relies on the controller's instructions and cannot provide independent support when the flight crew encounters an emergency.
[0010] 5. Document CN112363520A discloses an aircraft flight assisted piloting system and control method based on artificial intelligence technology. This method is more inclined towards the automation of flight operations and fails to provide immediate guidance in emergency situations. In an emergency, the flight crew requires a detailed response plan, not just automated operations. This method has limited processing capabilities for non-standard flight conditions and cannot provide the flight crew with a detailed response plan in an emergency. Emergency situations often exceed the scope of conventional operations and require more flexible and comprehensive support. This method may not provide sufficient support under extreme conditions, increasing flight risks.
[0011] 6. Document CN114299519A discloses an assisted flight method based on an electronic flight manual in XML format. This method mainly improves operational efficiency through visual assistance and does not involve voice interaction to adapt to emergency situations. In an emergency, the flight crew may not be able to distract themselves to view the screen and require more direct and quick guidance. This method lacks a mechanism for quickly obtaining information in an emergency and cannot meet the flight crew's needs to quickly obtain a disposal plan in a high-pressure environment. Decisions in emergencies need to be quick and accurate, and AR technology has limitations in this regard. The application scenario of this method is relatively single and is mainly suitable for routine flight operations. It cannot cover all needs in emergency situations.
[0012] 7. Document US20220147905A1 discloses a system and method for assisting crew members in executing release messages. This technology focuses on the optimization of daily operating processes and is not specifically designed for rapid response mechanisms in emergencies. In an emergency, the flight crew needs immediate guidance and support, not just process optimization. This technology lacks the function of providing immediate guidance to the flight crew in an emergency and cannot meet the needs of rapid decision-making in an emergency. Decisions in emergencies need to be quick and accurate, and this technology cannot provide sufficient support. This technology is mainly suitable for ground operations and pre-takeoff preparations. It has limited support for emergencies during flight and cannot fully cover the needs of the flight crew in various situations. Summary of the Invention
[0013] The embodiments of this specification provide a method and device for intelligently querying a flight manual based on voice recognition, so as to solve the technical problem of how to improve the effect and efficiency of flight manual query.
[0014] To solve the above technical problems, the embodiments of this specification provide the following technical solutions:
[0015] The embodiment of this specification provides a method for intelligently querying a flight manual based on speech recognition, the method comprising:
[0016] Acquiring target voice information and optimizing recognition of the target voice information;
[0017] Identify key information on the results obtained after optimized identification, and generate text instructions based on the key information identification results;
[0018] Searching for a corresponding target solution in the flight manual according to the text instruction and displaying the target solution;
[0019] The step of optimizing the recognition of the target voice information includes:
[0020] According to the pronunciation situation of the target voice information, an optimization recognition operation corresponding to the pronunciation situation is performed.
[0021] Preferably, for the pronunciation situation of the target voice information, performing an optimization recognition operation corresponding to the pronunciation situation includes:
[0022] The target speech information is detected to see if there is any unclear accent using a preset accent recognition model, and the detected unclear accent is corrected.
[0023] Preferably, for the pronunciation situation of the target voice information, performing an optimization recognition operation corresponding to the pronunciation situation includes:
[0024] It is detected whether the target voice information has a too fast speaking speed, and when it is detected that the speaking speed is too fast, the recognition accuracy of the short-duration voice segment is increased.
[0025] Preferably, detecting whether the target voice information is spoken too fast includes:
[0026] Extract audio features of the target voice information within a preset duration, compare the audio features with a standard pronunciation pattern, and determine whether the target voice information is spoken too fast.
[0027] Preferably, for the pronunciation situation of the target voice information, performing an optimization recognition operation corresponding to the pronunciation situation includes:
[0028] The target speech information is detected to determine whether it has a local accent, and when a local accent is detected, a preset local accent database is used to improve recognition accuracy in the case of a local accent.
[0029] Preferably, for the pronunciation situation of the target voice information, performing an optimization recognition operation corresponding to the pronunciation situation includes:
[0030] It is detected whether the target voice information has a low volume, and when the low volume is detected, automatic gain control is performed to improve the recognition accuracy in the low volume situation.
[0031] Preferably, for the pronunciation situation of the target voice information, performing an optimization recognition operation corresponding to the pronunciation situation includes:
[0032] Whether the target voice information has a large change in speech speed, and when it is detected that there is a large change in speech speed, dynamic time adjustment is performed to improve the recognition accuracy in the case of large change in speech speed.
[0033] Preferably, the target voice information is obtained by performing noise suppression processing on the original voice information.
[0034] Preferably, identifying key information of the results obtained after optimized identification includes:
[0035] Identify whether the optimized recognition results contain common flight terms and / or emergency instructions.
[0036] The embodiment of this specification provides a flight manual intelligent query device based on voice recognition, the device comprising:
[0037] An optimization recognition module is used to obtain target voice information and optimize the recognition of the target voice information;
[0038] A key information recognition module is used to identify key information from the results obtained after optimization and recognition, and generate text instructions based on the key information recognition results;
[0039] a solution query module, configured to query the flight manual for a corresponding target solution according to the text instruction and display the target solution;
[0040] The step of optimizing the recognition of the target voice information includes:
[0041] According to the pronunciation situation of the target voice information, an optimization recognition operation corresponding to the pronunciation situation is performed.
[0042] At least one of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects:
[0043] A user-customized flight crew voice recognition method that takes into account universality has been adopted to achieve real-time and accurate recognition and response of voice input, effectively improving the effect and efficiency of flight manual queries, and significantly improving flight safety and operational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following briefly describes the drawings required for use in the embodiments of this specification or the prior art description. Obviously, the following only describes the drawings required for use in some embodiments of this application. For those of ordinary skill in the art, other drawings can be derived from these drawings without inventive effort.
[0045] Figure 1 This is a schematic diagram of the flight manual intelligent query process based on voice recognition in the first embodiment of this specification.
[0046] Figure 2 This is a schematic diagram of the architecture of the first embodiment of this specification.
[0047] Figure 3 This is an interactive diagram in the first embodiment of this specification. DETAILED DESCRIPTION
[0048] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments involved in the specific implementation methods are only part of the embodiments of this application, not all of the embodiments. All other embodiments obtained based on the embodiments in the specific implementation methods by those skilled in the art without making creative work should fall within the scope of protection of this application.
[0049] The first embodiment of this specification (hereinafter referred to as "Embodiment 1") provides a method for intelligent querying a flight manual based on voice recognition. The execution subject of Embodiment 1 includes but is not limited to a terminal or a server or an operating system or an application, that is, the execution subject can be diverse and can be set, used or transformed as needed. In addition, a third-party application can also assist the execution subject in executing Embodiment 1. For example, the method for intelligent querying a flight manual based on voice recognition in Embodiment 1 can be executed by a server, and a corresponding application can be installed on a terminal (the terminal can be held by a user). Data can be transmitted between the terminal or the application and the server, thereby assisting the server in executing the method for intelligent querying a flight manual based on voice recognition in Embodiment 1.
[0050] refer to Figure 1 The flight manual intelligent query based on voice recognition provided in the first embodiment includes:
[0051] S101: Acquire target voice information and optimize recognition of the target voice information;
[0052] In Example 1, target voice information can be obtained. The target voice information can refer to voice emitted by aircraft-related personnel (including but not limited to crew members, the same below) (for example, relevant personnel emit voice and the execution subject of Example 1 receives it as target voice information), or it can be voice-related information emitted by aircraft-related personnel and generated after specific processing (for example, relevant personnel emit voice and the execution subject of Example 1 performs specific processing on the voice to generate target voice information). The target voice information can be voice or other forms, which are not limited to Example 1.
[0053] Preferably, the speech of the relevant personnel can be used as the original speech information, and noise suppression processing is performed on the original speech information to obtain the target speech information. Specifically, during the noise suppression process, the original speech information is first estimated and background noise is subtracted through spectral subtraction technology. This process operates primarily in the frequency domain, estimating the noise level by analyzing the spectrum of silent segments and removing this estimated value from the overall speech signal spectrum. In addition, an adaptive filtering method, namely the least mean square (LMS) algorithm, is used to dynamically adjust the filter parameters based on the input signal to reduce noise.
[0054] In the first embodiment, the target voice information is optimized for recognition, wherein the optimization for recognition of the target voice information may include: performing an optimization recognition operation corresponding to the pronunciation of the target voice information.
[0055] In actual situations, the pronunciation of relevant personnel may have various situations, such as unclear stress, rapid enunciation, local accent, low volume and large changes in speaking speed, etc. These situations can also be called non-standard pronunciation. In response to various situations that may occur in pronunciation, Example 1 performs corresponding optimized recognition operations. The following is a detailed description:
[0056] 1. Unclear accent
[0057] In embodiment 1, an optimized recognition operation corresponding to the pronunciation situation of the target voice information is performed, which may include: detecting whether the target voice information has unclear accent through a preset accent recognition model, and correcting the detected unclear accent.
[0058] Specifically, the preset accent recognition model is not limited to addressing unclear accents; it can also serve as part of a speech recognition system to assist in addressing other types of pronunciation issues. This model is trained or constructed based on deep learning frameworks, including but not limited to convolutional neural networks (CNNs) and recurrent neural networks (RNNs), particularly long short-term memory (LSTM) and gated recurrent units (GRUs). These architectures excel at capturing time-dependent features in audio signals. Specifically, the accent recognition model is trained using a large dataset containing diverse accent patterns. By analyzing features such as the spectrogram of the audio signal, it learns to distinguish between normal pronunciation and ambiguous accents. In practice, when the target speech signal is detected to have ambiguous accent, the accent recognition model can infer the most likely correct pronunciation based on contextual information and make appropriate corrections. It is worth noting that while this accent recognition model is primarily designed for unclear accents, it can also provide support for other pronunciation scenarios, including those described below, such as rapid enunciation and regional accents, by improving overall recognition accuracy through improved feature extraction methods. This means that the accent recognition model can be applied to other pronunciation scenarios as needed.
[0059] 2. Speaking too fast or too fast
[0060] In Example 1, an optimized recognition operation corresponding to the pronunciation situation of the target voice information is performed, which may include: detecting whether the target voice information has a fast speaking speed, and increasing the recognition accuracy of short-duration voice segments when it is detected that the speaking speed is too fast.
[0061] Among them, detecting whether the target voice information is spoken too fast may include: extracting audio features within a preset duration of the target voice information (the preset duration is generally short, such as a few milliseconds), comparing the audio features with a standard pronunciation pattern, and determining whether the target voice information is spoken too fast.
[0062] Specifically, when a fast speaking speed is detected, an algorithm based on Dynamic Time Warping (DTW) technology can be used to handle the fast speaking speed. DTW is an algorithm that measures the similarity between two time series and is particularly suitable for handling misalignment problems on the time axis caused by changes in speaking speed. Through the DTW algorithm, Example 1 can more accurately capture the speech characteristics of the target speech information at a fast speaking speed, thereby improving recognition accuracy.
[0063] 3. Regional accent
[0064] In the first embodiment, an optimized recognition operation corresponding to the pronunciation situation of the target voice information is performed, which may include: detecting whether the target voice information has a local accent, and when a local accent is detected, using a preset local accent database to improve the recognition accuracy in the case of a local accent.
[0065] Specifically, local accent samples from various regions can be collected in advance, training data for multiple local accents can be integrated into a comprehensive database. This comprehensive database can then be used to better understand and recognize pronunciations with local characteristics. For example, when a pilot speaks with a certain local accent, optimized recognition operations can compare the target speech information with samples in the comprehensive database to identify the pilot's local accent, thereby improving the recognition accuracy of the target speech information.
[0066] Preferably, after successfully identifying or detecting the local accent in the target voice information, in order to further improve the recognition rate, Example 1 adjusts the parameter settings of the voice recognition engine according to the characteristics of the local accent. Utilizing adaptive training technology, data specific to the local accent is added to the original general voice recognition model for fine-tuning to obtain a new voice recognition model, so that the newly obtained voice recognition model is more adaptable to the current pronunciation habits with local characteristics. In addition, personalized adjustment of the acoustic model is introduced, and the parameter configuration of the acoustic model is optimized according to the characteristics of different dialects collected in the local accent database, so that it is more sensitive to the sound characteristics of the current specific dialect. This approach not only improves the recognition accuracy, but also enhances the robustness and user experience of Example 1. Specifically, this personalized adjustment involves adjusting the feature extraction layer or decoding strategy in the acoustic model to better capture and understand the uniqueness of the local accent.
[0067] 4. Low volume
[0068] In Example 1, an optimized recognition operation corresponding to the pronunciation situation of the target voice information is performed, which may include: detecting whether the target voice information has a low volume situation (for example, judging whether there is a low volume situation based on the decibel number of the target voice information), and when a low volume situation is detected, improving the recognition accuracy in the low volume situation through automatic gain control.
[0069] In other words, automatic gain control is used for low-volume speech. This method amplifies low-volume sound signals without affecting other audio characteristics, thereby improving speech recognition in low-volume conditions. Specifically, this method dynamically adjusts the gain of the audio signal of the target voice information, ensuring that even whispered speech can be accurately captured and recognized.
[0070] For example, when a pilot speaks in a low voice in a noisy environment or for health reasons, optimizing the recognition operation can automatically enhance these low-volume sound signals in the target voice information, thereby improving the recognition rate.
[0071] 5. Large changes in speaking speed
[0072] In embodiment one, an optimized recognition operation corresponding to the pronunciation situation of the target voice information is performed, which may include: whether the target voice information has a large change in speech speed, and when a large change in speech speed is detected, dynamic time adjustment is performed to improve the recognition accuracy in the case of large change in speech speed.
[0073] Specifically, Dynamic Time Warping (DTW) is used to address large variations in speech rate. Specifically, the DTW algorithm adjusts the time axis of the target speech at different speech rates, ensuring accurate matching to the pre-set speech model even with large variations in speech rate. The DTW algorithm adapts to varying speech rates by calculating the optimal matching path between two sequences.
[0074] For example, a pilot may speak very quickly in an emergency situation, but more slowly in normal situations. By optimizing the recognition operation, the DTW algorithm can be used to adjust the timeline to ensure that each word is correctly recognized.
[0075] Through the above, we achieve recognition of the target voice information and perform corresponding optimized recognition operations for various pronunciation conditions in the target voice information. After optimized recognition, the optimized recognition results are optimized voice features, providing high-quality input for subsequent voice recognition and ensuring that voice commands can be accurately recognized.
[0076] It has been verified in practice that the recognition effect of the target voice information can be effectively improved through the corresponding optimization recognition operation. The specific verification results are shown in Table 1:
[0077]
[0078] Table 1
[0079] As shown in Table 1, by optimizing recognition operations for different pronunciation scenarios, the recognition rate of target speech information has been significantly improved. For example, by introducing an additional accent recognition model, the recognition rate for speech with unclear accents increased by 15%. This proves that optimized recognition ensures accurate recognition of target speech information under various complex pronunciation conditions, significantly improving speech recognition performance and efficiency.
[0080] S103: performing key information recognition on the result obtained after the optimized recognition, and generating a text instruction according to the key information recognition result;
[0081] In the first embodiment, key information is identified based on the optimized recognition results. This identification may include determining whether the optimized recognition results contain common flight terms and / or emergency instructions. This key information includes, but is not limited to, common flight terms and / or emergency instructions.
[0082] Taking common flight terms and emergency instructions as an example, in order to further improve the recognition ability of common flight terms and emergency instruction vocabulary, Example 1 adopts a method of adding specific flight terminology training data and performing specific recognition training on common flight terms and emergency instruction vocabulary. The specific approach is to expand the training set containing various types of flight professional terminology through actual flight data and simulated flight data, and especially strengthen the learning of special vocabulary in the aviation field. In this way, special training on common flight terms and emergency instruction vocabulary can significantly improve the recognition accuracy of complex or uncommon vocabulary.
[0083] For example, aviation terminology contains many proper nouns and technical terms (such as "ICE DETECTED" and "FIRE WARNING") that are not common in everyday language but are very important in flight operations. By increasing the training data for these specific terms, we can more accurately recognize these terms in real-world application scenarios.
[0084] Practical verification has shown that specific recognition training for commonly used flight terms and emergency instruction vocabulary can effectively improve the recognition of key information, including commonly used flight terms and emergency instructions. The specific verification results are shown in Table 2:
[0085]
[0086]
[0087] Table 2
[0088] As shown in Table 2, specific recognition training for commonly used flight terms and emergency instructions significantly improved the recognition rate of key information, including these terms. For example, by increasing training data for specific flight terms, the recognition rate for special terms increased by 25%. Training on the command "ICE DETECTED" increased the recognition rate for voice recognition of this command by 30%, and training on the command "FIRE WARNING" increased the recognition rate for voice recognition of this command by 32%. This demonstrates that specific recognition training for commonly used flight terms and emergency instructions significantly improves the speed and accuracy of voice recognition in emergency situations, ensuring that voice recognition can respond quickly in emergencies and significantly improving the effectiveness and efficiency of voice recognition.
[0089] After the above-mentioned key information identification, the key information identification results (such as the identified common flight terms and / or emergency instructions) can be generated into text instructions.
[0090] The execution body of embodiment one may have corresponding modules to respectively perform the above-mentioned optimization recognition and key information recognition, and the modules may be a speech recognition framework based on Mel-frequency cepstral coefficients (MFCC) feature extraction and Gaussian mixture model-hidden Markov model (GMM-HMM).
[0091] Among them, the module for performing optimized recognition (hereinafter referred to as the optimized recognition module) can be located at the beginning of the voice input, receive the voice instructions of the relevant personnel, perform various preprocessing on the voice signal including the aforementioned noise suppression, and deeply customize it for various possible pronunciation situations, ensuring a high recognition rate through optimized recognition.
[0092] The module for performing critical information recognition (hereinafter referred to as the CRI module) combines actual and simulated flight data to conduct intensive training specifically targeting critical information. This ensures rapid speech recognition response in emergency situations, significantly improving both the speed and accuracy of speech recognition in these situations. The CRI module receives optimized speech features from the optimized recognition module and further identifies and parses voice commands, ensuring rapid and accurate recognition of flight crew voice commands. Its output is the recognized text command, providing accurate input for subsequent flight manual queries.
[0093] Each model involved in Example 1 can be constructed in advance and deployed on a corresponding system or platform as part of Example 1, and the system or platform can be used to execute Example 1; or Example 1 can be written as an application, and each model runs as a partial module or function of the application.
[0094] S105: Querying a corresponding target solution in the flight manual according to the text instruction, and displaying the target solution.
[0095] In Example 1, after generating the text instruction, the corresponding target solution can be queried in the flight manual according to the text instruction (the query of the flight manual can be implemented through a system or application that can perform flight manual retrieval query, and these systems or applications can store the content of the flight manual), and the target solution can be displayed.
[0096] Specifically, natural language processing can be used to parse text commands to determine the true intent of the person speaking. A query command is then generated based on the analysis results. The query command is then passed to the flight manual for querying, resulting in the corresponding target solution. This solution is then passed to the front-end for display. The front-end display can dynamically adjust the displayed content based on the query results, ensuring that the information presented is both intuitive and easy to understand.
[0097] The following is an example to further illustrate the content of the first embodiment:
[0098] refer to Figure 2 , embodiment 1 can be executed by a corresponding on-board system.
[0099] Specifically, the flight crew issues voice commands, which are collected through a microphone. After pre-processing such as noise suppression, the collected voice is optimized for recognition as target voice information to ensure that the voice content can be accurately parsed later. The optimized recognition results are then processed through key information recognition to ensure that the voice commands can be quickly and accurately converted into text commands. The text commands are parsed through natural language processing to determine the true intentions of the person issuing the voice commands, and query commands are generated based on the parsed results. The query commands are passed to the flight manual for query, thereby finding the corresponding disposal plan (belonging to the target plan), and the found disposal plan is passed to the front-end human-computer interaction interface for display, including displaying it in a form that is easy for the flight crew to understand, so that the disposal plan in the flight manual can be given in a timely and accurate manner. The dialogue between the flight crew and the system is realized through the human-computer interaction interface, which is a direct channel for interaction between the flight crew and the system.
[0100] The first embodiment can be applied to an electronic flight bag, so the execution subject of the first embodiment can be an electronic flight bag or a device or system deployed with an electronic flight bag.
[0101] The feasibility of the first embodiment has been verified, as follows:
[0102] The method provided in the first embodiment has been implemented in a desktop simulation environment, for example Figure 3 To verify the feasibility and effectiveness of Example 1, a large number of experimental tests were conducted. The following are the specific content and results of the experiments, including key indicators such as word error rate (WER), recognition accuracy, and performance under different signal-to-noise ratios (SNR).
[0103] 1. Experimental Setup
[0104] 1.1. Test environment:
[0105] ●Simulated flight environment: Use professional audio software to simulate various noises during flight, including engine noise, wind noise, electronic equipment noise, etc.
[0106] Actual flight environment: Voice data recorded in a real cockpit flight environment, covering different flight phases and emergency situations.
[0107] 1.2. Test data:
[0108] Simulation Data: Contains 20,000 simulated voice commands, covering commonly used terms and emergency commands for flight crews.
[0109] ●Real-world data: Contains 10,000 actual in-flight voice commands, covering various flight phases and emergency situations.
[0110] 1.3. Test indicators:
[0111] Word Error Rate (WER / CER): This is an important indicator for measuring the performance of a speech recognition system, indicating the proportion of incorrect characters in the recognition results. The calculation formula is:
[0112]
[0113] Where S represents the number of incorrectly replaced characters, D represents the number of incorrectly deleted characters, and I represents the number of incorrectly inserted characters; N represents the total number of characters in the reference text. Lower WER values indicate better performance of the speech recognition system. For example, a system with a WER of 10% means that 1 out of every 10 characters is incorrectly recognized.
[0114] ●Recognition accuracy: refers to the ratio of characters correctly recognized by the speech recognition system to the total number of characters.
[0115] The calculation formula is:
[0116]
[0117] The higher the recognition accuracy, the better the performance of the speech recognition system. For example, if a system has a recognition accuracy of 90%, it means that 9 out of every 10 characters are recognized correctly.
[0118] Performance under different signal-to-noise ratios (SNRs): A measure of signal strength relative to the strength of background noise, usually expressed in decibels (dB). In speech recognition, the signal-to-noise ratio reflects the relative strength of the speech signal and background noise. A high signal-to-noise ratio (such as 30dB) indicates that the background noise is low and the speech signal is strong. The system should exhibit high recognition accuracy and low word error rate. A low signal-to-noise ratio (such as 0dB) indicates that the background noise is high and the speech signal is weak. The system's performance may decline, but it still needs to maintain a certain recognition accuracy and word error rate. By testing at different signal-to-noise ratios, the robustness and applicability of the system can be comprehensively evaluated to ensure that it can work effectively in various practical application scenarios.
[0119] 2. Test results
[0120]
[0121]
[0122] Table 3: Word error rate at different signal-to-noise ratios
[0123] Table 3 shows the word error rate (WER) for simulated and real data at different signal-to-noise ratios. As the SNR increases, the WER gradually decreases. At a 30dB SNR, the WER for the simulated and real data drops to 3.2% and 5.4%, respectively. This demonstrates that in a high SNR environment, the WER of Example 1 is significantly reduced, significantly improving performance.
[0124] Signal-to-noise ratio (dB) Accuracy of simulated data (%) Actual data accuracy (%) 0 64.8 61.5 5 71.9 68.6 10 79.7 76.3 15 85.8 82.5 20 90.2 87.9 25 93.7 91.5 30 96.8 94.6
[0125] Table 4: Recognition accuracy under different signal-to-noise ratios
[0126] Table 4 shows the recognition accuracy of simulated and real-world data at different signal-to-noise ratios. As the signal-to-noise ratio increases, the recognition accuracy gradually improves. At a signal-to-noise ratio of 30 dB, the recognition accuracy for the simulated and real-world data reaches 96.8% and 94.6%, respectively. This demonstrates that in high-SNR environments, the recognition accuracy of Example 1 significantly improves, demonstrating excellent performance.
[0127] 3. Test Conclusion
[0128] The above experimental results show that Example 1 exhibits excellent performance under different SNR environments. In particular, under high SNR conditions (20dB and above), Example 1 significantly reduces word error rate and improves recognition accuracy, meeting the requirements of practical applications.
[0129] Specifically:
[0130] ●Word error rate: When the signal-to-noise ratio is 20dB, the word error rates of simulated data and actual data are 9.8% and 12.1% respectively; when the signal-to-noise ratio is 30dB, the word error rates drop to 3.2% and 5.4% respectively.
[0131] Recognition accuracy: When the signal-to-noise ratio (SNR) is 20dB, the recognition accuracy for simulated data and real data is 90.2% and 87.9%, respectively. When the SNR is 30dB, the recognition accuracy reaches 96.8% and 94.6%, respectively.
[0132] These results demonstrate that Example 1 is highly practical and reliable in actual flight environments, significantly improving the efficiency of flight crews in acquiring and executing emergency response plans, reducing operational errors, and enhancing flight safety. Therefore, Example 1 is feasible for deployment on aircraft.
[0133] The first embodiment can achieve the following beneficial effects:
[0134] Example 1 adopts a customized flight manual query method that takes into account universality. By optimizing recognition customization, it achieves efficient and accurate recognition of specific pronunciation situations. Through specific training for key information recognition, it achieves efficient and accurate recognition of key information, including common flight terms and emergency instructions. Example 1 optimizes the flight manual query process through intelligent means, achieving real-time and accurate recognition and response to flight crew voice input. It can quickly and accurately help relevant personnel (including flight crew) obtain information in the flight manual in emergency situations, significantly improving flight safety and operational efficiency, and effectively supporting the flight crew's rapid decision-making and action in emergency situations.
[0135] Among them, the existing speech recognition technology has the problem of low recognition accuracy when dealing with the non-standard pronunciation of pilots (such as unclear stress, too fast pronunciation, local accent, low volume and large changes in speaking speed, etc.). Example 1 performs deep customization and optimization for various pronunciation habits to achieve corresponding optimized recognition operations. For example, additional stress recognition models are introduced, the recognition accuracy of short-duration voice segments is increased, training data of multiple local accents is integrated, automatic gain control and dynamic time warping (DTW) technology are used. This deep customization method can adapt to the diverse pronunciation habits of various groups of people, including pilots, and ensure that voice commands can be accurately recognized under various complex conditions, significantly improving the accuracy of speech recognition under complex pronunciation conditions, and improving speech recognition effect and efficiency.
[0136] In the aviation field, common flight terms and emergency command vocabulary (such as "ICE DETECTED", "FIRE WARNING", etc.) are highly professional and important. However, the existing system has limited recognition capabilities for these terms, especially in emergency situations. Example 1 combines actual flight data and simulated flight data to conduct special intensive training on various key information, which significantly improves the response speed and accuracy of speech recognition in emergency situations, and improves the effect and efficiency of speech recognition. This special training method ensures that relevant personnel (such as flight crews) can obtain accurate command recognition results in a timely manner, thereby improving flight safety and operational efficiency.
[0137] Example 1 can be integrated into an electronic flight bag, and a dialogue between relevant personnel and the subject used to execute Example 1 can be realized through a human-computer interaction interface, and ultimately the handling plan in the flight manual can be provided in a timely and accurate manner, thereby greatly improving the user experience and operational convenience, and improving flight safety and operational efficiency. It is particularly suitable for flight manual inquiries of commercial aircraft.
[0138] A second embodiment of this specification provides a flight manual intelligent query device based on voice recognition corresponding to the method described in the first embodiment, the device comprising:
[0139] An optimization recognition module is used to obtain target voice information and optimize the recognition of the target voice information;
[0140] A key information recognition module is used to identify key information from the results obtained after optimization and recognition, and generate text instructions based on the key information recognition results;
[0141] a solution query module, configured to query the flight manual for a corresponding target solution according to the text instruction and display the target solution;
[0142] The step of optimizing the recognition of the target voice information includes:
[0143] According to the pronunciation situation of the target voice information, an optimization recognition operation corresponding to the pronunciation situation is performed.
[0144] Preferably, for the pronunciation situation of the target voice information, performing an optimization recognition operation corresponding to the pronunciation situation includes:
[0145] The target speech information is detected to see if there is any unclear accent using a preset accent recognition model, and the detected unclear accent is corrected.
[0146] Preferably, for the pronunciation situation of the target voice information, performing an optimization recognition operation corresponding to the pronunciation situation includes:
[0147] It is detected whether the target voice information has a too fast speaking speed, and when it is detected that the speaking speed is too fast, the recognition accuracy of the short-duration voice segment is increased.
[0148] Preferably, detecting whether the target voice information is spoken too fast includes:
[0149] Extract audio features of the target voice information within a preset duration, compare the audio features with a standard pronunciation pattern, and determine whether the target voice information is spoken too fast.
[0150] Preferably, for the pronunciation situation of the target voice information, performing an optimization recognition operation corresponding to the pronunciation situation includes:
[0151] The target speech information is detected to determine whether it has a local accent, and when a local accent is detected, a preset local accent database is used to improve recognition accuracy in the case of a local accent.
[0152] Preferably, for the pronunciation situation of the target voice information, performing an optimization recognition operation corresponding to the pronunciation situation includes:
[0153] It is detected whether the target voice information has a low volume, and when the low volume is detected, automatic gain control is performed to improve the recognition accuracy in the low volume situation.
[0154] Preferably, for the pronunciation situation of the target voice information, performing an optimization recognition operation corresponding to the pronunciation situation includes:
[0155] Whether the target voice information has a large change in speech speed, and when it is detected that there is a large change in speech speed, dynamic time adjustment is performed to improve the recognition accuracy in the case of large change in speech speed.
[0156] Preferably, the target voice information is obtained by performing noise suppression processing on the original voice information.
[0157] Preferably, identifying key information of the results obtained after optimized identification includes:
[0158] Identify whether the optimized recognition results contain common flight terms and / or emergency instructions.
[0159] The contents not described in detail in the first and second embodiments can be referred to each other. The second embodiment can achieve the same beneficial effects as the first embodiment. The above embodiments can be used in combination.
[0160] The foregoing is merely an embodiment of the present invention and is not intended to limit the present application. For those skilled in the art, various modifications and variations may be made to the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included within the scope of the claims of the present application.
Claims
1. A flight manual intelligent query method based on speech recognition, characterized in that: The method comprises: Acquiring target voice information and optimizing recognition of the target voice information; Identify key information on the results obtained after optimized identification, and generate text instructions based on the key information identification results; Searching for a corresponding target solution in the flight manual according to the text instruction and displaying the target solution; The step of optimizing the recognition of the target voice information includes: According to the pronunciation situation of the target voice information, an optimization recognition operation corresponding to the pronunciation situation is performed.
2. The method according to claim 1, wherein According to the pronunciation of the target voice information, an optimization recognition operation corresponding to the pronunciation is performed, including: The target speech information is detected to see if there is any unclear accent using a preset accent recognition model, and the detected unclear accent is corrected.
3. The method according to claim 1, wherein According to the pronunciation of the target voice information, an optimization recognition operation corresponding to the pronunciation is performed, including: It is detected whether the target voice information has a too fast speaking speed, and when it is detected that the speaking speed is too fast, the recognition accuracy of the short-duration voice segment is increased.
4. The method according to claim 3, wherein Detecting whether the target voice information is spoken too fast includes: Extract audio features of the target voice information within a preset duration, compare the audio features with a standard pronunciation pattern, and determine whether the target voice information is spoken too fast.
5. The method according to claim 1, wherein According to the pronunciation of the target voice information, an optimization recognition operation corresponding to the pronunciation is performed, including: The target speech information is detected to determine whether it has a local accent, and when a local accent is detected, a preset local accent database is used to improve recognition accuracy in the case of a local accent.
6. The method according to claim 1, wherein According to the pronunciation of the target voice information, an optimization recognition operation corresponding to the pronunciation is performed, including: It is detected whether the target voice information has a low volume, and when the low volume is detected, automatic gain control is performed to improve the recognition accuracy in the low volume situation.
7. The method according to claim 1, wherein According to the pronunciation of the target voice information, an optimization recognition operation corresponding to the pronunciation is performed, including: Whether the target voice information has a large change in speech speed, and when it is detected that there is a large change in speech speed, dynamic time adjustment is performed to improve the recognition accuracy in the case of large change in speech speed.
8. The method according to claim 1, wherein The target voice information is obtained by performing noise suppression processing on the original voice information.
9. The method according to any one of claims 1 to 8, characterized in that Identification of key information from the optimized identification results includes: Identify whether the optimized recognition results contain common flight terms and / or emergency instructions.
10. A flight manual intelligent query device based on voice recognition, characterized in that: The device comprises: An optimization recognition module is used to obtain target voice information and optimize the recognition of the target voice information; A key information recognition module is used to identify key information from the results obtained after optimization and recognition, and generate text instructions based on the key information recognition results; a solution query module, configured to query the flight manual for a corresponding target solution according to the text instruction and display the target solution; The step of optimizing the recognition of the target voice information includes: According to the pronunciation situation of the target voice information, an optimization recognition operation corresponding to the pronunciation situation is performed.
Citation Information
Patent Citations
Aircraft flight auxiliary driving system based on artificial intelligence technology and control method
CN112363520A
Auxiliary flight method based on XML format electronic flight manual
CN114299519A
Aircraft air conflict detection method based on control voice semantic analysis
CN118072563A
System and method for assisting flight crew with the execution of clearance messages
US20220147905A1
Autonomous detect and avoid from speech recognition and analysis
US20240201696A1