Voice instruction calling handshake verification method and system

By combining speech recognition and voiceprint feature analysis with flight status recognition and standardization verification, the risks and unclear responsibilities associated with manual listening in aviation flights have been resolved. This has enabled automated verification and accountability for critical commands, thereby improving flight safety and reliability.

CN121789671APending Publication Date: 2026-04-03CHINA EASTERN TECH APPL RES & DEV CENT CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies rely on human hearing and interpretation of voice commands during flight, which carries risks such as missed hearing, mishearing, and delayed response. Traditional voice recognition systems have low recognition rates in noisy cockpit environments, lack role recognition leading to unclear responsibilities, lack automatic verification of command standardization, and fail to achieve intelligent matching based on flight status.

Method used

Employing technologies such as speech recognition, status recognition, text matching, role recognition, and instruction conformity judgment, and through speech preprocessing, voiceprint feature analysis, and multi-dimensional real-time parameter fusion, it achieves automatic recognition and verification of key instructions, including speech acquisition, preprocessing, text conversion, conformity verification, role binding, and anomaly alarm.

Benefits of technology

It improves flight safety, reduces human error in communication, ensures the authenticity, correctness, and standardization of instructions, clarifies the responsible parties, and enhances the reliability and security of instruction interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789671A_ABST
    Figure CN121789671A_ABST
Patent Text Reader

Abstract

The invention discloses a voice instruction shouting handshake verification method and system, and aims to realize automatic identification and verification of key instructions in flight, reduce man-made communication errors and improve flight safety. The method comprises the following steps: acquiring an original voice signal when an object executes a task; preprocessing the original voice signal to obtain a structured voice segment; converting the voice segment into text information, matching the text information with a standard instruction library, and preliminarily judging whether the text information belongs to a target instruction; the recognized text information is screened; performing format accuracy and normalization verification on the recognized text information; dynamically associating the voice instruction with the current task execution state, and verifying the time sequence rationality and scene consistency of the instruction; the speakers of different roles are distinguished based on voiceprint features, and accurate binding of the instruction and the responsibility subject is achieved; and finally verifying the authenticity, correctness and normalization of the instruction, triggering an alarm or prompt when the instruction is abnormal, and outputting a structured instruction text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of intelligent human-computer interaction, specifically to a multi-dimensional recognition, verification, and responsibility binding technology for cockpit voice commands. It is applicable to the application scenarios of automatic listening, standardization verification, role attribution determination, and anomaly warning of pilots' standardized voice commands during key phases such as taxiing, takeoff, climb, cruise, approach, landing, and runway of civil aviation and general aviation aircraft. Background Technology

[0002] During flight, standardized call-outs between pilots in the cockpit are a fundamental aspect of ensuring flight safety. For example, during takeoff, the co-pilot will issue key takeoff instructions, which the captain must confirm and execute accordingly; similarly, there are standardized call-outs during approach and landing. The accurate transmission and correct response to these voice instructions directly affect the standardization and safety of flight operations.

[0003] Currently, the transmission and verification of such critical instructions rely entirely on manual listening and subjective judgment. Pilots rely on their hearing to receive the other party's announcements and judge, based on their own experience, whether the content is correct, complete, and in accordance with the requirements of the current flight phase. The entire process lacks automated assistance and no technical system has been introduced to verify the authenticity, standardization, or timing of the announcements in real time.

[0004] However, this method, which relies entirely on manual operation, carries significant risks. First, the cockpit environment is noisy, including engine noise, airflow sounds, alarm tones, and interference from radio communications, which can easily lead to missed hearings, mishearings, or delayed reactions. Second, traditional general-purpose speech recognition systems have low recognition rates in high-noise, multi-speaker, and highly technical cockpit environments, making it difficult to meet the stringent accuracy requirements of aviation safety. Third, existing methods do not differentiate between the speakers, failing to assign specific instructions to specific roles such as captain, first officer, or observer, resulting in blurred boundaries of responsibility and hindering post-incident review and training evaluation. Fourth, there is a lack of automated verification mechanisms for the standardization of the instructions themselves—such as whether standard terminology is used, whether parameters are complete, whether the word order is correct, and whether vague expressions (such as "more or less" or "probably") are included. Fifth, the flight process has obvious phased characteristics (such as taxiing, takeoff, climb, cruise, descent, approach, landing, and runway), and different standard callout requirements correspond to different phases. However, existing technology has failed to dynamically associate the content of voice commands with real-time flight status (such as airspeed, altitude, vertical speed, flap / landing gear position, engine thrust, etc.), and cannot determine whether a certain command is reasonable or necessary at the current phase.

[0005] In summary, the existing technology has the following problems: (1) It relies entirely on human hearing, which poses risks such as missed hearing, mishearing, and delayed response; (2) The cockpit environment is noisy, and traditional voice recognition systems are difficult to achieve an effective recognition rate; (3) The different voices spoken by different roles are significantly different, and the lack of role recognition makes it difficult to distinguish responsibilities; (4) The standardization of instructions lacks automatic verification, such as "flight altitude, heading, speed" which are prone to incompleteness or incorrect sequence; (5) The flight phase is dynamic and changeable, and intelligent instruction matching is not combined with the flight status. These problems together restrict the safety and reliability of cockpit voice interaction. Summary of the Invention

[0006] The following provides a brief overview of one or more aspects to offer a basic understanding of them. This overview is not an exhaustive summary of all conceived aspects, nor is it intended to identify key or decisive elements of all aspects, nor to define the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form to prepare for the more detailed descriptions that follow.

[0007] The purpose of this invention is to solve the above-mentioned problems and provide a voice command handshake verification method and system. Through technologies such as voice recognition, status recognition, text matching, role recognition, and command standardization judgment, it realizes the automatic identification and verification of key commands in flight, reduces human communication errors, and improves flight safety.

[0008] The technical solution of this invention is as follows: This invention discloses a voice command handshake verification method, the method comprising: Step S1: Collect the raw audio signal of the subject while performing the task; Step S2: Preprocess the acquired raw speech signal to obtain structured speech segments; Step S3: Convert the structured speech fragments obtained in step S2 into text information and match them with the standard instruction library to preliminarily determine whether they belong to the target instruction; Step S4: Filter the text information identified in step S3. If the current speech is not a key instruction, no further processing will be performed. If the current speech is a key instruction, the subsequent processing in step S5 will be performed. Step S5: Verify the format accuracy and standardization of the text information identified in Step S3; Step S6: Dynamically associate the recognized voice commands with the current task execution status to verify the timing rationality and scenario consistency of the commands; Step S7: Distinguish speakers of different roles based on voiceprint features to achieve precise binding of instructions with responsible parties; Step S8: Perform final verification of the authenticity, correctness and standardization of the instructions, and trigger alarms or prompts when abnormalities occur, and finally output structured instruction text.

[0009] According to an embodiment of the voice command handshake verification method of the present invention, in step S2, the preprocessing includes high-frequency and low-frequency noise filtering, voice detection and silence removal, segmentation processing, normalization and feature extraction.

[0010] According to an embodiment of the voice command handshake verification method of the present invention, in step S4, a dual-channel structure of acoustics and text is adopted. The first sub-model of convolutional neural network / convolutional neural network and long short-term memory network is used to process the acoustic features of the speech signal and determine whether the speech has specific command voiceprint features. The second sub-model of semantic classification model is used to analyze the text content after speech-to-text conversion and determine whether the semantics are related to the standard command. Then, the confidence outputs of the two sub-models are fused to classify the speech into command speech and non-command speech.

[0011] According to an embodiment of the voice command handshake verification method of the present invention, in step S4, the process of detecting whether the voice belongs to a key command further includes introducing a keyword dictionary, SOP parameter detection, command semantic similarity judgment, and stage restriction strategy.

[0012] According to an embodiment of the voice command handshake verification method of the present invention, in step S5, the text output by voice recognition is subjected to multi-level parsing, including keyword parsing, word order judgment, command semantic deconstruction, and semantic similarity matching; The format accuracy and standardization verification includes the verification of items such as instruction structure, content, word order, auxiliary words, and password strength.

[0013] According to an embodiment of the voice command handshake verification method of the present invention, in step S6, the processing of the dynamic association between the voice command and the current flight state further includes: By integrating multi-dimensional real-time parameters from airborne sensors and flight control systems, a comprehensive perception of the current flight status can be constructed. By combining a rule base and a deep learning classification model based on multi-label classification, dynamic association between voice commands and the current flight status is achieved, including hard rule matching, vector matching of command semantics and stages, and a multi-label flight status classification network.

[0014] This invention also discloses a voice command-based handshake verification system, the system comprising: The voice acquisition module is used to acquire the raw voice signal of the object when performing the task; The speech preprocessing module is used to preprocess the acquired raw speech signal to obtain structured speech segments; The speech-to-text and instruction recognition module is used to convert the structured speech segments obtained by the speech preprocessing module into text information and match them with the standard instruction library to initially determine whether they belong to the target instruction. The non-command speech detection module is used to filter the text information recognized by the speech-to-text and command recognition modules. If the current speech is not a key command, no further processing will be performed. If the current speech is a key command, it will be processed by the subsequent command standardization judgment module. The instruction standardization judgment module is used to verify the format accuracy and standardization of the text information recognized by the speech-to-text and instruction recognition modules; The content status classification module is used to dynamically associate the recognized voice commands with the current task execution status to verify the timing rationality and scenario consistency of the commands. The role recognition module is used to distinguish speakers of different roles based on voiceprint features, so as to accurately bind instructions with responsible parties. The verification and alarm module is used to perform final verification of the authenticity, correctness and standardization of the instructions, and to trigger alarms or prompts when an anomaly occurs, and finally outputs the structured instruction text.

[0015] The present invention also discloses an electronic device, the electronic device including a controller, the controller including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; When the processor executes the program stored in memory, it implements the steps of the voice command shouting handshake verification method as described above.

[0016] The present invention also discloses a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the voice command shouting handshake verification method as described above.

[0017] The present invention also discloses a computer program product, which, when executed by a processor, implements the steps of the voice command shouting handshake verification method as described above.

[0018] Compared with the prior art, the present invention has the following advantages: First, this invention employs a multi-module collaborative technical architecture, integrating speech recognition, text matching, instruction conformity assessment, status recognition, and role recognition, to establish a multi-dimensional verification mechanism for instruction broadcasts. This mechanism comprehensively verifies each broadcast from three dimensions: authenticity (whether it is a valid instruction), correctness (whether the content is accurate), and conformity (whether it complies with standard operating procedures, SOPs). This technical feature enables the system not only to identify what the pilot said but also to determine whether it was "said correctly," "said completely," and "said appropriately," thereby significantly improving the reliability and security of instruction interaction.

[0019] Secondly, this invention introduces a dynamic correlation mechanism between flight phases and command content. By fusing aircraft sensor data and flight control parameters (such as airspeed, altitude, vertical speed, flap / landing gear status, etc.), it identifies the current flight phase (such as taxiing, takeoff, approach, etc.) in real time and dynamically matches the set of allowed standard commands accordingly. This technical feature ensures that the system only accepts specific commands during reasonable flight phases. For example, during the approach phase, it focuses on verifying announcements such as "Gear Down" and "Flaps XXX," while not triggering such verifications during the cruise phase. This avoids misjudgments out of context and improves the contextual awareness and timing rationality of command verification.

[0020] Third, this invention achieves role-based voiceprint binding through voiceprint feature extraction and comparison technology. Reference voiceprints are pre-collected during pilots' onboarding or training phases to establish a "role-voiceprint feature" database. During flight, voiceprint features such as MFCC, Fbank, PLP, or X-vector are extracted from each speech segment. Combined with a sliding window stabilization algorithm and temporal consistency verification, the identities of speakers such as captain, first officer, and observer are accurately distinguished. This technology clearly assigns each instruction to a specific responsible person, forming a traceable chain of responsibility and effectively solving the problem of unclear responsibility in traditional manual monitoring.

[0021] Fourth, this invention includes a non-command speech detection module, employing a dual-channel acoustic and text structure: on one hand, it uses a CNN or CNN+LSTM model to determine whether the speech possesses command-related voiceprint features; on the other hand, it utilizes a semantic classification model (such as BERT) to determine whether its semantics are related to standard commands, and combines a keyword dictionary, SOP parameter detection, and stage restriction strategies for comprehensive judgment. This technical feature effectively filters out irrelevant content such as everyday conversations, interjections, coughs, and mechanical noise, ensuring that the system only processes truly critical operational commands, thereby improving the overall algorithm's accuracy and robustness.

[0022] Fifth, this invention integrates an automatic alarm function for erroneous commands. After verifying the authenticity, standardization, and correctness of the commands, if any missing commands, incorrect content, non-standard format, or inconsistencies with the current flight phase are detected, the system will automatically generate a prompt or alarm message. This technical feature provides additional safety redundancy for flight operations, allowing for timely intervention before human negligence or communication errors occur, helping to prevent potential risks and improve flight safety. Attached Figure Description

[0023] The above-described features and advantages of the present invention will be better understood after reading the following detailed description of embodiments of the present disclosure in conjunction with the accompanying drawings. In the drawings, components are not necessarily drawn to scale, and components having similar related characteristics or features may have the same or similar reference numerals.

[0024] Figure 1 A flowchart of an embodiment of the voice command handshake verification method of the present invention is shown.

[0025] Figure 2 A schematic diagram of an embodiment of the voice command-based handshake verification system of the present invention is shown.

[0026] Figure 3 A schematic diagram of the structure of an electronic device according to the present invention is shown. Detailed Implementation

[0027] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. It should be noted that the aspects described below with reference to the accompanying drawings and specific embodiments are merely exemplary and should not be construed as limiting the scope of protection of the present invention in any way.

[0028] Figure 1 The flowchart of an embodiment of the voice command handshake verification method of the present invention is shown. This embodiment uses an aircraft cockpit as an application scenario for illustration. This scenario is only for illustrative purposes, and the present invention is not limited thereto. It can also be extended to any application scenario with strict requirements for high reliability, strong standardization, multi-role collaboration, and closed-loop verification of key commands.

[0029] Step S1: Collect the raw speech signal of the object (in this example scenario, the object is the pilot).

[0030] In this example, the acquisition device is installed on the pilot's side of the cockpit and uses a high-sensitivity microphone in conjunction with a DSP noise reduction module to acquire speech after environmental noise suppression.

[0031] For pilot scenarios, the collected behaviors cover all flight phases, including taxiing, takeoff, climb, cruise, approach, landing, and runway, ensuring that the quality of the raw voice signal meets the requirements of subsequent processing.

[0032] Step S2: Preprocess the acquired raw speech signal to obtain structured speech segments, providing a standardized input format for subsequent model recognition. Preprocessing includes high-frequency and low-frequency noise filtering, VAD (Voice Activity Detection) speech detection and silence removal, segmentation (e.g., dividing the raw speech signal into 0.5–3 second segments), normalization, and feature extraction.

[0033] High-frequency and low-frequency noise filtering removes interference components in the speech signal that are not within the human voice frequency range, thereby improving the accuracy and robustness of subsequent speech recognition and analysis.

[0034] VAD speech detection and silence removal determine which parts of an audio clip contain valid speech and which parts are silence or non-speech noise, and then trim or filter them accordingly.

[0035] Segmentation processing involves dividing speech data into segments of fixed duration to fit speech recognition models.

[0036] Normalization primarily standardizes the amplitude (volume) or energy of the speech signal to eliminate signal strength fluctuations caused by factors such as the speaker's distance from the microphone, varying vocal intensity, and differences in device gain. Feature extraction extracts key information from the speech signal that characterizes its content and speaker characteristics, converting it into a numerical vector form more suitable for algorithm processing.

[0037] Step S3: Convert the structured speech fragments obtained in step S2 into text information and match them with the standard instruction library to preliminarily determine whether they belong to the target instruction.

[0038] The conversion of speech segments into text information is achieved using a deep speech recognition model.

[0039] Step S4: Filter the text information identified in Step S3. If the current voice is not a critical command, no further processing is performed. If the current voice is a critical command, proceed to Step S5. This step filters out voice unrelated to the task (e.g., flight operations) to ensure that subsequent processing targets are critical commands (e.g., call-outs).

[0040] Non-mission-related voice content includes, but is not limited to: daily conversations (such as “I’ll have some water later”, “The weather is nice”), chat phrases, interjections (“uh”, “um”, “okay”), coughing and breathing sounds, cabin mechanical noises, switch sounds, and voices of non-cockpit personnel (such as flight attendants and ground staff).

[0041] This step employs a dual-channel structure of acoustics and text: a first sub-model, such as a Convolutional Neural Network (CNN) or a CNN + Long Short-Term Memory (LSTM) network, processes the acoustic features of the speech signal to determine whether the speech possesses specific command voiceprint features; a second sub-model, such as a semantic classification model (BERT), analyzes the text content after speech-to-text conversion to determine whether the semantics are related to standard commands. The confidence outputs of the two sub-models are then merged, classifying the speech into "Command Speech" and "Non-Command Speech." For example, the confidence scores from the two sub-models are mapped to a common probability space and then merged according to a certain weight ratio, ultimately generating a decision result that reflects both the voiceprint features and semantic accuracy.

[0042] If the speech is determined to be non-command speech, it will not undergo subsequent command recognition processing and will not participate in subsequent normative, role, and stage judgments. It can be recorded as background noise samples to improve the algorithm's robustness. To avoid misinterpreting colloquial commands as ordinary dialogue, a keyword dictionary, SOP parameter detection, command semantic similarity judgment, and stage restriction strategies are introduced (e.g., speech during takeoff is highly likely to be a command). For example, "Okay, flap 15" is colloquial but will be recognized as a command. If it is a command, the command recognition and verification process begins; otherwise, subsequent processing terminates.

[0043] For example, the keyword dictionary processing includes a built-in dictionary extracted from aviation standard operating procedures (SOPs) and historical flight data. This dictionary contains core terms from all typical instructions, such as "V1," "Rotate," "Flaps," "Gear," "Heading," "Set," "Checked," and "Stabilized." This dictionary not only covers English terms but also supports corresponding Chinese expressions (such as "flaps," "landing gear," and "heading"). When the text output by speech recognition contains one or more keywords, even if the sentence structure is incomplete or contains colloquial expressions (such as "uh, flaps fifteen"), it will be considered a potential instruction, triggering further verification. The keyword dictionary can be dynamically loaded by aircraft type, airline, or flight phase to improve adaptability.

[0044] For example, in SOP parameter detection, many standard instructions contain parameters in specific formats, such as numbers, angles, and speed values. Rule template matching is used to detect whether these parameters conform to SOP requirements. For instance, in "Flaps XX", "XX" should be a valid flap setting (e.g., 5, 15, 30); in "Heading XXX", "XXX" should be an integer between 0 and 360; and "Speed ​​XXX" should be a reasonable airspeed range. If a combination of parameters conforming to this structure exists in the identified text, even if the statement prefix is ​​a colloquial expression (e.g., "Okay, heading 270"), it will still be considered a valid instruction. This mechanism effectively prevents missed detections due to pilots' habitual use of language (e.g., adding interjections like "okay," "OK," or "um").

[0045] For example, in Instruction Semantic Similarity Judgment, a semantic embedding-based method is used to calculate the similarity in vector space between the identified text and each instruction in the standard instruction library. Specifically, this involves: using pre-trained language models (such as BERT or Sentence-BERT) to encode colloquial statements and standard instructions into high-dimensional vectors; calculating the cosine similarity or Euclidean distance between the two; and considering the similarity to be the same semantic instruction if it exceeds a set threshold. For instance, different expressions such as "flaps up to 15," "flaps 15," and "Flaps one five" are close to each other in the semantic vector space and can be uniformly mapped to the standard instruction "Flaps 15," thus avoiding misjudgment as non-instructions due to differences in expression.

[0046] For example, in the Phase-Based Prioritization Strategy, the flight process is divided into several distinct phases (such as taxiing, takeoff, climb, approach, and landing), each with its own set of typical instructions. The system automatically determines the current phase based on real-time flight parameters (such as airspeed, altitude, flap position, and landing gear status) and adjusts the prior probability of instruction decisions accordingly. For instance, during the takeoff roll phase (airspeed close to V1), any speech containing keywords such as "V1," "Rotate," or "climbing," even if incomplete or containing interjections, will be assigned a high instruction probability; while during the cruise phase, similar speech may be downgraded. This context-aware mechanism significantly improves the sensitivity to instructions in critical phases, effectively preventing the omission of crucial calls due to colloquial expressions or background noise interference.

[0047] The four strategies mentioned above do not operate in isolation, but are integrated into the non-command speech detection process in a fusion manner: keywords trigger initial screening, SOP parameters provide structural constraints, semantic similarity ensures robustness of recognition under diverse expressions, and stage restriction strategies inject flight context. These four strategies work synergistically to ensure that this step can filter out non-command content such as everyday conversations and environmental noise, while accurately capturing the pilot's natural, concise, and even accented commands during actual operations, thus supporting subsequent standardization, role, and stage verification processes.

[0048] Step S5: Verify the format accuracy and standardization of the text information identified in step S3.

[0049] The verification process checks whether the pilot's announcements conform to the standard announcement format requirements of the SOP (Standard Operating Procedure) issued by the airline, flight team, or Civil Aviation Administration, thus ensuring the rigor of the operational procedures. The pre-built standard announcement format library includes: mandatory commands for each flight phase (e.g., "V1", "Rotate", "Positive Climb"); parameter-based command templates (e.g., "Heading XXX", "Flaps XX", "Speed ​​XXX"); and confirmation commands (e.g., "Checked", "Set", "Stabilized"). All commands are stored in a dual structure of rule templates and semantic embedding vectors to ensure both aircraft type compatibility and conversational processing capabilities.

[0050] The text output by speech recognition is subjected to the following multi-level parsing (keyword parsing, word order determination, instruction semantic deconstruction, and semantic similarity matching): Keyword parsing first scans the text for predefined instruction keywords, such as "Flaps," "Gear," "Heading," "V1," and "Rotate." These keywords are prerequisites for triggering subsequent parsing. If no keywords are detected, the text is directly judged as a non-instruction or invalid instruction. The keyword database supports multiple languages ​​(such as Chinese and English) and can be customized by aircraft type and airline.

[0051] Word order judgment: Aviation instructions have a highly fixed word order structure, such as "Flaps 15" instead of "15 Flaps", and "Gear Up" instead of "Up Gear". Built-in grammar rule templates are used to verify whether the relative positions of keywords and their parameters conform to the standard format. Expressions that do not conform to the word order rules, even if semantically correct (such as "fifteen flaps"), may be marked as formatting errors to reinforce operational rigor.

[0052] Instruction semantic deconstruction (such as subject, parameters, actions), decomposes the instruction into structured semantic components, typically including: subject (usually implicit, defaulting to the current operator or system), action / operation type (such as "Set", "Check", "Confirm"), object (such as "Flaps", "Gear", "Speed"), parameter value (such as "15", "Down", "270"). For example, "Flaps fifteen" is deconstructed as: object = "Flaps", parameter = "15"; "Gear Down" is deconstructed as: object = "Gear", parameter = "Down". This deconstruction enables instructions in different languages or expressions to be mapped to a unified internal representation.

[0053] Semantic similarity matching, to be compatible with colloquial, dialectal or Chinese-English mixed expressions, calculates the similarity between the deconstructed semantic vectors and the embedded vectors in the standard instruction library. Each instruction in the standard instruction library is stored in a dual form of "rule template + semantic embedded vector". For example: "Flaps one five" (transcription of English pronunciation), "襟翼十五" (Chinese transcription), "Flaps 15" (numerical form) are close to each other in the semantic space and can be unified and merged into the standard instruction "Flaps 15" through cosine similarity matching. This mechanism effectively solves the problem of instruction fragmentation caused by speech recognition errors or differences in expression habits.

[0054] The content of format accuracy and standardization verification includes the verification of items such as instruction structure, content, word order, auxiliary words, and password strength. If the format accuracy and standardization verification matches successfully, it is marked as a correct instruction; if the format accuracy and standardization verification fails, it is marked as incorrect and a prompt or alarm message is generated.

[0055] Step S6: Dynamically associate the recognized voice instruction with the current flight state, and verify the timing rationality and scenario consistency of the instruction.

[0056] By integrating multi-dimensional real-time parameters from airborne sensors and the flight control system (FCS), a comprehensive perception of the current flight status can be constructed. These key parameters include: Indicated Airspeed (IAS): used to distinguish between taxiing (<50 knots), takeoff acceleration (>80 knots), and cruise (typically above 250–300 knots); Altitude Pressure (ALT): used in conjunction with climb / descent rate to determine whether it is climbing (e.g., 0–10,000 feet), cruise (level flight), or approach (<3,000 feet); Vertical Speed ​​(V / S): positive values ​​indicate climbing, negative values ​​indicate descent, and values ​​close to zero may indicate cruise or level flight approach; Landing gear status (retracted / extended) and flap position (e.g., 0°, 15°, 30°): key indicators for judging configuration adjustments after takeoff, approach preparation, and landing; Engine thrust status (e.g., N1 speed, throttle position): high thrust during takeoff, stable during cruise, and gradually reduced throttle during approach; Track changes (e.g., heading stability, glide slope deviation, ILS signal acquisition): used to identify whether it has entered the instrument approach or visual landing phase. These parameters, after being time-aligned and filtered, serve as input features for the flight phase classification model.

[0057] Then, by combining a rule base and a deep learning classification model based on multi-label classification, the dynamic association between voice commands and the current flight status is realized, including hard rule matching, vector matching of command semantics and stages, and a multi-label flight status classification network.

[0058] Rule-Based Matching: Based on Civil Aviation Standard Operating Procedures (SOPs) and aircraft type manuals, a series of deterministic logical rules are preset. For example: if IAS ≥ 80 knots and V / S > 500 fpm and flaps ≥ 5° → it is determined to be the "takeoff" phase; if ALT < 1,000 feet and landing gear deployed and flaps ≥ 30° → it is determined to be the "landing" phase; if the altitude is stable at cruising altitude ± 200 feet and V / S ≈ 0 → it is determined to be "cruising". In addition, hard rules also include command triggering conditions, such as: "80 knots, V1, Rotate, etc. must be called out during the takeoff phase". If they are not detected, a "critical command missing" alarm may be triggered.

[0059] The semantics of the command are matched with the vectors of the flight phase. The currently recognized voice command text is encoded into a semantic vector (e.g., using Sentence-BERT) and its similarity is calculated with the typical command semantic prototype vectors corresponding to each flight phase. For example, the semantic center height of "Gear Down" is close to that of the "Approach" phase, but far from that of the "Climb" phase. This matching can help correct ambiguities that may arise from relying solely on parameter judgment (e.g., low-altitude flyby vs. actual approach).

[0060] Multi-label flight state classification networks (such as Transformer Encoder): To handle complex, continuous, and overlapping flight states (such as the transition zone between "approach" and "landing"), an end-to-end deep neural network is deployed. A typical structure is a multi-label classifier based on Transformer Encoder, characterized by: Input: A time-series flight control parameter window (such as IAS, ALT, V / S, etc. for the past 30 seconds); Output: Probability distribution of multiple stages corresponding to each time step (because stages may overlap, such as "descent + approach"); Training data: derived from a large amount of real flight QAR (Quick Access Recorder) data, with precise stage labels; This multi-label flight state classification network can capture the nonlinear relationships and long-term dependencies between parameters, adapting to differences in flight profiles under different aircraft types, airports, and weather conditions. It runs in parallel with a rule base, and the final decision is obtained through weighted fusion (e.g., rule priority, model confidence weighting).

[0061] The processing of verifying the timing rationality and scenario consistency of commands includes: once the current flight phase is determined, the system initiates a targeted command verification strategy. For example: During the "approach" phase, the following commands will be monitored to ensure they are correctly issued: "Flaps XXX" (Check if the flaps are set according to the procedure); “Gear Down” (confirms landing gear has been lowered); "Landing Checklist Complete"

[0062] If these critical instructions are missing, delayed, or formatted incorrectly, they will be marked as “Phase-Instruction Inconsistency” and an alarm will be generated (such as “No landing gear drop call detected during approach phase”).

[0063] If it is in the "takeoff" phase, check whether "80 knots", "V1", "Rotate", "Positive Climb", etc. appear in sequence; If you are in the "taxi" phase, pay attention to ground operation commands such as "Taxi Checklist" and "Clear Left / Right".

[0064] This context-aware instruction verification mechanism ensures that it not only identifies "what was said", but also judges "whether it should be said at this time", thereby significantly improving the accuracy of flight operation compliance monitoring.

[0065] Step S7: Based on voiceprint features, distinguish speakers of different roles to achieve precise binding of instructions with responsible entities.

[0066] Based on voiceprint features, the system automatically distinguishes between roles such as captain, co-pilot, observer, or instructor, and binds the content of the announcements to the corresponding responsible parties, forming a chain of responsibility.

[0067] During the pilot's onboarding, training, or flight preparation phases: collect the pilot's reference voiceprint; generate the pilot's voiceprint vector (e.g., 192-dimensional, 256-dimensional); and establish a "role-voiceprint feature" database.

[0068] During flight, voiceprint features are extracted from speech segments, including MFCC (Mel-Frequency Cepstral Coefficients), Fbank (Filter Bank Features), PLP (Perceptual Linear Prediction), and a deep voiceprint encoder X-vector. Speaker identity is determined based on feature vector similarity (cosine similarity, PLDA (Probabilistic Linear Discriminant Analysis) score, etc.). To prevent role jumps caused by transient noise, a sliding window stabilization algorithm, temporal consistency verification, and voiceprint fusion strategy are incorporated to ensure consistency of the speaker's role in continuous speech, thereby achieving reliable traceability.

[0069] Step S8: Perform final verification of the authenticity, correctness and standardization of the instructions, and trigger alarms or prompts when abnormalities occur, and finally output structured instruction text.

[0070] The system automatically assesses the identified commands: whether they are correct, conform to the current flight procedure, are standardized, and whether any key commands are missing. If normal, a normal call tag is recorded; if abnormal, an automatic prompt or alarm is issued, and a violation tag is recorded.

[0071] This step integrates the outputs of steps S5 (standardization), S6 (state consistency), and S7 (role binding) to form a triple verification loop for the "authenticity, correctness, and standardization" of instructions. It can automatically generate a structured text report that displays the instruction recognition status, the results of erroneous instruction comparison, and data statistics.

[0072] Figure 2 The principle of one embodiment of the voice command-based handshake verification system of the present invention is illustrated. See also Figure 2 The system in this embodiment includes: a voice acquisition module, a voice preprocessing module, a voice-to-text and command recognition module, a non-command voice detection module, a command standardization judgment module, a content status classification module, a role recognition module, and a verification and alarm module.

[0073] The voice acquisition module is used to acquire the raw voice signal of the object performing the task. The internal processing of the voice acquisition module and... Figure 1 The steps S1 shown are the same and will not be repeated here.

[0074] The speech preprocessing module is used to preprocess the acquired raw speech signal to obtain structured speech segments. The internal processing of the speech preprocessing module and... Figure 1 The steps shown in step S2 are the same and will not be repeated here.

[0075] The speech-to-text and command recognition module converts structured speech segments obtained from the speech preprocessing module into text information and matches them against a standard command library to initially determine whether they belong to the target command. The internal processing of the speech-to-text and command recognition module... Figure 1 Step S3 is the same as shown, and will not be repeated here.

[0076] The non-command speech detection module filters the text information recognized by the speech-to-text and command recognition modules. If the current speech is not a key command, no further processing is performed; if the current speech is a key command, it proceeds to the subsequent command conformity judgment module. The internal processing of the non-command speech detection module and... Figure 1 Step S4 is the same as shown, and will not be repeated here.

[0077] The instruction standardization judgment module is used to verify the format accuracy and standardization of the text information recognized by the speech-to-text and instruction recognition modules. The internal processing of the instruction standardization judgment module and... Figure 1 Step S5 is the same as shown, and will not be repeated here.

[0078] The content status classification module dynamically associates recognized voice commands with the current task execution status, verifying the timing rationality and scenario consistency of the commands. The internal processing of the content status classification module... Figure 1 Step S6 is the same as shown, and will not be repeated here.

[0079] The role recognition module distinguishes speakers of different roles based on voiceprint features, enabling precise binding of instructions to responsible parties. The internal processing of the role recognition module and... Figure 1 Step S7 is the same as shown, and will not be repeated here.

[0080] The verification and alarm module performs final verification of the authenticity, correctness, and standardization of instructions, triggering alarms or prompts in case of anomalies, and ultimately outputting structured instruction text. The internal processing of the verification and alarm module... Figure 1 The steps shown in step S8 are the same and will not be repeated here.

[0081] In addition, such as Figure 3 As shown, the present invention also discloses an electronic device, which includes a controller, a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus.

[0082] Memory is used to store computer programs; the processor is used to execute the programs stored in memory, enabling functions such as... Figure 1 The steps of the voice command handshake verification method embodiment shown are as follows.

[0083] Furthermore, the present invention also discloses a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements as follows: Figure 1 The steps of the voice command handshake verification method embodiment shown are as follows.

[0084] Furthermore, this invention also discloses a computer program product, which, when executed by a processor, implements the following: Figure 1 The steps of the voice command handshake verification method embodiment shown are illustrated.

[0085] Although the methods described above are illustrated and depicted as a series of actions for the sake of simplicity, it should be understood and appreciated that these methods are not limited by the order of the actions, as some actions may occur in a different order and / or concurrently with other actions from the illustrations and descriptions herein or not illustrated and described herein but which may be understood by those skilled in the art, according to one or more embodiments.

[0086] Those skilled in the art will further appreciate that the various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, the various illustrative components, blocks, modules, circuits, and steps are described above in a generalized manner in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of the invention.

[0087] The various illustrative logic blocks, modules, and circuits described in conjunction with the embodiments disclosed herein can be implemented or performed using a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general-purpose processor may be a microprocessor, but in alternatives, it may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration.

[0088] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to a processor such that the processor can read and write information to / from the storage medium. In an alternative, the storage medium may be integrated into the processor. The processor and storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In an alternative, the processor and storage medium may reside as discrete components in the user terminal.

[0089] In one or more exemplary embodiments, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software as a computer program product, the functionality may be stored or transmitted as one or more instructions or code on or through a computer-readable medium. A computer-readable medium includes both computer storage media and communication media, encompassing any medium that facilitates the transfer of a computer program from one location to another. A storage medium may be any available medium accessible to a computer. By way of example and not limitation, such a computer-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and is accessible to a computer. Any connection is also legitimately referred to as a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of a medium. As used in this article, disk and disc include compact discs (CDs), laser discs, optical discs, digital multi-purpose discs (DVDs), floppy disks, and Blu-ray discs. Disks typically reproduce data magnetically, while discs reproduce data optically using lasers. Combinations of these should also be included within the scope of computer-readable media.

[0090] The prior description of this disclosure is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to this disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not intended to be limited to the examples and designs described herein, but should be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A voice command-based handshake verification method, characterized in that, The methods include: Step S1: Collect the raw audio signal of the subject while performing the task; Step S2: Preprocess the acquired raw speech signal to obtain structured speech segments; Step S3: Convert the structured speech fragments obtained in step S2 into text information and match them with the standard instruction library to preliminarily determine whether they belong to the target instruction; Step S4: Filter the text information identified in step S3. If the current speech is not a key instruction, no further processing will be performed. If the current speech is a key instruction, the subsequent processing in step S5 will be performed. Step S5: Verify the format accuracy and standardization of the text information identified in Step S3; Step S6: Dynamically associate the recognized voice commands with the current task execution status to verify the timing rationality and scenario consistency of the commands; Step S7: Distinguish speakers of different roles based on voiceprint features to achieve precise binding of instructions with responsible parties; Step S8: Perform final verification of the authenticity, correctness and standardization of the instructions, and trigger alarms or prompts when abnormalities occur, and finally output structured instruction text.

2. The voice command handshake verification method according to claim 1, characterized in that, In step S2, preprocessing includes high-frequency and low-frequency noise filtering, speech detection and silence removal, segmentation, normalization, and feature extraction.

3. The voice command handshake verification method according to claim 1, characterized in that, In step S4, a dual-channel structure of acoustics and text is adopted, and the first sub-model of convolutional neural network / convolutional neural network and long short-term memory network is used to process the acoustic features of speech signal and determine whether the speech has specific command voiceprint features. The second sub-model of the semantic classification model is used to analyze the text content after speech-to-text conversion, determine whether the semantics are related to standard instructions, and then combine the confidence outputs of the two sub-models to classify it into instruction speech and non-instruction speech.

4. The voice command handshake verification method according to claim 1, characterized in that, In step S4, the process of detecting whether the speech belongs to a key instruction also includes introducing a keyword dictionary, SOP parameter detection, instruction semantic similarity judgment, and stage restriction strategy.

5. The voice command handshake verification method according to claim 1, characterized in that, In step S5, the text output by speech recognition is parsed at multiple levels, including keyword parsing, word order judgment, instruction semantic deconstruction, and semantic similarity matching. The format accuracy and standardization verification includes the verification of items such as instruction structure, content, word order, auxiliary words, and password strength.

6. The voice command handshake verification method according to claim 1, characterized in that, In step S6, the processing of dynamically associating voice commands with the current flight status further includes: By integrating multi-dimensional real-time parameters from airborne sensors and flight control systems, a comprehensive perception of the current flight status can be constructed. By combining a rule base and a deep learning classification model based on multi-label classification, dynamic association between voice commands and the current flight status is achieved, including hard rule matching, vector matching of command semantics and stages, and a multi-label flight status classification network.

7. A voice command-based handshake verification system, characterized in that, The system includes: The voice acquisition module is used to acquire the raw voice signal of the object when performing the task; The speech preprocessing module is used to preprocess the acquired raw speech signal to obtain structured speech segments; The speech-to-text and instruction recognition module is used to convert the structured speech segments obtained by the speech preprocessing module into text information and match them with the standard instruction library to initially determine whether they belong to the target instruction. The non-command speech detection module is used to filter the text information recognized by the speech-to-text and command recognition modules. If the current speech is not a key command, no further processing will be performed. If the current speech is a key command, it will be processed by the subsequent command standardization judgment module. The instruction standardization judgment module is used to verify the format accuracy and standardization of the text information recognized by the speech-to-text and instruction recognition modules; The content status classification module is used to dynamically associate the recognized voice commands with the current task execution status to verify the timing rationality and scenario consistency of the commands. The role recognition module is used to distinguish speakers of different roles based on voiceprint features, so as to accurately bind instructions with responsible parties. The verification and alarm module is used to perform final verification of the authenticity, correctness and standardization of the instructions, and to trigger alarms or prompts when an anomaly occurs, and finally outputs the structured instruction text.

8. An electronic device, characterized in that, The electronic device includes a controller, which includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus. Memory, used to store computer programs; When a processor executes a program stored in memory, it implements the steps of the voice command shouting handshake verification method as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the voice command shouting handshake verification method as described in any one of claims 1-6.

10. A computer program product, characterized in that, When the program is executed by the processor, it implements the steps of the voice command shouting handshake verification method as described in any one of claims 1-6.