Pilot training operation evaluation method based on multi-modal time sequence fusion

By combining multimodal temporal fusion and causal logic topology graphs, the problem of insufficient causal constraints in pilot training assessment is solved, and stable and accurate assessment is achieved in dynamic environments.

CN121786764AActive Publication Date: 2026-04-03HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS +1
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-05
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies in pilot training and evaluation rely on single-dimensional sensor signal acquisition and static logic analysis, making it difficult to establish physical deterministic causal constraints between multi-source heterogeneous data streams in dynamic simulation training environments, leading to logical gaps and false alarms in the evaluation.

Method used

By acquiring multidimensional data and using a unified clock signal for timestamp calibration, multimodal feature vectors are extracted. Combined with cross-modal attention mechanisms and operational causal logic topology graphs, physical state boundary parameters are calculated in real time, and logical bias operators are applied to increase the decision weights, thereby achieving cross-dimensional logical arbitration.

Benefits of technology

To ensure the logical determinism of evaluation results in complex environments, eliminate evaluation interruptions caused by single-modal failures, improve identification stability, provide logical diagnostic functions, and avoid false alarms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121786764A_ABST
    Figure CN121786764A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electrical digital data processing, and discloses a pilot training operation evaluation method based on multi-modal time sequence fusion, which comprises the following steps: acquiring a video stream, an audio signal and a multi-dimensional operation state parameter sequence of a target object in an operation evolution process, and realizing frame-level time alignment; extracting space time action, instruction semantics and operation control trend feature vectors to generate multi-modal feature vectors; generating fusion features by using a cross-modal attention mechanism; constructing an operation causal logic topological graph containing physical boundary constraints; according to the method, the physical state boundary parameters of the controlled object are calculated, a logic bias operator is applied on the basis of the fusion features, the judgment weight of the semantic corresponding to the target operation node is adjusted and increased, logic blind compensation is achieved through a data reconstruction mechanism driven by the physical causal law, and the certainty of an evaluation result in a complex environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a pilot training operation evaluation method based on multimodal temporal fusion, belonging to the field of electronic digital data processing technology. Background Technology

[0002] Pilot training and assessment is a crucial component of the civil aviation training system. Existing digital data processing relies on single-dimensional sensor signal acquisition and static logic analysis. The main technical direction is to integrate cockpit video spatial action features, audio command semantic features, and fast access recorder parameters for joint assessment. For example, Chinese invention patent CN119861614B discloses an automated assisted driving method based on multimodal information fusion, which uses a multimodal Transformer model to classify and assess pilot behavior. This belongs to the big data training probability mapping mechanism. In a high-dynamic simulation training environment, such algorithms rely on statistical correlation of appearance features, making it difficult to establish physical deterministic causal constraints between multi-source heterogeneous data streams. Video action recognition is distorted by drastic changes in light and shadow or human occlusion, and speech semantics drifts due to high-noise background interference. The system relies on linear fusion based on statistical weights, leading to logical gaps or false alarms in judgment.

[0003] In dynamic simulation training environments, processing architectures based on multimodal overlays face semantic confidence mismatch. Changes in cockpit lighting, human posture occlusion, or high-noise background noise can lead to random signal loss and feature drift in the temporal domain of visual and acoustic features. To maintain data synchronization, a large amount of redundant computing resources are allocated for time alignment, which reduces the stability of the signal processing architecture to environmental noise. Sacrificing stability for increased feature dimensions causes the evaluation program to produce logical gaps when multi-source signals are locally damaged. Simply increasing the sampling frequency or deepening the feature completion model structure is limited by the hard constraints of the processor's real-time processing capabilities and cannot eliminate semantic conflicts between asynchronous data streams. Conventional weighted linear alignment algorithms produce judgment biases when processing nonlinear control sequences, making it difficult to establish a definite causal mapping between apparent features and flight physical intentions.

[0004] Therefore, how to construct an electrical digital data processing architecture that uses the laws of flight physics to deterministically calibrate random characteristic errors and realizes cross-dimensional logical arbitration in an asynchronous time-series environment has become the technical problem to be solved by this invention. Summary of the Invention

[0005] To address the problems mentioned in the background art, the technical solution of the present invention is as follows: A pilot training operation evaluation method based on multimodal temporal fusion, comprising the following steps:

[0006] Step S101: Obtain the video stream, audio signal, and multi-dimensional operating state parameter sequence of the target object during the operation evolution process; use a unified clock signal to timestamp the video stream, audio signal, and multi-dimensional operating state parameter sequence to achieve time alignment of each modal data at the frame level.

[0007] Step S102: Use a video processing model to extract spatial-temporal action feature vectors, use a speech recognition model to extract instruction semantic feature vectors, use a time series analysis model to extract operation control trend feature vectors of the target object, and combine the feature vectors into a multimodal feature vector.

[0008] Step S103: Use a cross-modal attention mechanism to perform weighted fusion of multimodal feature vectors to generate fused features that include the original confidence scores of each operation semantic category;

[0009] Step S104: Construct an operation causal logic topology graph containing physical boundary constraints. The operation causal logic topology graph consists of multiple operation nodes corresponding to the baseline operation procedure, and each operation node is associated with a physical state trigger vector.

[0010] Step S105: Calculate the physical state boundary parameters of the controlled object in real time based on the multidimensional running state parameter sequence; if the physical state boundary parameters match the physical state trigger vector of the target operation node, apply a logical bias operator on the fused features to increase the judgment weight of the corresponding operation semantics of the target operation node.

[0011] Step S106: Calculate the weighted scores of the identified operation events and the operation causal logic topology graph in terms of time consistency, sequence consistency, and parameter deviation, and output an evaluation report.

[0012] Preferably, applying the logical bias operator in step S105 includes: extracting the first-order state change rate, actuator displacement, and motion vector parameters from the multi-dimensional operating state parameter sequence in real time; mapping the extracted parameters to a preset physical energy state space to determine the current temporal state mode of the controlled object; when there is a preset logical correlation between the temporal state mode and the target operation node, calculating the probability correction increment for the target operation node using a nonlinear activation function, and accumulating the probability correction increment to the original confidence score to increase the judgment weight.

[0013] Preferably, the operation nodes in the causal logic topology graph are connected by directed arcs; the directed arcs are used to represent the sequential constraints of the operation actions; the physical state trigger vector contains the control quantity threshold range required to represent the operation compliance and the physical quantity change trend operator.

[0014] Preferably, the weighted fusion of multimodal feature vectors in step S103 includes: calculating the feature mutual information of the video stream, audio signal, and multidimensional running state parameter sequence within the current time window; dynamically allocating the attention weight of each modality according to the feature mutual information; and adjusting the attention weight of the corresponding state feedback signal in the multidimensional running state parameter sequence if the recognition confidence of the target action in the video stream is lower than a preset threshold.

[0015] Preferably, the process of increasing the attention weight follows the following weight adjustment formula: ;in, The corrected operation semantics determination weights, For the original fusion weights, To monitor the obtained physical state boundary parameter values ​​in real time, The physical trigger reference value corresponding to the target operation node. To characterize the variance factor of the sampling fluctuation of the multidimensional operating state parameter sequence, and and They have the same dimensions.

[0016] Preferably, the calculation of the weighted score in step S106 includes: calculating the path offset distance between the identified operation event sequence and the baseline operation procedure using a dynamic time warping algorithm; determining the corresponding deviation penalty coefficient based on the logical depth of each operation node in the operation causal logic topology graph; and multiplying the path offset distance and the deviation penalty coefficient to obtain a quantitative score representing the operation accuracy.

[0017] Preferably, the evaluation report traces the identified abnormal operations through an operational causal logic topology diagram to identify the physical logic correlation defects that cause operational deviations, and outputs improvement suggestions for the energy state management logic of the controlled object.

[0018] Preferably, step S101, which achieves time alignment of each modal data at the frame level, includes: obtaining the inherent sampling frequency of each modal sensor, wherein the sampling frequency range is [missing information]. to The Lagrange interpolation algorithm is used to upsample the data sequence of the low sampling rate mode so that its temporal resolution is aligned with the video frame rate of the highest sampling rate mode.

[0019] Preferably, the real-time calculation of the physical state boundary parameters of the controlled object in step S105 includes: constructing a real-time state profile model of the controlled object based on a multi-dimensional operating state parameter sequence; and determining the logical necessity of the target operating node being activated under the current physical environment by calculating the spatiotemporal logical coupling entropy between the real-time state profile model and the physical state trigger vector of the target operating node.

[0020] Preferably, after step S106, the method further includes: real-time monitoring of the numerical changes of physical state boundary parameters; if the slope of the change of physical state boundary parameters exceeds a preset mutation threshold, the operation causal logic topology diagram is switched to the emergency logic processing mode corresponding to the mutation threshold, and the weighted score is updated according to the constraint rules in the emergency logic processing mode.

[0021] Compared with the prior art, the beneficial effects of the present invention are:

[0022] 1. In pilot training operation evaluation, an operation causal logic topology network containing physical state trigger vectors is constructed. The flight parameter data stream with strong causal attributes is used as the logical calibration benchmark. Dynamic arbitration is performed on the video recognition features and audio semantic features with strong randomness. When the video acquisition channel loses instantaneous features due to sudden changes in cockpit lighting or human posture occlusion, the aircraft's physical energy state and attitude change pulses are used to apply logical bias to candidate operation nodes in the feature fusion layer. The closed-loop verification logic of the deterministic physical law is used to compensate for the uncertain sensor signals, eliminating the evaluation logic interruption caused by single-mode sensor failure, ensuring the continuity of data stream processing, and improving the logical determinism of the evaluation results in complex electromagnetic and light and shadow environments.

[0023] 2. An asynchronous arbitration method based on flight profile state is adopted to calculate the mutual information between flight parameter feature vectors and standard operating procedure topology nodes in real time. This allows the system to dynamically activate corresponding evaluation weights based on the current flight physical environment, eliminating the rigid dependence of operation recognition logic on absolute timestamps or fixed operation sequences. This enables the system to identify non-standard but logically valid personalized operation paths. The evaluation process has evolved from mechanical comparison of apparent action features to causal verification based on the coupling of operation intent and physical feedback. This avoids false alarms caused by fine-tuning of operation timing and improves the stability of the electronic digital processing system in recognizing nonlinear control behaviors.

[0024] 3. Multimodal temporal features are mapped to a high-dimensional feature vector space. A cross-modal attention mechanism is used to calculate the spatiotemporal coupling entropy between voice commands, hand gestures, and control surface feedback. The deep coupling mechanism of multi-dimensional features can capture operational delays or command execution deviations that cannot be represented by a single modality. By calculating time consistency, sequence consistency, and parameter deviation weighted indices, the discrete operation event processing process is converted into a continuous execution consistency score. The evaluation mechanism based on causal chains enables the evaluation report to trace back to the physical deviation root cause through a logical topology diagram, providing pilots with improvement suggestions pointing to the operation mechanism and realizing a functional leap from simple data recording to logical diagnosis. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of the evaluation method for multimodal temporal fusion and physical logic constraints of the present invention;

[0026] Figure 2 This is a bar chart comparing the recognition probabilities of four typical flight operations before and after the present invention.

[0027] Figure 3 This is a timing logic interaction diagram of the multimodal data acquisition and frame-level time alignment processing of the present invention. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] This invention provides a pilot training operation evaluation method based on multimodal temporal fusion. Through multi-source heterogeneous data acquisition and synchronous processing, a causal correlation model between pilot operational behavior and flight state parameters is established to achieve quantitative evaluation of the training process. The evaluation method includes acquiring video streams, audio signals, and multi-dimensional operational state parameter sequences of the target object during the operational evolution process; using a unified clock signal to timestamp each modal data to achieve frame-level time alignment; and extracting spatial-temporal action, command semantics, and operational control trend feature vectors to generate multimodal feature vectors. Utilizing cross-modal attention mechanisms to process multimodal feature vectors Perform weighted fusion and generate fusion features By constructing an operational causal logic topology graph containing physical boundary constraints, the real-time physical state boundary parameters of the controlled object are determined, and then the features are fused. A logical bias operator is applied to increase the judgment weight of specific operation semantics. Finally, the path offset distance between the identified operation event sequence and the baseline operation procedure is calculated, and an evaluation report is output. Video streams, audio signals, and multi-dimensional operational state parameter sequences of the target object during the operation evolution process are acquired. A unified clock signal is used to timestamp the video streams, audio signals, and multi-dimensional operational state parameter sequences, achieving frame-level time alignment of each modality's data. The video stream is acquired through multiple high-definition cameras deployed in the cockpit, with a sampling frequency of... to Within the specified range; the inherent sampling frequency of each modal sensor is obtained, and the data sequence of the low sampling rate mode is upsampled using the Lagrange interpolation algorithm to align its temporal resolution with that of the highest sampling rate mode; spatial-temporal action feature vectors are extracted using a video processing model, instruction semantic feature vectors are extracted using a speech recognition model, and operation control trend feature vectors in the multidimensional operating state parameter sequence are extracted using a time series analysis model, and the feature vectors are combined into a multimodal feature vector. Multimodal feature vectors Represented as: ,in, For video action feature vectors, For voice command feature vectors, This is the feature vector of flight state parameters.

[0030] Utilizing cross-modal attention mechanisms for multimodal feature vectors Perform weighted fusion to generate fusion features Calculate the feature mutual information of the video stream, audio signal, and multidimensional running state parameter sequence within the current time window, and dynamically allocate the attention weights of each modality based on the feature mutual information. If the recognition confidence of the target action in the video stream is lower than a preset threshold, adjust the attention weights of the corresponding state feedback signals in the multidimensional running state parameter sequence. The adjustment process follows the following weight correction formula: ,in, The corrected operation semantics determination weights, For the original fusion weights, To monitor the obtained physical state boundary parameter values ​​in real time, The physical trigger reference value corresponding to the target operation node. To characterize the variance factor of the sampling fluctuation of the multidimensional operating state parameter sequence, and and They have the same dimensions; when calculating the mutual information of features, the system obtains the video action feature vector sequence and the flight state parameter feature vector sequence within the current sliding time window, calculates the joint probability distribution of the two sets of sequences and their respective marginal probability distributions, and determines the mutual information value for different modal channels based on the distribution entropy of each modal data stream. When a component of a certain modality shows a stronger causal correlation with other modalities, the arbitration unit maps the mutual information value to the dynamic attention weight of that modal channel through a nonlinear normalization function, realizing adaptive weight allocation across modal information; an operational causal logic topology graph containing physical boundary constraints is constructed. The operational causal logic topology graph consists of multiple operational nodes corresponding to the baseline operational procedures, and each operational node is associated with a physical state trigger vector; the operational nodes in the operational causal logic topology graph are connected by directed arc segments, which are used to represent the sequential temporal constraint relationship of operational actions; the physical state trigger vector contains the control quantity threshold range required to represent operational compliance and the physical quantity change trend operator.

[0031] The physical state boundary parameters of the controlled object are calculated in real time based on the multidimensional operating state parameter sequence; a real-time state profile model of the controlled object is constructed based on the multidimensional operating state parameter sequence, and the spatiotemporal logical coupling entropy between the real-time state profile model and the physical state trigger vector of the target operation node is calculated to determine the logical necessity of the target operation node being activated under the current physical environment; if the physical state boundary parameters match the physical state trigger vector of the target operation node, then in the feature fusion... A logical bias operator is applied to increase the judgment weight of the operation semantics corresponding to the target operation node. Specifically, when applying the logical bias operator, the first-order state change rate, actuator displacement, and motion vector parameters in the multi-dimensional operating state parameter sequence are extracted in real time. The extracted parameters are mapped to a preset physical energy state space to determine the current temporal state mode of the controlled object. When there is a preset logical correlation between the temporal state mode and the target operation node, a nonlinear activation function is used to calculate the probability correction increment for the target operation node, and the probability correction increment is accumulated to the original confidence score. The weighted scores of the identified operation events and the operation causal logic topology graph in terms of time consistency, sequence consistency, and parameter deviation are calculated and output. The evaluation report is generated. When calculating the weighted score, a dynamic time warping algorithm is used to calculate the path offset distance between the identified operation event sequence and the baseline operation procedure. Based on the logical depth of each operation node in the operation causal logic topology graph and the corresponding deviation penalty coefficient, the path offset distance and the deviation penalty coefficient are multiplied to obtain a quantitative score representing the operation accuracy. The evaluation report traces the identified abnormal operations through the operation causal logic topology graph, identifying the physical logic association defects that cause the operation deviation. During the generation of the evaluation report, the arbitration unit performs backtracking calculations on nodes whose path offset distance exceeds a preset deviation threshold through the operation causal logic topology graph, identifying the specific physical quantity deviation components that cause the judgment deviation. If continuous deviations are detected... If the physical parameter value corresponding to a sampling point exceeds the envelope range defined by the physical state trigger vector, the system will map the identified physical attribute defect to the preset manipulation defect classification library and output an evaluation report containing the conclusion of the energy state imbalance of the controlled object and targeted operation correction guidance.

[0032] Example 1: In a nighttime approach and landing training scenario, affected by high-frequency flickering light interference in the cockpit, the feature vector of the pilot's operation of the landing gear handle in the video stream is affected. The confidence level of identification is determined by decay to Due to the dramatic changes in lighting, the operational semantic features extracted by the video processing unit drift in the local time domain. If conventional linear weight allocation logic is used, the failure of video modal weights will prevent the system from accurately determining the critical operation node of landing gear deployment.

[0033] The system detected flight parameter feature vectors Rate of atmospheric pressure drop in the middle When the landing gear is in the glide slope capture zone and the landing gear status indicator changes state, the cross-modal attention unit recognizes a confidence level below a preset threshold based on the video. The determination result calls the logic bias operator calculation program and uses the formula Increase the decision weight of the node corresponding to the landing phase in the causal logic topology graph of the operation. The corrected operation semantics determination weights, For the original fusion weights, To monitor the obtained physical state boundary parameter values ​​in real time, The physical trigger reference value corresponding to the target operation node. To characterize the variance factor of the sampling fluctuation of the multidimensional operating state parameter sequence, this weight correction value directly affects the fused features. The calculation process changes the semantic determination probability of the corresponding landing gear deployment operation from... Upgraded to This process utilizes the physical causality of aircraft operation to calibrate random sensing errors in video modalities, establishing a deterministic causal mapping between landing gear state parameters and operational intentions. In cases where visual modal features are lost, the system uses physical logic bias and multimodal temporal fusion mechanisms to generate a model consistent with the pilot's operational intentions. Figure 1 Fusion characteristics The algorithm uses dynamic time warping to calculate the path offset distance between the identified operation event sequence and the baseline operation procedure. The final evaluation report points the deviation to the energy state management logic of the controlled object through the operation causal logic topology diagram.

[0034] Example 2: In the evaluation test of a nighttime low-visibility takeoff and climb scenario constructed using a flight simulation platform, the functional specifications of the flight simulation platform require the control input output accuracy to be no less than [missing information]. The degree, and the sampling frequency of the physical state parameters is fixed at . The experiment aims to verify the recognition stability of the operational causal logic topology graph, which includes physical boundary constraints, in the face of asynchronous sensor noise and partial visual occlusion. Data sources include raw joystick displacement and elevator angle data generated by the simulator, and the average superimposed signal-to-noise ratio. The system uses Gaussian pulse interference to simulate sensor signal jitter in a real cockpit under airflow turbulence; the sampling period is set to... The purpose of this parameter setting is to capture the pilot's millisecond-level response characteristics at critical operational nodes, while reducing the processing load.

[0035] The experimental design includes an experimental group constructed using the method of this invention, and a control group that removes the operational causal logic topology graph, retaining only the multimodal feature weights. The system's evaluation accuracy is tested as video stream confidence decreases by injecting gradient-increasing noise interference into the input signal source; the data reflects the effect of increasing interference intensity... Upgraded to During the process, because the method of the present invention utilizes the physical state boundary parameters calculated in real time... and physical trigger baseline value Generate logical bias operators and corrected operation semantics judgment weights. Follow the formula Compensate for damaged visual feature components. The corrected operation semantics determination weights, For the original fusion weights, The physical state boundary parameter values ​​obtained in real time are in units of , The physical trigger baseline value corresponding to the target operation node, in units of , The variance factor characterizes the sampling fluctuation of the multidimensional operating state parameter sequence, with units of . This ensured that the accuracy of the experimental group's assessment remained at [percentage missing]. In summary, the control group, lacking an arbitration mechanism based on physical causality, has a lower identification accuracy when the interference intensity exceeds a certain threshold. Later dropped to the following.

[0036] Table 1: Results of Multimodal Evaluation Stability Gradient Test

[0037]

[0038] The conclusions drawn from the data in Table 1 show that the solution of the present invention has evaluation determinism in complex environments. Its mechanism lies in the cross-domain causal logic arbitration mechanism, which changes the dependence of the evaluation model on the presentation of appearance features. By transforming the physical inertial constraints in asynchronous time-series data into logical biases in the multimodal fusion process, the system realizes semantic reconstruction using aircraft state parameters with causal attributes when a single mode fails.

[0039] Example 3: This example combines Figures 1 to 3 This section describes a pilot training operation evaluation method based on multimodal temporal fusion, such as... Figure 1As shown, the process begins with the acquisition of multi-source heterogeneous data. After acquiring video streams, audio signals, and multi-dimensional operational state parameter sequences, the system enters the frame-level time alignment S101 stage, using a unified clock signal for timestamp calibration. The process consists of two main lines. The first line is multi-modal feature extraction and combination S102, which extracts spatial-temporal action, command semantics, and operation control trend feature vectors, and generates fused features containing the original confidence scores through cross-modal attention weighted fusion S103. The other line executes the real-time calculation of physical state boundary parameters S105 in parallel. This step calculates the controlled object based on the multi-dimensional operational state parameter sequence. The system takes parameters and combines them with the node information of physical boundary constraints and benchmark operation procedures constructed in the operation causal logic topology diagram S104. If the physical parameters match the trigger vector, the logic bias operator is triggered in step S105. The judgment weight of the target node is increased based on the fused features. Finally, the system integrates the above fused feature information and performs multi-dimensional consistency weighted score calculation S106 based on the logic depth and deviation penalty coefficient benchmark provided by the operation causal logic topology diagram. The system quantifies the scores of time, sequence consistency and parameter deviation dimensions, and performs the evaluation report output step to generate a report containing operation event identification and scoring.

[0040] like Figure 2 As shown, the chart presents a comparison of the recognition probabilities of four typical flight operations before and after the application of this method, presented in bar chart form. The horizontal axis represents the operation type, arranged in order of landing gear operation, flap operation, throttle operation, and heading operation. The vertical axis represents the recognition probability in percentage. The legend is distinguished by different fill textures: horizontal stripes represent the recognition probability before application, and diagonal stripes represent the recognition probability after application. In the four key operations of landing gear, flaps, throttle, and heading, the height of the diagonal stripe bars representing the recognition probability after application is higher than that of the horizontal stripe bars representing the recognition probability before application. Figure 3 As shown, this timing diagram reveals the data interaction and processing logic between various hardware components, including the cockpit camera, audio acquisition unit, QAR recorder, clock synchronization unit, interpolation processor, and data buffer. The process begins with the data output of each sensor. The cockpit camera outputs video frames at a frequency of 30-60Hz, the audio acquisition unit outputs audio sampling data, and the QAR recorder outputs flight status parameters at a frequency of 50Hz. All three signals are sent to the clock synchronization unit simultaneously. The clock synchronization unit performs the operation of acquiring the inherent sampling frequency of each mode, generates a unified timestamp, and transmits the calibrated data stream to the interpolation processor. The interpolation processor then identifies the mode with the highest sampling rate and uses the Lagrange interpolation upsampling algorithm to process the low sampling rate data to achieve time resolution alignment. Finally, the aligned video frame sequence, audio sequence, and parameter sequence are written into the data buffer, and frame-level time alignment is completed.

[0041] Example 4: In a scenario where the hardware performance of the video acquisition module degrades due to thermal drift of the photosensitive unit caused by continuous operation, the operational feature vector extracted from the video stream is... The spatial coordinates exhibit systematic deviations, causing the confidence level for recognizing flap handle actuation actions to decrease from the normal level. Descending to Because the confidence level of the sensor data is lower than the preset threshold This leads to a risk of interruption in the operational semantics determination of the evaluation logic due to feature distortion; the system arbitration unit determines the variance factor in the weight correction formula. The online parameter calibration procedure obtains the preceding sequence of the current time step. Multidimensional operating state parameter sequence The sampled values ​​are used to calculate the statistical fluctuation of the rate of change of air pressure altitude within that time period using the discrete variance formula. This statistical fluctuation value is then used as the variance factor. The current iteration input.

[0042] Synchronous extraction of multidimensional operating state parameter sequences The indicated airspeed and altitude change rate are defined as physical state boundary parameters. ,calculate The physical trigger reference value of the corresponding flap adjustment node in the causal logic topology diagram. The absolute deviation; the logical bias operator is invoked to execute the weight compensation procedure, and the correction process follows the formula. ,in, The corrected operation semantics determination weights, For the original fusion weights, The physical state boundary parameter values ​​obtained in real time are in units of , The physical trigger baseline value corresponding to the target operation node, in units of , The variance factor characterizes the sampling fluctuation of the multidimensional operating state parameter sequence, with units of . ; through the logical bias operator in feature fusion The probability correction increment for flap handle actuation is superimposed on top, making the final recognition probability of this operation node increase from... Upgraded to .

[0043] Example 5: In a flight training center scenario where the system is deployed with heterogeneous avionics data acquisition terminals, the system executes a pre-calibration procedure for sensor latency consistency. The central control unit sends synchronization test level signals to the video stream acquisition interface, audio capture unit, and multi-dimensional operational status parameter sequence bus. It detects the difference in the original timestamps of the synchronization test level signals arriving in each modal data buffer and calculates the hardware latency reference value for different modal transmission links. The unit is Using linear regression algorithm to The arrival time of a group of consecutive sampling pulses is fitted to determine the dynamic drift compensation amount of the video modality relative to the multidimensional operating state parameter sequence bus. The unit is , calculate The clock calibration operator written to the data buffer enables subsequent fusion features. The calculation process is based on a time delay deviation of less than Synchronous spatiotemporal axis operation; when the system is applied to non-standard handling procedure evaluation scenarios for specific machine models, the system executes an operation causal logic topology graph parameter filling procedure based on offline data-driven operation; by acquiring The dataset includes flight operation quality monitoring data for landing maneuvers under complex weather conditions as initial state samples. A Hidden Markov Model (HMM) is used to partition the state space of the multidimensional operational state parameter sequence. The transition probabilities of elevator angle, throttle position, and control surface displacement under specific physical energy states are extracted. These calculated transition probabilities are then filled into the directed arc segment weight parameters of the operational causal logic topology graph, and combined with the statistical distribution variance of the physical quantity change trend operator. Determine the sensitive range of the physical state trigger vector for each operation node.

[0044] In scenarios where the evaluation system is deployed in the cockpit of a specific type of flight simulator, to address video coordinate reference errors caused by installation axis deviation of the acquisition module and ambient light background brightness shift, the system employs a spatial calibration method. This involves identifying visual feature anchor points at physical locations on the control panel and throttle handle, calculating their pixel offsets in the video coordinate system, and then using a homography transformation matrix to calculate the video motion feature vector. normalization operator This ensures that the coordinate deviation of the calibrated eigenvector in three-dimensional Euclidean space is within the range of... The system utilizes historical flight operation quality monitoring data to determine the boundary parameters of the operational causal logic topology; it analyzes multiple sets of standard approach and landing profile data to calculate the statistical mean of the pressure-altitude descent rate during the glide path capture phase, and sets this mean as the physical trigger reference value for the target operational node. The unit is Simultaneously calculate the statistical fluctuation range of the rate of change of air pressure altitude in the sample, and use the variance formula. Determine the variance factor The initial reference value, in units of ; For the first The physical parameter values ​​of each sampling point, in units of ; This represents the average physical parameter value within the sampling period, in units of... ; The total number of samples.

[0045] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0046] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A pilot training operation evaluation method based on multimodal temporal fusion, characterized in that, Includes the following steps: Step S101: Obtain the video stream, audio signal, and multi-dimensional operating state parameter sequence of the target object during the operation evolution process; use a unified clock signal to timestamp the video stream, audio signal, and multi-dimensional operating state parameter sequence to achieve time alignment of each modal data at the frame level. Step S102: Use a video processing model to extract spatial-temporal action feature vectors, use a speech recognition model to extract instruction semantic feature vectors, use a time series analysis model to extract operation control trend feature vectors of the target object, and combine the feature vectors into a multimodal feature vector. Step S103: Use a cross-modal attention mechanism to perform weighted fusion of multimodal feature vectors to generate fused features that include the original confidence scores of each operation semantic category; Step S104: Construct an operation causal logic topology graph containing physical boundary constraints. The operation causal logic topology graph consists of multiple operation nodes corresponding to the baseline operation procedure, and each operation node is associated with a physical state trigger vector. Step S105: Calculate the physical state boundary parameters of the controlled object in real time based on the multidimensional running state parameter sequence; if the physical state boundary parameters match the physical state trigger vector of the target operation node, apply a logical bias operator on the fused features to increase the judgment weight of the corresponding operation semantics of the target operation node. Step S106: Calculate the weighted scores of the identified operation events and the operation causal logic topology graph in terms of time consistency, sequence consistency, and parameter deviation, and output an evaluation report.

2. The pilot training operation evaluation method based on multimodal temporal fusion according to claim 1, characterized in that, The application of the logical bias operator in step S105 includes: extracting the first-order state change rate, actuator displacement, and motion vector parameters from the multi-dimensional operating state parameter sequence in real time; mapping the extracted parameters to a preset physical energy state space to determine the current temporal state mode of the controlled object; when there is a preset logical correlation between the temporal state mode and the target operation node, calculating the probability correction increment for the target operation node using a nonlinear activation function, and accumulating the probability correction increment to the original confidence score.

3. The pilot training operation evaluation method based on multimodal temporal fusion according to claim 1, characterized in that, In the causal logic topology graph, the operation nodes are connected by directed arcs; the directed arcs are used to represent the temporal constraints of the operation actions. The physical state trigger vector contains the control quantity threshold range required to characterize operational compliance, as well as the physical quantity change trend operator.

4. The pilot training operation evaluation method based on multimodal temporal fusion according to claim 1, characterized in that, Step S103, which involves weighted fusion of multimodal feature vectors, includes: calculating the mutual information of features in the video stream, audio signal, and multidimensional running state parameter sequence within the current time window; dynamically allocating attention weights for each modality based on the mutual information of features; and adjusting the attention weights of the corresponding state feedback signals in the multidimensional running state parameter sequence if the recognition confidence of the target action in the video stream is lower than a preset threshold.

5. The pilot training operation evaluation method based on multimodal temporal fusion according to claim 4, characterized in that, The process of increasing attention weight follows the weight adjustment formula: ;in, The corrected operation semantics determination weights, For the original fusion weights, To monitor the obtained physical state boundary parameter values ​​in real time, The physical trigger reference value corresponding to the target operation node. To characterize the variance factor of the sampling fluctuation of the multidimensional operating state parameter sequence, and and They have the same dimensions.

6. The pilot training operation evaluation method based on multimodal temporal fusion according to claim 1, characterized in that, Step S106 involves calculating the weighted score by: using a dynamic time warping algorithm to calculate the path offset distance between the identified operation event sequence and the baseline operation procedure; determining the corresponding deviation penalty coefficient based on the logical depth of each operation node in the operation causal logic topology graph; and multiplying the path offset distance and the deviation penalty coefficient to obtain a quantitative score representing the operation accuracy.

7. The pilot training operation evaluation method based on multimodal temporal fusion according to claim 1, characterized in that, The assessment report traces the identified anomalous operations through the operation of a causal logic topology diagram to identify the physical logic defects that cause operational deviations, and outputs improvement suggestions for the energy state management logic of the controlled object.

8. The pilot training operation evaluation method based on multimodal temporal fusion according to claim 1, characterized in that, Step S101, which achieves time alignment of each modal data at the frame level, includes: obtaining the inherent sampling frequency of each modal sensor, with a sampling frequency range of [missing information]. to The Lagrange interpolation algorithm is used to upsample the data sequence of the low sampling rate mode so that its temporal resolution is aligned with the video frame rate of the highest sampling rate mode.

9. The pilot training operation evaluation method based on multimodal temporal fusion according to claim 1, characterized in that, The real-time calculation of the physical state boundary parameters of the controlled object in step S105 includes: constructing a real-time state profile model of the controlled object based on a multi-dimensional operating state parameter sequence; and determining the logical necessity of the target operating node being activated under the current physical environment by calculating the spatiotemporal logical coupling entropy between the real-time state profile model and the physical state trigger vector of the target operating node.

10. The pilot training operation evaluation method based on multimodal temporal fusion according to claim 1, characterized in that, Step S106 and thereafter includes: real-time monitoring of the numerical changes of physical state boundary parameters; if the slope of the change of physical state boundary parameters exceeds a preset mutation threshold, the operation causal logic topology diagram is switched to the emergency logic processing mode corresponding to the mutation threshold, and the weighted score is updated according to the constraint rules in the emergency logic processing mode.

Citation Information

Patent Citations

  • Automated assisted driving method based on multi-modal information fusion

    CN119861614B

  • Pilot manipulation ability evaluation method and system based on multi-modal data and online learning driving

    CN120296663A

  • Flight simulator multi-modal data-based feature fusion model construction method

    CN120673206A

  • Pilot dynamic evaluation method, system and equipment based on TEM model and storage medium

    CN120911788A

  • Intelligent model construction method and system for student training data analysis

    CN121257673A