Non-intrusive experimental evaluation system and method for pilot-human trust behavior

By collecting multi-dimensional data for synchronous matching and comprehensive analysis, a comprehensive trust behavior index is generated, which solves the problem of human-machine trust assessment results deviation in existing technologies, realizes accurate identification and dynamic assessment of pilot trust status, and supports flight training and airborne system design.

CN122153275APending Publication Date: 2026-06-05AIR FORCE MEDICAL CENT PLA +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
AIR FORCE MEDICAL CENT PLA
Filing Date
2026-01-16
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing human-machine trust assessment methods lack the time-synchronous correlation of subjective and objective data and the fusion verification of multi-source data, resulting in deviations between assessment results and actual trust behavior. This makes it difficult to accurately identify the imbalance of pilots' trust status in emergency situations and fails to provide accurate data support for flight training optimization and intelligent design of airborne systems.

Method used

By collecting pilots' eye-tracking data, operational command data, and flight instrument parameter data, and combining subjective assessments and expert evaluations, multi-dimensional data mapping and synchronous matching are performed to generate a comprehensive trust behavior index, enabling a non-intrusive assessment of pilots' human-machine trust behavior.

Benefits of technology

It enables accurate identification and dynamic quantitative assessment of behavioral characteristics such as excessive trust, insufficient trust, and self-perception bias among pilots, providing reliable data support for flight training optimization and intelligent design of airborne systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122153275A_ABST
    Figure CN122153275A_ABST
Patent Text Reader

Abstract

The application provides a non-intrusive experimental evaluation system and method for pilot human-machine trust behavior, and relates to the technical field of human-computer interaction, which comprises the following steps: collecting a plurality of original data points including eye movement trajectory data points, operation instruction data points and flight instrument parameter data points of a pilot to form an objective behavior data set; receiving a plurality of subjective evaluation data points of the pilot based on the objective behavior data set to form a subjective evaluation data set; receiving a plurality of expert evaluation data points of a flight expert according to the subjective evaluation data set and the objective behavior data set to form an expert evaluation data set; and synchronously matching each data point in the objective behavior data set, the subjective evaluation data set and the expert evaluation data set according to a flight task time node to generate a multi-dimensional behavior mapping data set with time corresponding relationship. The application can realize accurate identification, dynamic evaluation and quantitative analysis of pilot human-machine trust behavior.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-computer interaction technology, and in particular to a non-invasive experimental evaluation system and method for pilot human-computer trust behavior. Background Technology

[0002] With the widespread application of semi-automatic and automated systems and artificial intelligence technologies on aircraft, the frequency of human-machine interaction has increased significantly. Human-machine trust has become a key factor affecting pilots' cognitive decision-making and flight safety. In the future trend of manned and unmanned collaboration, the pilot's trust level with automated systems (over-trust or under-trust) is directly related to the smooth execution of flight missions. Once the trust level is unbalanced, it may lead to problems such as delayed response to abnormal system information and operational decision-making errors, which may in turn cause major safety accidents.

[0003] Currently, human-machine trust assessment mainly relies on subjective reports (such as questionnaires) or single-dimensional objective data collection, which has the following technical shortcomings: It lacks a mechanism for the synchronous correlation of subjective and objective data and the fusion and verification of multi-source data. For example, existing assessment methods either rely solely on pilots' subjective feedback, failing to accurately reflect dynamic trust changes in actual operation; or they only collect objective data such as eye movements and flight parameters, without effectively linking them with subjective cognition and expert evaluation, leading to discrepancies between assessment results and actual trust behavior. This deficiency can lead to several adverse consequences: it is difficult to accurately identify pilots' trust imbalances in key scenarios such as emergencies and flight illusions; for example, it is impossible to capture the implicit behavior of pilots ignoring instrument anomalies due to excessive trust in the system through subjective questionnaires; at the same time, the assessment results lack comprehensiveness and reliability, failing to provide accurate data support for flight training optimization and intelligent design of airborne systems, and making it difficult to effectively intervene in and regulate human-machine trust states. Summary of the Invention

[0004] The technical problem to be solved by this invention is to provide a non-invasive experimental evaluation system and method for pilot human-machine trust behavior, which can realize accurate identification, dynamic evaluation and quantitative analysis of pilot human-machine trust behavior, and provide support for flight training optimization, intelligent design of airborne systems and flight safety assurance.

[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0006] Firstly, a non-invasive experimental evaluation method for pilot-machine trust behavior, the method comprising:

[0007] Collect multiple raw data points, including pilot eye movement trajectory data points, operation command data points, and flight instrument parameter data points, to form an objective behavioral data set;

[0008] Based on the objective behavioral data set, multiple subjective evaluation data points from the pilot are received to form a subjective evaluation data set;

[0009] Based on the subjective evaluation data set and the objective behavior data set, multiple expert evaluation data points from flight experts are received to form an expert evaluation data set;

[0010] Each data point in the objective behavior data set, subjective evaluation data set, and expert evaluation data set is synchronously matched according to the flight mission time nodes to generate a multi-dimensional behavior mapping data set with time correspondence.

[0011] By using a multi-dimensional behavior mapping dataset, the visual attention feature data points and flight operation response data points are correlated and processed to form data points for the consistency verification of subjective and objective evaluation results.

[0012] Based on the data points of the consistency verification results of subjective and objective assessments, the data points of the trust behavior questionnaire, the self-assessment data points of the behavior, and the expert assessment data points are standardized, transformed, and weighted to generate a comprehensive trust behavior index data point.

[0013] Based on the comprehensive trust behavior index data points, the standardized trust behavior indicator data points are compared and analyzed with the objective flight parameter data points to identify and output the behavioral characteristic assessment results data points that represent excessive trust, insufficient trust, or self-cognition bias, thus completing a non-intrusive comprehensive assessment of the pilot's human-machine trust behavior.

[0014] Furthermore, based on the objective behavioral data set, multiple subjective evaluation data points from the pilots are received to form a subjective evaluation data set, including:

[0015] Based on the objective behavioral data set, the trust behavior questionnaire data points, which are filled out by the pilots before entering the flight simulator and are used to characterize the pilots' preset level of trust in the airborne automation capabilities and their behavioral tendencies, are received and recorded. At the same time, the trust behavior questionnaire data points are used as the first type of subjective evaluation data points.

[0016] The system receives and records the self-evaluation data points of the pilots' behavior after completing all simulated flight missions. These self-evaluation data points include the pilots' self-evaluation data on their visual attention allocation, actual operational behavior, and operational risk perception in various mission scenarios. The self-evaluation data points are also used as the second type of subjective evaluation data points.

[0017] The first type of subjective evaluation data points and the second type of subjective evaluation data points are combined to form a subjective evaluation data set.

[0018] Furthermore, based on the subjective assessment data set and the objective behavioral data set, multiple expert evaluation data points from flight experts are received to form an expert evaluation data set, including:

[0019] It receives objective behavioral data sets and provides flight experts with a synchronized observation interface for eye-tracking data points and flight instrument parameter data points.

[0020] Based on the synchronous observation interface, the system receives real-time expert evaluation data points generated by flight experts during the simulated flight mission. These data points include flight experts' evaluation data on the pilot's visual attention characteristics, operational response characteristics, and operational risk level in various mission scenarios.

[0021] All expert real-time evaluation data points are aggregated to form an expert evaluation data set.

[0022] Furthermore, for each data point in the objective behavior data set, subjective evaluation data set, and expert evaluation data set, synchronous matching processing is performed according to the flight mission time nodes to generate a multi-dimensional behavior mapping data set with time correspondence, including:

[0023] Acquire an objective behavior data set and mark the corresponding first experimental timestamp for each eye-tracking data point, operation command data point, and flight instrument parameter data point in the objective behavior data set.

[0024] Receive the expert evaluation data set and mark the second experimental timestamp corresponding to the observation and evaluation time for each real-time self-evaluation data point of the expert in the expert evaluation data set;

[0025] Obtain behavioral self-evaluation data points from the subjective evaluation dataset, and associate all behavioral self-evaluation data points with corresponding scenario time periods based on the specific simulated flight mission scenarios evaluated by the behavioral self-evaluation data points.

[0026] Based on the first experimental timestamp, the second experimental timestamp, and the scene time period identifier, data points from the objective behavior data set, the subjective evaluation data set, and the expert evaluation data set, which correspond to the same flight mission phase, are aligned and integrated to generate a multi-dimensional behavior mapping data set.

[0027] Furthermore, through a multi-dimensional behavior mapping dataset, correlation analysis is performed on visual attention feature data points and flight operation response data points to form data points for verifying the consistency of subjective and objective evaluations, including:

[0028] From the multi-dimensional behavior mapping dataset, visual attention feature data points and flight operation response data points corresponding to the same flight mission phase are extracted; the visual attention feature data points are derived from eye track data points, and the flight operation response data points are derived from operation command data points and the flight instrument parameter change data points they cause.

[0029] Based on the extracted visual attention feature data points and flight operation response data points, correlation analysis index data points are calculated and generated; wherein the correlation analysis index data points include at least the attention transition time interval data points from the appearance of key flight information to the pilot's first gaze at the area, and the operation response delay data points from the pilot's stable gaze at the key information to the execution of the corresponding operation command.

[0030] The calculated correlation analysis index data points are compared and analyzed with the self-evaluation data points of behavior extracted from the multi-dimensional behavior mapping data set that correspond to the same flight mission phase, to generate data points for the consistency verification of subjective and objective evaluation.

[0031] Furthermore, based on the data points from the consistency verification results of subjective and objective assessments, the trust behavior questionnaire data points, self-assessment data points, and expert peer assessment data points are standardized, transformed, and weighted to generate a comprehensive trust behavior index data point, including:

[0032] Based on the data points of the consistency verification results of subjective and objective assessments, trust behavior questionnaire data points and self-assessment data points are extracted from the subjective assessment data set, and expert peer assessment data points are extracted from the expert evaluation data set to form the original assessment data point set.

[0033] The original set of assessment data points is preprocessed to generate a preprocessed set of assessment data points. This preprocessed set consists of trust behavior questionnaire data points, self-assessment data points, and expert peer assessment data points for each individual pilot, forming a multi-dimensional data point set within the assessment space. The convex hull of the data point set is calculated to identify data points inside the convex hull as valid core data points, while data points outside the convex hull are identified as marginal or outlier data points. Marginal or outlier data points are then filtered or adjusted to form the preprocessed set of assessment data points.

[0034] The original evaluation value data points from various evaluation sources in the preprocessed evaluation data point set are standardized and transformed to generate corresponding standardized subscale score data points.

[0035] Based on preset weighting coefficients, all standardized subscale score data points are weighted and summed to generate a comprehensive trust behavior index data point.

[0036] Furthermore, based on the comprehensive trust behavior index data points, the standardized trust behavior indicator data points are compared and analyzed with the objective flight parameter data points to identify and output the behavioral characteristic assessment results data points representing excessive trust, insufficient trust, or self-perception bias. This completes a non-invasive comprehensive assessment of the pilot's human-machine trust behavior, including:

[0037] It receives comprehensive trust behavior index data points, standardized subscale score data points as standardized trust behavior indicator data points, and key scenario objective flight parameter data points extracted from multi-dimensional behavior mapping data sets;

[0038] Call the preset trust state discrimination threshold data points and the preset behavior pattern rule base data points;

[0039] Based on the comprehensive trust behavior index data points, and compared with the preset trust status judgment threshold data points, a preliminary trust level classification is performed to generate initial trust classification data points;

[0040] Based on standardized subscale score data points and key scenario objective flight parameter data points, combined with initial trust classification data points, and according to the preset behavior pattern rule base data points, behavior pattern matching and cross-validation analysis are performed to generate detailed trust status discrimination data points.

[0041] By integrating the initial trust classification data points with the detailed judgment data points of trust status, the system generates and outputs behavioral characteristic assessment result data points that represent excessive trust, insufficient trust, or self-perception bias.

[0042] Secondly, a non-invasive experimental evaluation system for pilot-machine trust behavior includes:

[0043] The data point acquisition module is used to collect multiple raw data points, including pilot's eye movement trajectory data points, operation command data points, and flight instrument parameter data points, to form an objective behavior data set; based on the objective behavior data set, it receives multiple subjective evaluation data points from the pilot to form a subjective evaluation data set; and based on the subjective evaluation data set and the objective behavior data set, it receives multiple expert evaluation data points from flight experts to form an expert evaluation data set.

[0044] The synchronization matching module is used to perform synchronous matching processing on each data point in the objective behavior data set, subjective evaluation data set, and expert evaluation data set according to the flight mission time nodes, and generate a multi-dimensional behavior mapping data set with time correspondence.

[0045] The correlation analysis module is used to perform correlation analysis on visual attention feature data points and flight operation response data points through a multi-dimensional behavior mapping dataset, forming data points for verifying the consistency between subjective and objective evaluations.

[0046] The conversion and calculation module is used to standardize and weight the trust behavior questionnaire data points, self-assessment data points, and expert assessment data points based on the consistency verification results of subjective and objective evaluations, and generate comprehensive trust behavior index data points.

[0047] The identification and output module is used to compare and analyze standardized trust behavior index data points with objective flight parameter data points based on comprehensive trust behavior index data points, identify and output behavioral characteristic assessment result data points that represent excessive trust, insufficient trust or self-cognition bias, and complete a non-intrusive comprehensive assessment of pilot human-machine trust behavior.

[0048] The above-described solution of the present invention has at least the following beneficial effects:

[0049] Because it employs multi-source objective data acquisition, time synchronization matching of subjective and objective evaluation data, convex hull preprocessing and standardized weighted calculation, and cross-validation of behavioral patterns, it effectively overcomes the technical shortcomings of traditional evaluations, such as the lack of fusion verification of subjective and objective data and the difficulty in accurately capturing dynamic trust behavior. This enables accurate identification and dynamic quantitative evaluation of behavioral characteristics such as excessive trust, insufficient trust, and self-perception bias in pilots, providing reliable data support for flight training optimization, intelligent design of airborne systems, and flight safety assurance. Attached Figure Description

[0050] Figure 1 This is a flowchart illustrating a non-invasive experimental evaluation method for pilot-machine trust behavior provided in an embodiment of the present invention.

[0051] Figure 2 This is a schematic diagram of a non-invasive experimental evaluation system for pilot-machine trust behavior provided in an embodiment of the present invention. Detailed Implementation

[0052] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0053] like Figure 1 As shown, embodiments of the present invention propose a non-invasive experimental evaluation method for pilot-machine trust behavior, the method comprising the following steps:

[0054] Step 1: Collect multiple raw data points, including pilot's eye movement trajectory data points, operation command data points, and flight instrument parameter data points, to form an objective behavioral data set;

[0055] Step 2: Based on the objective behavioral data set, receive multiple subjective evaluation data points from the pilots to form a subjective evaluation data set;

[0056] Step 3: Based on the subjective evaluation data set and the objective behavior data set, receive multiple expert evaluation data points from flight experts to form an expert evaluation data set;

[0057] Step 4: Synchronously match each data point in the objective behavior data set, subjective evaluation data set, and expert evaluation data set according to the flight mission time nodes to generate a multi-dimensional behavior mapping data set with time correspondence.

[0058] Step 5: Through a multi-dimensional behavior mapping dataset, perform correlation analysis on visual attention feature data points and flight operation response data points to form data points for verification of consistency between subjective and objective evaluations.

[0059] Step 6: Based on the data points of the consistency verification results of subjective and objective assessments, standardize and weight the trust behavior questionnaire data points, self-assessment data points, and expert peer assessment data points to generate comprehensive trust behavior index data points.

[0060] Step 7: Based on the comprehensive trust behavior index data points, compare and analyze the standardized trust behavior index data points with the objective flight parameter data points to identify and output the behavioral characteristic assessment results data points that represent excessive trust, insufficient trust, or self-cognition bias, thus completing the non-intrusive comprehensive assessment of the pilot's human-machine trust behavior.

[0061] In this embodiment of the invention, multi-dimensional objective behavioral data, subjective assessments, and expert evaluations are collected and synchronized with flight mission time nodes to achieve effective correlation between data from different sources. The consistency between subjective and objective assessments is verified through correlation analysis of visual attention and operational response. A comprehensive trust behavior index is generated through standardized transformation and weighted calculation, accurately identifying behavioral characteristics such as excessive trust, insufficient trust, and self-perception bias in pilots. This method does not interfere with normal pilot operations, overcomes the limitations of traditional subjective assessments, and provides comprehensive and reliable evaluation results, offering strong support for flight training optimization, intelligent design of airborne systems, and flight safety assurance.

[0062] In a preferred embodiment of the present invention, step 1 above may include:

[0063] Step 1.1 involves real-time acquisition of the pilot's eye movement trajectory data points using a wearable eye tracker, real-time acquisition of the pilot's operational command data points through the flight simulator's operating interface, and real-time acquisition of flight instrument parameter data points. Specifically, this includes: using a high-fidelity flight simulator as the core hardware device; the simulator accurately recreates the real flight cockpit environment and operational logic, supporting various simulated flight missions such as takeoff, handling illusions, handling in-flight emergencies, and return landing; and recording core parameters during flight in real time; using a pre-calibrated wearable eye tracker to capture real-time information such as the pilot's pupil dynamic changes, gaze coordinates, gaze path, and gaze area during simulated flight, forming continuous eye movement trajectory data points; capturing real-time information such as the pilot's trigger actions, operational force, and operational range on the joystick, throttle, buttons, and other operating components through the flight simulator's control interface, generating corresponding operational command data points; and simultaneously acquiring various core parameters displayed on the flight instruments, such as flight altitude, flight speed, heading angle, bank angle, angle of attack, and Mach number, during simulated flight, forming flight instrument parameter data points.

[0064] Step 1.2 involves performing real-time timestamp marking and data integrity verification on the real-time acquired eye-tracking data points, operation command data points, and flight instrument parameter data points. This generates verified eye-tracking data points, operation command data points, and flight instrument parameter data points with unified timestamps. Specifically, based on the unified timing benchmark after the flight simulator starts and enters the formal experimental state, a unique timestamp is marked for each real-time acquired eye-tracking data point, operation command data point, and flight instrument parameter data point to ensure that the time recording standards of the three types of data are consistent. Subsequently, integrity verification is performed on the three types of data to check for missing keyframes in the eye-tracking data, broken records in the operation commands, and abnormal or blank values ​​in the flight instrument parameters. Data points with missing or abnormal values ​​are marked or preliminarily filtered, ultimately generating verified eye-tracking data points, operation command data points, and flight instrument parameter data points with unified timestamps and valid data.

[0065] Step 2.3 involves summarizing the verified eye-tracking data points, operational command data points, and flight instrument parameter data points with unified timestamps to form an objective behavior data set. Specifically, this includes: using the unified timestamp as the core index, systematically summarizing the verified eye-tracking data points, operational command data points, and flight instrument parameter data points according to the chronological order of the experiment; during the summarization process, binding the three types of data corresponding to the same timestamp one by one to form complete data entries containing time stamps, eye-tracking behavior, operational actions, and flight status, ensuring a precise correlation between the pilot's visual behavior and operational behavior at each experimental moment and the corresponding flight parameters; subsequently, integrating and storing all bound structured data entries in chronological order to construct a complete and coherent objective behavior data set, which clearly presents the pilot's behavior and status changes throughout the simulated flight.

[0066] In this embodiment of the invention, multi-dimensional objective data such as pilot eye movement trajectories, operation commands, and flight instrument parameters are comprehensively collected. Then, the time synchronization of various types of data is achieved by marking them with a unified timestamp. Data validity is ensured through data integrity verification. Finally, the verified multi-source data is summarized according to time correlation to form an objective behavior data set. This not only ensures the comprehensiveness of data coverage and the consistency of time dimensions, but also improves the reliability and structure of the data. It provides a foundation for the synchronous matching and subjective-objective fusion analysis of subjective evaluation data and expert evaluation data, and effectively supports the dynamic and accurate assessment of pilot human-machine trust behavior.

[0067] In a preferred embodiment of the present invention, step 2 above may include:

[0068] Step 2.1: Based on the objective behavioral data set, receive and record the trust behavior questionnaire data points filled out by the pilots before entering the flight simulator. These data points characterize the pilots' pre-set trust level and behavioral tendencies towards the onboard automation capabilities. Simultaneously, these trust behavior questionnaire data points are used as the first type of subjective evaluation data points. Specifically, this includes: before the formal experiment begins, preparing and verifying a standardized human-machine trust behavior evaluation questionnaire to ensure it contains both behavioral intentions and potential behaviors, with no missing items; before the pilots enter the flight simulator, clearly informing them of the purpose of filling out the questionnaire and the information confidentiality rules, explaining that the behavioral intentions section should revolve around pre-set scenarios such as routine flight and emergency handling, requiring them to fill out their subjective trust tendencies and behavioral plans towards the onboard automation system; and the potential behaviors section should combine past flight experience to predict the explicit operational reactions to the onboard automation system in actual flight situations; the pilots independently fill out the questionnaire in an independent environment, with questions answered promptly without interfering with their independent judgment; the questionnaires are collected on-site after completion and checked one by one. After ensuring completeness and logical consistency, and confirming accuracy, the responses are recorded according to unified rules, generating trust behavior questionnaire data points. Each data point is then associated with a unique pilot identifier, the completion time, and the attributes of the first type of subjective assessment data point, thus completing the collection and organization of preset trust-related subjective data. Finally, it should be noted that the aforementioned airborne automation system is the core auxiliary operating system carried by the aircraft, integrating semi-automatic control, intelligent decision support, and other technologies, covering key aspects of the entire flight process. Its core functions include automatically adjusting core parameters such as flight altitude, speed, and heading to assist in completing routine tasks such as takeoff, cruise, and landing. It can also identify in-flight emergencies, flight illusions, and other abnormal scenarios, issuing warnings and providing standardized handling suggestions or partially autonomous handling operations. By reducing repetitive pilot operations and improving mission execution efficiency, the airborne automation system becomes the core carrier of human-machine interaction. Its operational reliability and functional adaptability directly affect the pilot's trust tendency and are the core object of the questionnaire assessment of the pilot's preset trust level in this step.

[0069] Step 2.2: Receive and record the behavioral self-assessment data points submitted by the pilots after completing all simulated flight missions. These data points include the pilots' self-evaluation data on their visual attention allocation, actual operational behavior, and operational risk perception in various mission scenarios. Simultaneously, these behavioral self-assessment data points are used as the second type of subjective evaluation data points. Specifically, after completing all simulated flight missions, including takeoff, illusion handling, in-flight emergency handling, and return landing, the pilots first confirm that all flight data has been fully recorded, and then provide a pilot self-assessment form consistent with the flight expert's evaluation indicators. The three core dimensions of the self-assessment form are clearly defined: the visual attention level must correspond to each mission scenario, and the evaluation focuses on the cockpit. The timeliness, appropriateness, and duration of attention to internal instruments are assessed. At the operational level, the consistency between actions and instrument information, the speed of response after focusing on the instruments, and the degree of compliance with prescribed actions are reviewed. At the risk perception level, the risk level of one's own operational behavior is objectively evaluated. A brief prompt is provided regarding the sequence of each task scenario in the experiment to aid in accurate matching, without guiding the evaluation direction. Pilots independently complete the form based on their actual operational process and real-time experience. After completion, the completeness is checked, and once confirmed to be complete, the self-evaluation results are recorded according to unified rules, generating behavioral self-evaluation data points. These data points are then linked to the pilot's unique identifier, the experiment completion time, and the attributes of the second type of subjective evaluation data points, completing the collection and organization of relevant subjective data for post-experimental self-evaluation.

[0070] Step 2.3 involves aggregating the first and second categories of subjective evaluation data points to form a subjective evaluation dataset. This includes: initiating the aggregation process for both categories of subjective evaluation data using the pilot's unique identifier as the core association basis; firstly, extracting the first and second categories of subjective evaluation data points according to the identifier, and supplementing each category with complete labels such as experimental batch, experimental date, data source type (preset trust or post-event self-evaluation), and corresponding completion time; then establishing data associations and constructing a unified data framework containing fields such as pilot's unique identifier, experimental batch, data source, evaluation dimension, specific indicator score, and completion time; integrating the two categories of data for the same pilot according to the evaluation dimension to form a complete subjective evaluation data entry for a single individual; comprehensively verifying all integrated entries, identifying and correcting any anomalies such as missing data, incorrect identifier correspondence, or label confusion; and storing all data entries in a standardized format after verification to ultimately form a complete and standardized subjective evaluation dataset, ensuring that the two categories of subjective data can be quickly associated and matched with objective data and expert evaluation data through the core identifier.

[0071] In a preferred embodiment of the present invention, step 3 above may include:

[0072] Step 3.1 involves receiving an objective behavioral data set and providing flight experts with a synchronized observation interface for eye-tracking data points and flight instrument parameter data points. Specifically, this includes: first, receiving an objective behavioral data set that has been timestamped and verified for integrity. This set contains continuous eye-tracking data points, including pupil changes, gaze coordinates, observation paths, operational command data points, and flight instrument parameter data points, including flight altitude, speed, heading, bank angle, angle of attack, and Mach number. Based on a unified timing benchmark, a real-time data synchronization and association mechanism is established, precisely binding eye-tracking data corresponding to the same timestamp with flight instrument parameters to ensure complete alignment in time, generating a synchronized data stream. A dedicated synchronized observation interface is then built in a visual format, divided into three core functional areas: an eye-tracking dynamic display area, presenting the pilot's eye movement trajectory and gaze focus in real time; a real-time instrument data area, synchronously refreshing various flight parameters and abnormal warning information; and a mission scenario identification area, clearly indicating the type of flight mission currently being executed, such as takeoff, illusion handling, and emergency handling, allowing flight experts to intuitively and clearly see the real-time correspondence between the pilot's visual attention behavior and the concurrent flight status.

[0073] Step 3.2: Based on the synchronous observation interface, receive real-time expert evaluation data points generated by flight experts during the simulated flight mission. These data points include evaluations of the pilot's visual attention characteristics, operational response characteristics, and operational risk levels in various mission scenarios. Specifically, after the simulated flight mission officially begins, the synchronous observation interface updates data in real-time as the mission progresses. An evaluation team of no fewer than three flight experts tracks the dynamic process of each mission scenario based on the eye-tracking trajectory and flight instrument parameter linkage information displayed on the interface. For different stages such as takeoff, various random illusion scenarios, various random special situation scenarios, and return landing, real-time evaluations must be conducted strictly according to preset unified evaluation indicators: the timeliness of visual attention focus on instruments, and the appearance of key information. Whether the focus is rapid and reasonable, whether the gaze path conforms to the task logic and gaze duration, and whether the information processing is appropriate; the operational response characteristics revolve around the consistency between the operation and instrument information, whether the operation matches the instrument display, the speed of the action response after focusing on the instrument, whether the corresponding operation is executed in a timely manner, the degree of compliance with the prescribed actions, and whether the standard operating procedures are followed; the operational risk level is comprehensively judged based on the stability of the flight status and the safety of handling anomalies; at the end of each mission scenario or at key event nodes, such as when the airborne automation system issues a warning or when parameters fluctuate abnormally, the respective evaluation results are immediately recorded, generating real-time expert evaluation data points that include the pilot's unique identifier, experimental batch, current mission scenario, evaluation dimension, specific indicator score, evaluation timestamp, and expert number, and ensuring that each evaluation result is accurately traceable.

[0074] Step 3.3 involves aggregating all real-time expert evaluation data points to form an expert evaluation data set. This includes: initiating the aggregation process for real-time expert evaluation data points based on the pilot's unique identifier and the experimental batch; firstly, classifying all real-time expert evaluation data points according to the pilot's unique identifier, and then reorganizing the evaluation data for the same pilot according to the experimental batch and mission scenario; supplementing each data point with complete tag information, including the evaluation expert number, corresponding mission scenario identifier, and evaluation timestamp, to ensure data attribute integrity; and then conducting data consistency verification: on the one hand, verifying the evaluations of the same expert for the same pilot in the same scenario. The evaluation results are checked for logical contradictions. On the other hand, the evaluation indicators of different experts for the same pilot in the same scenario are cross-compared to ensure that the evaluation direction is consistent and there is no obvious deviation. Any abnormal data is promptly communicated with the corresponding experts for verification and correction. After verification, all other evaluation data items are integrated according to a standardized data structure. The fields of each item clearly include the pilot's unique identifier, experimental batch, mission scenario, evaluation dimension, specific indicator score, evaluation timestamp, and expert number. Finally, a standardized, complete and traceable set of expert evaluation data is formed to ensure that it can achieve accurate matching in the time dimension through the core identifier and the set of objective behavioral data and subjective evaluation data.

[0075] In this embodiment of the invention, by receiving and merging objective behavioral datasets to build a synchronous observation interface for eye-tracking trajectories and flight instrument parameters, a comprehensive and real-time evaluation basis is provided for flight experts. Then, during the simulated flight mission, flight experts generate real-time expert evaluation data points covering visual attention features, operational response features, and operational risk levels based on observations. Finally, these data are summarized to form a standardized set of expert evaluation data. This not only achieves dynamic synchronization between expert evaluation and the flight process, ensuring the timeliness and relevance of the evaluation, but also allows expert evaluation to form a multi-source linkage with objective behavioral data and subjective evaluation data. This effectively makes up for the lack of professional perspective verification in traditional evaluations, improves the objectivity, comprehensiveness, and accuracy of pilot-machine trust behavior evaluation, and lays a reliable expert evaluation foundation for multi-dimensional data fusion analysis and trust status identification.

[0076] In a preferred embodiment of the present invention, step 4 above may include:

[0077] Step 4.1: Obtain the objective behavior data set and mark each eye-tracking data point, operation command data point, and flight instrument parameter data point in the objective behavior data set with its corresponding first experimental timestamp. Specifically, this includes: first, obtaining the objective behavior data set that has undergone data integrity verification and structured integration; the objective behavior data set contains continuously collected eye-tracking data points, including pupil changes, gaze point coordinates, and observation path; operation command data points, including joystick actions, throttle adjustments, and button triggers; and flight instrument parameter data points, including flight altitude, speed, heading, bank angle, angle of attack, and Mach number; before initiating the timestamp marking process. First, confirm that the unified timing benchmark used in the experiment has been calibrated to ensure that the timing accuracy reaches the millisecond level and remains stable throughout the process. Then, according to the data acquisition sequence, mark the corresponding first experimental timestamp for each data point: the timestamp of the eye track data point precisely corresponds to the instant when the eye tracker acquires the data, the timestamp of the operation command data point is synchronized with the actual moment when the pilot triggers the operation component, and the timestamp of the flight instrument parameter data point is completely consistent with the moment when the instrument parameters are refreshed in real time. After marking, attach a timing benchmark calibration record to each timestamp, including calibration time and calibration deviation value, to ensure that the time recording standards of the three types of data are unified and the source is traceable.

[0078] Step 4.2: Receive the expert evaluation data set and mark each real-time peer-evaluation data point in the expert evaluation data set with a second experimental timestamp corresponding to the observation and evaluation time. Specifically, this includes: receiving the expert evaluation data set that has been classified, checked for consistency, and supplemented with labels. The expert evaluation data set contains all real-time peer-evaluation data points entered by no fewer than three flight experts during the simulated flight mission, and each data point is associated with attributes such as the pilot's unique identifier, experimental batch, and mission scenario; initiating the second experimental timestamp marking process based on the system's unified timing benchmark that is completely consistent with the first experimental timestamp; capturing the precise moment when each expert submits their evaluation result in real time, and using this moment as the core value of the second experimental timestamp for the corresponding real-time peer-evaluation data point, ensuring that the time accuracy remains consistent with the first experimental timestamp during the marking process; simultaneously, binding the second experimental timestamp with the corresponding expert real-time peer-evaluation data point's expert number, mission scenario identifier, evaluation dimension, and other attributes to form a correlation between timestamp, evaluation attributes, and evaluation content, ensuring that the expert evaluation behavior is accurately aligned with the corresponding flight process moment and can effectively connect with the time dimension of objective behavioral data.

[0079] Step 4.3: Obtain the behavioral self-evaluation data points from the subjective evaluation dataset. Simultaneously, based on the specific simulated flight mission scenarios evaluated by each behavioral self-evaluation data point, associate all behavioral self-evaluation data points with corresponding scenario time periods. Specifically, this includes: obtaining the subjective evaluation dataset that has undergone structured integration and integrity verification; selecting the second type of subjective evaluation data points, namely behavioral self-evaluation data points; this type of data contains the pilot's self-evaluation of visual attention allocation, actual operational behavior, and operational risk perception under various mission scenarios; and retrieving the mission execution logs automatically recorded during the experiment, which clearly record takeoff, various random illusion fields, etc. The specific start and end times of all simulated flight mission scenarios, including various random special situations and return landing, are based on a unified timing benchmark. For each behavioral self-evaluation data point, the scenario keywords in its evaluation description, such as illusion handling and special situation response, are analyzed. Combined with the scenario recall prompts recorded by the pilot, the data point is accurately matched with the specific scenario in the mission execution log. After the matching is confirmed, a corresponding scenario time period identifier is associated with the behavioral self-evaluation data point, including the scenario name, scenario start timestamp, and scenario end timestamp, ensuring that each behavioral self-evaluation data point can be clearly associated with a specific flight mission stage.

[0080] Step 4.4: Based on the first experimental timestamp, the second experimental timestamp, and the scene time period identifier, data points from the objective behavior data set, subjective evaluation data set, and expert evaluation data set that correspond to the same flight mission phase are aligned and integrated to generate a multi-dimensional behavior mapping data set. Specifically, this includes: First, based on the pre-set mission flow and mission execution log, the entire simulated flight experiment is divided into independent flight mission phases such as takeoff, various random illusion scenarios, various random special situation scenarios, and return landing, clearly defining the time range boundaries of each phase, using a unified timing benchmark timestamp as the basis; then, the multi-source data alignment and integration process is initiated: Within each flight mission phase, all data points in the objective behavior data set whose first experimental timestamp falls within the time range of that phase are first selected; then, expert real-time peer evaluation data points in the expert evaluation data set whose second experimental timestamp falls within the time range of that phase are matched; simultaneously, the self-evaluation data of behavior in the subjective evaluation data set that completely corresponds to the scene time period identifier of that phase is extracted. The data points were grouped into three categories based on pilot unique identifier, experimental batch, and mission phase. Each group was then sorted chronologically, and structured integration entries were constructed. Each entry included core fields such as pilot unique identifier, experimental batch, mission phase name, phase time range, objective behavioral data (including the first experimental timestamp and specific values ​​for each data point), expert evaluation data (including the second experimental timestamp, expert number, and evaluation score for each data point), and behavioral self-evaluation data (including scene time period identifier and specific self-evaluation content). After integration, data correlation verification was performed on each entry: checking the rationality of objective data timestamps and expert evaluation timestamps within the same phase, confirming that the scene time period corresponding to the behavioral self-evaluation data was completely consistent with the phase time range, and investigating for time conflicts, incorrect scene correspondence, data missing, etc. Any abnormal data discovered was promptly verified and corrected. After verification, all structured integration entries were stored uniformly according to the flight mission phase sequence, generating a complete and coherent multi-dimensional behavioral mapping data set.

[0081] In a preferred embodiment of the present invention, step 5 above may include:

[0082] Step 5.1: Extract visual attention feature data points and flight operation response data points corresponding to the same flight mission phase from the multi-dimensional behavior mapping dataset. The visual attention feature data points originate from eye-tracking data points, and the flight operation response data points originate from operation command data points and the resulting changes in flight instrument parameters. Specifically, this includes: first, retrieving the multi-dimensional behavior mapping dataset that has already been integrated from multiple sources. This dataset is structured and stored according to flight mission phases such as takeoff, various random illusion scenarios, various random special situation scenarios, and return landing. Each phase is associated with complete objective behavior data, subjective evaluation data, and expert evaluation data. For each independent flight mission phase, first locate the corresponding objective behavior data module, and then extract two types of core data from it step by step: one is visual attention features... The data points are derived directly from eye-tracking data, specifically including the amplitude of pupil contraction or dilation, the precise coordinates of the fixation point on the instrument panel, the path of the line of sight between different instruments, and the distribution of fixation areas in each instrument, comprehensively reflecting the pilot's visual attention state. The second is the flight operation response data points, which consist of two parts. One part is the pilot's operation command data, including the push and pull angles of the control stick, the adjustment range of the throttle, the trigger type and trigger sequence of function buttons, and other operation details. The other part is the flight instrument parameter change data triggered by the operation command, including the specific fluctuation values ​​and trends of parameters such as flight altitude, speed, heading, bank angle, and angle of attack before and after the operation. This ensures that both types of data accurately correspond to the same flight mission phase and resonate with the evaluation indicators (visual attention, operational behavior) of that phase.

[0083] Step 5.2: Based on the extracted visual attention feature data points and flight operation response data points, calculate and generate correlation analysis index data points. These correlation analysis index data points include at least the attention transition time interval data points from the appearance of key flight information to the pilot's first gaze at the area, and the operation response delay data points from the pilot's stable gaze at the key information to the execution of the corresponding operation command. Specifically, this includes: firstly, extracting all key flight information within the current flight mission phase from the background records of the simulated flight mission experimental program. This information includes abnormal warning signals, flight parameter change prompts, special situation trigger commands, illusion inducement markers, etc., and each key piece of information is associated with a corresponding first experimental timestamp, i.e., the time when the information appeared; based on the extracted visual attention feature data points, by analyzing the trajectory of the gaze point coordinate change, locating the pilot's first shift of gaze to the key flight information. For the instrument panel area, the time difference is calculated by subtracting the first experimental timestamp of the critical information occurrence time from the first experimental timestamp of that time, generating attention transition time interval data points. Then, a stable gaze judgment criterion is set: when the gaze point remains continuously in the instrument panel area corresponding to the critical information, and the coordinate fluctuation range does not exceed a preset threshold within 300 milliseconds, it is determined to be a stable gaze state, and the start time of this state is recorded, corresponding to the first experimental timestamp. Next, the time when the pilot executes the corresponding operation command for that critical information is extracted from the flight operation response data points, i.e., the first experimental timestamp of the operation command data point. The time difference is calculated by subtracting the stable gaze start time from the operation command execution time, generating operation response delay data points. Both types of correlation analysis indicator data points must be bound to the corresponding flight mission phase name, critical information type, and first experimental timestamp to ensure data traceability.

[0084] Step 5.3 involves comparing the calculated correlation analysis index data points with the self-evaluation data points of the same flight mission phase extracted from the multi-dimensional behavior mapping data set. This generates data points verifying the consistency between subjective and objective assessments. Specifically, this includes: first, extracting the corresponding self-evaluation data points from the multi-dimensional behavior mapping data set based on the name of the current flight mission phase. These data points are consistent with the flight expert's peer-evaluation indicators, covering core dimensions such as the timeliness of attention to instruments at the visual level and the speed of action response at the operational level. This includes the pilot's self-evaluation results and specific scores for these dimensions. A pre-determined reasonable time range standard is established by no fewer than three flight experts, considering different flight mission phases and key flight information types. Experts base their decisions on their extensive flight experience and refer to the typical behavioral characteristics of pilots within a reasonable trust range, namely, timely attention and rapid response when key information appears. This clarifies the attention shift and operational response in various scenarios. The reasonable range standard serves as an important reference for evaluating trust behavior. The generated attention switching time interval data points are compared with the self-assessment's indicator of the timeliness of instrument attention to determine if the actual time falls within the preset reasonable range, matching the degree of consistency with the self-assessment results. Operational response delay data points are compared with the self-assessment's action response speed indicator, analyzing the consistency between the actual situation and the self-assessment description based on the reasonable delay range. During comparison, if the actual data is within the reasonable range and consistent with the self-assessment results, it is marked as completely consistent; if it is within the reasonable range but has a slight deviation, and the deviation does not exceed a preset threshold, it is marked as basically consistent; if it exceeds the reasonable range or has a large deviation, it is marked as inconsistent and the specific deviation value is recorded, while also associating whether there is a trust mismatch in this stage, i.e., potential characteristics of excessive or insufficient trust. Finally, data points for verifying the consistency of subjective and objective evaluations are generated, including the flight mission stage name, correlation analysis indicator type, actual data value, self-assessment results, consistency level, and any deviation values.

[0085] In a preferred embodiment of the present invention, step 6 above may include:

[0086] Step 6.1: Based on the data points from the consistency verification results of subjective and objective assessments, extract trust behavior questionnaire data points and self-assessment data points from the subjective assessment dataset, and simultaneously extract expert peer assessment data points from the expert evaluation dataset to form the original assessment data point set. Specifically, this includes: first, retrieving the generated data points from the consistency verification results of subjective and objective assessments, whereby these data points clearly record the degree of consistency and deviation between the correlation analysis indicators and self-assessment results for the same pilot in each flight mission phase; and then, using the pilot's unique identifier and experimental batch as the core correlation basis, accurately extracting the corresponding pilot's trust behavior questionnaire data points from the subjective assessment dataset, i.e., the [data point number]. The system consists of two types of subjective assessment data points: one for subjective evaluation and the other for behavioral self-assessment. Both types of data include scores for various assessment indicators across visual, operational, and risk levels. Simultaneously, expert peer assessment data points for the pilot are extracted from the expert evaluation dataset. These expert peer assessment data points are provided by at least three flight experts based on eye-tracking and operational behaviors, and also cover the core indicator scores consistent with the subjective assessment. The three types of extracted data points are then bound to the pilot's unique identifier and experimental batch to ensure that all data points correspond to the same assessment object and experimental scenario. This results in a final set of raw assessment data points containing questionnaire, self-assessment, and peer assessment data.

[0087] Step 6.2 involves preprocessing the original assessment data set to generate a preprocessed assessment data set. This preprocessed set comprises the trust behavior questionnaire data points, self-assessment data points, and expert peer assessment data points for each individual pilot, forming a multidimensional data set to construct the data point set within the assessment space. The convex hull of the data point set is calculated to identify and differentiate the data distribution characteristics: data points located outside the convex hull and significantly deviating from the group distribution in multidimensional features are identified as potential marginal or outlier data points; while data points located inside the convex hull and conforming to the overall distribution pattern are identified as valid core data points; marginal or outlier data points are further classified as... The data points are filtered or adjusted to form a pre-processed evaluation data point set. Specifically, this includes: taking each individual pilot as an independent unit, initiating a data integration process: multi-dimensionally integrating the trust behavior questionnaire data points, self-evaluation data points, and expert peer evaluation data points from their original evaluation data point set to construct a multi-dimensional data point for each pilot; the dimensions of this multi-dimensional data point strictly correspond to a unified evaluation indicator system, namely 3 indicators at the visual level, 3 indicators at the operational level, and 1 indicator at the risk level, for a total of 7 dimensions, each corresponding to a specific original score value; based on all the valid multi-dimensional data points of the pilots that pass the screening, a complete evaluation system is constructed within the pre-set 7-dimensional evaluation space. The data point set is used to evaluate each coordinate axis of the evaluation space, which corresponds to an evaluation index. The coordinate values ​​of the data points are the raw scores of the corresponding indexes. By traversing all data points, the convex hull boundary of the data point set is calculated and drawn. The convex hull boundary is formed by connecting the extreme value data points of each dimension's score in the data point set. It can define the distribution range of most data points. Data points located inside the convex hull, with a balanced distribution of scores and conforming to the overall evaluation pattern, are identified as effective core data points. Data points located outside the convex hull, with one or more scores significantly deviating from the overall distribution, or with large differences from other data points, are identified as marginal or anomalous data points. For marginal or anomalous data... The data points are reviewed in conjunction with the consistency verification results of subjective and objective assessments and the experimental process logs. If the deviation originates from non-trust factors such as pilot operation errors, data entry errors, or temporary equipment malfunctions, it is adjusted according to the reasonable distribution range of similar data or directly removed. If the deviation reflects the pilot's true trust state (such as extreme over-trust or extreme under-trust), the data point is retained and marked with abnormal trust characteristics. After reviewing and processing all data points, all valid core data points and confirmed retained marginal or abnormal data points are integrated to form a preprocessed evaluation data point set, ensuring that the data in the set is both authentic and valid and covers various trust states.

[0088] Step 6.3 involves standardizing the original evaluation data points from various evaluation sources in the preprocessed evaluation data point set to generate corresponding standardized subscale score data points. Specifically, this includes: first, clarifying the subscale composition of the three evaluation sources in the preprocessed evaluation data point set: the trust behavior questionnaire data points constitute the questionnaire subscale, containing 7 core evaluation indicators; the self-evaluation data points constitute the self-evaluation subscale, containing 7 core evaluation indicators; and the expert peer evaluation data points constitute the peer evaluation subscale, containing 7 core evaluation indicators. For each subscale, the scores of all valid pilots for the corresponding indicators are summarized, and the mean and standard deviation of the scores for each indicator within each subscale are calculated. These statistical values ​​are used as the benchmark parameters for the standardization transformation of that subscale, ensuring that the benchmark parameters reflect the overall score distribution characteristics. Using a unified standardized transformation method, the raw scores of each pilot for each indicator in various subscales are uniformly transformed according to the degree of deviation and standard deviation from the corresponding indicator mean. This ensures that the transformed scores are distributed within a preset standard range, eliminating dimensional differences between different indicators and subscales. After completing the standardized transformation of individual indicators, the standardized scores of the seven indicators within the same subscale are summed to obtain the comprehensive standardized score for each pilot corresponding to the three types of subscales. This generates standardized subscale score data points for the questionnaire, self-rated standardized subscale score data points, and peer-rated standardized subscale score data points. Each standardized subscale score data point is associated with the pilot's unique identifier, experimental batch, subscale name, standardized scores of each indicator, and comprehensive subscale score.

[0089] Step 6.4: Based on the preset weighting coefficients, perform weighted summation on all standardized subscale scores to generate a comprehensive trust behavior index. This involves: First, pre-setting weighting coefficients: A weighting assessment group composed of no fewer than three flight experts assigns weighted scores to the three subscales based on the reliability, importance, and evaluation criteria of the assessment data. The expert peer-rating subscale, considered the gold standard for trust behavior evaluation, has the highest reliability and a relatively high weighting. The trust behavior questionnaire subscale reflects the pilot's pre-set trust tendency, while the self-rating subscale reflects the pilot's actual post-event feelings. The weights of both are reasonably allocated based on their contribution to the comprehensive trust assessment. The assessment group first scores independently, then coordinates... Commercial calibration eliminates extreme biases and ultimately determines the fixed weight coefficients for the three subscales: questionnaire, self-report, and peer-report, with the sum of the three weight coefficients being 1. Based on the pilot's unique identifier, the corresponding standardized subscale score data points are retrieved. The comprehensive standardized score of each subscale is multiplied by the corresponding preset weight coefficient to obtain the weighted score of the three subscales. Then, the weighted scores of the three subscales are summed to obtain the pilot's final comprehensive score, which generates the comprehensive trust behavior index data point. This data point must completely record the pilot's unique identifier, experimental batch, original scores of the three subscales, standardized scores of the three subscales, each weight coefficient, weighted score, and final comprehensive index value.

[0090] In this embodiment of the invention, by using the consistency verification results of subjective and objective assessments as the screening criterion, three types of core data—trust behavior questionnaires, self-assessments, and expert evaluations—are accurately extracted to form an original assessment set. Then, convex hull analysis is used to identify effective core data and eliminate or adjust marginal outliers, ensuring the authenticity and reliability of the assessment data. Subsequently, the preprocessed data undergoes standardization transformation to eliminate dimensional differences between different assessment tools, ensuring horizontal comparability of various scores. Finally, weighted summation is performed based on preset weight coefficients to generate a comprehensive trust behavior index data point. This process not only reasonably highlights the importance of different assessment sources (especially expert evaluations as the gold standard) but also achieves multi-source fusion and complementarity of subjective assessment data. This process not only overcomes the limitations of traditional single subjective assessments but also reduces data noise and errors through a progressive data processing flow. The resulting comprehensive trust behavior index can comprehensively and accurately represent the overall human-machine trust level of pilots.

[0091] In a preferred embodiment of the present invention, step 7 above may include:

[0092] Step 7.1 involves receiving the comprehensive trust behavior index data points, the standardized subscale score data points (which serve as standardized trust behavior indicator data points), and the key scenario objective flight parameter data points extracted from the multi-dimensional behavior mapping data set. Specifically, this includes: first retrieving the generated comprehensive trust behavior index data points, which contain the pilot's unique identifier, experimental batch, scores for various subscales, and the final comprehensive index value; simultaneously extracting the generated standardized subscale score data points, including questionnaire standardized subscale scores, self-rating standardized subscale scores, and peer-rating standardized subscale scores, and then... As standardized trust behavior indicator data points, objective flight parameter data points corresponding to key scenarios are selected from a multi-dimensional behavior mapping data set. Key scenarios specifically refer to stages where trust status is prone to mismatch, such as illusion handling scenarios and emergency handling scenarios. Objective flight parameters include real-time fluctuation data such as flight altitude, speed, heading, bank angle, angle of attack, and Mach number under the scenario, as well as the parameter change range caused by operation commands. The three types of data are bound one by one according to the pilot's unique identifier, experimental batch, and key scenario name to ensure that all data accurately correspond to the same evaluation object and experimental scenario, forming a complete set of data to be judged.

[0093] Step 7.2 involves calling preset trust state discrimination threshold data points and preset behavioral pattern rule base data points. Specifically, this includes: calling trust state discrimination threshold data points pre-constructed based on historical experimental data and domain expert experience. These thresholds are pre-determined by no fewer than three flight experts in conjunction with a large amount of measured data from experienced pilots and reasonable trust interval characteristics. They clearly define threshold ranges for three levels: insufficient trust, reasonable trust, and excessive trust. Each range corresponds to a specific numerical range of the comprehensive trust behavior index and is adapted to the complexity of the flight mission. Simultaneously, it calls behavioral pattern rule base data points pre-established based on expert knowledge and historical analysis. This rule base is constructed based on trust mismatch characteristics and includes judgment rules for three types of behavioral patterns: excessive trust, insufficient trust, and self-perception bias. The excessive trust pattern rules cover features such as prolonged staring at a single instrument, delayed response to abnormal instrument information, and inconsistency between operation and instrument display. The insufficient trust pattern rules include features such as frequent switching of gaze points, excessive adjustment of operation, and ignoring effective system prompts. The self-perception bias pattern rules focus on significant deviations between self-assessment results and other-assessment results, and objective data, where the deviation exceeds a reasonable range.

[0094] Step 7.3: Based on the comprehensive trust behavior index data points, and comparing them with the preset trust status discrimination threshold data points, a preliminary trust level classification is performed to generate initial trust classification data points. Specifically, this includes: comparing each pilot's comprehensive trust behavior index data point one by one with the preset trust status discrimination threshold data points as the basis for judgment; if the comprehensive trust behavior index is lower than the lower limit of the discrimination threshold, it indicates that the pilot's overall trust level has not reached a reasonable standard, and it is initially judged as insufficient trust, generating a corresponding initial trust classification data point; if the comprehensive trust behavior index is higher than the upper limit of the discrimination threshold, it indicates that the pilot's overall trust level exceeds the actual capability matching range of the airborne automation system, and it is initially judged as excessive trust, generating a corresponding initial trust classification data point; if the comprehensive trust behavior index is in the middle range of the discrimination threshold, it indicates that the overall trust level is basically reasonable, and it is initially judged as reasonable trust, generating a corresponding initial trust classification data point; each initial trust classification data point is associated with the pilot's unique identifier, experimental batch, specific value of the comprehensive trust behavior index, and preliminary classification result to ensure that the classification is traceable.

[0095] Step 7.4: Based on the standardized subscale score data points and key scenario objective flight parameter data points, combined with the initial trust classification data points, and according to the preset behavioral pattern rule base data points, conduct behavioral pattern matching and cross-validation analysis to generate refined trust status discrimination data points. Specifically, this includes: for each pilot's initial trust classification data points, combining their standardized subscale score data points and key scenario objective flight parameter data points, conducting behavioral pattern matching and cross-validation; if the initial classification is over-trust, compare with the over-trust pattern rules in the rule base to check whether the key scenario objective parameters have characteristics such as response delays exceeding the reasonable range and inconsistencies between operation and instrument information, while simultaneously analyzing the standardized subscale... If the self-assessment score is significantly higher than the peer assessment score, and the initial classification is insufficient trust, check whether the objective parameters have excessive fluctuations due to frequent operations, or whether the frequency of fixation switching exceeds the reasonable standard. At the same time, compare whether the self-assessment score is significantly lower than the peer assessment score. If the initial classification is reasonable trust, focus on checking the consistency between the self-assessment, peer assessment, and objective data. If the deviation of the three exceeds the judgment standard for self-perception bias in the rule base, it is preliminarily determined that there is self-perception bias. During the verification process, each feature matching result needs to be associated with specific data support to finally generate detailed judgment data points for trust status, clarify whether the initial classification is accurate, whether there is a detailed classification (such as excessive trust accompanied by self-perception bias), and the specific judgment basis.

[0096] Step 7.5 integrates the initial trust classification data points with the refined trust status judgment data points to generate and output behavioral characteristic assessment result data points representing excessive trust, insufficient trust, or self-perception bias. Specifically, this includes: structurally integrating the initial trust classification data points with the refined trust status judgment data points; if the refined judgment does not adjust the initial classification, then the behavioral characteristic basis in the refined judgment is supplemented based on the initial classification; if the refined judgment corrects the initial classification or adds a self-perception bias label, then the comparison process of the initial classification is linked to the refined judgment result; finally, behavioral characteristic assessment result data points are generated, which clearly represent the pilot's trust status as excessive trust, insufficient trust, or self-perception bias, and record the corresponding behavioral characteristics in detail: characteristics of excessive trust include excessive gaze duration on a single instrument in critical scenarios, delayed operation response, and inconsistency between operation and instrument information; characteristics of insufficient trust include frequent gaze switching in critical scenarios, excessive operation leading to parameter fluctuations, and ignoring effective system prompts; characteristics of self-perception bias include self-assessment results and other-assessment results, and actual behavioral deviations reflected by objective flight parameters exceeding the reasonable range; outputting the assessment result data points in a standardized format.

[0097] In the specific implementation of the embodiments of the present invention, the following processes and methods are also included:

[0098] This embodiment first uses a weighted method to construct a comprehensive trust index as its core, combining subjective assessment data and objective flight data to build a multi-source synthesized comprehensive trust behavior index. Through modeling, it achieves future trust behavior prediction, providing a precise and predictable technical solution for assessing pilot-machine trust status. The experimental system consists of hardware, software, and assessment tools. The hardware includes a flight simulator, a wearable eye tracker, and an operation console. The flight simulator simulates flight operations and records parameters; the eye tracker collects eye movement data such as pupil changes and fixation points; and the console is responsible for task settings and data export. The software is a program for simulating dynamic flight operation tasks, supporting takeoff, at least two random illusion scenarios, at least two random special situation scenarios, and return landing tasks. It can automatically record objective parameters such as flight altitude, speed, and heading. The assessment tools include a human-machine trust behavior assessment questionnaire, a flight expert evaluation form, and a pilot self-assessment form. The questionnaire covers both behavioral intentions and potential behaviors, with both external and self-assessments using visual, operational, and risk-related indicators as core assessment indicators. The participants included experienced pilots who were being evaluated, testers responsible for equipment calibration and experimental assistance, flight commanders responsible for simulator operation and mission switching, and no fewer than three flight experts who served as the gold standard for evaluation.

[0099] The experimental procedure was carried out in an orderly manner according to the following steps: Before the experiment, the pilot had to complete the human-machine trust behavior assessment questionnaire. The experimenter explained in detail the functions of each instrument on the flight simulator and the experimental task details, including the operating procedures, methods and precautions for takeoff, illusion handling, emergency handling, return landing, etc., to ensure that the pilot understood the experimental requirements. Then, the experimenter had the pilot wear a wearable eye tracker to perform precise calibration of the fixation point. By having the pilot fixate on key instruments in the simulator in sequence, such as the altimeter, speedometer and heading instrument, the accuracy of the equipment data acquisition was calibrated to ensure data accuracy. After the calibration was completed, the experimenter left the simulator and closed the cabin door. The flight commander started the simulator reset procedure to put all hardware and software into normal standby state. During the formal experimental phase, pilots completed various tasks in a random order. These included illusion scenarios such as instrument illusions and flight attitude illusions, and emergency scenarios such as engine failures and avionics system anomalies. Throughout the experiment, the software automatically recorded flight parameters in real time, and the eye tracker continuously collected eye movement data, which was then wirelessly transmitted to the control console display. Flight experts observed changes in eye movement and instrument data in real time on the display and provided peer evaluations of the pilots' performance in each scenario. If pilots encountered problems with equipment operation or mission comprehension, they could communicate with the flight commander via the built-in communication device to ensure the smooth progress of the experiment. After the experiment, pilots completed a self-evaluation form based on their operational experience and feelings. The experimenters collected questionnaire scores, peer evaluation scores, self-evaluation scores, and all recorded objective flight parameters and eye movement data to complete the initial data collection.

[0100] The data preprocessing stage begins with core data extraction. Three types of subjective data for each pilot are extracted from the collected data: trust behavior questionnaire scores, flight expert peer assessment scores, and pilot self-assessment scores. All three types of data include seven core indicators: visual aspect (timeliness of attention, reasonableness, and gaze duration); operational aspect (operational consistency, response speed, and compliance); and risk aspect (operational risk level). Standardization is then performed. Scores for each indicator for all pilots are aggregated for each of the three types of data, and the mean and standard deviation of each indicator are calculated as standardization benchmark parameters. The Z-score standardization method is used to adjust the dimensions of each pilot's raw scores for each indicator in each subscale based on their deviation from the mean and the size of the standard deviation. Specifically, the raw score is subtracted from the mean and divided by the standard deviation, converting it into standardized scores distributed in the interval {-3, 3}, namely Z1, Z2, and Z3, corresponding to the questionnaire, peer assessment, and self-assessment, respectively. This completely eliminates the dimensional differences between different assessment tools, ensuring the horizontal comparability of the three types of scores and laying the data foundation for subsequent weighted calculations and modeling.

[0101] The weighting coefficients are jointly determined by a weighting assessment panel composed of no fewer than three flight experts participating in the peer review. Before the assessment, the experts need to fully discuss the reliability and importance of the three subscales in conjunction with the evaluation logic of this invention and actual flight scenarios: the flight expert peer review score is based on real-time observed eye-tracking data (such as the gaze point transition trajectory in illusion scenarios) and operational behavior (such as operational response during emergency handling), directly reflecting real trust behavior. As the gold standard for evaluation, it has the highest reliability, so its weighting is set at 0.45; the trust behavior questionnaire score reflects the pilot's pre-experimental pre-existing trust tendency, influencing their initial operational decisions, and its weighting is set at 0.3; the pilot's self-evaluation score reflects the actual feeling after the experiment, which can supplement the subjective cognitive dimension, and its weighting is set at 0.25, ultimately ensuring that the sum of w1, w2, and w3 is 1. The calculation of the comprehensive trust behavior index is based on standardized scores and weighting coefficients, using a weighted summation formula, i.e., comprehensive trust behavior index = Tcomprehensive = w1Z1 + w2Z2 + w3Z3. This formula intuitively quantifies the overall human-machine trust level of each pilot.

[0102] To ensure the reliability of the index, a targeted effectiveness test is required: Two core scenarios, illusion handling and emergency handling, are selected. Objective behavioral indicators are extracted for each scenario: gaze transition time and stable gaze duration in illusion scenarios; operational response delay and flight parameter fluctuation amplitude in emergency scenarios. A correlation analysis is then performed between the comprehensive trust index and these objective indicators. If the index falls within the expert-preset reasonable range (4 to 6 points out of 10), and the pilot's operational response delay in emergency handling is less than 1.5 seconds and the flight parameter fluctuation amplitude is less than 5%, and the correlation analysis shows a significant correlation between the index and the objective indicators (correlation coefficient r ≥ 0.6), then the index effectively represents real trust behavior. Simultaneously, the highly correlated objective indicators selected in the effectiveness test are identified as scenario-based feature variables for the prediction model, achieving seamless integration of weighted construction and modeling.

[0103] The construction and training of the predictive model are based on a reliable comprehensive trust index and highly relevant scenario-based features, proceeding systematically in stages. First, feature preprocessing is performed: the comprehensive trust behavior index and selected scenario-based feature variables are integrated into the model input feature set. Continuous variables in the feature set, such as response delay and fluctuation amplitude, are normalized and mapped to the {0, 1} interval. The 3σ principle is used to identify and remove extreme outliers caused by equipment failures, and then the outliers are replaced with the median of the feature to ensure feature quality. Next, data is partitioned: complete data from 80 pilots at different flight experience levels (junior, intermediate, and advanced) are collected and divided into a 7:2:1 ratio: 56 pilots for training, 16 for validation, and 8 for testing. Each dataset includes input features and corresponding trust status labels, jointly labeled by three flight experts based on the comprehensive trust index and objective behavioral performance, categorized into three types: excessive trust, reasonable trust, and insufficient trust.

[0104] For the model, a random forest classification model was chosen, and the core parameters were initialized based on domain experience: 120 decision trees, a tree depth limit of 10 layers, and a minimum number of sample splits of 5. This ensured both the capture of complex feature relationships and the avoidance of overfitting. An iterative optimization strategy was employed during training: After the first round of training, validation set evaluation showed that the over-trust recognition accuracy was only 65%, lower than the preset standard of 75%. Analysis revealed that the feature weight for operational response delay in engine failure scenarios was insufficient. In the second round of training, the weight of this feature was increased by 20%, and the number of decision trees was increased to 150. After retraining, the over-trust accuracy improved to 78%, but the recall for under-trust was only 70%. The tree depth was further optimized to 12 layers to supplement the gaze switching frequency feature in instrument illusion scenarios. After 5 rounds of iteration, the model achieved an overall accuracy of 88%, precision of 86%, and recall of 87% on the validation set, meeting the performance requirements. Finally, a generalization ability test was conducted: test set data showed that the model had a prediction accuracy of 90% for intermediate pilots and no less than 85% for primary and advanced pilots. Moreover, the prediction effect was stable in engine failure and instrument illusion scenarios, proving that the model has good versatility and scenario adaptability.

[0105] After successful model validation, two core results are output: First, each pilot's comprehensive trust behavior index and corresponding trust status determination, which, combined with preset thresholds, clarifies whether the pilot is over-trusting, reasonably trusting, or under-trusting. For example, a pilot with a comprehensive trust index of 8.2 out of 10, and whose operational response delay exceeds 1.8 seconds during emergency handling, is judged to have over-trusted behavior. Second, the trained trust prediction model and future trust behavior risk warnings. For example, by inputting a pilot's historical comprehensive trust index and gaze transition characteristics from past illusion handling scenarios into the model, the probability of over-trust leading to operational delays in similar engine failure emergencies in the future can be predicted. These results can be directly applied to targeted flight training interventions, such as strengthening system anomaly identification and rational operation training in emergency scenarios for over-trusted pilots, and conducting automated system reliability verification and operational confidence building training for under-trusted pilots. Simultaneously, it provides data support for the intelligent optimization design of aircraft avionics systems. For example, for pilots with low comprehensive trust indices, the visual recognition of emergency warning instruments can be optimized, helping to build a human-machine interaction system that better matches the pilots' trust behavior characteristics and ensure flight safety.

[0106] like Figure 2 As shown, embodiments of the present invention also provide a non-invasive experimental evaluation system for pilot-machine trust behavior, comprising:

[0107] The data point acquisition module is used to collect multiple raw data points, including pilot's eye movement trajectory data points, operation command data points, and flight instrument parameter data points, to form an objective behavior data set; based on the objective behavior data set, it receives multiple subjective evaluation data points from the pilot to form a subjective evaluation data set; and based on the subjective evaluation data set and the objective behavior data set, it receives multiple expert evaluation data points from flight experts to form an expert evaluation data set.

[0108] The synchronization matching module is used to perform synchronous matching processing on each data point in the objective behavior data set, subjective evaluation data set, and expert evaluation data set according to the flight mission time nodes, and generate a multi-dimensional behavior mapping data set with time correspondence.

[0109] The correlation analysis module is used to perform correlation analysis on visual attention feature data points and flight operation response data points through a multi-dimensional behavior mapping dataset, forming data points for verifying the consistency between subjective and objective evaluations.

[0110] The conversion and calculation module is used to standardize and weight the trust behavior questionnaire data points, self-assessment data points, and expert assessment data points based on the consistency verification results of subjective and objective evaluations, and generate comprehensive trust behavior index data points.

[0111] The identification and output module is used to compare and analyze standardized trust behavior index data points with objective flight parameter data points based on comprehensive trust behavior index data points, identify and output behavioral characteristic assessment result data points that represent excessive trust, insufficient trust or self-cognition bias, and complete a non-intrusive comprehensive assessment of pilot human-machine trust behavior.

[0112] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A non-invasive experimental evaluation method for pilot-machine trust behavior, characterized in that, The method includes: Collect multiple raw data points, including pilot eye movement trajectory data points, operation command data points, and flight instrument parameter data points, to form an objective behavioral data set; Based on the objective behavioral data set, multiple subjective evaluation data points from the pilot are received to form a subjective evaluation data set; Based on the subjective evaluation data set and the objective behavior data set, multiple expert evaluation data points from flight experts are received to form an expert evaluation data set; Each data point in the objective behavior data set, subjective evaluation data set, and expert evaluation data set is synchronously matched according to the flight mission time nodes to generate a multi-dimensional behavior mapping data set with time correspondence. By using a multi-dimensional behavior mapping dataset, the visual attention feature data points and flight operation response data points are correlated and processed to form data points for the consistency verification of subjective and objective evaluation results. Based on the data points of the consistency verification results of subjective and objective assessments, the data points of the trust behavior questionnaire, the self-assessment data points of the behavior, and the expert assessment data points are standardized, transformed, and weighted to generate a comprehensive trust behavior index data point. Based on the comprehensive trust behavior index data points, the standardized trust behavior indicator data points are compared and analyzed with the objective flight parameter data points to identify and output the behavioral characteristic assessment results data points that represent excessive trust, insufficient trust, or self-cognition bias, thus completing a non-intrusive comprehensive assessment of the pilot's human-machine trust behavior.

2. The non-invasive experimental evaluation method for pilot-machine trust behavior according to claim 1, characterized in that, Based on the objective behavioral data set, multiple subjective evaluation data points from pilots are received to form a subjective evaluation data set, including: Based on the objective behavioral data set, the trust behavior questionnaire data points, which are filled out by the pilots before entering the flight simulator and are used to characterize the pilots' preset level of trust in the airborne automation capabilities and their behavioral tendencies, are received and recorded. At the same time, the trust behavior questionnaire data points are used as the first type of subjective evaluation data points. The system receives and records the self-evaluation data points of the pilots' behavior after completing all simulated flight missions. These self-evaluation data points include the pilots' self-evaluation data on their visual attention allocation, actual operational behavior, and operational risk perception in various mission scenarios. The self-evaluation data points are also used as the second type of subjective evaluation data points. The first type of subjective evaluation data points and the second type of subjective evaluation data points are combined to form a subjective evaluation data set.

3. The non-invasive experimental evaluation method for pilot-machine trust behavior according to claim 2, characterized in that, Based on the subjective assessment data set and the objective behavioral data set, multiple expert evaluation data points from flight experts are received to form an expert evaluation data set, including: It receives objective behavioral data sets and provides flight experts with a synchronized observation interface for eye-tracking data points and flight instrument parameter data points. Based on the synchronous observation interface, the system receives real-time expert evaluation data points generated by flight experts during the simulated flight mission. These data points include flight experts' evaluation data on the pilot's visual attention characteristics, operational response characteristics, and operational risk level in various mission scenarios. All expert real-time evaluation data points are aggregated to form an expert evaluation data set.

4. The non-invasive experimental evaluation method for pilot-machine trust behavior according to claim 3, characterized in that, For each data point in the objective behavior data set, subjective evaluation data set, and expert evaluation data set, synchronous matching processing is performed according to the flight mission time nodes to generate a multi-dimensional behavior mapping data set with time correspondence, including: Acquire an objective behavior data set and mark the corresponding first experimental timestamp for each eye-tracking data point, operation command data point, and flight instrument parameter data point in the objective behavior data set. Receive the expert evaluation data set and mark the second experimental timestamp corresponding to the observation and evaluation time for each real-time self-evaluation data point of the expert in the expert evaluation data set; Obtain behavioral self-evaluation data points from the subjective evaluation dataset, and associate all behavioral self-evaluation data points with corresponding scenario time periods based on the specific simulated flight mission scenarios evaluated by the behavioral self-evaluation data points. Based on the first experimental timestamp, the second experimental timestamp, and the scene time period identifier, data points from the objective behavior data set, the subjective evaluation data set, and the expert evaluation data set, which correspond to the same flight mission phase, are aligned and integrated to generate a multi-dimensional behavior mapping data set.

5. The non-invasive experimental evaluation method for pilot-machine trust behavior according to claim 4, characterized in that, By using a multi-dimensional behavior mapping dataset, correlation analysis is performed on visual attention feature data points and flight operation response data points to form data points for verifying the consistency of subjective and objective evaluations, including: From the multi-dimensional behavior mapping dataset, visual attention feature data points and flight operation response data points corresponding to the same flight mission phase are extracted; the visual attention feature data points are derived from eye track data points, and the flight operation response data points are derived from operation command data points and the flight instrument parameter change data points they cause. Based on the extracted visual attention feature data points and flight operation response data points, correlation analysis index data points are calculated and generated; wherein the correlation analysis index data points include at least the attention transition time interval data points from the appearance of key flight information to the pilot's first gaze at the area, and the operation response delay data points from the pilot's stable gaze at the key information to the execution of the corresponding operation command. The calculated correlation analysis index data points are compared and analyzed with the self-evaluation data points of behavior extracted from the multi-dimensional behavior mapping data set that correspond to the same flight mission phase, to generate data points for the consistency verification of subjective and objective evaluation.

6. The non-invasive experimental evaluation method for pilot-machine trust behavior according to claim 5, characterized in that, Based on the data points from the consistency verification results of subjective and objective assessments, the trust behavior questionnaire data points, self-assessment data points, and expert peer assessment data points are standardized, transformed, and weighted to generate a comprehensive trust behavior index data point, including: Based on the data points of the consistency verification results of subjective and objective assessments, trust behavior questionnaire data points and self-assessment data points are extracted from the subjective assessment data set, and expert peer assessment data points are extracted from the expert evaluation data set to form the original assessment data point set. The original set of assessment data points is preprocessed to generate a preprocessed set of assessment data points. This preprocessed set consists of trust behavior questionnaire data points, self-assessment data points, and expert peer assessment data points for each individual pilot, forming a multi-dimensional data point set within the assessment space. The convex hull of the data point set is calculated to identify data points inside the convex hull as valid core data points, while data points outside the convex hull are identified as marginal or outlier data points. Marginal or outlier data points are then filtered or adjusted to form the preprocessed set of assessment data points. The original evaluation value data points from various evaluation sources in the preprocessed evaluation data point set are standardized and transformed to generate corresponding standardized subscale score data points. Based on preset weighting coefficients, all standardized subscale score data points are weighted and summed to generate a comprehensive trust behavior index data point.

7. The non-invasive experimental evaluation method for pilot-machine trust behavior according to claim 6, characterized in that, Based on the comprehensive trust behavior index data points, standardized trust behavior indicator data points are compared and analyzed with objective flight parameter data points to identify and output behavioral characteristic assessment results data points representing excessive trust, insufficient trust, or self-perception bias. This completes a non-intrusive comprehensive assessment of pilot-machine trust behavior, including: It receives comprehensive trust behavior index data points, standardized subscale score data points as standardized trust behavior indicator data points, and key scenario objective flight parameter data points extracted from multi-dimensional behavior mapping data sets; Call the preset trust status judgment threshold data points and the preset behavior pattern rule base data points; Based on the comprehensive trust behavior index data points, and compared with the preset trust status judgment threshold data points, a preliminary trust level classification is performed to generate initial trust classification data points; Based on standardized subscale score data points and key scenario objective flight parameter data points, combined with initial trust classification data points, and according to the preset behavior pattern rule base data points, behavior pattern matching and cross-validation analysis are performed to generate detailed trust status discrimination data points. By integrating the initial trust classification data points with the detailed judgment data points of trust status, we can generate and output the behavioral characteristic assessment results data points that represent excessive trust, insufficient trust, or self-perception bias.

8. A non-invasive experimental evaluation system for pilot-machine trust behavior, the system implementing the method as described in any one of claims 1 to 7, characterized in that, include: The data point acquisition module is used to collect multiple raw data points, including pilot eye movement trajectory data points, operation command data points, and flight instrument parameter data points, to form an objective behavioral data set; Based on the objective behavioral data set, multiple subjective evaluation data points from pilots are received to form a subjective evaluation data set; based on the subjective evaluation data set and the objective behavioral data set, multiple expert evaluation data points from flight experts are received to form an expert evaluation data set. The synchronization matching module is used to perform synchronous matching processing on each data point in the objective behavior data set, subjective evaluation data set, and expert evaluation data set according to the flight mission time nodes, and generate a multi-dimensional behavior mapping data set with time correspondence. The correlation analysis module is used to perform correlation analysis on visual attention feature data points and flight operation response data points through a multi-dimensional behavior mapping dataset, forming data points for verifying the consistency between subjective and objective evaluations. The conversion and calculation module is used to standardize and weight the trust behavior questionnaire data points, self-assessment data points, and expert assessment data points based on the consistency verification results of subjective and objective evaluations, and generate comprehensive trust behavior index data points. The identification and output module is used to compare and analyze standardized trust behavior index data points with objective flight parameter data points based on comprehensive trust behavior index data points, identify and output behavioral characteristic assessment result data points that represent excessive trust, insufficient trust or self-cognition bias, and complete a non-intrusive comprehensive assessment of pilot human-machine trust behavior.