A multi-dimensional evaluation method for civil aviation flight trainees
Through the multi-dimensional evaluation method of calculating dynamic weights with Chi-square test and Phi correlation coefficient, the problem of insufficient distinction in the existing flight training evaluation system is solved, and the refined evaluation and dynamic management of trainee abilities is realized, which improves the structural and interpretability of the scoring results.
Patent Information
- Application Number
- CN202510754865.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-06-06
AI Technical Summary
The existing flight training evaluation system cannot accurately reflect the distinction between key checkpoints, the score results are lacking in structure, it is difficult to identify the shortcomings of trainees' abilities, and the classification method of the evaluation results is single, so it cannot adapt to the dynamic stratified management of trainees at different levels.
A multi-dimensional evaluation method based on chi-square test and Phi correlation coefficient is adopted. By calculating the relationship between the checkpoint and the final evaluation category of the trainees, dynamic weights are obtained, and a scoring mechanism based on checkpoint failure marks and weighted deductions is constructed to achieve a refined quantitative evaluation of the trainees' training performance.
It realizes dynamic response ability to students' operational errors, improves the structural and interpretability of the scoring results, supports personalized training feedback and differentiated teaching strategies, and improves the management efficiency and intelligence level of the training system.
Smart Images

Figure CN120296360B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of flight training assessment and skill classification, and in particular to a multi-dimensional assessment method for civil aviation flight trainees. Background Art
[0002] Current skill training and competency assessment systems generally use linear statistical methods based on scoring tables to assess student performance. These systems typically use fixed weights to accumulate scores or use project completion rates as the evaluation basis. This results in different checkpoints contributing evenly to the final score, making it difficult to reflect the actual contribution of each checkpoint to differentiate competency.
[0003] During the grading process, failed checkpoints are often simply marked or assigned a uniform deduction. This approach lacks a tiered representation of the importance of failed items, resulting in some high-risk, high-weight operational errors not being fully reflected in the final results, making it difficult to truly reflect the student's ability gaps.
[0004] Scoring results are often presented as an overall score, lacking structured analysis of the individual scores. This makes it difficult for system managers to identify specific problem areas, limiting the implementation of personalized training feedback and differentiated teaching strategies.
[0005] Assessment results are often categorized into binary or single-level intervals, lacking a continuous grading mechanism. This setup fails to adapt to the dynamic tiered management needs of students at different levels, limiting the ability to optimize training paths based on data.
[0006] In order to achieve more targeted and structured training evaluation, a comprehensive scoring and classification model based on behavioral differences and task differentiation is needed to improve the adaptability and feedback efficiency of the training evaluation system in complex task environments. Summary of the Invention
[0007] In response to the shortcomings of the existing technology, the present invention provides a multi-dimensional assessment method for civil aviation flight trainees, which solves the problems in the existing assessment method that the key checkpoint discrimination cannot be accurately reflected, the scoring results lack structure, and the granularity of the trainee ability classification is insufficient.
[0008] To achieve the above objectives, the present invention is implemented through the following technical solutions: a multi-dimensional assessment method for civil aviation flight trainees, comprising the following steps:
[0009] S1. Collecting original evaluation data of each checkpoint during flight training, where the evaluation result of each checkpoint is a binary value of pass or fail;
[0010] S2. Perform a chi-square test on the relationship between the data at each checkpoint and the student's final assessment category (category A or category B), and calculate the chi-square value for each checkpoint;
[0011] S3. Calculate the Phi correlation coefficient between each checkpoint and the student's final classification based on the chi-square value;
[0012] S4. normalize the Phi correlation coefficient in each dimension to obtain a normalized Phi coefficient;
[0013] S5. Calculate the category discrimination of each checkpoint, defined as the difference in pass rates between Category A and Category B students at that checkpoint;
[0014] S6. Multiply the normalized Phi coefficient by the corresponding category discrimination to obtain the dynamic weight of each checkpoint;
[0015] S7. Calculate the score of each checkpoint based on the basic score, dynamic weight, and student performance of each checkpoint;
[0016] S8. Add up the scores of all failed checkpoints and subtract them from 100 to get the student's final total score;
[0017] S9. Determine the comprehensive ability level of the students based on the final score, and implement a multi-dimensional classification assessment of flight students.
[0018] Preferably, the calculation formula of the Phi correlation coefficient is:
[0019]
[0020] in, is the Phi correlation coefficient, is the chi-square value, is the total number of samples.
[0021] Preferably, the normalized Phi correlation coefficient is calculated as follows:
[0022] ;
[0023] in, Indicates the The normalized Phi coefficient of each checkpoint in its dimension, Indicates the The raw Phi correlation coefficient of the checkpoints, Indicates the sum of the Phi correlation coefficients of all checkpoints in the same dimension. is the index of the checkpoint within the dimension, is the number of checkpoints within the dimension.
[0024] Preferably, the calculation formula of the category discrimination CD is:
[0025] ;
[0026] in, and These are the pass rates of Category A and Category B students at a certain checkpoint.
[0027] Preferably, the dynamic weight of each checkpoint Calculated by the following formula:
[0028]
[0029] in, Indicates the The dynamic weight of the checkpoints, Indicates the The normalized Phi coefficient of each checkpoint, Representative The class discrimination of each checkpoint.
[0030] Preferably, checkpoint score The calculation formula is:
[0031] ;
[0032] in, Representative The score or potential penalty for each checkpoint, Representative The dynamic weight of the checkpoints, It is the basic score of the dimension, and 100 is used to standardize the score to 100 points.
[0033] Preferably, the final score calculation formula for the student is:
[0034] ;
[0035] in, Represents the student's final total score. Represents the student's initial full score or basic score, For checkpoints The failure indicator variable, For checkpoints score.
[0036] Preferably, the dimensions include flight trajectory management, communication, procedure execution, problem solving, and situational awareness core competency dimensions.
[0037] Preferably, the evaluation method is applicable to any flight training stage, and the data used by the evaluation method includes data samples from the 9-hour initial training and 13-hour intensive training stages.
[0038] The present invention provides a multi-dimensional assessment method for civil aviation flight trainees. It has the following beneficial effects:
[0039] 1. This invention utilizes a total score calculation method based on a fusion of checkpoint failure marking and weighted deductions, enabling dynamic responsiveness to student operational errors within the scoring mechanism. Compared to existing approaches that simply accumulate passing items or score based on percentages, this solution clearly quantifies the specific impact of failed items, addressing the lack of differentiation for failed items found in traditional methods.
[0040] 2. This invention constructs a weighting model based on "normalized Phi coefficient × category discrimination × base score" and embeds it into the final scoring logic, achieving an organic integration of capability differentiation and evaluation weighting. Traditional solutions often fail to account for differences in checkpoint discrimination capabilities, resulting in skewed and lacking structure in scoring results. This solution overcomes the technical shortcoming of such insensitive scoring.
[0041] 3. The "Total Score Model," constructed by superimposing failure items, offers strong interpretability. Managers can clearly identify student deduction points and their impact. Compared to existing methods that often employ black-box models or fuzzy scoring rules, this solution effectively addresses the difficulty in tracing training evaluation results, hindering corrective guidance.
[0042] 4. The multi-level classification mechanism based on score ranges provided by this invention supports subsequent integration with modules such as automatic tiered training and training feedback delivery. Unlike existing solutions that only provide a single judgment of whether or not a target has been met, this mechanism offers greater adaptability and scalability, significantly improving the management efficiency and intelligence level of the training system. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 This is the improved scoring flow chart of the present invention;
[0044] Figure 2 This is the original scoring flow chart;
[0045] Figure 3 This is the original score chart for the 9-hour flight training;
[0046] Figure 4 Adjusted score chart for 9 hours of flight training;
[0047] Figure 5 This is the original score chart for the 13-hour flight training;
[0048] Figure 6 Adjusted score chart for 13 hours of flight training. DETAILED DESCRIPTION
[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the specification of the present invention. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0050] To more accurately reflect the comprehensive performance of flight trainees during training, this paper provides a multi-dimensional assessment method for civil aviation flight trainees based on statistical correlation analysis. This method combines the correlation between checkpoint performance and final classification (A / B) to scientifically assess and dynamically assign weights to trainees' key competencies. The following describes the implementation of this method in detail, combining specific steps.
[0051] Please see the attached Figure 1 -Attached Figure 6 The embodiment of the present invention provides a multi-dimensional assessment method for civil aviation flight trainees, comprising the following steps:
[0052] S1. Collecting raw evaluation data for each checkpoint during flight training. The evaluation results of the checkpoints are binary values of pass or fail.
[0053] To ensure that the scoring mechanism truly and objectively reflects the student's performance during training, high-quality raw data collection is essential for flight trainee competency assessment. This step, as the foundation of the overall assessment method, directly determines the accuracy and applicability of subsequent correlation analysis and weight calculation.
[0054] As a prerequisite, the collected data should cover all operational behavior nodes with evaluation significance in the flight training mission, and must include the paired data relationship between these nodes and the students' assessment results to support subsequent multi-dimensional modeling work based on statistics.
[0055] In the present invention, the core goal of step S1 is to construct an original data structure system for evaluation and analysis. The data system must meet the following requirements: on the one hand, it must be able to fully reflect all performance details of each student during the training process; on the other hand, it must have a standardized data organization format to facilitate the direct calling and processing of the statistical analysis model.
[0056] In this embodiment, the raw data collection objects are the checkpoint results and final evaluation results of the tasks performed by the trainees at fixed flight training time points.
[0057] Checkpoints are various observation points preset in the flight training mission process based on the Civil Aviation Competency-Based Training and Assessment (CBTA) framework. Each checkpoint corresponds to a skill or behavioral requirement.
[0058] Each student pilot undergoes multiple phases of training, each with several checkpoints. For example, during the 9-hour initial training, each student will go through 22 phases with a total of 108 checkpoints; during the 13-hour intensive training, the number will reach 36 phases with a total of 158 checkpoints.
[0059] It should be noted that each checkpoint in the training is clearly defined as a binary state of "whether it meets the standard". The following encoding rules are used in this invention:
[0060] If the trainee's performance on this checkpoint meets the predetermined operational standards, it is scored as 0 (pass);
[0061] If the performance does not meet the assessment criteria, it is scored as 1 (failed).
[0062] To support the Phi correlation coefficient analysis method used in this invention, all checkpoint data must be binary variables. This provides the data basis for the subsequent construction of a 2×2 contingency table and the statistical independence test.
[0063] In this embodiment, in addition to collecting the trainee's performance data at each checkpoint, it is also necessary to simultaneously record the final evaluation results of each trainee in this stage of training. This result is also in the form of binary classification:
[0064] Category A: indicates that the training has met the standards and the candidate is qualified for subsequent training;
[0065] Category B: Indicates that the training does not meet the standards and there is a risk of grounding.
[0066] In one possible implementation, each sample record needs to contain the following fields:
[0067] Student's unique ID (student_id, used for data collation, not involved in scoring);
[0068] Number of each checkpoint (e.g. FPM.2, SAW.1, etc., refer to CBTA code specifications);
[0069] The pass status of each checkpoint (0 / 1);
[0070] The student's final classification result (A or B);
[0071] The core capability dimension to which the checkpoint belongs (e.g., FPM, COM, PRO, etc.);
[0072] The flight phase (such as engine start, cruise flight, landing, etc.).
[0073] As an option, to ensure the standardization of the data structure, the present invention uses a table structure to organize all data into a training record table. Each row represents a complete training performance sample of a student, and each column is the evaluation result of a checkpoint and the final assessment label.
[0074] Specifically, the following are typical data structure styles:
[0075] Table 1:
[0076]
[0077] As you can understand, each "0 / 1" mark in the table indicates whether the student's performance at the corresponding checkpoint met the standard, while the "Assessment Results" provides a comprehensive evaluation of the overall performance during the training phase. This structure supports both human readers and machine programs in performing statistical analysis on large sample data.
[0078] In some embodiments, dimension attribution information of each checkpoint may also be recorded to support subsequent normalization processing and dimension-level decentralized control.
[0079] For example, in a 9-hour training session, 108 checkpoints can be mapped to the following five capability dimensions:
[0080] Table 2: 9-hour dimension details
[0081]
[0082] During the 13-hour training, a sixth dimension, KNO (Knowledge Application, 10 items in total), was added. The above mapping information can be found in the following table:
[0083] Table 3: 13-hour dimension details
[0084]
[0085] As a further technical treatment, the following standardization and cleaning steps are performed on the collected data in the present invention:
[0086] Delete all fields containing student identification information (such as name, gender, ID card, etc.) to comply with regulations;
[0087] Use missing value processing algorithms to eliminate incomplete records;
[0088] The final assessment result labels are standardized into binary variables, with category A being 1 and category B being 0 or vice versa, depending on the implementation;
[0089] Ensure that all checkpoint data types are Boolean or numeric 0 / 1 to facilitate subsequent statistics.
[0090] Data processing and statistical analysis (Phi correlation coefficient) in step S1
[0091] After data collection was complete, the Phi correlation coefficient was calculated to analyze the correlation between performance at each checkpoint. The following table shows the Phi correlation coefficient analysis results for the 9-hour and 13-hour flight training phases. The table shows the degree of correlation between each checkpoint and the trainee's final assessment.
[0092] Table 4: Phi correlation coefficient table for 9-hour flight training
[0093]
[0094] Table 4 shows that some checkpoints, such as Checkpoint 20.1.5 (SAW) and Checkpoint 18.4.2 (FPM), have high Phi coefficients of 0.408 and 0.34, respectively, indicating that these checkpoints have a strong discriminatory effect on the students' final classification. Checkpoints with high Phi coefficients are mainly concentrated in the SAW and FPM dimensions, indicating that students' information management and manual flight skills in complex environments have a significant impact on their final performance.
[0095] Table 5: Phi correlation coefficient table for 13 hours flight training
[0096]
[0097] Table 5 shows that checkpoints 21.3 (FPM) and 32.2.1 (PRO) have high Phi coefficients of 0.298 and 0.275, respectively, indicating that these checkpoints have a strong indicative effect on trainee classification. In the 13-hour dataset, some checkpoints in the PSD and PRO dimensions also exhibit high Phi coefficients, reflecting the increasing importance of decision-making and rule execution skills in longer training sessions.
[0098] The Phi correlation coefficient table above shows the correlation between different checkpoints and the final assessment results. For example, the Phi value of checkpoint 20.1.5 (SAW dimension) in a 9-hour flight training is 0.408, indicating a strong correlation with the trainee's final assessment results.
[0099] In summary, step S1, by collecting standardized and structured flight training checkpoint data, provides the data foundation for the dynamic weighted scoring method based on the Phi correlation coefficient and class discrimination. With complete checkpoint records and corresponding classification labels, subsequent steps can be used to conduct checkpoint analysis and scoring modeling.
[0100] S2. Perform a chi-square test on the relationship between the data at each checkpoint and the student's final assessment category (category A or category B), and calculate the chi-square value for each checkpoint;
[0101] After completing flight training data collection in step S1, the main purpose of step 2 is to quantify the correlation between each checkpoint and the trainee's final assessment result (Category A or B) through chi-square tests and Phi correlation coefficients. The Phi correlation coefficient is a statistic that measures the strength of the relationship between binary variables and can help analyze the importance of each checkpoint in flight training evaluation.
[0102] Chi-square test and Phi correlation coefficient formula: The chi-square test is used to detect the difference between the observed frequency and the expected frequency. If the difference is significant, it means that there is a certain correlation between the variables. The calculation formula of the chi-square test is:
[0103]
[0104] in: is the chi-square value, is the observed frequency, that is, the frequency in the actual sample data; is the expected frequency, the frequency calculated based on the independence assumption.
[0105] In binary data, the contingency table will show the distribution of students who passed and failed each checkpoint in Category A and Category B. The specific 2×2 contingency table format is as follows:
[0106] Table 6:
[0107]
[0108] According to this contingency table, the expected frequency The calculation formula is as follows:
[0109]
[0110]
[0111]
[0112] Among them, a represents the number of students who passed the specific checkpoint among those who were finally assessed as Class A; b represents the number of students who passed the specific checkpoint among those who were finally assessed as Class B; c represents the number of students who failed the specific checkpoint among those who were finally assessed as Class A; d represents the number of students who failed the specific checkpoint among those who were finally assessed as Class B.
[0113] , , , is the expected frequency calculated based on these observed frequencies: corresponds to the expected frequency of cell a; corresponds to the expected frequency of cell b; corresponds to the expected frequency of cell c; corresponds to the expected frequency of cell d.
[0114] After obtaining the chi-square value, we can further calculate the Phi correlation coefficient to measure the strength of the relationship between two variables. The formula for the Phi correlation coefficient is:
[0115] ;
[0116] in: is the Phi correlation coefficient, is the chi-square value, is the total number of samples.
[0117] The Phi correlation coefficient value ranges from -1 to 1. The closer the value is to 1, the stronger the correlation between the checkpoint and the student's final assessment result; the closer the value is to 0, the weaker the correlation.
[0118] Implementation of Step 2
[0119] Data collation and contingency table construction: Based on the student assessment data collected in step S1, construct a 2×2 contingency table corresponding to each checkpoint, recording the distribution of passed and failed students in Class A and Class B students.
[0120] Calculation of Chi-square value: For each checkpoint, use the Chi-square test formula to calculate the Chi-square value between it and the student's final assessment result. .
[0121] Calculation of Phi correlation coefficient: Based on the chi-square value, the Phi correlation coefficient is calculated to quantify the strength of the relationship between the checkpoint and the student's final assessment result.
[0122] Statistical analysis and result presentation: The chi-square value and Phi correlation coefficient calculation results are organized into a table to show the correlation strength of each checkpoint, helping to further analyze the impact of each checkpoint on the student's assessment results.
[0123] To improve computational efficiency, especially when faced with a large amount of checkpoint data, the following scaling solutions can be considered:
[0124] Data preprocessing and distribution optimization: Before calculating the Phi correlation coefficient, the data is preprocessed and normalized to remove outliers and improve the accuracy of the chi-square test.
[0125] Use parallel computing: When dealing with large amounts of data, you can use parallel computing to speed up the calculation of the chi-square test and Phi correlation coefficient, improving efficiency.
[0126] Apply cross-validation: In order to verify the reliability of the Phi correlation coefficient, the robustness of the model can be checked by cross-validation.
[0127] Therefore, in step 2, we quantified the correlation between each checkpoint and the student's final assessment result by calculating the Phi correlation coefficient using a chi-square test. This process effectively assesses the importance of each checkpoint in flight training and provides a scientific basis for student scoring. Through chi-square tests and Phi correlation coefficient analysis, we can identify the checkpoints that have the greatest impact on the student's final assessment result, thereby further optimizing the training evaluation system.
[0128] S3. Calculate the Phi correlation coefficient between each checkpoint and the student's final classification based on the chi-square value; normalize the Phi correlation coefficient within each dimension to obtain the normalized Phi coefficient;
[0129] In step 2, the Phi correlation coefficient for each checkpoint was calculated using a chi-square test. To eliminate the impact of differences in the number of checkpoints in each dimension and avoid scoring imbalances, this step normalizes the Phi coefficients within each dimension, ensuring that the sum of the Phi coefficients for each dimension is 1. This normalization standardizes the relative importance of checkpoints within each dimension, providing a more impartial basis for the subsequent weighted scoring model.
[0130] To avoid scoring imbalance due to the uneven number of checkpoints in each dimension, the normalization method is used to standardize the Phi values in each dimension so that their sum is 1. The normalization formula is as follows:
[0131] ;
[0132] in, Indicates the The normalized Phi coefficient of each checkpoint in its dimension, Indicates the The raw Phi correlation coefficient of the checkpoints, Indicates the sum of the Phi correlation coefficients of all checkpoints in the same dimension. is the index of the checkpoint within the dimension, is the number of checkpoints within the dimension.
[0133] Normalized Phi coefficient It indicates the importance of the checkpoint in the dimension to which it belongs, and the sum of the normalized Phi coefficients of all checkpoints in the dimension is 1. In this way, the normalized Phi coefficient of any checkpoint can intuitively reflect its relative contribution in the dimension.
[0134] Implementation of step 3:
[0135] Determine the Phi coefficient of the checkpoint in each dimension;
[0136] Based on the calculation results of step 2, list the Phi values of all checkpoints in each dimension Calculate the sum of the Phi coefficients in each dimension;
[0137] Sum the Phi values in each dimension to get the sum of the Phi coefficients of that dimension: .
[0138] Normalization: Normalize the Phi coefficient of each checkpoint and use Formula 3 to convert the Phi coefficient of each checkpoint into a normalized Phi coefficient.
[0139] Structured display: The normalized Phi coefficients are organized into a table to show the relative importance of each checkpoint in each dimension.
[0140] Example: Phi coefficient normalization of 13 hours of flight training data
[0141] Table 7: Normalized Phi coefficients for 13 hours of flight training
[0142]
[0143] Table 7 shows that the scoring method using Phi coefficient adjustment significantly improves the discrimination of students' true abilities compared to the traditional equally weighted scoring method. For example, checkpoint 30.3.3, with a normalized Phi value of 0.73, saw its final score increase from a raw score of 6.7 to 26.8, demonstrating strong discriminatory power. Similarly, checkpoint 32.2.1 saw its score increase to 12.8 after normalization, further demonstrating its effectiveness in distinguishing between Category A and Category B students. Some checkpoints with lower raw scores but higher Phi coefficients (such as Checkpoint 21.3) saw significant score increases after weight adjustment, demonstrating the improved scoring method's emphasis on key skill points. Conversely, checkpoints with lower category discrimination (such as Checkpoint 5.1) saw smaller score increases even with higher Phi normalization coefficients, avoiding overemphasis on ineffective discriminatory points. Overall, the improved scoring method effectively highlights key ability indicators, improves discrimination of students' true abilities, and provides a more scientific and reasonable scoring standard for flight student selection and training.
[0144] To ensure the computational efficiency and accuracy of this step, the following optimization measures can be considered:
[0145] Parallel Computing: The Phi normalization calculation for each dimension is processed in parallel to improve computational efficiency, especially when processing large amounts of data.
[0146] Visualization tools: To better display the normalized Phi coefficient, corresponding visualization tools can be developed to display the results graphically, allowing analysts to quickly identify important checkpoints.
[0147] Weighted normalization: For cases where certain dimensions are more important, a weighting factor can be introduced so that the normalized Phi coefficient of a specific dimension has a higher weight in the overall score.
[0148] Step 3 normalizes the Phi coefficients within each dimension, ensuring a fair comparison of the importance of checkpoints within each dimension and preventing the impact of varying numbers of checkpoints across dimensions on the scoring results. The resulting Phi coefficients can be used as input to subsequent scoring models, providing a scientific basis for further training and evaluation.
[0149] S4. Calculate the category discrimination of each checkpoint, defined as the difference in pass rates between Category A and Category B students at that checkpoint;
[0150] In the assessment method of this invention, the category discrimination (CD) metric is introduced to enhance the assessment model's ability to distinguish between different student categories. This metric aims to reflect the effectiveness of a checkpoint in distinguishing between different student categories by measuring the difference in pass rates between Category A and Category B students at the same checkpoint. A higher CD indicates a stronger ability of the checkpoint to distinguish between students. Therefore, calculating CD is a key step in the assessment model of this invention, ensuring that the model more sensitively identifies differences in ability among students.
[0151] Alternatively, category discrimination is measured by comparing the pass rates of Category A and Category B students at each checkpoint. By calculating the category discrimination for each checkpoint, we can assess the effectiveness of different assessment dimensions in distinguishing student categories, thus providing a basis for the subsequent weighted scoring model.
[0152] Category discrimination calculation formula:
[0153] In this embodiment, the category discrimination (CD) is calculated using the following formula:
[0154] ;
[0155] in: Indicates the The pass rate of Category A students at each checkpoint; Indicates the The pass rate of Category B students at each checkpoint; Indicates the The class discrimination of each checkpoint.
[0156] Specifically, category discrimination is quantified by comparing the pass rates of Category A and Category B students at the same checkpoint. The absolute value in the formula represents the difference in pass rates between the two categories at that checkpoint. A larger value indicates a stronger ability of that checkpoint to distinguish between Category A and Category B students.
[0157] Calculation steps for category discrimination:
[0158] In one possible implementation, the calculation process of step 4 includes the following key steps:
[0159] First, we need to obtain the pass rate data of Class A and Class B students at each checkpoint. This data usually comes from the students' actual assessment performance at each checkpoint. Specifically, calculate the pass rate of Class A students at a certain checkpoint. and the pass rate of Category B students , which can be achieved by:
[0160]
[0161] It should be noted that the definition of Category A and Category B students is determined based on specific assessment criteria and may be divided according to the students' ability level, training time or other criteria.
[0162] Next, calculate the category discrimination , which is the absolute value of the pass rate difference shown in the formula. For each checkpoint, by comparing the pass rate difference between category A and category B students, the category discrimination of each checkpoint is obtained.
[0163] In this embodiment, the calculated results of category discrimination are used to evaluate the importance of each checkpoint in differentiating students' abilities. Specifically, a checkpoint with a higher category discrimination has a stronger effect on differentiating the abilities of Category A and Category B students, reflecting the contribution of that checkpoint to the assessment model.
[0164] In some embodiments, each evaluation dimension may be weighted based on the degree of category discrimination. For example, checkpoints with higher category discrimination may be given a higher weight so that these checkpoints with stronger discrimination are given more weight in the comprehensive scoring.
[0165] In order to further improve the discrimination of the evaluation model, in an extended embodiment, if the category discrimination of some checkpoints is low, that is, If the value is small, it means that the checkpoint is not effective in distinguishing between type A and type B students. In this case, you can optimize it in the following ways:
[0166] Adjust checkpoint settings: For checkpoints with low category discrimination, consider adjusting their evaluation criteria or adding new dimensions to improve the discrimination of the checkpoint.
[0167] Add new category differentiation checkpoints: Some new checkpoints can be added that can provide stronger differentiation information when distinguishing between Category A and Category B students.
[0168] It is important to understand that by optimizing checkpoints with low category discrimination, the sensitivity of the entire evaluation model in distinguishing different student categories can be improved, thereby improving the accuracy and fairness of student scoring.
[0169] For example, in a flight training assessment, if a checkpoint (e.g., simulated flight operation) does not clearly differentiate between beginners (Category A) and experienced trainees (Category B), the category discrimination of the checkpoint is In this case, you can consider increasing the category differentiation of this checkpoint by adding more differentiated items such as flight operation skills and emergency response.
[0170] In another possible implementation, if the category differentiation in a certain training dimension (such as flight planning and management capabilities) is low, more detailed evaluation criteria can be introduced, such as improving the category differentiation of this dimension based on the trainees' decision-making performance in specific situations.
[0171] By calculating the category discrimination, we can clarify the effectiveness of each checkpoint in distinguishing student categories, thereby optimizing the sensitivity and accuracy of the evaluation model. The introduction of category discrimination not only enhances the discrimination power of the evaluation model, but also provides a scientific basis for subsequent scoring weighting and optimization of evaluation dimensions.
[0172] S5. Multiply the normalized Phi coefficient by the corresponding category discrimination to obtain the dynamic weight of each checkpoint;
[0173] In the assessment method of this invention, step 5 aims to further optimize the role of checkpoints in the scoring model by calculating the final checkpoint weights. The final weights are a comprehensive consideration of each checkpoint's normalized Phi coefficient and class discrimination (CD). The goal is to ensure that each checkpoint fairly and effectively reflects the student's ability level during the assessment process.
[0174] In the previous steps, we calculated the normalized Phi coefficient and class discrimination for each checkpoint. The former measures the relative importance of the checkpoint within the evaluation dimension, while the latter measures the effectiveness of the checkpoint in distinguishing between different student categories. The core of Step 5 is to combine these two factors and calculate the final weight for each checkpoint through a weighted approach. In this way, the present invention can ensure that the scoring model more accurately reflects the student's ability performance and has stronger discriminatory power.
[0175] In this example, the final weight of the checkpoint It is determined by two main factors:
[0176] Normalized Phi coefficient , which reflects the weight of the checkpoint within the evaluation dimension;
[0177] Category discrimination , which measures the effectiveness of the checkpoint in distinguishing different categories of learners.
[0178] Alternatively, the final weights are calculated as follows:
[0179]
[0180] in: It is Normalized Phi coefficient of each checkpoint; It is Category discrimination of each checkpoint; and It is the weight coefficient for adjusting the relative importance of the normalized Phi coefficient and category discrimination.
[0181] Specifically, the weight coefficient and It can be adjusted according to the needs of actual application. For example, when the ability differentiation of the assessment dimension has a greater impact on the student's performance, it can be appropriately increased. On the contrary, if the normalized Phi coefficient has a greater impact on the evaluation results, you can increase The weight of .
[0182] In this embodiment, the process of calculating the final weight of each checkpoint includes the following steps:
[0183] First, calculate the normalized Phi coefficient for each checkpoint , this process has been completed in step 3. The normalized Phi coefficient reflects the relative importance of the checkpoint within the dimension. The calculation method is to normalize the original Phi coefficient so that the sum of the Phi coefficients of all checkpoints in each dimension is 1.
[0184] Then, the class discrimination of each checkpoint is calculated , this process has been completed in step 4. Category discrimination reflects the effect of this checkpoint in distinguishing between category A and category B students, and the calculation formula is:
[0185]
[0186] in, and The A and B students are The pass rate at each checkpoint.
[0187] Next, the formula is used to combine the normalized Phi coefficient and the class discrimination to calculate the final weight of each checkpoint. and Adjustment, final weight It can be flexibly changed according to actual application requirements.
[0188] In one possible implementation, the weight coefficient can be adjusted according to the actual situation. and To better meet specific assessment objectives. For example, in some assessment scenarios (such as comprehensive ability assessment of trainees), category discrimination may be more important, while in other scenarios (such as assessment of specific skills), the normalized Phi coefficient may be more important.
[0189] Specifically, in some embodiments, when the class distinction is low, it is possible to consider increasing the weight of the normalized Phi coefficient to reduce the impact of the class distinction. Conversely, when the class distinction is high, the weight coefficient can be adjusted to Enhance its contribution to the final weight to better reflect the importance of the checkpoint in distinguishing the learner classes.
[0190] Assume that in a 9-hour flight training evaluation, some checkpoints (such as flight plans) have high normalized Phi coefficients but low class discrimination. In this case, the weight of the normalized Phi coefficient can be appropriately increased. , and reduce the weight of category discrimination , thereby placing more emphasis on the evaluation role of this checkpoint in the final scoring.
[0191] For example, in the flight operation skills assessment, certain skills have a strong role in distinguishing between Category A students (beginners) and Category B students (experienced ones), and the category discrimination is high. Therefore, the weight of the category discrimination can be increased to highlight the contribution of this checkpoint in distinguishing students' abilities.
[0192] By calculating the final weights of the checkpoints, the present invention effectively combines the normalized Phi coefficient with class discrimination, improving the discriminatory power and accuracy of the scoring model. The calculation of the final weights not only allows for flexible adjustment based on actual assessment needs but also provides a more scientific basis for subsequent comprehensive student scoring. This weighting approach provides significant flexibility for optimizing and adjusting the assessment model, adapting to the needs of different student groups and assessment objectives, thereby improving the fairness and accuracy of the assessment.
[0193] S6. Calculate the score for each checkpoint based on its base score, dynamic weight, and the student's performance. Add up the scores of all failed checkpoints and subtract them from 100 to obtain the student's final total score.
[0194] In the previous steps, we calculated the normalized Phi coefficient, class discrimination, and final weight, providing a quantitative basis for the evaluation contribution of each checkpoint. To further optimize the student evaluation process and ensure the fairness and rationality of the scores for each checkpoint, this step aims to calculate the final score for each checkpoint.
[0195] The core idea of this step is to calculate the score for each checkpoint by combining its normalized Phi coefficient, category discrimination, and base score. This process ensures that each checkpoint fully reflects its role in differentiating students' abilities when assessing them, and adjusts the weights to make the final score more representative.
[0196] Checkpoint score calculation formula:
[0197] In this embodiment, the score calculation of the checkpoint is achieved through the following two steps:
[0198] Calculate the weight of the checkpoint The weight of each checkpoint is calculated based on its normalized Phi coefficient and category discrimination The specific calculation formula is:
[0199]
[0200] in: Indicates the The dynamic weight of the checkpoints, Indicates the The normalized Phi coefficient of each checkpoint, Representative Category discrimination of each checkpoint;
[0201] It is The class discrimination of each checkpoint.
[0202] This calculation method combines the importance of the checkpoint in the assessment dimension (reflected by the Phi coefficient) and its contribution to the differentiation of students' abilities (reflected by category discrimination) to derive the weight of each checkpoint.
[0203] Calculate the final score of the checkpoint Calculate the weight of each checkpoint Finally, you need to combine the basic score of the checkpoint To calculate the final score The calculation formula is as follows:
[0204] ;
[0205] in: Representative The score or potential penalty for each checkpoint, It is The weight of each checkpoint; is the base score for this checkpoint; 100 is used to normalize the score to a 100-point scale.
[0206] It should be noted that the basic score It is determined based on the preliminary assessment results at each checkpoint and is usually calculated dynamically based on the student's performance at that checkpoint.
[0207] How the computational process is implemented
[0208] In this embodiment, the process of calculating the score of each checkpoint includes the following steps:
[0209] First, according to the normalized Phi coefficient obtained in step 5 and category discrimination , calculate the weight of each checkpoint using the formula .
[0210] Next, combine the base scores for each checkpoint , calculate the final score of the checkpoint according to the formula . Basic score It is usually based on the student's operation or test results at the checkpoint. In some embodiments, the basic score can also be dynamically adjusted according to factors such as the student's training time and actual performance.
[0211] Final score The calculation result not only takes into account the weight of each checkpoint, but also incorporates the students' actual ability performance (reflected by the basic score) into the assessment, making the assessment result more accurate and reasonable.
[0212] Basic score It is a quantification of the student's actual performance at each checkpoint, usually determined in the following ways:
[0213] Static score: For certain standardized checkpoints (such as basic flight operations), a fixed basic score can be preset.
[0214] Dynamic Scores: Based on the student's real-time performance during training, the basic score can be dynamically calculated. For example, during flight training, the basic score can be dynamically adjusted based on the student's time, accuracy, and other indicators of the task completed.
[0215] Assume that in a flight training assessment, the scores for Checkpoint 1 (e.g., takeoff maneuvers) and Checkpoint 2 (e.g., landing maneuvers) are calculated as follows:
[0216] The normalized Phi coefficient of checkpoint 1 is 0.8, the class discrimination is 0.6, and the base score is 80;
[0217] Checkpoint 2 has a normalized Phi coefficient of 0.5, a class discrimination of 0.9, and a base score of 90.
[0218] First, calculate the weight of the checkpoint using the formula:
[0219] ;
[0220] ;
[0221] Then, the final score is calculated according to the formula:
[0222]
[0223] In this case, although the basic score of Checkpoint 2 is higher, its final score is slightly lower than that of Checkpoint 1 due to its lower normalized Phi coefficient. This shows that Checkpoint 1 plays a more prominent role in the assessment of trainees' abilities.
[0224] In order to further improve the accuracy and fairness of the evaluation model, the weight coefficient or the calculation method of the basic score can be adjusted according to the specific application scenario. For example:
[0225] Adjusting the base score: When certain checkpoints are less effective in differentiating students of different abilities, the scoring can be optimized by adjusting the weight of the base score.
[0226] Dynamically adjust weights: In different groups of students, the weights of the normalized Phi coefficient and category discrimination can be dynamically adjusted based on the students' overall performance, thereby improving the adaptability of the model.
[0227] By calculating checkpoint scores, the present invention provides a method that comprehensively considers the importance, discrimination, and actual performance of each checkpoint in assessing student ability. The final score not only accurately reflects the actual assessment effect of each checkpoint but can also be adjusted based on the student's specific performance, ensuring the accuracy and fairness of the assessment results. Furthermore, the score calculation method of the present invention is highly adaptable and can be optimized and adjusted according to the needs of different assessment scenarios, further enhancing the effectiveness of the assessment model.
[0228] S7. Determine the comprehensive ability level of the students based on the final score, and implement a multi-dimensional classification assessment of flight students.
[0229] After quantifying the scores for each checkpoint, a method for calculating a total score based on checkpoint failures is needed to systematically assess the student's overall training performance. This step is closely aligned with Step 6 above. Its core goal is to construct a total score function model, combining the failure information and score weights of each checkpoint, to form an assessment mechanism that can be used to categorize student ability levels.
[0230] It should be noted that this mechanism relies on the checkpoint score calculated in step 6 , and introduce the checkpoint failure identification variable By establishing a "failure-deduction" mapping relationship, we can achieve a systematic rollback of total scores based on individual operational errors. This method can accurately reflect students' training deficiencies in key areas and form a graded classification result to support subsequent training recommendations and tiered development strategies.
[0231] In this embodiment, the student's total score is calculated according to the following functional relationship:
[0232] ;
[0233] in, Represents the student's final total score. Represents the student's initial full score or basic score, For checkpoints The failure identification variable of , whose value is:
[0234] ;
[0235] For checkpoints The specific calculation method is detailed in step 6. Its value comes from the product of the normalized Phi coefficient and the category discrimination and then combined with the basic score, that is:
[0236] ;
[0237] The above formula comprehensively reflects the weight of the checkpoint in the ability evaluation system while retaining the differences in training performance. The product relationship between the two scores is based on the weighted deduction of scores for failed items, which constitutes a structured scoring method of "weighted deduction based on failed items".
[0238] In this embodiment, the students' performance is classified based on the Total Score calculation results.
[0239] In the standard classification mode, a fixed threshold is introduced to divide the score range into different ability level ranges. The exemplary classification rules are as follows:
[0240] when When , the trainees are classified as Class A, indicating a high degree of training achievement;
[0241] when When the students are divided into Class B, they are considered to have shortcomings in their abilities or insufficient performance.
[0242] Specifically, the present invention is not limited to a two-level classification model, and can also be expanded to a multi-level classification scheme according to application requirements in an actual evaluation system. For example, as an option, a more detailed interval division can be set as follows:
[0243] : Category S (Excellent Performance);
[0244] : Class A (meets standards);
[0245] : Category B (to be strengthened);
[0246] : Category C (unqualified).
[0247] The specific values of the classification boundaries can be adjusted according to the score distribution curve of a large sample of students collected during the system debugging phase, or they can be set based on the normal distribution mean or percentile method to adapt to the complexity and tolerance requirements of different training tasks.
[0248] In one possible implementation, a classification mechanism is combined with a training feedback module for linkage processing.
[0249] For example, if the system identifies a student with a Total Score of 78 and their primary failure points are concentrated in low-weighted checkpoints, they can be marked as "critically qualified" through the "marginal tolerance mechanism" and allowed to proceed to the re-evaluation stage. Alternatively, the re-evaluation can be manually confirmed by the training administrator based on the specific failure points and the recommended classification can be adjusted.
[0250] It's important to note that this mechanism, fundamentally based on a "failure item plus rebate" strategy, offers excellent scalability and fault tolerance. As an extension, additional evaluation dimensions (such as task completion time and response latency) can be introduced and incorporated into the basic score calculation in Formula 7 via a linear fusion coefficient, making the classification results more comprehensive across dimensions.
[0251] For example, the following is an application case of a specific calculation process:
[0252] The total number of checkpoints a student participated in the assessment is The failed checkpoints are items 3, 5, and 9, with corresponding scores: S3=4.2, S5=3.8, and S9=2.7 respectively.
[0253] Substitute into the formula for calculation: ;
[0254] Then the student can be classified as Class A according to the standard classification rules, indicating that the training requirements have been basically met.
[0255] It can be understood that the integrated scoring-classification mechanism of the present invention takes into account both scoring transparency and evaluation accuracy, and effectively integrates the capability characterization of multi-dimensional weights while being based on key checkpoint failure statistics.
[0256] The above method is particularly suitable for comprehensive capability evaluation scenarios such as flight training, industrial skills assessment, and medical simulation, which require the combination of process performance and error analysis. It is also suitable for the basic logical interface for the subsequent construction of training plan scheduling and feedback push modules.
[0257] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A multi-dimensional assessment method for civil aviation flight trainees, characterized by: The following steps are involved: S1. Collecting original evaluation data of each checkpoint during flight training, where the evaluation result of each checkpoint is a binary value of pass or fail; S2. The final assessment category of the trainee is either A or B. A chi-square test is performed on the relationship between the data of each checkpoint and the final assessment category of the trainee, and the chi-square value of each checkpoint is calculated; S3. Calculate the Phi correlation coefficient between each checkpoint and the student's final classification based on the chi-square value; S4. normalize the Phi correlation coefficient in each dimension to obtain a normalized Phi coefficient; S5. Calculate the category discrimination of each checkpoint, defined as the difference in pass rates between Category A and Category B students at that checkpoint; S6. Multiply the normalized Phi coefficient by the corresponding category discrimination to obtain the dynamic weight of each checkpoint; S7. Calculate the score of each checkpoint based on the basic score, dynamic weight, and student performance of each checkpoint; S8. Add up the scores of all failed checkpoints and subtract them from 100 to get the student's final total score; S9. Determine the comprehensive ability level of the students based on the final score, and implement a multi-dimensional classification assessment of flight students.
2. The multi-dimensional assessment method for civil aviation flight trainees according to claim 1, characterized in that: The calculation formula of the Phi correlation coefficient is: ; in, is the Phi correlation coefficient, is the chi-square value, is the total number of samples.
3. The multi-dimensional assessment method for civil aviation flight trainees according to claim 1, characterized in that: The normalized Phi correlation coefficient is calculated as follows: ; in, Indicates the The normalized Phi coefficient of each checkpoint in its dimension, Indicates the The raw Phi correlation coefficient of the checkpoints, Indicates the sum of the Phi correlation coefficients of all checkpoints in the same dimension. is the index of the checkpoint within the dimension, is the number of checkpoints within the dimension.
4. The multi-dimensional assessment method for civil aviation flight trainees according to claim 1, characterized in that: The calculation formula of the category discrimination CD is: ; in, and These are the pass rates of Category A and Category B students at a certain checkpoint.
5. The multi-dimensional assessment method for civil aviation flight trainees according to claim 3, characterized in that: Dynamic weight of each checkpoint Calculated by the following formula: ; in, Indicates the The dynamic weight of each checkpoint, Indicates the The normalized Phi coefficient of each checkpoint, Representative The class discrimination of each checkpoint.
6. The multi-dimensional assessment method for civil aviation flight trainees according to claim 5, characterized in that: Checkpoint score The calculation formula is: ; in, Representative The score or potential penalty for each checkpoint, Representative The dynamic weight of each checkpoint, It is the basic score of the dimension.
7. The multi-dimensional assessment method for civil aviation flight trainees according to claim 1, characterized in that: The final score of the students is calculated as follows: ; in, Represents the student's final total score. Represents the student's initial full score or basic score, For checkpoints The failure indicator variable, For checkpoints score.
8. The multi-dimensional assessment method for civil aviation flight trainees according to claim 1, characterized in that: The dimensions include flight trajectory management, communication, procedure execution, problem solving, and situational awareness core competency dimensions.
9. The multi-dimensional assessment method for civil aviation flight trainees according to claim 1, characterized in that: The evaluation method is applicable to any flight training stage, and the data used by the evaluation method includes data samples from the 9-hour initial training and 13-hour intensive training stages.
Citation Information
Patent Citations
Hyperspectral anomaly detection method based on multilevel tensor prior constraint
CN114331976A
VENN model-based competency evaluation standard making method
CN117057673A