Electronic confrontation training evaluation method based on data analysis
By classifying training data and applying descriptive, validation and data mining analysis methods, the diversified problems of electronic adversarial training evaluation are solved, and the flexibility and accuracy of training evaluation are improved, and the multi-dimensional evaluation needs of complex scenarios are adapted to the needs of multi-dimensional evaluation in complex scenarios.
Patent Information
- Application Number
- CN202510353479.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
AI Technical Summary
The existing electronic adversarial training evaluation methods based on data analysis lack systematic guidance on diversified training evaluation issues, making it difficult to support multi-dimensional evaluation of complex scenarios, failing to fully utilize training data to tap potential value, and the evaluation results are insufficiently reliable.
By classifying training data, it is used to describe, verify, and data mining analysis, and combining descriptive statistics of the overall data, non-parametric tests and data mining methods for small data sets, a flexible data analysis framework is built to carry out all-round and in-depth applications of small-scale data for electronic adversarial training.
It realizes objective and accurate evaluation of training evaluation, improves the flexibility, comprehensiveness and accuracy of evaluation, adapts to the complexity and dynamic uncertainty of electronic confrontation training, and supports efficient solution to diversified training evaluation problems.
Smart Images

Figure CN120296453A_ABST
Abstract
Description
Technical Field:
[0001] The present invention belongs to the technical field of electronic countermeasure training, and mainly relates to an electronic countermeasure training evaluation method based on data analysis. Background Art:
[0002] Electronic countermeasure training evaluation, as a key link for comprehensively understanding, value analysis, and diagnostic evaluation of various elements and activities in training, runs through the entire training process. It has extremely important practical significance in discovering and solving training problems, promoting the implementation of actual combat training, and testing the effectiveness of training.
[0003] Different from traditional training evaluation which mainly focuses on the construction of index systems, the establishment of index weights, and the construction of evaluation method models, electronic countermeasure training evaluation based on data analysis takes training data as the core driver, emphasizes the in-depth application of data analysis methods in the whole process and multi-dimensional training evaluation. At the same time, this method regards the training results obtained by traditional evaluation methods as the key components of training data, and conducts research on the application of data analysis methods on this basis, which is a beneficial supplement and expansion to traditional training evaluation methods.
[0004] The construction of electronic countermeasure training scenarios is difficult, usually relying on training bases for implementation. Limited by the training duration of the bases and resource guarantee conditions, it is difficult to obtain training data. In addition, the process of electronic countermeasure training in a complex electromagnetic environment is affected by many complex and uncertain factors, and electronic countermeasure training data often presents small sample data or small data sets that do not satisfy the normal distribution. Therefore, the electronic countermeasure training evaluation method based on data analysis focuses on the analysis and application of such training data.
[0005] Currently, the application of electronic countermeasure training evaluation methods based on data analysis mainly has the following deficiencies: First, it is often limited to the application of a certain analysis algorithm to a certain type of problem in training evaluation, resulting in a lack of systematic data analysis method guidance for diverse training evaluation topics, making it difficult to support the complex scenarios and multi-dimensional training evaluation of electronic countermeasure training, and the training evaluation benefits are difficult to fully demonstrate; Second, in terms of data utilization, it fails to comprehensively and integratively utilize various types of data such as training condition data, trainee data, and training result data in training, making it difficult to deeply explore the deep laws and potential values of the data, and the decision-making advantages of the data cannot be effectively exerted; Third, it ignores the complex characteristics of electronic countermeasure training data and the requirements of analysis methods for the statistical characteristics of training data such as dimension, quantity, and distribution, resulting in the situation that although a certain algorithm is used in data analysis, the training data does not meet the usage prerequisites of the algorithm, thereby reducing the reliability of the evaluation results.
[0006] "Evaluation of the Effectiveness of Base Training for Electronic Countermeasure Forces Based on TOPSIS Method" constructs an evaluation index system for training effectiveness according to the requirements and characteristics of base training for electronic countermeasure forces, determines the weights of evaluation indicators by combining the entropy weight method, constructs an evaluation model using the TOPSIS method to achieve quantitative evaluation of training effectiveness, and finally obtains the training effectiveness score. "Evaluation Method for the Training Effect of Shipborne Electronic Countermeasures Based on Effectiveness" proposes an idea for constructing an operational effectiveness evaluation index system and a method for expressing training effects of "decomposed based on task requirements" according to the characteristics of shipborne electronic countermeasure professional training. "Evaluation of the Training Effect of the Operational Application of Electronic Countermeasure Equipment Based on Fuzzy Comprehensive Evaluation" starts from the construction requirements of training evaluation indicators, establishes an evaluation index system for the training effect of the operational application of electronic countermeasure equipment, and uses the fuzzy comprehensive evaluation method to establish an evaluation model for the training effect of the operational application of electronic countermeasure equipment. However, the above methods focus on elaborating a comprehensive training evaluation algorithm model, and have not carried out systematic analysis and application research on small sample data or small data sets in electronic countermeasures, and at the same time, they cannot effectively solve the diverse problems in electronic countermeasure training evaluation. Summary of the Invention:
[0007] In order to overcome the above deficiencies, the present invention provides an electronic countermeasure training evaluation method based on data analysis.
[0008] The technical solution adopted by the present invention to solve its technical problems:
[0009] An electronic countermeasure training evaluation method based on data analysis includes the following steps:
[0010] S1. Classification of evaluation issues based on training data analysis;
[0011] The classification of the evaluation issues includes training evaluation issues based on descriptive analysis, training evaluation issues based on confirmatory analysis, and training evaluation issues based on data mining;
[0012] S2. Collection and collation of training data;
[0013] S21. Data collection
[0014] Training data can be divided into three different types: training condition data, trainee data, and training result data; according to the determined classification of evaluation issues, collect electronic countermeasure training data targeted, and the collected training data includes one or more of the above types;
[0015] S22. Data collation
[0016] Training data collation aims at reorganizing and comprehensively integrating different data items according to business logic for different types of collected data based on evaluation topics, and preprocessing the constructed data set through data cleaning, so as to form a data set for specific evaluation topics;
[0017] S3. Matching of data analysis methods;
[0018] By identifying the statistical characteristics of training data, further select appropriate data analysis methods according to the characteristics of the data;
[0019] S4. Result generation and application
[0020] Rely on the data analysis software platform to conduct data analysis, and based on the data analysis results, achieve multi-dimensional evaluation and estimation of military training, providing a basis for accurately evaluating training quality and optimizing training strategies.
[0021] The training evaluation topic based on descriptive analysis is used to analyze or compare the training results of trainees belonging to one or more units;
[0022] The training evaluation topic based on confirmatory analysis is used to verify whether the training effect of the training object during a certain training period meets the expectations, providing a basis for adjusting subsequent training plans; or verifying the effectiveness of training in improving the electronic countermeasure capabilities of multiple training objects, and determining the advantages and disadvantages of training methods;
[0023] Or verifying the differences in competition results among multiple competition objects in a competition, providing support for ranking the places of multiple objects;
[0024] The training evaluation topic based on data mining is used to divide multiple trainees into different groups according to the similarity of training effects, providing a basis for optimizing training design; or studying the correlation between the situations of different trainees and training results, providing a scientific basis for analyzing the training effects of trainees;
[0025] Or distinguishing the training difficulty levels according to the settings of training conditions, so as to comprehensively reflect the training results and training effectiveness of the training object under specific training difficulties.
[0026] The training condition data includes data items such as training subjects, terrain, weather, meteorology, background electromagnetic environment, and electromagnetic confrontation intensity;
[0027] The training object data refers to the basic situations of trainees, including data items such as training units, trainees, educational levels, professional years, post years, operation years, number of major drills, number of base trainings, and post training time;
[0028] The training result data includes data items such as training objects, training time, training subjects, organization methods, and training results;
[0029] Among them, the training object data item in the training result data includes two forms: trainees and training units. The training of trainees is carried out in the equipment operation training stage, and the training of training units is carried out in the tactical confrontation training stage;
[0030] The values of the organizational mode data item in the training result data include daily training, assessment and competition. Daily training focuses on the improvement of the trainees' stage-by-stage ability level during the training process; assessment focuses on the mandatory and normative inspection of the trainees under the specific assessment content and standards; competition focuses on the comparison of the ability levels of multiple trainees under the same competitive conditions.
[0031] In step S2, the data set for a specific assessment topic includes:
[0032] For evaluation topics that analyze or compare the training results of trainees from one or more units, the collated data set is the assessment results of the personnel from one or more training units;
[0033] For the evaluation of whether the training effect of a trainee during a certain period of training meets the expectations, the collated data set is the daily training results of a trainee during a specific period of time;
[0034] In order to evaluate the effectiveness of verification training in improving the electronic warfare capabilities of multiple trainees, the collated data set is the assessment results of multiple trainees before and after training;
[0035] In order to evaluate the differences in competition results between multiple competitors in the verification competition, the collated data set is the competition results of multiple rounds of multiple training objects;
[0036] For the evaluation topic of dividing multiple trainees into different groups according to the similarity of training effects, the sorted data set is the historical data of the assessment results of multiple subjects of multiple trainees;
[0037] For the evaluation topic of studying the correlation between different trainee conditions and training results, the collated data set is an integration of different trainee conditions and trainee assessment results in a unit;
[0038] Regarding the evaluation issue of distinguishing the difficulty levels of training according to the setting of training conditions, the collated data set includes the training condition data in previous trainings and the corresponding difficulty level classification, as well as the training condition data required to determine the training difficulty level.
[0039] In step S3, the training data is divided into samples and populations; the statistical characteristics of the samples include sample type, sample size, and distribution of the population represented by the samples;
[0040] The sample types include single sample, two samples and multiple samples;
[0041] The sample capacity includes large sample data and small sample data;
[0042] The distribution of the population represented by the sample includes known distribution data and unknown distribution data.
[0043] The data analysis methods in step S3 include descriptive statistical analysis of overall data, non-parametric test methods for small samples, and data mining methods for small data sets.
[0044] The descriptive statistical analysis of the overall data is to identify the assessment results of the personnel of one or more training units as overall data, and comprehensively use histograms, data concentration trends, deviation trends, and data distribution measurements to achieve characteristic descriptive analysis of the overall data. The data concentration trends include mean, median, and mode; the deviation trends include standard deviation, variance, and dispersion coefficient; the data distribution measurements include kurtosis and skewness.
[0045] There are three nonparametric test methods for small samples, namely:
[0046] 1) The daily training results of a certain training object in a specific period are identified as a single sample variable, and the Cox-Stuart test is used to implement the trend test of a single variable. The idea of the sign test is used to count the number of times the data rises and falls, and then the corresponding probability distribution is used to determine whether the original hypothesis can be rejected at a given significance level, thereby inferring whether there is a significant trend change in the training effect. This method has no strict requirements on the data distribution form and data sample capacity;
[0047] 2) The assessment results before and after training of multiple trainees are identified as two paired samples, and the Wilcoxon test is used to implement the significant difference test of the two paired samples. For the two paired samples, the difference of each pair of data is first calculated, and then the absolute value of the difference is sorted and assigned a rank, considering the positive and negative signs of the difference, by comparing the positive rank sum and negative rank sum statistics, according to the specific probability distribution, to determine whether there is a significant difference in the overall distribution represented by the two groups of paired data. This method has no strict requirements on the data distribution form and data sample capacity;
[0048] 3) Based on the competition results of multiple rounds of multiple trainees, multiple groups of paired samples are identified, and the Friedman test is used to realize the significant difference analysis of multiple groups of related samples. The observed values of each sample under different treatments are first sorted, and the rank sum of each sample is calculated. By analyzing the difference in the rank sum, it is inferred whether there is a significant difference in the medians of multiple paired samples. This method is suitable for multiple groups of related samples and the data does not require normal distribution. There is no strict requirement on the data sample capacity.
[0049] The data mining method for small data sets is applicable to small data sets, which refers to data sets with a data record volume of dozens to hundreds and an observation dimension of less than ten. There are three data mining methods for small data sets, namely:
[0050] 1) For the historical data of the assessment results of multiple subjects of multiple trainees, it is identified as a data set of multidimensional variables describing multiple objects, and the k-means algorithm is used to implement cluster analysis. First, K initial cluster centers are randomly determined, and then the distances from each data point to these cluster centers are calculated. The data points are assigned to the category where the nearest cluster center is located. Then the mean of each category is recalculated as the new cluster center. The above steps are repeated until the cluster center no longer changes significantly or the set number of iterations is reached. Through this iterative process, the data is divided into different categories, so that the data in the same category has a high similarity and the data between different categories have a large difference;
[0051] 2) The integration of different trainee conditions and trainee assessment results of a certain unit is identified as transactional data, and the Apriori algorithm is used to mine association rules of transactional data sets. The data is first discretized, and based on the prior properties of frequent item sets, all frequent single item sets are first found, and then higher-level frequent item sets are gradually generated through connection and pruning operations. Finally, the association rules are selected according to the minimum support and confidence.
[0052] 3) The training condition data and the corresponding difficulty level divisions in previous trainings are identified as multidimensional data sets with classification labels. Assuming that each feature condition is independent of each other under given category conditions, naive Bayes is used to achieve data classification. Based on Bayes' theorem and the assumption of feature condition independence, the posterior probability of the sample under each category is calculated, and the category with the largest probability is selected as the prediction result. When the training data is limited and it is difficult to obtain a comprehensive combination, the existing qualitative variable characteristics and empirical knowledge can be used to infer and evaluate the training difficulty.
[0053] There are seven situations for generating and applying the result of step S4, which are:
[0054] 1) Generation and application of training evaluation results based on descriptive analysis. Analyze and compare the training assessment results of personnel from one or more training units. With the help of the data analysis platform, generate detailed data statistical charts to describe the central trend, deviation trend and distribution characteristics of the data, and intuitively present the ability level and differences of each object, so as to provide a basis for subsequent training resource allocation, training difficulty adjustment and experience sharing, aiming to promote the balanced improvement of the overall training level and make the training more targeted;
[0055] 2) Generation and Application of Training Evaluation Results Based on Confirmatory Analysis
[0056] 21) Trend Test of a Single Training Data Sample
[0057] For the analysis of the change trend of the daily training results of a single training unit in a specific training course, the Cox-Stuart trend test is carried out with the help of the data analysis platform to provide a basis for judging whether the training effect has been significantly improved, and then assist in deciding whether to adjust the training course and carry out higher-level training, so as to ensure that the training content and difficulty match the improvement of the trainees' abilities and reasonably plan the training progression path;
[0058] 22) Difference Analysis of Two Training Data Samples
[0059] For the comparison of the assessment result differences of the same batch of trainees before and after training, the Wilcoxon signed-rank test is carried out with the help of the data analysis platform to test the effectiveness of the training. Based on the analysis results, successful experiences are summarized, and the key factors leading to the improvement of the training effect are sorted out to provide scientific guidance for the improvement and perfection of the same type of training in the future and promote the continuous improvement of the training quality;
[0060] 23) Difference Analysis of Multiple Training Data Samples
[0061] It is used to compare and rank the competition results of multiple training units in the competition scenario. The Friedman test is carried out with the help of the data analysis platform to provide an objective basis for the evaluation of the competition results;
[0062] 3) Generation and Application of Training Evaluation Results Based on Data Mining
[0063] 31) Cluster Analysis of Training Data
[0064] For the historical assessment results of multiple trainees of the same type, the k-means cluster analysis is realized by using the data analysis platform. For the training levels of different categories of objects, the training design is optimized pertinently to improve the pertinence and efficiency of the training;
[0065] 32) Mining of Association Rules of Training Data
[0066] For the analysis of the association relationship between different personnel conditions and assessment results, the Apriori algorithm analysis is realized by using the data analysis platform. The relevant analysis conclusions provide scientific references for personnel selection, training arrangement and training resource allocation;
[0067] 33) Bayesian Inference of Training Data
[0068] In view of the differences between the training conditions and the historical training conditions, the data analysis platform is used to perform Bayesian inference on the difficulty of the current training conditions. On this basis, the actual effectiveness and quality of the training are accurately measured in combination with the training results.
[0069] Due to the adoption of the technical solution as described above, the present invention has the following advantages:
[0070] The present invention constructs a flexible and highly adaptable systematic data analysis framework, fully considering the high complexity, dynamic uncertainty and derived diverse training evaluation problems in electronic countermeasure training. Through a rich variety of data analysis methods and means, focusing on the small-scale data of electronic countermeasure training, the all-round, in-depth and high-efficiency application of training data is achieved based on the precise matching of algorithms, realizing the objective and accurate evaluation of training evaluation, so as to effectively cope with complex and changeable training scenarios and improve the flexibility, comprehensiveness and accuracy of training evaluation. With the continuous expansion of the depth and breadth of electronic countermeasure training evaluation and the extensive application of data analysis, the training evaluation method based on data analysis has broad development prospects. Brief Description of the Drawings:
[0071] Figure 1 is a flowchart of the present invention;
[0072] Figure 2 is a histogram of the assessment results of three units in Example 1;
[0073] Figure 3 is a time series diagram of training data in Example 2;
[0074] Figure 4 is a Friedman analysis of variance table in Example 3;
[0075] Figure 5 is a window schematic diagram of interactive multiple comparisons;
[0076] Figure 6 is a clustering result silhouette diagram in Example 4;
[0077] Figure 7 is a result diagram of the association rule analysis of influencing factors in Example 5. Detailed Embodiments:
[0078] To comprehensively demonstrate the effectiveness and flexibility of this evaluation method in different training scenarios and evaluation topics, the following representative cases are specifically selected for analysis, and these cases cover seven different specific scenarios.
[0079] (1) Application Example of Training Evaluation Based on Descriptive Analysis
[0080] 1. Evaluation Topic
[0081] During a certain base training, joint training for the same specialty was carried out for multiple units. The focus was on conducting equipment operation assessments on the trainees belonging to three training units in the training area, and it was intended to conduct a comparative analysis of the assessment results of the personnel belonging to different training units.
[0082] 2. Data Collection and Sorting
[0083] A data set was constructed and analyzed according to the assessment results of the personnel belonging to the training units. During the training, the training situations of the 3 participating units were statistically analyzed. The statistical data reflected the assessment results of all the trainees in the participating units, and all the assessment results were obtained under the same training conditions.
[0084] 3. Data Analysis Algorithm Matching
[0085] It was identified as overall data, and exploratory analysis of the training data was intended to be achieved by comprehensively using methods such as central tendency, dispersion tendency, data distribution, and histograms.
[0086] 4. Result Generation and Application
[0087] The algorithm implementation of data analysis was achieved through Excel and Spss. The charts of data analysis are as follows:
[0088] Table 1 Descriptive Statistics of the Assessment Results of the Personnel Belonging to the Three Units
[0089]
[0090] Based on the data in Table 1, comparative analysis can be carried out from three perspectives: central tendency, dispersion tendency, and data distribution of the data:
[0091] (1) In terms of central tendency, looking at the mean, median, and mode, all three values of Unit 3 are small, indicating that in terms of the average level, the ability of Unit 3 is below that of the other two units.
[0092] (2) In terms of dispersion tendency, since the mean and standard deviation are different, only the coefficient of variation can be looked at. The coefficient of variation of Unit 2 is the smallest, indicating that its data changes the least, meaning there are neither particularly excellent nor particularly poor ones.
[0093] (3) In terms of data distribution, the histogram of data distribution shows the distribution patterns of the assessment results of each unit. It can be seen from the figure that the distribution of the assessment results of Unit 3 is wider, which is mutually confirmed with the data distribution characteristics calculated through kurtosis and skewness, further judging that the distribution of the assessment results of the personnel in Unit 3 is wider.
[0094] Combining the three, Unit 3 is generally weaker (although there is a maximum value). Unit 1 and Unit 2 still need to be compared by combining quantile analysis.
[0095] Comparison Results of Quantiles of Assessment Results of Three Units - Percentiles
[0096]
[0097] As can be seen from Table 2, after the 25% quantile, Unit 1 has always been larger than Unit 2. Unit 2 has always been between the other two units with little fluctuation, which also verifies that its coefficient of variation is the smallest. After the 75% quantile, the assessment results of Unit 3 are significantly higher than those of other units, indicating that some of its personnel have high-level combat capabilities, but the overall distribution has a large dispersion (coefficient of variation 0.407).
[0098] Based on the above analysis, the following suggestions can be obtained:
[0099] (1) Unit 1 has a relatively high average ability, and the level in the middle part exceeds that of other units, indicating that its overall ability level is relatively high. It is recommended to appropriately increase the training difficulty for Unit 1 in subsequent training. At the same time, encourage Unit 1 to carry out experience sharing activities within the unit, organize excellent training methods into a booklet for reference by other units, so as to drive the improvement of the overall training level.
[0100] (2) Unit 2 has an average level in the middle, with small fluctuations and relatively balanced ability levels. It is recommended to maintain the existing training strategy in subsequent training, and at the same time communicate with Unit 1 to discover and improve its own deficiencies through comparison.
[0101] (3) Unit 3 has a relatively weak average ability and only exceeds other units at low and high levels, indicating that there are great differences in the adaptability of its personnel to complex electromagnetic environments. It is recommended to increase the training duration and basic skill training in subsequent training. Stratified and grouped training should be carried out according to the personnel's ability levels. For personnel with weak foundations, carry out special basic strengthening, including the explanation of electromagnetic theory knowledge and the training of basic equipment operation specifications; for high-level personnel, carry out advanced training courses.
[0102] (2) Application Example of Training Evaluation Based on Verification
[0103] 1. Trend Test of a Single Training Data Sample
[0104] (1) Assessment Issue
[0105] During a certain base daily training, a unit carried out tactical confrontation training in typical scenarios, aiming to test whether the training level of the trained unit has been significantly improved during the training period, and further provide a basis for whether to adjust the training courses.
[0106] (2) Data Collection and Sorting
[0107] Analyze the changing trend of the training results of the trained unit multiple times, and construct a data set according to the business logic of the daily training results of the trained unit within a certain time sequence during the training period, such asFigure 3 The following is a time series graph of the training data of the trainees in a certain training. The horizontal axis represents the date, and the vertical axis represents the daily training results. It is difficult to intuitively judge the change trend of the training data from the above graph, and the data does not satisfy a certain specific distribution. Regarding the problem of training effect trend detection.
[0108] (3) Data analysis algorithm matching
[0109] Considering that the training data is single-sample data that does not satisfy the normal distribution and the sample size is small, the Cox-Stuart test in non-parametric tests is adopted. This method can test the overall trend of a single sample and has no requirements for the distribution form of the sample.
[0110] (4) Result generation and application
[0111] The specific analysis process is as follows:
[0112] The first step: First, make the null hypothesis, that is, there is no upward trend in the training effect of the trainees. The alternative hypothesis is that there is an upward trend in the training effect.
[0113] The second step: Use the R software to call the pbinom statement for the Cox-Stuart trend test, and get p = 0.01757813.
[0114] The third step: Compare with the given significance level and make a statistical inference: For the given significance level a, if a ≥ 0.02, then reject the null hypothesis and consider that the training effect of the trainees has a significant upward trend.
[0115] It is recommended to further carry out high-level tactical training subjects, such as coordinating and grouping with other specialties under the threat of multiple types of electromagnetic targets, efficiently sharing target information, and real-time collaborative interference strategies to improve the collaborative combat ability.
[0116] 2. Difference analysis of two training data samples
[0117] (1) Evaluation issues
[0118] During a certain base training, the same-field training was carried out for 10 units. According to the training plan, a pre-training assessment was conducted at the beginning of the training and a post-training assessment was carried out at the end of the training. It is intended to test whether this training is effective in improving the overall training level of the participating units through the differences in the assessment results of the 10 units before and after the training.
[0119] (2) Training data collection and collation
[0120] Collect and collate the assessment results of the 10 participating units, and construct a data set according to the business logic of the participating units and their relevant assessment results before and after the training. The assessment results of different units before and after the training are shown in Table 3 below.
[0121] Table 3 Data of Different Trainees Before and After Training
[0122] Trainees Before training After training Absolute difference Sign of difference Rank 1 0.65 0.52 0.13 + 4 2 0.25 0.52 0.27 - 8 3 0.48 0.45 0.03 + 1 4 0.35 0.57 0.22 - 7 5 0.42 0.58 0.16 - 5 6 0.66 0.95 0.29 - 10 7 0.52 0.42 0.1 + 2 8 0.25 0.42 0.17 - 6 9 0.38 0.5 0.12 - 3 10 0.2 0.48 0.28 - 9
[0123] (3) Data Analysis Algorithm Matching
[0124] As shown in the above table, the first column represents the numbers of different trainees. The two training data obtained before and after training in each row of the table are typical paired samples. Since the training data are two related samples, with a small sample size and not meeting the normal distribution, Wilcoxon in non-parametric tests can be used. The Wilcoxon signed-rank test is a test for two paired data samples, used to infer which of the different treatments that generate the two groups of data is better, and has no requirements for the distribution form and quantity of the samples.
[0125] (4) Result Generation and Application
[0126] The specific analysis process is as follows:
[0127] The first step: Make the null hypothesis: The change trend of the overall assessment result has no significant improvement. It can be directly seen from the table that the number of negative signs is more than that of positive signs. Therefore, the alternative hypothesis is that the overall change trend has a significant improvement.
[0128] The second step: Use Matlab to call the signrank function for the signed-rank sum test, and the calculated result is p = 0.0371.
[0129] The third step: Draw a conclusion. That is, for the given significance level a, if a ≥ 0.04, then reject the original hypothesis and consider that the training effect has a significant improvement.
[0130] Combined with the above conclusion in the training summary stage, deeply analyze the key factors leading to the improvement of the training effect, such as the improvement of training methods, the optimal allocation of training resources, the improvement of personnel quality, etc. Further sort out and summarize these successful experiences, so as to provide a scientific basis and guidance for future training.
[0131] 3. Difference Analysis of Multiple Training Data Samples
[0132] (1) Evaluation Issues
[0133] During a certain base training, a combat competition was carried out for multiple units of the same type, and the combat training subjects of four training units were focused on. It is intended to compare the competition results of different trainees during the training to further provide a basis for the ranking of the competition.
[0134] (2) Data Collection and Sorting
[0135] Construct a dataset according to the business logic of the training units and their related competition results. As shown in Table 4.
[0136] Table 4 Training data of four units
[0137]
[0138]
[0139] (3) Data analysis algorithm matching
[0140] This training dataset is a set of paired samples with a small amount of data, and the data distribution does not satisfy the normal distribution. The Friedman test in non-parametric tests can be used to statistically infer whether there are significant differences in the training effects of different trainees. The Friedman test is a test for multiple paired samples, used to infer whether there are significant differences in the medians of multiple paired samples, and there is no requirement for the data to be normally distributed.
[0141] (4) Result generation and application
[0142] First step: Make the original hypothesis, that is, there is no difference in the competition results of the four units, and take the significance level a = 0.05. The Matlab statistical toolbox provides the Friedman function for the test, which can generate an analysis of variance table. In Matlab, call this function for simulation analysis, and the result is as Figure 4 shown. From the result, p = 0.0434 is obtained, indicating that the original hypothesis can be rejected at the significance level of 0.05, and it is considered that there are significant differences in the assessment results of the four units.
[0143] Second step: Matlab provides the multcompare function for multiple comparisons. This function can generate an intuitive interactive graphic window. In the window, click on any selected line segment, and the selected line segment will turn blue. If other line segments are red, it indicates that the difference from the selected line segment is significant. Call the multcompare function in Matlab for multiple comparisons, and the final result is as follows Figure 5 shown. It can be seen that at the significance level of 0.05, the difference in the competition results between the first unit and the fourth unit is significant, and the differences between the remaining units are not significant.
[0144] Taken together, based on the existing data and analysis results, a rough ranking is made as follows: Unit 1 wins the first place, Unit 2 and Unit 3 respectively win the second or third place (since the difference between them is not significant, the ranking can be fine-tuned according to other comprehensive factors), and Unit 4 wins the fourth place.
[0145] (3) Application example of training evaluation based on data mining
[0146] 1. Cluster Analysis of Training Data
[0147] (1) Evaluation of Issues
[0148] During a certain base training, joint training for the same specialty was carried out for multiple units of the same type. Before task planning, a training plan was formulated based on the ability characteristics of 8 units. It was intended to reasonably distinguish the training units through the training data of 3 training courses of 8 units in the past, so as to optimize the training design pertinently and improve the pertinence and training efficiency of training.
[0149] (2) Data Collection and Sorting
[0150] Construct a data set according to the business logic of the assessment results of different training courses of different training units, as shown in the following table.
[0151] Table 5 Cluster Analysis of Training Data
[0152] Trainee serial number Subject A Subject B Subject C 1 0.47 0.69 0.60 2 0.72 0.48 0.58 3 0.75 0.58 0.58 4 1.00 0.53 0.57 5 0.40 0.81 0.41 6 0.40 0.42 0.68 7 0.39 0.45 0.26 8 0.82 0.29 0.45
[0153] (3) Matching of Data Analysis Algorithms
[0154] Since the data in training targets a relatively small data set, and considering that the influence degree of feature dimensions on clustering is roughly the same, the k-means algorithm is used to implement cluster analysis, and the K-means clustering method of the partitioning method is adopted to realize the clustering of training levels.
[0155] (4) Result Generation and Application
[0156] Use Matlab to implement k-means cluster analysis, and its analysis process is as follows:
[0157] The first step: First, read the data, standardize the data, and then select the initial condensation points for clustering. It is calculated that the ability levels of 8 units can be divided into three groups, as shown in Table 6. Among them, the units classified into the same group tend to be similar in terms of training level in a certain sense.
[0158] Table 6 Cluster Results of Training Levels
[0159] Unit serial number Horizontal clustering 1 2 2 1 3 1 4 1 5 2 6 1 7 1 8 3
[0160] The second step: Draw a silhouette plot according to the clustering results, and observe whether the clustering is reasonable from the silhouette plot.
[0161] The value range of the silhouette value S(i) is [-1, 1]. The larger the value of S(i), the more reasonable the clustering of the i-th point. When S(i)<0 , it indicates that the clustering of the i-th point is unreasonable. The silhouette plot of the above example is as Figure 6As shown, when these data are divided into three categories, the silhouette value of each observation is positive and above 0.2, indicating that it is appropriate to divide these observations into three categories.
[0162] By observing the clustering results, it can be found that the first category of units (unit numbers 2, 3, 4, 6, 7) has relatively good overall balance. From the results of the three training subjects given, there are no obvious short boards in these units for each subject; the comprehensive level is moderately above average, and the numerical values of all results are generally in a relatively high range. In the subsequent training design, appropriately adjust the complexity of the environment to increase the training difficulty of each subject.
[0163] For the second category of units (unit numbers 1, 5), subject B has outstanding advantages. It can be clearly seen that the results of subject B are significantly higher than those of other units. Relatively speaking, the results of subject A and subject C are slightly inferior. In the subsequent training design, focus on making up for the short boards of subjects A and C, and sort out the successful training experience for subject B and share it with other units.
[0164] For the third category of units (unit number 8), the result imbalance is significant. This unit has significant advantages in subject A, and subjects B and C need to be strengthened, showing an obvious imbalance. In the future, it is necessary to focus on training the weak subjects B and C.
[0165] 2. Mining Association Rules of Training Data
[0166] (1) Evaluation Issues
[0167] A certain unit has participated in base training in multiple batches over the years, and has focused on carrying out basic training on interference capabilities for its affiliated personnel. It is intended to analyze and find the correlation between different situations of personnel and interference effects, so as to provide a scientific basis for the analysis of assessment results.
[0168] (2) Data Collection and Sorting
[0169] In order to summarize the hidden laws in the training data, a data set is constructed according to the business logic of the basic situation of personnel and their relevant assessment results, recording 8 data items such as educational level (x1), professional years (x2), post years (x3), operation years (x4), number of major drills (x5), number of base trainings (x6), post training time (x7) and interference effect (y). Among them, the data items (x1~x7) can be regarded as the influencing factor set of the interference effect (y), and the interference effect (y) is reflected by the assessment results of the interference subject.
[0170] (3) Data Analysis Algorithm Matching
[0171] Identified as transactional data, considering that there is no hierarchical relationship among the training data variables and the dataset size is small, the Apriori algorithm is used in this example. The Apriori algorithm is the most commonly used and classic algorithm for mining frequent item sets in association rules, which can discover interesting associations or correlations among item sets in a large amount of data.
[0172] (4) Result generation and application
[0173] The Apriori algorithm analysis is implemented based on R software. The analysis process is as follows:
[0174] Step 1: Discretize the above original data according to certain rules. The discretization rules are shown in Table 7, and the discretized results are shown in Table 8.
[0175] Table 7 Data discretization results
[0176]
[0177] Table 8 Data instances for association rule analysis of influencing factors
[0178]
[0179] Step 2: The obtained data has 8 variables. Since there is no hierarchical relationship among the variables, the Apriori algorithm is used to perform association analysis on the training data. Apriori is the most commonly used and classic algorithm for mining frequent item sets in association rules. Its core idea is to generate candidate items and their support degrees through connection, and then generate frequent item sets through pruning, so as to discover interesting associations or correlations among item sets in a large amount of data. Analysis process: Use R software to read the preprocessed data and call the association rule package for association rule mining. To ensure the practicality of the association rules, combined with the actual problems concerned in data analysis, the minimum support degree of the rules is specified as 10%, the minimum confidence degree is 60%, the lift is greater than 1, and the antecedent of the generated rules is the influencing factor, and the consequent of the rules is the assessment result.
[0180] Table 9 Association rule analysis results
[0181]
[0182] The following three association rules are obtained from Table 9:
[0183] Rule 1: There is a 71% confidence level to believe that the assessment result is good when the number of times a person is trained in the base is large, and the applicability of this association rule is 26%;
[0184] Rule 2: There is an 80% confidence level to believe that the assessment result is good when a person has a high educational level and a long professional tenure, and the applicability of this association rule is 21%;
[0185] Rule 3: There is an 83% confidence level to believe that the assessment results are good when the personnel have a high cultural level and a large number of major drills. The applicability of this association rule is 26%.
[0186] Step 3: Rule visualization. As Figure 7 shown, the antecedent and consequent of the association rule are represented by a polyline with an arrow from left to right. The thickness of the polyline represents the magnitude of the rule support, and the depth of the gray scale represents the height of the lift.
[0187] Based on the association rules between the personnel status and the assessment results, optimize in terms of personnel selection and training. For example, for personnel with fewer base training times, their base training arrangements can be appropriately increased. Given the positive impact of the base training times on the results, if training resources permit, the frequency of base training can be appropriately increased, the content and scale of base training can be expanded, and at the same time, the curriculum settings of base training can be optimized to make it more targeted and effective; for personnel with a high cultural level, focus on training them on professional positions and increase their opportunities to participate in major drills to improve their assessment results and combat capabilities. At the same time, establish long-term tracking and feedback on the personnel status and assessment results. Through long-term tracking, analyze which factors play a key role in the long-term development of personnel and which association rules will change in applicability at different stages. Further, adjust the personnel training path according to the feedback information from these long-term trackings.
[0188] 3. Bayesian Inference of Training Data
[0189] (1) Evaluation Issues
[0190] In a certain base training, considering that there are differences in the training condition settings compared with previous training conditions, it is planned to classify the difficulty level of the training conditions this time to accurately analyze the assessment results of the training objects under specific conditions.
[0191] (2) Data Collection and Sorting
[0192] Combined with the characteristics of electronic countermeasure technology, terrain, weather, meteorology, background electromagnetic environment, and electromagnetic countermeasure intensity are regarded as the key focuses of attention and identified as the key uncertain factors of training conditions. A conditional factor set X = (X1, X2, X3, X4, X5) is set. Different states are taken for X according to the meaning of the variables. X1 represents terrain variables, with 1 representing mountains, 2 representing hills, 3 representing plateaus, and 4 representing cities; X2 represents weather variables, with 1 representing daytime and 2 representing night; X3 represents climate variables, with 1 representing sunny, 2 representing rainy, 3 representing snowy, and 4 representing foggy; X4 represents the background electromagnetic environment, with 1 representing simple, 2 representing medium, and 3 representing complex; X5 represents the intensity of electromagnetic countermeasures, with 1 representing low intensity, 2 representing medium intensity, and 3 representing high intensity. C represents the training difficulty, with 1 representing primary difficulty, 2 representing intermediate difficulty, and 3 representing advanced difficulty. The training condition data and difficulty levels of previous trainings are used as the constructed data set, as shown in Table 10.
[0193] Table 10 Partial Data Examples
[0194] <![CDATA[X1]]> <![CDATA[X2]]> <![CDATA[X3]]> <![CDATA[X4]]> <![CDATA[X5]]> C 2 1 1 3 2 3 1 1 2 3 3 2 4 2 1 1 1 1 2 1 2 2 2 2 1 2 1 1 1 2 3 1 3 1 1 2 …… …… …… …… …… ……
[0195] (3) Data Analysis Algorithm Matching
[0196] The above data set is identified as a multi-dimensional data set with classification labels. Due to limited actual training time and resources, it is difficult to obtain comprehensive combinations of different factors under different conditions and difficult to obtain comprehensive data. At the same time, considering the small scale of the data set, a Bayesian classifier is created and trained for the above data set, and it is assumed that each feature condition is independent given the category condition.
[0197] (4) Result Generation and Application
[0198] Based on the python software platform, the NB classification program is implemented. First, necessary libraries and functions are imported, the data set is loaded and divided into a training set and a test set. Then, a MultinomialNB classifier is created and trained. On this basis, the trained classifier is used for prediction and the performance of the model is evaluated, that is, the accuracy is calculated and the classification report is printed. Finally, for the data with a given unknown category, the trained classifier is used to infer the new data and the probability of each category is output. According to the results shown by the code running, the overall accuracy of the model is 0.93, indicating that the classifier has good performance overall. The specific classification report is shown in Table 11 below. The report results further confirm the classification performance. For Bayesian inference on a given set of data (2, 1, 3, 2, 2), the posterior probabilities of the subject difficulty are 0.285, 0.572, and 0.143 respectively. According to the maximum posterior criterion, medium is selected as the class label for this instance, which is consistent with the actual judgment.
[0199] Further combine the training difficulty level with the training results to objectively and accurately measure the actual effectiveness and quality of the training. In addition, in the long-term tracking and analysis of the training effect, incorporate the training difficulty level as a key variable into the consideration. By establishing a long-term training database, record the training results under different difficulty levels.
[0200] Table 11 Classification Report Table
[0201] precision recall f1-score support class1 1.00 1.00 1.00 5 class2 1.00 0.83 0.91 6 class3 0.80 1.00 0.89 4 accuracy 0.93 15 macro avg 0.93 0.94 0.93 15 weighted avg 0.95 0.93 0.93 15
[0202] The parts not described in detail in the above content are prior arts, so no detailed description is provided.
Claims
1. An electronic countermeasure training evaluation method based on data analysis, characterized in that: It includes the following steps: S1. Evaluation issue classification based on training data analysis; The classification of the evaluation issues includes training evaluation issues based on descriptive analysis, training evaluation issues based on confirmatory analysis, and training evaluation issues based on data mining; S2. Training data collection and collation; S21. Data collection The training data can be divided into three different types: training condition data, trainee data, and training result data; According to the determined evaluation issue classification, electronic countermeasure training data is collected specifically, and the collected training data includes one or more of the above types; S22. Data collation The training data collation is carried out for different types of collected data. According to the evaluation issues, different data items are recombined and comprehensively integrated according to the business logic, and the constructed data set is preprocessed through data cleaning, so as to form a data set for specific evaluation issues; S3. Matching of data analysis methods; By identifying the statistical characteristics of the training data, a suitable data analysis method is further selected according to the characteristics of the data; S4. Result generation and application Relying on the data analysis software platform for data analysis, multi-dimensional evaluation and estimation of military training are realized based on the data analysis results, providing a basis for accurately evaluating training quality and optimizing training strategies.
2. The electronic countermeasure training evaluation method based on data analysis according to claim 1, characterized in that: The training evaluation issue based on descriptive analysis is used to analyze or compare the training results of trainees belonging to one or more units; The training evaluation issue based on confirmatory analysis is used to verify whether the training effect of the trainee meets the expectations during a certain training period, providing a basis for adjusting the subsequent training plan; or verifying the effectiveness of the training in improving the electronic countermeasure capabilities of multiple trainees, and judging the quality of the training method; Or verifying the difference in competition results among multiple competition participants in the competition, providing support for ranking the places among multiple participants; The training evaluation issue based on data mining is used to divide multiple trainees into different groups according to the similarity of training effects, providing a basis for optimizing training design; Or studying the correlation between the conditions of different trainees and training results, providing a scientific basis for analyzing the training effects of trainees; Or distinguishing the training difficulty levels according to the settings of training conditions, so as to comprehensively reflect the training results and training effectiveness of trainees under specific training difficulties.
3. The electronic countermeasure training evaluation method based on data analysis according to claim 1, wherein: The training condition data includes data items such as training subjects, terrain, weather, meteorology, background electromagnetic environment, and electromagnetic confrontation intensity; The trainee data refers to the basic situation of the trainees, including data items such as the training unit, trainees, educational level, professional years, post years, operation years, number of major drills, number of base trainings, and post training time; The training result data includes data items such as trainees, training time, training subjects, organization methods, and training results; Among them, the trainee data item in the training result data includes two forms: trainees and training units. The training of trainees is carried out in the equipment operation training stage, and the training of training units is carried out in the tactical confrontation training stage; The values of the organizational mode data item in the training result data include daily training, assessment and competition. Daily training focuses on the improvement of the trainees' stage-by-stage ability level during the training process; assessment focuses on the mandatory and normative inspection of the trainees under the specific assessment content and standards; competition focuses on the comparison of the ability levels of multiple trainees under the same competitive conditions.
4. The electronic countermeasure training evaluation method based on data analysis according to claim 3, characterized in that: In step S2, the data set for a specific evaluation topic includes: an evaluation topic for analyzing or comparing the training results of trainees belonging to one or more units, and the collated data set is the assessment results of the personnel belonging to one or more training units; For the evaluation of whether the training effect of a trainee during a certain period of training meets the expectations, the collated data set is the daily training results of a trainee during a specific period of time; In order to evaluate the effectiveness of verification training in improving the electronic warfare capabilities of multiple trainees, the collated data set is the assessment results of multiple trainees before and after training; In order to evaluate the differences in competition results between multiple competitors in the verification competition, the collated data set is the competition results of multiple rounds of multiple training objects; For the evaluation topic of dividing multiple trainees into different groups according to the similarity of training effects, the sorted data set is the historical data of the assessment results of multiple subjects of multiple trainees; For the evaluation topic of studying the correlation between different trainee conditions and training results, the collated data set is an integration of different trainee conditions and trainee assessment results in a unit; Regarding the evaluation issue of distinguishing the difficulty levels of training according to the setting of training conditions, the collated data set includes the training condition data in previous trainings and the corresponding difficulty level classification, as well as the training condition data required to determine the training difficulty level.
5. The electronic countermeasure training evaluation method based on data analysis according to claim 1, characterized in that: In step S3, the training data is divided into samples and populations; the statistical characteristics of the samples include sample type, sample size, and distribution of the population represented by the samples; The sample types include single sample, two samples and multiple samples; The sample capacity includes large sample data and small sample data; The distribution of the population represented by the sample includes known distribution data and unknown distribution data.
6. The method for evaluating electronic countermeasure training based on data analysis according to claim 5, wherein: The data analysis methods in step S3 include descriptive statistical analysis of overall data, non-parametric test methods for small samples, and data mining methods for small data sets.
7. An electronic countermeasure training evaluation method based on data analysis according to claim 6, characterized in that: The descriptive statistical analysis of the overall data is to identify the assessment results of the personnel of one or more training units as overall data, and comprehensively use histograms, data concentration trends, deviation trends, and data distribution measurements to achieve characteristic descriptive analysis of the overall data.
8. The electronic countermeasure training evaluation method based on data analysis according to claim 6, characterized in that: There are three nonparametric test methods for small samples, namely: 1) The daily training results of a certain trainee in a specific period of time are identified as a single sample variable, and the Cox-Stuart test is used to implement the trend test of a single variable. This method has no strict requirements on the distribution form and quantity of data; 2) For the assessment results before and after training of multiple trainees, they are identified as two paired samples. The Wilcoxon test is used to achieve the significance difference test of the two paired samples. This method has no strict requirements for the data distribution form and the data sample size; 3) For the competition results of multiple trainees in multiple rounds, they are identified as multiple groups of paired samples. The Friedman test is used to achieve the significance difference analysis of multiple groups of related samples, which is applicable to the situation of multiple groups of related samples and has no requirement for normal distribution of data. This method has no strict requirements for the data sample size.
9. An electronic countermeasure training evaluation method based on data analysis according to claim 6, characterized in that: The data mining method for the small dataset is applicable to the small dataset, where the small dataset refers to the data record quantity ranging from dozens to hundreds, and the observation dimension is within ten; There are three data mining methods for the small dataset, which are respectively: 1) For the historical data of the assessment results of multiple subjects of multiple trainees, it is identified as a dataset describing multiple objects with multi-dimensional variables, and the k-means algorithm is used to achieve clustering analysis; 2) For the integration of the status of different trainees and the assessment results of trainees in a certain unit, it is identified as transactional data, and the Apriori algorithm is used to achieve the association rule mining of the transaction dataset; 3) For the training condition data and the corresponding difficulty level division in previous trainings, it is identified as a multi-dimensional dataset with classification labels. Assuming that each feature condition is independent under the given category condition, the naive Bayes is used to achieve data classification.
10. The electronic countermeasure training evaluation method based on data analysis according to claim 1, characterized in that: There are seven situations for the result generation and application in step S4, which are respectively: 1) Generation and application of training evaluation results based on descriptive analysis For the analysis and comparison of the training assessment results of the personnel belonging to one or more training units, detailed data statistical charts are generated with the help of the data analysis platform, realizing the description of the central tendency, dispersion tendency and distribution characteristics of the data, intuitively presenting the ability levels and differences of each object, so as to provide a basis for subsequent training resource allocation, training difficulty adjustment and experience sharing, aiming to promote the balanced improvement of the overall training level and make the training more targeted; 2) Generation and application of training evaluation results based on confirmatory analysis 21) Trend test of a single training data sample For the analysis of the change trend of the daily training results of a single training unit in a specific training subject, the Cox-Stuart trend test is carried out with the help of the data analysis platform to provide a basis for judging whether the training effect is significantly improved, and then assist in deciding whether to adjust the training subject and carry out higher-level training, ensuring that the training content and difficulty match the ability improvement of the trainees, and reasonably planning the training progression path; 22) Difference analysis of two training data samples For the comparison of the assessment result differences of the same batch of trainees before and after training, the Wilcoxon signed-rank test is carried out with the help of the data analysis platform to test the effectiveness of the training. Based on the analysis results, successful experiences are summarized, and the key factors leading to the improvement of the training effect are sorted out, providing scientific guidance for the improvement and perfection of the same type of training in the future, and promoting the continuous improvement of the training quality; 23) Difference analysis of multiple training data samples Used to compare and rank the competition results of multiple training units in the combat competition scenario, conduct Friedman tests with the help of the data analysis platform, and provide an objective basis for the evaluation of competition results; 3) Generation and application of training evaluation results based on data mining 31) Cluster analysis of training data For the historical assessment results of multiple trainees of the same type, use the data analysis platform to implement k-means cluster analysis, and optimize the training design for different categories of trainees according to their training levels, so as to improve the pertinence and effectiveness of training; 32) Mining of association rules for training data For the analysis of the association relationship between different personnel conditions and assessment results, use the data analysis platform to implement Apriori algorithm analysis, and the relevant analysis conclusions provide scientific references for personnel selection, training arrangement, and training resource allocation; 33) Bayesian inference of training data For the differences between the current training conditions and the historical training conditions, use the data analysis platform to conduct Bayesian inference on the difficulty of the current training conditions, and on this basis, accurately measure the actual effectiveness and quality of training in combination with the training results.