Learning method for passive fault tolerance and active error correction of multi-dimensional sample features
By using the learning methods of passive fault tolerance and active error correction in machine learning in machine learning, the abnormal data is processed, and the problem of serious impact of outliers in the existing technology is solved, and the fault tolerance and accuracy of the learning method is improved.
Patent Information
- Application Number
- CN202510486947.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-06-06
AI Technical Summary
Existing machine learning methods lack effective fault tolerance when dealing with outliers, resulting in severe impact on learning processes and results, especially in big data environments.
A learning method for passive fault tolerance and active error correction of multidimensional sample features is proposed. By dividing the samples into training sets and test sets, passive fault tolerance learning and active error correction testing are adopted, and abnormal data is handled using statistical features and fault tolerance mechanisms.
The fault tolerance of learning methods in the presence of exception samples is realized, learning errors caused by abnormal data are avoided, and the reliability and accuracy of machine learning are improved.
Smart Images

Figure CN120106249A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine fault-tolerant learning, and in particular to a learning method for passive fault tolerance and active error correction of multi-dimensional sample features. Background Art
[0002] Machine learning has been a hot topic in the field of artificial intelligence in recent years. Since the first international seminar on machine learning was held at Carnegie Mellon University in the United States in 1980, machine learning research has been booming around the world. In 1983, Michalski, Carbonell and Mitchell edited Machine Learning: An Artificial Intelligence Approach, which divided machine learning research into different types such as "learning from examples", "learning in problem solving and planning", "learning through observation and discovery" and "learning from instructions", and systematically summarized the progress of machine learning research at that time; in 1984, the international journal Machine Learning was founded, which strongly promoted the rapid development of machine learning research; in 1990, Carbonell published Machine Learning: Paradigms and Methods, which was subsequently widely used in different technical fields such as natural language understanding, non-monotonic reasoning, machine vision, and pattern recognition. In the past 30 years, machine learning has become one of the important ways to realize artificial intelligence, especially in the environment of big data. How to effectively use information and obtain hidden, effective and understandable knowledge from it has become the key to affecting the application of intelligent technology.
[0003] Feature learning, also known as representation learning, is one of the key areas of machine learning, especially in deep learning. The purpose of feature learning is to automatically learn effective feature representations from sample data, which can better help subsequent machine learning algorithms (such as classification, regression, clustering, etc.) to process tasks.
[0004] Feature learning, whether learning features from examples or discovering features through observation, cannot be separated from data or samples. However, due to the influence of various uncertain factors such as sampling errors, recording errors or environmental mutations, the sample data used for machine learning inevitably has outliers, especially big data. The existence of abnormal samples will not only affect the learning process but also the learning results, and even lead to serious deviations or errors. As pointed out by Widmann et al (2018) and Santoyo et al (2017), for machine learning, wild values or outliers may be a serious problem when training algorithms, which will seriously affect the reliability of the learning algorithm and affect the correctness of the learning results. Therefore, existing learning methods must clean the sample data before performing machine learning on the sample data, which is a very laborious task that affects the real-time nature of learning.
[0005] Unfortunately, there is no universal solution for identifying outliers, let alone a machine learning solution with fault tolerance. Some of the existing learning models and abnormal data cleaning methods require the number of outliers or the proportion of abnormal data to be known in advance, while others require the statistical distribution model and model parameters that the sample population obeys to be known in advance. Even so, it is difficult to avoid the adverse effects of abnormal data. For example, multivariate linear regression and K-means clustering lack the ability to tolerate abnormal data. Even a small amount of abnormal data in the sample will lead to regression offset and clustering bias. Summary of the invention
[0006] The purpose of the present invention is to provide a learning method for passive fault tolerance and active error correction of multi-dimensional sample features. It is a machine learning method with fault tolerance that is not adversely affected by abnormal data in the sample set. It will not "learn wrong" knowledge due to abnormal samples that may exist in the training sample set, nor will it give wrong test conclusions due to abnormal samples that may exist in the test sample set.
[0007] To achieve the above object, the present invention provides a learning method for passive error tolerance and active error correction of multi-dimensional sample features, comprising the following steps:
[0008] Step S1: Divide the samples into training set S I and the test set S II ;
[0009] Step S2, passive fault-tolerant learning of statistical characteristics of multi-dimensional samples of impurity training set;
[0010] Step S3, using the fault tolerance test and outlier correction of the test set with impurities;
[0011] Step S4: Determine the test result.
[0012] Preferably, in step S1, the training set S I For training and testing the learning model II Used for testing and validating learning models.
[0013] Preferably, in step S2, for the training set S formed by n p-dimensional samples I ={X i ∈R p , i=1,2,...,n}, where S represents the sample set, X i represents the i-th sample, R p represents p-dimensional real number space; training set S I The statistical characteristic of is the mathematical expectation value μ∈R of the matrix p and the scatter matrix ∑∈R p×p , including parameters, among which, Indicates the number of combinations;
[0014] A passive fault-tolerant learning method based on multi-dimensional sample statistical characteristics is established, specifically, an M-learning mechanism is constructed:
[0015]
[0016] Where, d 2 (X,μ,Σ)=(X-μ) τ Σ -1 (X-μ); {φ μ ,φ η} is in R + = is a real function pair on [0, +∞), μ is the first-order moment of the sample, i.e., the mathematical expectation, ∑ represents the covariance of the mother population corresponding to the sample, or the second-order moment, and ∧ represents the statistical estimator.
[0017] Preferably, in step S2, a binary group of the following form is constructed:
[0018]
[0019] In the formula, r 0 and c are adjustable parameters, the default value is r 0 =3p,c=3r 0 ;
[0020] Perform machine learning on the multidimensional feature quantity (μ, Σ) of the sample, denoted by X i ∈R p Component form (x 1 (t i ),...,x p (t i )) τ and median operator , construct an iterative calculation method:
[0021]
[0022] And select the initial value of the iteration:
[0023]
[0024] Through iterative calculation, the optimal passive fault tolerance learning is performed on the multi-dimensional feature quantity (μ, Σ) of the sample to obtain the optimal fault tolerance estimator
[0025] Preferably, in step S3, an elimination strategy or an avoidance strategy is adopted to convert the test set S II Partitioned into support set S ⅡS and the exclusion set S ⅡA Two parts, construct the repulsion index:
[0026]
[0027] By I IIA The size of determines the degree of fit between the learning model and the test set samples;
[0028] Exclusion set determination:
[0029] For any sample X i ∈S II , when it follows the normal distribution N(μ II ,Σ II ) is a valid sample, then Therefore, the sample mean and scatter matrix estimates obtained by learning the training set are Construct a confidence ellipsoid:
[0030]
[0031] In the formula is 2 -(1-α)×100% quantile of the distribution;
[0032] According to formula (6), for the test set S II The support set S ⅡS and the exclusion set S ⅡA Division:
[0033] S IIS =S II ∩Q II ,S IIA =S II -S IIS (7);
[0034] Get the optimal fault-tolerant learning of multivariate sample feature quantity (μ, Σ)
[0035] The optimal fault-tolerant learning algorithm is given a test set S II The compliance ratio I IIS and the rejection ratio I IIA They are:
[0036]
[0037] Preferably, in step S3, the influence support set S ⅡS and the exclusion set S ⅡA The important factor for division is the learning result of the training set feature parameters With the test set S Ⅱ Whether the sample mean and scatter matrix are consistent;
[0038] For two sets of samples from the training set and the test set and Two groups of samples construct the null hypothesis H 0 :
[0039] H 0 :μ=μ II ,Σ=Σ II (9);
[0040] remember and The sample sizes are n and m, and the means are X a and Construct the null hypothesis H 0 Likelihood ratio test statistic for :
[0041]
[0042] Constructing Negative Domains
[0043] When the statistic When , we accept the null hypothesis H with (1-α)×100% confidence. 0 .
[0044] Preferably, in step S3, the training set is recorded and test set The sample covariance matrices are and The likelihood ratio test statistic is obtained as follows:
[0045]
[0046] Error-tolerant statistics using covariance matrix and Substituting the sample covariance in equation (11), we obtain the modified λ which is tolerant to abnormal data. 4 Test statistic:
[0047]
[0048] When the statistic If , we accept the null hypothesis H 0 .
[0049] Preferably, in step S3, an abnormal data error correction and repair algorithm is established in the test set to correct any X i ∈S IIA , establish a fault-tolerant proportional compression method:
[0050]
[0051] Any X i ∈S IIA Both Therefore, the compression transformation formula (13) compresses all samples in the exclusion set back to the set S IIS In the process, error correction and repair are performed on abnormal samples.
[0052] Therefore, the present invention adopts the above-mentioned learning method of passive fault tolerance and active error correction of multi-dimensional sample features, and the beneficial effects are as follows:
[0053] The present invention establishes a fault-tolerant machine learning method that is not adversely affected by abnormal data in the sample set. It will not "learn wrong" knowledge due to abnormal samples that may exist in the training sample set, nor will it give wrong test conclusions due to abnormal samples that may exist in the test sample set.
[0054] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is an overall flow chart of an embodiment of a learning method for passive error tolerance and active error correction of multi-dimensional sample features of the present invention;
[0056] Figure 2 It is a schematic diagram of sample dispersion and error ellipsoid for two situations, without and with outliers, of an embodiment of a learning method for passive fault tolerance and active error correction of multi-dimensional sample features of the present invention, wherein (a) is the learning result of 100 samples in a training set without outliers, (b) is the conventional learning result when the training set has 4 outliers, (c) is the fault-tolerant learning result when the training set has 4 outliers, and (d) is the fault-tolerant learning result when the test set has 4 outliers. DETAILED DESCRIPTION
[0057] The technical solution of the present invention is further described below through the accompanying drawings and embodiments.
[0058] Unless otherwise defined, technical or scientific terms used in the present invention shall have the common meanings understood by one having ordinary skills in the field to which the present invention belongs.
[0059] A learning method for passive error tolerance and active error correction of multi-dimensional sample features, comprising the following steps:
[0060] Step S1: Considering the existing machine learning process, the samples are usually divided into training set S I and the test set S II Two parts, training set S I For training and testing set S of learning model II Therefore, the present invention specifically includes two parts: a passive fault-tolerant learning step S2 for the training set and an active fault-tolerant learning step S3 based on the test set.
[0061] Step S2: Passive fault-tolerant learning of statistical characteristics of multi-dimensional samples of impurity training set; for the training set S formed by n p-dimensional samples I ={X i ∈R p , i=1,2,...,n}, where S represents the sample set, X i represents the i-th sample, R p represents p-dimensional real number space; training set S I The important statistical feature of the matrix is the mathematical expectation value μ∈R p and the scatter matrix ∑∈R p×p , including parameters, among which, Indicates the number of combinations.
[0062] The present invention establishes a passive fault-tolerant learning method for multi-dimensional sample statistical features. The advantage of this method is that it can effectively suppress the adverse effects of abnormal data and improve the reliability of feature learning. Specifically, an M-learning mechanism is constructed:
[0063]
[0064] Where, d 2 (X,μ,Σ)=(X-μ) τ Σ -1 (X-μ); {φ μ ,φ η} is in R + = is a real function pair on [0, +∞), μ is the first-order moment of the sample, i.e., the mathematical expectation, ∑ represents the covariance of the mother population corresponding to the sample, or the second-order moment, and ∧ represents the statistical estimator.
[0065] The M-learning mechanism is a collection of different learning methods, including at least the commonly used maximum likelihood method and second-order moment estimation method. μ ,φ η Different selection methods can produce different learning results.
[0066] However, both the maximum likelihood estimation and the second-order moment estimation lack the ability to tolerate abnormal samples. I There are abnormal samples in the above learning model, and the estimated value of the feature quantity (μ, Σ) given by It will deviate from the true value and form erroneous knowledge cognition.
[0067] In order to select a learning mechanism with fault tolerance from the M-learning mechanism, the present invention constructs a tuple of the following form:
[0068]
[0069] In the formula, r 0 and c are adjustable parameters, the default value is r 0 =3p,c=3r 0 .
[0070] Based on the M-learning mechanism of the function binary shown in formula (2), the training set S I The abnormal information brought by the abnormal samples in the is fault-tolerant. When actually conducting machine learning on the multidimensional feature quantity (μ, Σ) of the sample, let X i ∈R p Component form (x 1 (t i ),...,x p (t i )) τ and median operator Construct an iterative calculation method:
[0071]
[0072] And select the initial value of the iteration:
[0073]
[0074] Through iterative calculation, the optimal passive fault-tolerant learning of the sample multidimensional feature quantity (μ, Σ) can be achieved to obtain the optimal fault-tolerant estimator
[0075] The advantage of the above learning method is that in the training set S I When it contains outliers or abnormal data, the fault-tolerant learning of these features can ablate the set S IThe adverse effects of outliers and abnormal data that may be included ensure that the learning results will not be distorted by the influence of outliers and abnormal data.
[0076] Step S3, using the fault tolerance test and outlier correction of the test set with impurities;
[0077] The acquisition of a high-quality learning algorithm is inseparable from both the training set and the test set. The training set is mainly used to train the model and determine the parameters to form a knowledge learning method; the test set is used to test the learning method and its generalization ability. The test set plays an important role in the formation of the learning model and the application and promotion of the learning method.
[0078] If there are wrong samples in the test set, such as outliers or spots, and the wrong samples are used to test the machine learning method, the test results are likely to be biased. To solve this problem, a elimination strategy or avoidance strategy is adopted to reduce the test set S II Partitioned into support set S ⅡS and the exclusion set S ⅡA Two parts, construct the repulsion index:
[0079]
[0080] By I IIA The size of determines the degree of fit between the learning model and the test set samples;
[0081] Exclusion set determination:
[0082] For any sample X i ∈S II , when it follows the normal distribution N(μ II ,Σ II ) is a valid sample, then Therefore, the sample mean and scatter matrix estimates obtained by learning the training set are Construct a confidence ellipsoid:
[0083]
[0084] In the formula is 2 -(1-α)×100% quantile of the distribution;
[0085] According to formula (6), we can implement the test set S II The support set S ⅡS and the exclusion set S ⅡA Division:
[0086] S IIS =S II ∩Q II ,SIIA =S II -S IIS (7);
[0087] And get the optimal fault-tolerant learning of multivariate sample feature quantity (μ, Σ)
[0088] The optimal fault-tolerant learning algorithm is characterized by that given a test set S II The compliance ratio I IIS and the rejection ratio I IIA They are:
[0089]
[0090] Affects the above support set S ⅡS and the exclusion set S ⅡA The important factor for division is the learning result of the training set feature parameters With the test set S Ⅱ Whether the sample mean and scatter matrix are consistent.
[0091] To this end, for two sets of samples from the training set and the test set and Two groups of samples construct the null hypothesis H 0 :
[0092] H 0 :μ=μ II ,Σ=Σ II (9);
[0093] remember and The sample sizes are n and m, and the means are X a and Construct the null hypothesis H 0 Likelihood ratio test statistic for :
[0094]
[0095] Constructing Negative Domains
[0096] When the statistic When , we accept the null hypothesis H with (1-α)×100% confidence. 0 .
[0097] Remember the training set and test set The sample covariance matrices are Σ a and The likelihood ratio test statistic can be obtained as follows:
[0098]
[0099] Error-tolerant statistics using covariance matrix and Substituting the sample covariance in equation (11), we can obtain the modified λ that is tolerant to abnormal data. 4 Test statistic:
[0100]
[0101] When the statistic If the null hypothesis H 0 .
[0102] From the test set S Ⅱ Eliminate the exclusion set S ⅡA This will lead to a continuous reduction in the test set and information loss. Considering the important role of the test set in the machine learning process, an abnormal data correction and repair algorithm for the test set is established. Specifically, for any X i ∈S IIA , establish a fault-tolerant proportional compression method:
[0103]
[0104] Any X i ∈S IIA Both Therefore, the compression transformation formula (13) can compress all samples in the exclusion set back to the set S IIS In the process, error correction and repair are performed on abnormal samples.
[0105] It should be noted that the error correction and repair method can only be used for the set S IIA Number of samples in |S IIA |Significantly less than |S IIS |The situation.
[0106] Step S4: Determine the test result.
[0107] Specific implementation method:
[0108] For active fault-tolerant learning of multi-dimensional sample feature quantities (μ, Σ), since abnormal samples such as whether there are outliers and the number of outliers are usually unknown, we do not know the proportion nor which samples they occur in. Therefore, it is necessary to implement feature fault-tolerant learning and fault-tolerant testing of learning algorithms in two stages. The implementation process is divided into two stages, such as Figure 1 shown.
[0109] In the first stage, passive fault-tolerant learning is achieved through four steps:
[0110] Step 1. All samples participate in learning.
[0111] Using the training set S I All samples of , the feature quantity (μ, Σ) is preliminarily learned, the specific method is:
[0112]
[0113] Step 2: Euclidean metric sorting of all samples.
[0114] For the training set S I All samples of Euclidean metric of And sort from small to large to get the sample partial order set {X (i) ,....,X (n)};
[0115] Step 3: 90% of the samples participate in learning.
[0116] Using the 90% points on the left side of the partial order relation, record the real number rounding operator as “[]”, and perform feature learning in a similar way to formula (14):
[0117]
[0118] Step 4: Correction of learning results. Return to Step 2 and use replace Repeat Step 2 and Step 3 to obtain the corrected passive fault-tolerant learning result.
[0119] In the second stage, active fault-tolerant testing of the passive fault-tolerant learning method is implemented through five steps:
[0120] Step 5: Use the sample mean and scatter matrix estimates learned from the training set Construct a confidence ellipsoid:
[0121]
[0122] The constant α in the formula is 0.05 or 0.01.
[0123] Step 6. Test set S Ⅱ Partitioned into support set S ⅡS and the exclusion set S ⅡA Two parts:
[0124] S IIS =S II ∩Q II ,S IIA =S II -S IIS
[0125] Step 7: Construct the null hypothesis H 0 :μ=μ II ,Σ=Σ II , construct the null hypothesis H 0 The likelihood ratio test statistic λ 4 and Negation Domain
[0126] Step8, when support set S ⅡS When the number of sample data meets the test requirements, the fault-tolerant statistic is used and Replace λ 4 The sample covariance in the expression is used to calculate the modified λ that is tolerant to abnormal data such as outliers. 4 Test statistic:
[0127]
[0128] Step 9, when the support set S ⅡS When the number of sample data in does not meet the test requirements, the exclusion set S ⅡA Sample X i ∈S IIA , through the fault-tolerant proportional compression method:
[0129]
[0130] Put all samples in the exclusion set back to the set S by compression ⅡS Then calculate the test statistic λ 4 .
[0131] Step 10, test result determination: When the statistic When accepting the null hypothesis H 0 , the feature fault-tolerant learning method passed the test.
[0132] Example
[0133] Specifically, according to Figure 1 And the above implementation process is implemented, the learning results obtained are as follows Figure 2 shown.
[0134] Select 100 groups of two-dimensional normal samples, the distribution is as follows Figure 2 As shown in (a), the asterisk in the figure is the machine learning result of the sample feature quantity, and the ellipse is the sample scatter domain; 4 sample points in the above 100 groups of samples are biased to form 4 outlier points. Figure 2 (b) shows the sample scatter diagram with outliers and the sample scatter error elliptic curve drawn by the conventional learning method; Figure 2 (c) shows the Figure 2The 100 groups of samples with 4 abnormal data shown in (b) are used as the objects, and the feature quantities and scattered error ellipse curves are obtained through fault-tolerant machine learning.
[0135] contrast Figure 2 (a) Figure 2 (b) and Figure 2 It can be clearly seen from (c) that due to the influence of abnormal samples, Figure 2 The scattered error ellipse and Figure 2 Compared with the scatter domain shown in (a), there is a significant deformation, which shows that conventional machine learning algorithms lack fault tolerance for abnormal data; Figure 2 The scattered error ellipse shown in (c) is significantly better than Figure 2 The error ellipse shown in (b) is close to Figure 2 The error ellipse shown in (a) in Figure 3 proves that the fault-tolerant machine learning method can effectively overcome the adverse effects of abnormal samples and improve the accuracy and credibility of machine feature learning.
[0136] Figure 2 (d) in Figure 1 shows the above fault-tolerant learning method applied to an 80-point test set, which also contains 4 abnormal samples. Figure 2 As can be seen from (d) in the figure, fault-tolerant machine learning successfully passed the fault-tolerance test, and the error ellipse was hardly affected by abnormal samples.
[0137] Therefore, the present invention adopts the above-mentioned learning method of passive fault tolerance and active error correction of multi-dimensional sample features, establishes a set of concise and practical fault-tolerant design methods for multi-dimensional sample feature learning processes, and abnormal sample error correction methods for test sets, improves the passive fault tolerance and active fault tolerance capabilities of machine learning methods for abnormal samples, and realizes data cleaning of multi-dimensional samples and fault-tolerant extraction of feature information, so as to ensure the reliability of machine learning and the correctness of knowledge acquisition, and avoid the machine learning of incorrect knowledge or being misled by abnormal samples due to abnormal samples.
[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.
Claims
1. A learning method for passive error tolerance and active error correction of multi-dimensional sample features, characterized in that: The following steps are involved: Step S1: Divide the samples into training set S I and the test set S II ; Step S2, passive fault-tolerant learning of statistical characteristics of multi-dimensional samples of impurity training set; Step S3, using the fault tolerance test and outlier correction of the test set with impurities; Step S4: Determine the test result.
2. According to claim 1, a learning method for passive error tolerance and active error correction of multi-dimensional sample features is characterized in that: In step S1, the training set S I For training and testing the learning model II Used for testing and validating learning models.
3. The learning method of passive error tolerance and active error correction of multi-dimensional sample features according to claim 2 is characterized in that: In step S2, for the training set S formed by n p-dimensional samples I ={X i ∈R p , i=1,2,...,n}, where S represents the sample set, X i represents the i-th sample, R p represents p-dimensional real number space; training set S I The statistical characteristic of is the mathematical expectation value μ∈R of the matrix p and the scatter matrix ∑∈R p×p , including parameters, among which, Indicates the number of combinations; A passive fault-tolerant learning method based on multi-dimensional sample statistical characteristics is established, specifically, an M-learning mechanism is constructed: Where, d 2 (X,μ,Σ)=(X-μ) τ Σ -1 (X-μ); {φ μ ,φ η } is in R + = is a real function pair on [0, +∞), μ is the first-order moment of the sample, i.e., the mathematical expectation, ∑ represents the covariance of the mother population corresponding to the sample, or the second-order moment, and ∧ represents the statistical estimator.
4. The learning method of passive error tolerance and active error correction of multi-dimensional sample features according to claim 3 is characterized in that: In step S2, a tuple of the following form is constructed: Where r0 and c are adjustable parameters, and the default values are r0=3p, c=3r0; Perform machine learning on the multidimensional feature quantity (μ, Σ) of the sample, denoted by X i ∈R p Component form (x1(t i ),...,x p (t i )) τ and median operator Construct an iterative calculation method: And select the initial value of the iteration: Through iterative calculation, the optimal passive fault tolerance learning is performed on the multi-dimensional feature quantity (μ, Σ) of the sample to obtain the optimal fault tolerance estimator 5. The learning method of passive error tolerance and active error correction of multi-dimensional sample features according to claim 4 is characterized in that: In step S3, an elimination strategy or an avoidance strategy is adopted to convert the test set S II Partitioned into support set S ⅡS and the exclusion set S ⅡA Two parts, construct the repulsion index: By I IIA The size of determines the degree of fit between the learning model and the test set samples; Exclusion set determination: For any sample X i ∈S II , when it follows the normal distribution N(μ II ,Σ II ) is a valid sample, then Therefore, the sample mean and scatter matrix estimates obtained by learning the training set are Construct a confidence ellipsoid: In the formula is 2 -(1-α)×100% quantile of the distribution; According to formula (6), for the test set S II The support set S ⅡS and the exclusion set S ⅡA Division: S IIS =S II ∩Q II ,S IIA =S II -S IIS (7); Get the optimal fault-tolerant learning of multivariate sample feature quantity (μ, Σ) The optimal fault-tolerant learning algorithm is given a test set S II The compliance ratio I IIS and the rejection ratio I IIA They are:
6. A learning method for passive error tolerance and active error correction of multi-dimensional sample features according to claim 5, characterized in that: In step S3, the influence support set S ⅡS and the exclusion set S ⅡA The important factor for division is the learning result of the training set feature parameters With the test set S Ⅱ Whether the sample mean and scatter matrix are consistent; For two sets of samples from the training set and the test set and Two groups of samples construct the null hypothesis H0: H0: μ=μ II ,S=S II (9); remember and The sample sizes are n and m, and the means are and Construct the likelihood ratio test statistic for the null hypothesis H0: Constructing Negative Domains When the statistic When , the null hypothesis H0 is accepted with (1-α)×100% confidence.
7. The learning method of passive error tolerance and active error correction of multi-dimensional sample features according to claim 6 is characterized in that: In step S3, record the training set and test set The sample covariance matrices are and The likelihood ratio test statistic is obtained as follows: Error-tolerant statistics using covariance matrix and Replacing the sample covariance in equation (11), we obtain the modified λ4 test statistic that is tolerant to abnormal data: When the statistic , then accept the null hypothesis H0.
8. The method for learning passive error tolerance and active error correction of multi-dimensional sample features according to claim 7, characterized in that: In step S3, an abnormal data error correction and repair algorithm is established in the test set. i ∈S IIA , establish a fault-tolerant proportional compression method: Any X i ∈S IIA Both Therefore, the compression transformation formula (13) compresses all samples in the exclusion set back to the set S IIS In the process, error correction and repair are performed on abnormal samples.