Abnormal sample identification method and device, storage medium and system

By constructing multiple single-model exception sample recognition models, calculating consensus mean and variance, and setting thresholds to identify abnormal samples, the problems of missed and false positives in the existing technology are solved, and the accuracy of identification of abnormal samples and the robustness of the model are improved.

CN120354291APending Publication Date: 2025-07-22HEILONGJIANG BAYI AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410040489.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing abnormal sample identification methods have missed and false positives, which leads to interference in model prediction performance and the accuracy needs to be improved.

Method used

Multiple single-model exception sample recognition models are constructed. By calculating the mean and variance of sample residuals, we obtain the consensus mean vector and consensus variance vector, set the threshold to identify the abnormal samples, and remove samples that meet the threshold conditions.

Benefits of technology

Identifying abnormal samples through multi-model consensus improves the correlation between near-infrared spectrogram and chemical values, eliminates interference to model prediction performance, and significantly improves the detection accuracy and robustness of abnormal samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354291A_ABST
    Figure CN120354291A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of data processing, and provides an abnormal sample recognition method, device, storage medium and system.The method comprises the steps that firstly, a modeling set is used for establishing a plurality of single-model abnormal sample recognition regression models, the regression models are used for predicting samples of a prediction set, and the prediction residual error of each sample is calculated; each modeling method is repeated for multiple times, and it is guaranteed that multiple prediction residuals of each sample can be obtained. And calculating a mean value and a variance of prediction residuals of each sample according to a modeling method, respectively selecting appropriate threshold values for the variance and the mean value according to a multi-model consensus method, and identifying the sample with the variance or the mean value of the prediction residuals greater than the threshold values as an abnormal sample. According to the scheme, the abnormal sample is identified in a multi-model consensus mode, so that the correlation between the near infrared spectrogram and the chemical value is increased, the interference on the prediction performance of the model is eliminated, and the accuracy and robustness of abnormal sample detection are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and particularly relates to an abnormal sample recognition method, device, storage medium and system. Background Art

[0002] Near-infrared spectroscopy is a non-destructive analysis method for analyzing the composition of substances. It uses the absorption characteristics of substances in the near-infrared band (usually 700 to 2500 nanometers) to determine the content of chemical components in samples. The principle of near-infrared spectroscopy is based on the absorption of light caused by molecular vibrations. Different functional groups or chemical bonds will exhibit specific absorption peaks in the near-infrared band, so that the composition of substances can be identified through spectral maps.

[0003] Near-infrared spectroscopy has the characteristics of being fast, accurate, non-destructive and diverse, so it is widely used in the fields of agriculture, food, medicine, chemical industry, environmental monitoring, etc. In near-infrared spectroscopy analysis, the presence of abnormal samples will seriously interfere with the prediction performance of the model. Therefore, the identification and elimination of abnormal samples are the basis and important links for establishing a robust model. By identifying and eliminating abnormal samples, interference can be reduced, and the accuracy and reliability of the model can be improved.

[0004] Common abnormal sample recognition methods at home and abroad include: one type is classical diagnosis, such as Mahalanobis distance discriminant method, spectral residual method and hat matrix method. When there are multiple singular samples in the sample set, the recognition efficiency of these diagnostic methods is often not ideal. Another type is robust regression methods, such as MVT (Multivariate Tri-mming), MVE (Minimum Volume Estimator), etc. These will smooth the data, which may lead to a decrease in the sensitivity to abnormal samples. Cao et al. proposed a method for identifying abnormal samples by combining Monte Carlo cross-validation and PLS, which can solve the above two problems, but it is found that there are "biases" and "preferences" in identifying abnormal samples.

[0005] Therefore, it can be seen that the existing analysis methods for abnormal samples have relatively prominent missed reports and false reports, and the capture accuracy of abnormal samples needs to be improved. Summary of the Invention

[0006] The purpose of the embodiments of the present application is to provide an abnormal sample recognition method, aiming to solve the problems that the existing analysis methods for abnormal samples have relatively prominent missed reports and false reports, and the capture accuracy of abnormal samples needs to be improved.

[0007] The embodiments of the present application are implemented as follows. An abnormal sample recognition method, the method includes:

[0008] Based on a plurality of spectral data matrices XN×K A sample data set is formed to construct q single - model abnormal sample recognition models, and each of the single - model abnormal sample recognition models is used to recognize abnormal samples; predict the q single - model abnormal sample recognition models respectively to obtain the sample residual matrix Z(q) of each single - model abnormal sample recognition model R×N and the mean vector M(q) of each sample residual 1×N and the variance vector V(q) 1×N , where q≥2, N represents the number of samples, K represents the number of spectral feature variables, and R is the number of repetitions; based on the mean vector M(q) of each sample residual of all single - model abnormal sample recognition models 1×N , obtain the consensus mean vector Mp 1×N ; based on the variance vector V(q) of each sample residual of all single - model abnormal sample recognition models 1×N , obtain the consensus variance vector Vp 1×N ; sort the values in the consensus mean vector Mp 1×N in descending order, delete the maximum value of the vector, and calculate the mean m of the top t values; sort the values in the consensus variance vector Vp 1×N in descending order, delete the maximum value of the vector, and calculate the mean v of the top t values in the vector to obtain the threshold pair (m, v); the samples that satisfy Mp 1×N ≥m and / or Vp 1×N ≥v are used as the finally recognized abnormal samples S

[0009] Another object of the embodiments of the present application is to provide an abnormal sample recognition device, and the abnormal sample recognition device includes:

[0010] A single - model abnormal sample recognition model construction module, which is used to construct q single - model abnormal sample recognition models based on a sample data set composed of several spectral data matrices X N×K Each of the single - model abnormal sample recognition models is used to recognize abnormal samples; a residual matrix, mean vector and variance vector acquisition module, which is used to predict the q single - model abnormal sample recognition models respectively to obtain the sample residual matrix Z(q) of each single - model abnormal sample recognition model R×N and the mean vector M(q) of each sample residual 1×N and the variance vector V(q) 1×N , where q≥2, N represents the number of samples, K represents the number of spectral feature variables, and R is the number of repetitions; a consensus mean vector and consensus variance vector acquisition module, which is used to obtain the consensus mean vector Mp based on the mean vector M(q) of each sample residual of all single - model abnormal sample recognition models 1×N , obtain the consensus mean vector Mp 1×N; Based on the variance vector V(q) of the residuals of all single - model anomaly sample recognition models 1×N , a consensus variance vector Vp is obtained 1×N ; A threshold pair acquisition module, which is used to sort the values in the consensus mean vector Mp 1×N in descending order, delete the maximum value of the vector, and calculate the mean m of the top t values; sort the values in the consensus variance vector Vp 1×N in descending order, delete the maximum value of the vector, and calculate the mean v of the top t values in the vector, obtaining the threshold pair (m, v); an anomaly sample acquisition module, which is used to take the samples that satisfy Mp 1×N ≥m and / or Vp 1×N ≥v as the finally identified anomaly samples S

[0011] Another object of the embodiments of the present application is to provide a computer - readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the processor executes the steps of an anomaly sample recognition method as described above

[0012] Another object of the embodiments of the present application is to provide an anomaly sample recognition system. When the system runs, it executes the steps of an anomaly sample recognition method as described above

[0013] An anomaly sample recognition method provided by the embodiments of the present application realizes the elimination of anomaly samples through multi - model consensus, and identifies anomaly samples through the method of multi - model consensus, so as to increase the correlation between near - infrared spectrograms and chemical values, eliminate the interference to the prediction performance of the model, and significantly improve the detection accuracy and robustness of anomaly samples Description of the Drawings

[0014] Figure 1 is an application environment diagram of an anomaly sample recognition method provided by the embodiments of the present application

[0015] Figure 2 is a step diagram of an anomaly sample recognition method provided by the embodiments of the present application

[0016] Figure 3 is a flowchart of an anomaly sample recognition method provided by the embodiments of the present application

[0017] Figure 4 is a schematic diagram of a residual matrix and its mean and variance vectors provided by the embodiments of the present application

[0018] Figure 5 is an original near - infrared spectrogram of corn provided by the embodiments of the present application

[0019] Figure 6The original near-infrared spectrogram of a kind of cereal provided by the embodiment of the present application;

[0020] Figure 7 The mean-variance diagram for identifying abnormal samples of corn data by combining the PLS, GPR, SVM, and BP modeling methods based on this solution provided by the embodiment of the present application;

[0021] Figure 8 The mean-variance diagram for identifying abnormal samples of starch data by combining the PLS, GPR, SVM, and BP modeling methods based on this solution provided by the embodiment of the present application;

[0022] Figure 9 The structural block diagram of an abnormal sample identification device provided by the embodiment of the present application;

[0023] Figure 10 The internal structural block diagram of a computer device in an embodiment. Detailed implementation manners

[0024] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0025] It can be understood that the terms "first", "second", etc. used in the present application can be used to describe various elements herein, but unless otherwise specified, these elements are not limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the present application, the first script can be called the second script, and similarly, the second script can be called the first script.

[0026] Figure 1 The application environment diagram of an abnormal sample identification method provided by the embodiment of the present application, as Figure 1 shown. In this application environment, it includes a data acquisition device 110 and a computer device 120.

[0027] The computer device 120 can be an independent physical server or terminal, or a server cluster composed of multiple physical servers. It can be a cloud server providing basic cloud computing services such as cloud servers, cloud databases, cloud storage, and CDN.

[0028] The data acquisition device 110 can be an infrared spectrum acquisition device, a spectral imaging device, an image acquisition device, an image processing or generation device, etc., but is not limited thereto. The data acquisition device 110 and the computer device 120 can be connected through a network, and the present application does not limit this here.

[0029] AsFigure 2 As shown, in one embodiment, an abnormal sample recognition method is proposed. In this embodiment, the method is mainly illustrated by applying it to the computer device 120 in the above Figure 1 . An abnormal sample recognition method may specifically include the following steps:

[0030] Step S201: Based on a sample data set composed of several spectral data matrices X N×K , construct q single-model abnormal sample recognition models, and each of the single-model abnormal sample recognition models is used to recognize abnormal samples.

[0031] Step S203: Respectively perform predictions on the q single-model abnormal sample recognition models to obtain the sample residual matrix Z(q) R×N , the mean vector M(q) of each sample residual 1×N and the variance vector V(q) 1×N , where q≥2, N represents the number of samples, K represents the number of spectral feature variables, and R is the number of repetitions.

[0032] Step S205: Based on the mean vector M(q) of each sample residual of all the single-model abnormal sample recognition models 1×N , obtain the consensus mean vector Mp 1×N ; based on the variance vector V(q) of each sample residual of all the single-model abnormal sample recognition models 1×N , obtain the consensus variance vector Vp 1×N .

[0033] Step S207: Sort the values in the consensus mean vector Mp 1×N in descending order, delete the maximum value of the vector, and calculate the mean m of the top t values; sort the values in the consensus variance vector Vp 1×N in descending order, delete the maximum value of the vector, and calculate the mean v of the top t values in the vector to obtain the threshold pair (m, v).

[0034] Step S209: Take the samples that satisfy Mp 1×N ≥m and / or Vp 1×N ≥v as the finally recognized abnormal samples S.

[0035] Those skilled in the art know that in this solution, q single-model abnormal samples are constructed, and q can be two or more. To balance accuracy and system calculation speed, preferably, q is selected as 4. This embodiment only takes 4 single-abnormal sample recognition models as an example. Using more or fewer single-abnormal sample recognition models can also achieve consensus recognition based on this solution, thereby improving the recognition accuracy of abnormal samples.

[0036] It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all fall within the protection scope of this application.

[0037] In one embodiment of this application, as Figure 3 shown, a flowchart of a method for identifying abnormal samples is provided. Based on this method, the accuracy and generalization ability of the model can be improved. This solution first preprocesses the spectral data set collected by the near-infrared spectrometer, then uses the Monte Carlo random sampling method to take 80% of the samples as the modeling set, and the remaining part as the prediction set. Several regression models are respectively built using the modeling set. These regression models are used to predict the samples in the prediction set, and the prediction residuals of each sample are calculated. Each modeling method is repeated multiple times to ensure that multiple prediction residuals of each sample can be obtained. The mean and variance of the prediction residuals of each sample are calculated respectively according to the modeling method. Based on the mean vector and variance vector of the residuals of all single-model abnormal sample recognition models, a consensus mean vector and a consensus variance vector are obtained. Appropriate thresholds are selected for the consensus variance and the consensus mean respectively, and the samples with the variance or mean of the prediction residuals greater than the threshold are identified as abnormal samples.

[0038] The spectral data matrix is represented by X N×K where each row in X represents a sample, and each column represents a spectral feature variable; the concentration vector is represented by Y N×1 where one value in Y corresponds to one row in X, representing the chemical concentration value of a sample. Among them, N is the number of samples, and K is the number of spectral feature variables.

[0039] An abnormal sample recognition method provided by an embodiment of this application realizes the elimination of abnormal samples through multi-model consensus, and identifies abnormal samples through the way of multi-model consensus, so as to increase the correlation between the near-infrared spectrogram and the chemical value, eliminate the interference to the prediction performance of the model, and significantly improve the elimination accuracy of abnormal samples.

[0040] In one embodiment of the embodiments of this application, the method for obtaining the spectral data matrix X N×K is as follows:

[0041] Obtain the original spectral data matrix; preprocess the original spectral data matrix to obtain the spectral data matrix, and the preprocessing is used to eliminate the abnormal data generated by factors such as particle size, surface scattering, and optical path transformation in the original data matrix.

[0042] In one embodiment of the embodiments of this application, based on a sample data set composed of several spectral data matrices X N×K q single-model abnormal sample recognition models are constructed, and then the sample residual matrix Z(q) of each of the single-model abnormal sample recognition models is obtainedR×N The method is as follows:

[0043] S1: Randomly select several data sets from the sample data set to construct the train-set data set, and use the sample data set that is not selected as the train-set data set as the predict-set data set.

[0044] S2: Based on the train-set data set, construct a single-model abnormal sample recognition model.

[0045] S3: Use the single-model abnormal sample recognition model to predict the predict-set data set, and calculate and collect the prediction residuals of each sample in the predict-set data set.

[0046] S4: Return to step S1, and repeat the above steps S1 to S3 several times until the preset number of repetitions is reached. Preferably, the number of repetitions is 200 times. Collect the prediction residual data of each sample in all the predict-set data sets generated during the repeated execution of the above steps, and then obtain the residual matrix Z(q) corresponding to each single-model abnormal sample recognition model. R×N 。

[0047] In an embodiment of the present application, the constructed single-model abnormal sample recognition models are 4, which are regression models established based on the train-set data set using PLS, GPR, SVM, and BP modeling methods respectively.

[0048] In the embodiments of the present application, PLS, GPR, SVM, and BP modeling methods can be used to construct a single-model abnormal sample recognition model, but it is not limited to only using these four methods, and other methods can also be used for modeling.

[0049] Among them, PLS refers to Partial Least Squares Regression, which is a multivariate statistical analysis method usually used to establish a linear relationship model between variables. GPR refers to Gaussian Process Regression, which is a non-parametric regression method. SVM refers to Support Vector Machine, which is a supervised learning model for classification and regression analysis. BP refers to Back Propagation Neural Network, which is an artificial neural network model often used to solve classification, regression and other problems.

[0050] In an embodiment of the present application, the mean vector M(q) of each sample residual 1×NWith the variance vector V(q) 1×N The method for obtaining it is as follows:

[0051] The prediction residuals of the train-set data set as the modeling set are counted as NaN. Ignoring the NaN values, the mean and variance of each column in the sample residual matrix Z(q) R×N are obtained to get the mean vector M(q) 1×N of the sample residuals and the variance vector V(q) 1×N .

[0052] In an embodiment of the present application, based on the mean vector M(q) 1×N of the sample residuals of all single-model anomaly recognition models, the consensus mean vector Mp 1×N is obtained; based on the variance vector V(q) 1×N of the sample residuals of all single-model anomaly recognition models, the consensus variance vector Vp 1×N is obtained. The method is as follows:

[0053] Calculate the percentage of each value in the mean vector M(q) 1×N in the sum of all N values to obtain the mean percentage vector Mp(q) 1×N . Calculate the average value of each sample in the mean percentage vector Mp(q) 1×N according to the following method to obtain the consensus mean vector Mp 1×N of the residuals:

[0054] Mp 1×N = [Mp(1) 1×N (i) + Mp(2) 1×N (i) + … + Mp(q) 1×N (i)] / q,

[0055] where i represents the i-th sample.

[0056] Calculate the percentage of each value in the variance vector V(q) 1×N in the sum of all N values to obtain the variance percentage vector Vp(q) 1×N . Calculate the average value of each sample in the variance percentage vector Vp(q) 1×N according to the following method to obtain the consensus variance vector Vp 1×V of the residuals:

[0057] Vp 1×N = [Vp(1) 1×N (i) + Vp(2) 1×N (i) + … + Vp(q) 1×N (i)] / q.

[0058] In the embodiments of the present application, in order to prevent inaccurate prediction by a single model and the influence of excessive prediction residuals on the selection of thresholds, since the prediction performances of the regression models are different, the percentage of each value in the mean vector and variance vector of each regression model in the total of N values is calculated, and the average value of each sample in the mean percentage vector of each regression model is calculated. Let Mp 1×N and Vp 1×N be sorted in descending order, and the maximum values in Mp 1×N and Vp 1×N are respectively deleted. After that, the mean and variance of the first t in the vector are taken to form a threshold pair, and the value range of t is 0 - 1. For example, when t takes 1 / 4, that is, the first 1 / 4 of the larger numbers in Vp 1×N are taken, and the mean and variance of these numbers are calculated, which is equivalent to taking the first 25% in Vp 1×N . If Vp 1×N = [1.5, 2, 2.5, 3.5, 4, 5, 5.5, 7.5, 9.5, 10, 12.5, 13.5, 23], after deleting the maximum value 23, the first 1 / 4 of the larger numbers (10, 12.5, 13.5) of the remaining 12 values are taken, and the average value of these four numbers is the variance threshold. The specific value can be determined according to the actual situation. The reason for deleting the maximum values in Mp 1×N and Vp 1×N is that the maximum values in Mp 1×N and Vp 1×N will make the finally calculated threshold too large, resulting in fewer abnormal samples being excluded, so that some abnormal samples may not be identified.

[0059] In an embodiment of the present application, taking the construction of 4 single - model abnormal sample recognition models as an example, the present solution includes the following steps:

[0060] Step 1: There are a spectral data matrix X N×K and a concentration vector Y N×1 , where N is the number of samples and K is the number of spectral data points. The spectral data matrix X N×K is pre - processed;

[0061] Step 2: Use the Monte Carlo random sampling method to take 80% of the samples as the modeling set train - set, and the remaining part as the prediction set predict - set.

[0062] Step 3: Respectively use modeling methods such as PLS, GPR, SVM, BP, etc. to establish regression models using the data of the modeling set train - set, use each regression model to predict the data of the prediction set predict - set, and calculate the prediction residuals of each sample in the predict - set. The prediction residuals of the samples in the modeling set train - set are recorded as NaN.

[0063] Step 4: Repeat Step 2 to Step 3 multiple times. Obtain the residual matrices Z1 R×N , Z2 R×N , Z3 R×N , Z4 R×N. , …, where R is the number of repetitions (at least 200 times) and N is the number of samples.

[0064] Step 5: Ignoring NaN values, find the mean and variance of each column of Z1 R×N , Z2 R×N , Z3 R×N , Z4 R×N , … respectively. Obtain the mean vectors M1 1×N , M2 1×N , M3 1×N , M4 1×N , … and the variance vectors V1 1×N , V2 1×N , V3 1×N , V4 1×N , …. ( Figure 2 )

[0065] Step 6: Find the percentage of each value in the mean vectors M1 1×N , M2 1×N , M3 1×N , M4 1×N , … in the sum of N values, obtaining the mean percentage vectors Mp1 1×N , Mp2 1×N , Mp3 1×N , Mp4 1×N , …. Find the percentage of each value in the variance vectors V1 1×N , V2 1×N , V3 1×N , V4 1×N , … in the sum of N values, obtaining the variance percentage vectors Vp1 1×N , Vp2 1×N , Vp3 1×N , Vp4 1×N , ….

[0066] Step 7: According to the mean percentage vectors Mp1 1×N , Mp2 1×N , Mp3 1×N , Mp4 1×N , … of each regression model, find the average value of each sample, i.e., Mp 1×N = (Mp1 1×N (i) + Mp2 1×N (i) + Mp3 1×N (i) + Mp4 1×N(i) + …) / x, where x is the number of regression models; according to the variance percentage vectors Vp1 of each regression model 1×N , Vp2 1×N , Vp3 1×N , Vp4 1×N , …, find the average value of each sample, that is, Vp 1×N = (Vp1 1×N (i) + Vp2 1×N (i) + Vp3 1×N (i) + Vp4 1×N (i) + …) / x, where x is the number of regression models.

[0067] Step 8: Sort the numbers in the consensus mean vector Mp 1×N in descending order, delete the maximum value of the vector, and obtain the vector Mp 1×(N-1) ; sort the numbers in the consensus variance vector Vp 1×N in descending order, delete the maximum value of the vector, and obtain the vector Vp 1×(N-1) . Select the mean value m of the top q numbers in the vector Mp 1×(N-1) and the mean value v of the top q numbers in the vector Vp 1×(N-1) . Obtain the threshold pair (m, v).

[0068] Step 9: Use the samples that satisfy Mp 1×N ≥ m and / or Vp 1×N ≥ v as the finally identified abnormal samples S.

[0069] In an embodiment of the present application, take the scheme provided in this embodiment and apply it to a set of near-infrared spectral data of corn samples and a set of near-infrared spectral data of cereal samples as an example. The present application is also applicable to the identification of abnormal samples before the near-infrared spectral modeling process of other types of samples.

[0070] The corn sample data used in this embodiment is the germination rate and spectral data of corn seeds collected by the research group of the researcher. The sample data includes two parts, the NIR spectral data of the samples and the germination rate data, with a total of 245 samples. As Figure 5 shown, the selected wavelength range of the sample spectrum is 1138 - 2080 nm, with a measurement value corresponding to every 2 nm, and the total number of wavelength points is 966. Figure 5 Each curve in

[0071] represents a row in the sample matrix X, that is, the spectral data of a sample. The present application uses the corn near-infrared spectrum as the independent variable and the germination rate as the dependent variable for abnormal sample identification and modeling prediction analysis to prove the effectiveness of this method. The method used in the present application is:

[0072] Step 2: Use the Monte Carlo random sampling method to select 80% of the samples as the modeling set, and the remaining part as the prediction set. The modeling set contains 196 samples, and the prediction set contains 49 samples. The spectral data of the modeling set is a 196×966 matrix train-set, and the spectral data of the validation set is a 49×966 matrix predict-set.

[0073] Step 3: Use the PLS, GPR, SVM, and BP modeling methods to establish regression models using the train-set data of the modeling set, use the regression models to predict the predict-set data, and calculate the prediction residuals of each sample in the predict-set. The prediction residuals of the samples in the train-set of the modeling set are NaN.

[0074] Step 4: Repeat steps 2 to 3 R (R = 1000) times. Obtain the residual matrix Z1 R×N , Z2 R×N , Z3 R×N , Z4 R×N , where R is the number of repetitions 1000, and N is the number of samples 245.

[0075] Step 5: Ignoring the NaN values, calculate the mean and variance of each column of Z1 R×N , Z2 R×N , Z3 R×N , Z4 R×N to obtain four groups of mean vectors M1 1×N , M2 1×N , M3 1×N , M4 1×N and variance vectors V1 1×N , V2 1×N , V3 1×N , V4 1×N . The sample mean-variance diagram is as shown in Figure 7 the figure.

[0076] Step 6: Calculate the percentage of each value in the mean vectors M1 1×N , M2 1×N , M3 1×N , M4 1×N ,... in the sum of N values to obtain the mean percentage vectors Mp1 1×N , Mp2 1×N , Mp3 1×N , Mp4 1×N ,.... Calculate the percentage of each value in the variance vectors V1 1×N , V2 1×N , V3 1×N , V4 1×NThe percentage of each value in … out of the total of N values is obtained to get the variance percentage vector Vp1 1×N , Vp2 1×N , Vp3 1×N , Vp4 1×N , ….

[0077] Step 7: Based on the mean percentage vectors Mp1 1×N , Mp2 1×N , Mp3 1×N , Mp4 1×N , …, find the average value of each sample, that is, Mp 1×N= (Mp1 1×N (i) + Mp2 1×N (i) + Mp3 1×N (i) + Mp4 1×N (i) + …) / x, where x is the number of regression models; Based on the variance percentage vectors Vp1 1×N , Vp2 1×N , Vp3 1×N , Vp4 1×N , …, find the average value of each sample, that is, Vp 1×N= (Vp1 1×N (i) + Vp2 1×N (i) + Vp3 1×N (i) + Vp4 1×N (i) + …) / x, x = 4;.

[0078] Step 8: Sort the numbers in the vector Mp 1×N in descending order and delete the maximum value of the vector to get the vector Mp 1×(N-1) ; Sort the numbers in the vector Vp 1×N in descending order and delete the maximum value of the vector to get the vector Vp 1×(N-1) . Select the mean value m of the top q numbers in the vector Mp 1×(N-1) and the mean value v of the top q numbers in the vector Vp 1×(N-1) . Obtain the threshold pair (0.8733, 0.8674).

[0079] Step 9: Samples that satisfy Mp 1×N ≥m and / or Vp 1×N ≥v are the finally identified abnormal samples S. The number of S is 35. The number of abnormal samples identified by the MCCV+PLS method is 37.

[0080] To illustrate the advantages of this application in abnormal sample recognition, under the same other conditions, models are established in three cases: (1) Using the method proposed in this application to delete abnormal samples, and then establishing models such as PLS, GPR, SVM, BP (denoted as MMCODM); (2) On the basis of SNV preprocessing, establishing PLS, GPR, SVM, BP models without removing samples (denoted as Non); (3) Using MCCV+PLS to remove abnormal samples, and then establishing PLS, GPR, SVM, BP models (denoted as MCCV+PLS). Compare the results of the above three cases to verify the effectiveness of the algorithm. The prediction ability of the model is mainly determined by the coefficient of determination of the modeling set (R c 2 ), the root mean square error of the modeling level (RMSEC), the coefficient of determination of the prediction set (R p 2 ), and the root mean square error of the prediction set (RMSEP) indicators. Among them, R c , R p The closer the values are to 1, the closer RMSEC and RMSEP are to 0, the better the fitting of the model, the higher the prediction accuracy, and the better the effect of abnormal sample removal.

[0081] The model parameters of the models established based on the three abnormal sample removal methods are shown in Table 1. From the overall performance (average model performance) of the four types of models, the performance of the models obtained based on the method of this application is better than that of the MCCV+PLS abnormal sample removal method and the method of not removing abnormal samples.

[0082] Table 1 Comparison of the effects of three abnormal sample removal results on the corn dataset

[0083]

[0084] In one embodiment, a set of publicly available near-infrared spectral data of grains from the website EigenVector is used. This dataset includes 80 grain samples measured by 3 different near-infrared spectrometers. The wavelength range of the sample spectra is 1100-2498 nm, as Figure 6 shown, the sampling interval is 2 nm, and there are a total of 700 wavelength points. The chemical properties include moisture, oil, protein, and starch values. In this example, the near-infrared spectrum measured by the instrument mp6 is selected, and the near-infrared spectrum of the grain is used as the independent variable and the starch content as the dependent variable for abnormal sample recognition and modeling prediction to illustrate the effectiveness of this method.

[0085] Step 1: Preprocess the original spectral data using the standard normal variate transformation (SNV) method.

[0086] Step 2: Use the Monte Carlo random sampling method to select 80% of the samples as the modeling set, and the remaining part as the prediction set. The modeling set contains 64 samples, and the prediction set contains 16 samples. The spectral data of the modeling set is a 64×700 matrix train-set, and the spectral data of the validation set is a 16×700 matrix predict-set.

[0087] Step 3: Use the PLS, GPR, SVM, and BP modeling methods to establish regression models using the modeling set data train-set, use the regression models to predict the prediction set data predict-set, and calculate the prediction residuals of each sample in predict-set. The prediction residuals of the samples in the modeling set train-set are NaN.

[0088] Step 4: Repeat steps 2 to 3 R (R = 1000) times. Obtain the residual matrix Z1 R×N , Z2 R×N , Z3 R×N , Z4 R×N , where R is the number of repetitions 1000, and N is the number of samples 80.

[0089] Step 5: Ignoring the NaN values, calculate the mean and variance of each column of Z1 R×N , Z2 R×N , Z3 R×N , Z4 R×N to obtain four groups of mean vectors M1 1×N , M2 1×N , M3 1×N , M4 1×N and variance vectors V1 1×N , V2 1×N , V3 1×N , V4 1×N . The sample mean-variance diagram is as shown in Figure 8 .

[0090] Step 6: Calculate the percentage of each value in the mean vectors M1 1×N , M2 1×N , M3 1×N , M4 1×N ,... in the sum of N values to obtain the mean percentage vectors Mp1 1×N , Mp2 1×N , Mp3 1×N , Mp4 1×N ,.... Calculate the percentage of each value in the variance vectors V1 1×N , V2 1×N , V3 1×N , V4 1×N ,... in the sum of N values to obtain the variance percentage vectors Vp1 1×N, Vp2 1×N , Vp3 1×N , Vp4 1×N , ….

[0091] Step 7: According to the mean percentage vectors Mp1 1×N , Mp2 1×N , Mp3 1×N , Mp4 1×N , …, find the average value of each sample, that is, Mp 1×N= (Mp1 1×N (i) + Mp2 1×N (i) + Mp3 1×N (i) + Mp4 1×N (i) + …) / x, where x is the number of regression models; According to the variance percentage vectors Vp1 1×N , Vp2 1×N , Vp3 1×N , Vp4 1×N , …, find the average value of each sample, that is, Vp 1×N= (Vp1 1×N (i) + Vp2 1×N (i) + Vp3 1×N (i) + Vp4 1×N (i) + …) / x, x = 4.

[0092] Step 8: Sort the numbers in the vector Mp 1×N in descending order, delete the maximum value of the vector, and get the vector Mp 1×(N- 1); Sort the numbers in the vector Vp 1×N in descending order, delete the maximum value of the vector, and get the vector Vp 1×(N-1) . Select the mean value m of the top q numbers in the vector Mp 1×(N-1) and the mean value v of the top q numbers in the vector Vp 1×(N-1) . Obtain the threshold pair (2.2368, 2.4059).

[0093] Step 9: Samples that satisfy Mp 1×N ≥m and / or Vp 1×N ≥v are the finally identified abnormal samples S. The number of S is 11. The number of abnormal samples identified by the MCCV+PLS method is 13.

[0094] The performance parameters of the 4 types of models built based on three abnormal sample rejection methods are shown in Table 2. From the overall performance (average model performance) of the four models, the performance of the model obtained based on the method of the present invention (MMCODM) is better than that of the traditional MCCV+PLS method and the method only after SNV preprocessing.

[0095] Table 2 Comparison of the effects of three abnormal sample rejection results on the cereal dataset

[0096]

[0097] In the embodiments of the present application, the existing spectral data matrix X N×K and the concentration vector Y N×1 can be used to preprocess the spectral data matrix X N×K . Among them, N is the number of samples, and K is the number of spectral data points. The data preprocessing methods can include standard normal transformation, multiplicative scatter correction, detrending algorithm, first derivative processing, second derivative processing, Savitzky-Golay smoothing processing, etc. The purpose is to eliminate the influence of particle size, surface scattering and optical path transformation on diffuse reflection during the acquisition of spectral data, and lay a foundation for the identification of abnormal samples.

[0098] In an embodiment of the present application, the Monte Carlo random sampling method can be used to take some samples as the modeling set train-set. Preferably, 80% is selected, and the remaining part is used as the prediction set predict-set. The PLS, GPR, SVM, and BP modeling methods can be used to establish a regression model using the data of the modeling set train-set, use each regression model to predict the data of the prediction set predict-set, and calculate the prediction residuals of each sample in the predict-set. The prediction residuals of the samples in the modeling set train-set are counted as NaN. By repeating the above steps given in this embodiment several times, preferably more than 200 times, the residual matrices Z1 R×N , Z2 R×N , Z3 R×N , Z4 R×N. ,... are obtained. Through the above method, multiple prediction residuals of each sample can be obtained, which is convenient for obtaining the mean and variance of the more accurate prediction residuals of each sample, so as to more accurately judge whether the sample is an abnormal sample.

[0099] As Figure 9 shown, in one embodiment, an abnormal sample identification device is provided. This abnormal sample identification device can be integrated into the above computer device 120, and specifically can include: a single-model abnormal sample identification model construction module 510, a residual matrix, a mean vector and a variance vector acquisition module 520, a consensus mean vector and a consensus variance vector acquisition module 530, a threshold pair acquisition module 540, and an abnormal sample acquisition module 550.

[0100] The single-model abnormal sample identification model construction module 510 is used to construct q single-model abnormal sample identification models based on a sample data set composed of a plurality of spectral data matrices X N×K , and each of the single-model abnormal sample identification models is used to identify abnormal samples.

[0101] The residual matrix, mean vector, and variance vector acquisition module 520 is configured to perform predictions on the q single-model abnormal sample recognition models respectively, and obtain the sample residual matrix Z(q) of each single-model abnormal sample recognition model R×N , the mean vector M(q) of each residual 1×N and the variance vector V(q) 1×N , where q≥2, N represents the number of samples, K represents the number of spectral feature variables, and R is the number of repetitions.

[0102] The consensus mean vector and consensus variance vector acquisition module 530 is configured to obtain the consensus mean vector Mp 1×N based on the mean vector M(q) of each residual of all single-model abnormal sample recognition models 1×N ; and obtain the consensus variance vector Vp 1×N based on the variance vector V(q) of each residual of all single-model abnormal sample recognition models 1×N .

[0103] The threshold pair acquisition module 540 is configured to sort the values in the consensus mean vector Mp 1×N in descending order, delete the maximum value of the vector, and calculate the mean m of the top t values therein; sort the values in the consensus variance vector Vp 1×N in descending order, delete the maximum value of the vector, and calculate the mean v of the top t values in the vector, to obtain the threshold pair (m, v).

[0104] The abnormal sample acquisition module 550 is configured to use the samples that satisfy Mp 1×N ≥m and / or Vp 1×N ≥v as the finally identified abnormal samples S.

[0105] In the embodiments of the present application, for the explanation of the above device, reference may be made to the steps of an abnormal sample recognition method in the foregoing text, which will not be elaborated here.

[0106] Through an abnormal sample recognition method provided by the embodiments of the present application, the elimination of abnormal samples through multi-model consensus is realized, and abnormal samples are identified through the multi-model consensus method, so as to increase the correlation between the near-infrared spectrogram and the chemical value, eliminate the interference to the prediction performance of the model, and significantly improve the elimination accuracy of abnormal samples.

[0107] Figure 10 shows the internal structure diagram of a computer device in an embodiment. The computer device may specifically be Figure 1 the device 120 in Figure 10As shown, the computer device includes a processor, a memory, a network interface, an input device, and a display screen connected via a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor can implement an abnormal sample recognition method. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor can execute an abnormal sample recognition method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0108] Those skilled in the art can understand that Figure 10 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0109] In one embodiment, an abnormal sample recognition device provided by this application can be implemented in the form of a computer program, and the computer program can run on a computer device such as Figure 10 shown. Each program module constituting the abnormal sample recognition device can be stored in the memory of the computer device. For example, Figure 9 the single-model abnormal sample recognition model construction module 510, the residual matrix, the mean vector and variance vector acquisition module 520, the consensus mean vector and consensus variance vector acquisition module 530, the threshold pair acquisition module 540, and the abnormal sample acquisition module 550 shown in. The computer program composed of each program module enables the processor to execute the steps in an abnormal sample recognition method of each embodiment of this application described in this specification.

[0110] For example, Figure 10 the computer device shown can execute step S201 through the single-model abnormal sample recognition model construction module 510 in an abnormal sample recognition device such as Figure 9 shown, and the 520 module executes step S203, and so on.

[0111] In one embodiment, a computer device is proposed. The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of an abnormal sample recognition method.

[0112] In the embodiments of the present application, for the explanation of the above device operation method, reference can be made to the steps of an abnormal sample recognition method in the foregoing text, which will not be elaborated herein.

[0113] Through an abnormal sample recognition method provided by the embodiments of the present application, the elimination of abnormal samples through multi-model consensus is realized. Abnormal samples are recognized through the consensus of multiple models, so as to increase the correlation between near-infrared spectrograms and chemical values, eliminate the interference with the prediction performance of the model, and significantly improve the elimination accuracy of abnormal samples.

[0114] In one embodiment, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the processor is caused to execute the steps of an abnormal sample recognition method as described above.

[0115] In the embodiments of the present application, for the explanation of the above method, reference can be made to the steps of an abnormal sample recognition method in the foregoing text, which will not be elaborated herein.

[0116] Through an abnormal sample recognition method provided by the embodiments of the present application, the elimination of abnormal samples through multi-model consensus is realized. Abnormal samples are recognized through the consensus of multiple models, so as to increase the correlation between near-infrared spectrograms and chemical values, eliminate the interference with the prediction performance of the model, and significantly improve the elimination accuracy of abnormal samples.

[0117] It should be understood that although the steps in the flowcharts of the embodiments of the present application are shown sequentially according to the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0118] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0119] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0120] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. An abnormal sample recognition method, characterized in that, The method includes: Based on a sample data set composed of several spectral data matrices X N×K q single-model abnormal sample recognition models are constructed, and each of the single-model abnormal sample recognition models is used to recognize abnormal samples; Predict the q single - model abnormal sample recognition models respectively to obtain the sample residual matrix Z(q) of each single - model abnormal sample recognition model R×N , the mean vector M(q) of each sample residual 1×N and the variance vector V(q) 1×N , where q≥2, N represents the number of samples, K represents the number of spectral feature variables, and R is the number of repetitions; The mean vector M(q) of the respective residuals of all single-model anomaly sample recognition models 1×N , the consensus mean vector Mp is obtained 1×N ; Based on the variance vector V(q) of the respective residuals of all single-model anomaly sample recognition models 1×N , the consensus variance vector Vp is obtained 1×N ; Sort the values in the consensus mean vector \(M_p\) 1×N in descending order, calculate the mean \(m\) of the top \(t\) values after deleting the maximum value of the vector; Sort the values in the consensus variance vector \(V_p\) 1×N in descending order, calculate the mean \(v\) of the top \(t\) values in the vector after deleting the maximum value of the vector, and obtain the threshold pair \((m, v)\); Samples that meet Mp 1×N ≥ m and / or Vp 1×N ≥ v are used as the finally identified abnormal samples S.

2. The abnormal sample recognition method according to claim 1, characterized in that The spectral data matrix X N×K is obtained by the following method: Obtain the original spectral data matrix; Preprocess the original spectral data matrix to obtain the spectral data matrix, The preprocessing is used to eliminate the abnormal data generated by factors such as particle size, surface scattering, and optical path transformation in the original data matrix.

3. The abnormal sample recognition method according to claim 1, wherein Based on a sample data set composed of several spectral data matrices X N×K q single-model abnormal sample recognition models are constructed, and then the sample residual matrix Z(q) of each of the single-model abnormal sample recognition models is obtained R×N The method is as follows: Randomly select several data sets from the sample data set to construct the train-set data set, and use the sample data set that is not selected as the train-set data set as the predict-set data set; Based on the train-set data set, construct a single-model abnormal sample recognition model; Use the single-model abnormal sample recognition model to predict the predict-set data set, and calculate and collect the prediction residuals of each sample in the predict-set data set; Return to the step of randomly selecting several data sets from the sample data set to construct the train-set data set, and using the sample data set that is not selected as the train-set data set as the predict-set data set. Repeat the above backtracking step until the preset number of repetitions is reached, and collect the prediction residual data of each sample in all the predict-set data sets generated in the repeated execution of the above steps, so as to obtain the residual matrix Z(q) corresponding to each single-model abnormal sample recognition model R×N 。 4. An abnormal sample recognition method according to claim 3, characterized in that, Four single-model abnormal sample recognition models are constructed, which are regression models established based on the train-set data set using PLS, GPR, SVM, and BP modeling methods respectively.

5. The abnormal sample recognition method according to claim 3, wherein The mean vector M(q) of the respective residuals 1×N and the variance vector V(q) 1×N are obtained as follows: Count the prediction residuals of the train-set data set used as the modeling set as NaN, Ignoring NaN values, calculate the mean and variance of each column in the sample residual matrix \(Z(q)\) respectively, to obtain the mean vector \(M(q)\) of each sample residual R×N and the variance vector \(V(q)\) of each column 1×N 1×N .​ 6. The abnormal sample recognition method according to claim 1, wherein The mean vector M(q) of the respective residuals of all single-model anomaly sample recognition models 1×N , a consensus mean vector Mp is obtained 1×N ; based on the variance vector V(q) of the respective residuals of all single-model anomaly sample recognition models 1×N , a consensus variance vector Vp is obtained 1×N The method is as follows: Calculate the mean vector M(q) 1×N The percentage of each value in 1×N out of the total sum of all N values is obtained to get the mean percentage vector Mp(q) 1×N The mean percentage vector Mp(q) is calculated according to the following method 1×N The average value of each sample in 1×N is obtained to get the consensus mean vector Mp 1×N :[[]]END]] Mp 1×N = [Mp(1) 1×N (i) + Mp(2) 1×N (i) + … + Mp(q) 1×N (i)] / q, where i represents the i-th sample; Calculate the variance vector V(q) 1×N The percentage of each value in the total sum of all N values is obtained to get the variance percentage vector Vp(q) 1×N , and calculate the variance percentage vector Vp(q) according to the following method 1×N The average value of each sample in is obtained to get the consensus variance vector Vp 1×N :[[]]END]] Vp 1×N = [Vp(1) 1×N (i) + Vp(2) 1×N (i) + … + Vp(q) 1×N (i)] / q。 7. An abnormal sample recognition device, characterized in that, The abnormal sample recognition device includes: A single-model abnormal sample recognition model construction module, which is used to construct q single-model abnormal sample recognition models based on a sample data set composed of several spectral data matrices X N×K Each of the single-model abnormal sample recognition models is used to recognize abnormal samples; A residual matrix, mean vector, and variance vector acquisition module for respectively performing predictions on the q single-model abnormal sample recognition models to obtain the sample residual matrix Z(q) of each of the single-model abnormal sample recognition models R×N , the mean vector M(q) of each sample residual 1×N and the variance vector V(q) 1×N , where q ≥ 2, N represents the number of samples, K represents the number of spectral feature variables, and R is the number of repetitions; A consensus mean vector and consensus variance vector obtaining module, which is used to obtain a consensus mean vector Mp based on the mean vector M(q) of the respective residuals of all single-model anomaly sample recognition models 1×N ; and obtain a consensus variance vector Vp based on the variance vector V(q) of the respective residuals of all single-model anomaly sample recognition models 1×N ; 1×N ; 1×N ; Threshold pair acquisition module, configured to sort the values in the consensus mean vector Mp 1×N in descending order, delete the maximum value of the vector, and calculate the mean m of the top t values therein; sort the values in the consensus variance vector Vp 1×N in descending order, delete the maximum value of the vector, and calculate the mean v of the top t values in the vector, to obtain a threshold pair (m, v); An abnormal sample acquisition module, which is used to use the samples that satisfy Mp 1×N ≥m and / or Vp 1×N ≥v as the finally identified abnormal sample S.

8. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the processor executes the steps of an abnormal sample recognition method according to any one of claims 1 to 6.

9. An abnormal sample recognition system, characterized in that, When the system is running, it executes the steps of an abnormal sample recognition method according to any one of claims 1 to 6.