Industrial detection classification model quantitative improvement method based on parameter estimation theory

By constructing a multi-sample verification set and calculating confidence intervals, the problem of the difference in conventional verification accuracy when the test set changes is solved, and more accurate evaluation of the overall accuracy of the classification model and robust decision support are achieved.

CN120180177APending Publication Date: 2025-06-20SHENZHEN TIEYUE ELECTRIC CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510153775.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Conventional verification accuracy can only reflect the performance of the classification model in a specific verification set. If the test set changes, the accuracy may vary greatly.

Method used

By constructing a verification set containing multiple samples, calculating sampling accuracy and average sampling errors, a confidence interval for the overall accuracy of the binary detection classification model is constructed to more accurately evaluate the overall accuracy of the model.

Benefits of technology

It provides a more accurate estimate of the overall accuracy of the model, helps to understand the uncertainty of model performance, provides more robust support for decision-making, and is suitable for a variety of binary classification models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180177A_ABST
    Figure CN120180177A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial detection classification model quantification improvement method based on a parameter estimation theory, and belongs to the technical field of data processing, and the method specifically comprises the following steps: constructing a verification set containing n samples based on a marked industrial scene data set, inputting each sample in the verification set into a trained binary detection classification model, and obtaining a binary detection classification model; the model outputs a prediction label, and the number of samples which are correctly predicted to be Posive and the number of samples which are wrongly predicted to be Posive in the prediction labels are counted; the sampling precision is calculated; according to repeated sampling, calculating an average sampling error based on the sampling precision; according to the sampling precision and the average sampling error, the confidence interval of the total precision of the binary detection classification model is constructed, the numerical value of the total precision of the model is expressed under the given confidence level, and the quantitative improvement effect on the industrial detection classification model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a quantization improvement method for an industrial detection classification model based on parameter estimation theory. Background Art

[0002] With the rapid development of Industry 4.0 and intelligent manufacturing, industrial detection and sensing technologies play an increasingly important role in modern industrial production. By using machine learning and artificial intelligence technologies, industrial detection systems can automatically identify and predict abnormal situations in the production process, thereby improving production efficiency and product quality. In these applications, binary detection classification models are widely used to detect the state of targets, such as determining whether a person wears a safety helmet, makes a phone call, smokes, or predicting whether equipment needs maintenance.

[0003] Binary detection classification is a common task in machine learning. Whether it is naive Bayes based on probability, clustering theory, or convolutional classification networks based on deep learning, they all need to be trained to output a classification model that maps the input vector to a Negative or Positive result label.

[0004] Commonly, the classification effect of the model is measured by precision. The definition of precision P is the ratio of the number of samples correctly inferred as the Positive label to the total number. In this case, the calculated precision P is based on the inference results of a specific number of samples, and the sum of these samples becomes the validation set. Precision is applicable to all fields of classification detection. Strictly speaking, precision can only reflect the performance of the classification model under a specific validation set. If the test set changes, the precision may vary greatly, that is, precision does not have the ability to evaluate the overall generalization. Summary of the Invention

[0005] The purpose of the present invention is to provide a quantization improvement method for an industrial detection classification model based on parameter estimation theory, and solve the following technical problems:

[0006] Conventional verification precision can only reflect the performance of the classification model under a specific validation set. If the test set changes, the precision may vary greatly.

[0007] The purpose of the present invention can be achieved by the following technical solutions:

[0008] A quantization improvement method for an industrial detection classification model based on parameter estimation theory, comprising the following steps:

[0009] Construct a validation set with n samples based on the labeled industrial scenario dataset, where n is a positive integer. Input each sample in the validation set into the trained binary detection and classification model, and the model outputs a predicted label.

[0010] Count the number of samples N that are correctly predicted as Positive among these predicted labels TP , and the number of samples N that are incorrectly predicted as Positive FP ; According to the number of samples N TP and the number of samples N FP , calculate the sampling precision P N ;

[0011] According to repeated sampling, calculate the average sampling error σ based on the sampling precision P N ; p ;

[0012] According to the sampling precision P N and the average sampling error σ p Construct the confidence interval of the overall precision of the binary detection and classification model , representing the value of the overall precision of the model at a given confidence level.

[0013] As a further solution of the present invention: The calculation formula for the average sampling error σ p is:

[0014]

[0015] Where P N is the sampling precision, PN = N TP / (N TP + N FP ).

[0016] As a further solution of the present invention: The construction formula for the confidence interval of the overall precision of the binary detection and classification model is:

[0017]

[0018] Where Z represents the probability degree.

[0019] As a further solution of the present invention: The samples in the validation set are randomly selected from the original dataset with the same distribution as the training set.

[0020] As a further solution of the present invention: Before the samples in the validation set are input into the model, preprocess the samples in the validation set, including data cleaning, missing value filling, and outlier handling.

[0021] As a further solution of the present invention: Before the samples in the validation set are input into the model, feature engineering processing is also performed, including feature selection and feature transformation.

[0022] As a further solution of the present invention: The binary detection and classification model is based on a logistic regression model, a support vector machine model, or a decision tree model.

[0023] As a further solution of the present invention: Evaluate the constructed confidence interval. If the width of the confidence interval is large, increase the number of samples in the validation set and recalculate the sampling precision and average sampling error to narrow the confidence interval to the set standard.

[0024] As a further solution of the present invention: Divide the samples in the validation set into multiple subsets, each subset corresponding to a different feature distribution or data source. Calculate the sampling precision and average sampling error for each subset respectively, and construct the corresponding confidence interval to evaluate the precision of the model under different feature distributions or data sources.

[0025] As a further solution of the present invention: During the construction process of the validation set, the stratified sampling method is adopted to ensure that the proportion of different category samples in the validation set is consistent with the proportion of the corresponding category samples in the original dataset.

[0026] Advantages of the present invention:

[0027] By constructing a validation set containing multiple samples, the method of the present invention can more comprehensively evaluate the performance of the classification model. Calculating the confidence interval using the sampling precision and average sampling error provides a more accurate estimate of the overall precision of the model. Introducing the concept of the confidence interval makes the evaluation of the model performance not limited to a single value, but provides a range that reflects the possible fluctuations of the model precision at a given confidence level. This helps to understand the uncertainty of the model performance and provides more robust support for decision-making. The present invention is applicable to a variety of common binary classification models, such as logistic regression models, support vector machine models, and decision tree models. This makes the method of the present invention have wide applicability and can be applied to different industrial detection and perception scenarios; through more accurate model evaluation, the present invention can help the industrial detection system more reliably identify and predict abnormal situations, thereby improving production efficiency. Brief description of the drawings

[0028] The present invention will be further described below with reference to the drawings.

[0029] Figure 1 It is a flow schematic diagram of the present invention. Detailed implementation manners

[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0031] Please refer to Figure 1 As shown, the present invention is an improved method for quantifying an industrial detection classification model based on parameter estimation theory, including:

[0032] Embodiment 1:

[0033] Constructing a validation set: In the present invention, the construction of the validation set is based on the labeled industrial scenario dataset. Specifically, first, samples with clear annotation information are selected from a large number of industrial scenario data. These annotation information usually includes the status of the target object (such as Positive or Negative) to ensure that the samples in the validation set can accurately reflect the actual industrial detection scenario. Then, randomly select n samples from these labeled samples to form a validation set, where n is a positive integer. This validation set will be used in the subsequent model evaluation process to ensure that the evaluation results are representative and reliable. Randomly draw n samples from the original dataset with the same distribution as the training set to form a validation set. This step ensures that the validation set is consistent with the training set in statistical characteristics, so that the evaluation results on the validation set can better reflect the performance of the model in actual applications. If the training set is extracted from a certain fixed industrial scenario, then the validation set should also be randomly drawn from the data of the same source to ensure its representativeness. During the construction of the validation set, the method of stratified sampling is adopted to ensure that the proportion of different category samples in the validation set is consistent with the proportion of corresponding category samples in the original dataset.

[0034] Model prediction and label statistics: Input each sample in the validation set into the trained binary detection classification model. Here, the binary detection classification model can be a logistic regression model, a support vector machine model, a decision tree model, etc. The model will output a predicted label for each sample. Then, count the number of samples N TP correctly predicted as Positive (positive class) among these predicted labels, and the number of samples N FP wrongly predicted as Positive. This step is the key to evaluating the model performance because it directly reflects the performance of the model on actual data.

[0035] Calculating Sampling Precision: Before calculating the sampling precision, preprocess the samples in the validation set, including data cleaning, missing value imputation, and outlier handling. These preprocessing steps help improve the quality of the data and the accuracy of subsequent analysis. Feature engineering is also carried out, including feature selection and feature transformation. After completing the preprocessing, calculate the sampling precision P TP and N FP values to calculate the sampling precision P N , and its calculation formula is P N = N TP / (N TP + N FP ). This step is an important part of evaluating the model's accuracy because only accurate data can lead to reliable conclusions.

[0036] When evaluating the performance of a classification model, simply calculating the sampling precision PN once may not fully reflect the reliability and stability of the model. To gain a deeper understanding of the model's performance on different datasets, we use the method of repeated sampling to calculate the average sampling error σ p . This step aims to obtain a more robust error estimate through multiple samplings and calculations, thereby more accurately evaluating the overall performance of the model.

[0037] Conduct multiple (e.g., k times) independent and random samplings from the original dataset, and each time draw n samples to form a new validation set. In this way, we can obtain k different validation sets, and each validation set represents a subset of the original dataset.

[0038] For each newly generated validation set, input each sample into the trained binary detection classification model. The model will output a predicted label for each sample. Then, count the sampling precision of these predicted labels, and based on the sampling precision values, we can calculate the average sampling error σ p . The calculation formula for the average sampling error is:

[0039]

[0040] When evaluating the performance of a classification model, we not only care about the average performance of the model but also hope to understand its stability and reliability under different conditions. To this end, we can use the sampling precision PN and the average sampling error σ p to construct a confidence interval for the overall precision of the binary detection classification model. This confidence interval can represent the possible range of the model's overall precision at a given confidence level, thus providing us with a more comprehensive view of the model's performance.

[0041] Specifically, the process of constructing the confidence interval is as follows:

[0042] Determine the confidence level: First, we need to select a confidence level, such as 95%. This confidence level represents our confidence in the overall accuracy estimate of the model. For a 95% confidence level, the corresponding Z-value (i.e., the probability degree) is approximately 1.96.

[0043] Calculate the confidence interval: Using the sampling precision P N and the average sampling error σ p , we can construct the confidence interval of the overall accuracy of the model. The construction formula of the confidence interval is:

[0044]

[0045] Evaluate the confidence interval: Evaluate the constructed confidence interval. If the width of the confidence interval is large, it indicates that there is a large uncertainty in our estimate of the overall accuracy of the model. In this case, we can increase the sample size of the validation set to improve the accuracy of the estimate. Specifically, we can recalculate the sampling precision and the average sampling error to narrow the confidence interval to the set standard.

[0046] The confidence interval (Confidence Interval, CI) is an important concept in statistics for estimating the range of population parameters (various precisions in the present invention). It is an interval calculated based on sample data, indicating the probability that the population parameter falls within this interval (referred to as the confidence level). The confidence interval consists of two parts: the estimated value and the error range. The estimated value is usually the sample statistic, and the error range in the present invention generally reflects the uncertainty caused by sampling randomness.

[0047] Subset analysis: To further evaluate the performance of the model under different feature distributions or data sources, we can divide the samples in the validation set into multiple subsets. Each subset corresponds to a different feature distribution or data source. Then, we calculate the sampling precision and the average sampling error for each subset respectively, and construct the corresponding confidence intervals. In this way, we can obtain detailed information about the performance of the model under different conditions, so as to more comprehensively evaluate the applicability and stability of the model.

[0048] Improvement theory:

[0049] 1. Adopting the classical statistical view, a specific validation set of the binary detection classification model is a random sampling of the population. Therefore, estimating the overall parameter accuracy from the accuracy of the finite sample parameters has more generalization significance.

[0050] 2. The precision is a percentage, and the "proportion" sampling theory in statistics is used to estimate the overall accuracy.

[0051] 3. Since it is an estimation of the population by samples, the accuracy of the estimated parameter is an interval and is associated with the probability degree (confidence level).

[0052] 4. The estimated population accuracy interval can better and more comprehensively reflect the generalization ability of the binary detection and classification model, and significantly improve the problem of the specificity deviation of the single-point accuracy for the validation set.

[0053] Example 2:

[0054] The above improved method using parameter estimation is not only applicable to the accuracy (Precision) index of the binary detection and classification model, but also applicable to other classification model evaluation indexes due to its theoretical generality. Including:

[0055] Recall (Recall, Sensitivity):

[0056]

[0057] Where: N TP is the number of samples for which the binary detection and classification model correctly infers the Positive label, and N FN is the number of samples for which the binary detection and classification model wrongly infers the Negtive label.

[0058] By calculating the recall rate, the performance of the model in identifying Positive samples can be evaluated, that is, how many samples that are actually Positive can the model correctly identify. This is very important in some application scenarios. A high recall rate means that more abnormal scenarios can be detected and the situation of missed detections can be reduced.

[0059] The recall rate and the precision rate (Precision) are two complementary indicators. Using them in combination can balance the tendency of the model between identifying positive samples and avoiding misjudging negative samples. For example, in some cases, it may be necessary to increase the recall rate to ensure that important samples are not missed, even if this may lead to some misjudgments.

[0060] Accuracy:

[0061]

[0062] Where: N TP is the number of samples for which the binary detection and classification model correctly infers the Positive label, and N FN is the number of samples for which the binary detection and classification model wrongly infers the Negtive label, and N TN is the number of samples for which the binary detection and classification model correctly infers the Negtive label, and N FP is the number of samples for which the binary detection and classification model wrongly infers the Positive label.

[0063] Accuracy is an indicator that measures the proportion of correctly classified samples in the total samples of the model, which can intuitively reflect the classification effect of the model on the overall data. For a dataset with balanced classes, accuracy is a simple and effective evaluation indicator, facilitating the understanding and comparison of the performance of different models.

[0064] During the model selection and optimization process, accuracy can be used as one of the important reference indicators. By comparing the accuracies of different models on the same dataset, a model with better performance can be selected, or when tuning the model parameters, adjustments can be made with the goal of improving accuracy.

[0065] Specificity:

[0066]

[0067] Where: N TN is the number of samples correctly inferred as the Negative label by the binary detection classification model, and N FP is the number of samples wrongly inferred as the Positive label by the binary detection classification model.

[0068] Specificity reflects the performance of the model in identifying Negative samples, that is, how many actual Negative samples the model can correctly identify. In fields such as industrial environmental monitoring, a high specificity means that the risk of misjudging the degree of environmental pollution can be reduced, avoiding unnecessary treatment measures.

[0069] Specificity and recall jointly focus on the classification effect of the model on different classes. Recall focuses on the identification of Positive samples, while specificity focuses on the identification of Negative samples. The combination of the two can more comprehensively evaluate the classification performance of the model, especially in the case of class imbalance.

[0070] F1 Score:

[0071]

[0072] Where: N TP is the number of samples correctly inferred as the Positive label by the binary detection classification model, and N FP is the number of samples wrongly inferred as the Positive label by the binary detection classification model, and N FN is the number of samples wrongly inferred as the Negative label by the binary detection classification model.

[0073] The F1 score is the harmonic mean of precision and recall, which can comprehensively consider the accuracy and integrity of the model when identifying positive samples. When the model needs to balance precision and recall, the F1 score provides a comprehensive evaluation metric, avoiding the bias that a single metric may bring.

[0074] In an imbalanced dataset, the F1 score can better reflect the performance of the model. Compared with accuracy, the F1 score is not affected by the difference in the number of classes and can more accurately evaluate the classification effect of the model on minority classes, helping to optimize the model to improve the recognition ability of minority classes.

[0075] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0076] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0077] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or modules can be in electrical, mechanical, or other forms.

[0078] In addition, in each embodiment of the present application, the various functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.

[0079] If the above functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0080] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0081] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A quantitative improvement method for industrial detection classification model based on parameter estimation theory, characterized in that: The following steps are involved: Based on the labeled industrial scene dataset, a validation set containing n samples is constructed, where n is a positive integer. Each sample in the validation set is input into the trained binary detection and classification model, and the model outputs a predicted label. Count the number of samples N that are correctly predicted as Positive among these predicted labels TP , and the number of samples N that are incorrectly predicted as Positive FP ; According to the sample size N TP and the sample size N FP , calculate the sampling precision P N ; According to repeated sampling, based on the sampling precision P N Calculate the average sampling error σ p ; According to the sampling precision P N and the average sampling error σ p Constructing the overall accuracy of the binary detection classification model The confidence interval of , which represents the numerical value of the overall accuracy of the binary detection classification model at a given confidence level.

2. According to claim 1, a quantitative improvement method for industrial detection classification model based on parameter estimation theory is characterized in that: Average sampling errorσ p The calculation formula is: Among them, P N is the sampling precision, P N =N TP / (N TP +N FP ).

3. According to claim 1, a quantitative improvement method for industrial detection classification model based on parameter estimation theory is characterized in that: The construction formula for the overall accuracy confidence interval of the binary detection classification model is: Where Z represents the probability.

4. According to claim 1, a quantitative improvement method for industrial detection classification model based on parameter estimation theory is characterized in that: The samples in the validation set are randomly selected from the original data set with the same distribution as the training set.

5. According to claim 1, a quantitative improvement method for industrial detection classification model based on parameter estimation theory is characterized in that: Before the samples in the validation set are input into the model, the samples in the validation set are preprocessed, including data cleaning, missing value filling and outlier processing.

6. The method for quantitatively improving an industrial detection classification model based on parameter estimation theory according to claim 1 is characterized in that: The samples in the validation set are also subjected to feature engineering before being input into the model. Including feature selection and feature transformation.

7. The method for quantitatively improving an industrial detection classification model based on parameter estimation theory according to claim 1 is characterized in that: The binary detection classification model is based on a logistic regression model, a support vector machine model or a decision tree model.

8. The method for quantitatively improving an industrial detection classification model based on parameter estimation theory according to claim 1 is characterized in that: The constructed confidence interval is evaluated. If the width of the confidence interval is greater than the set standard, the number of samples in the validation set is increased, and the sampling accuracy and average sampling error are recalculated to narrow the confidence interval to the set standard.

9. The method for quantitatively improving an industrial detection classification model based on parameter estimation theory according to claim 1 is characterized in that: The samples in the validation set are divided into multiple subsets, each subset corresponds to a different feature distribution or data source, the sampling accuracy and average sampling error are calculated for each subset, and the corresponding confidence interval is constructed to evaluate the accuracy of the model under different feature distributions or data sources.

10. The method for quantitatively improving an industrial detection classification model based on parameter estimation theory according to claim 1 is characterized in that: During the construction of the validation set, a stratified sampling method is adopted to ensure that the proportion of samples of different categories in the validation set is consistent with the proportion of samples of corresponding categories in the original data set.