Part quality grade classification method, device, equipment, medium and product based on uncertain data pattern classification
By employing a genetic algorithm with improved fitness functions and an evidence K-nearest neighbor model, the method addresses inefficiencies and inaccuracies in classifying component quality levels using uncertain data, achieving improved efficiency and accuracy.
Patent Information
- Application Number
- CN202510264674.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-07
AI Technical Summary
Existing uncertain data processing and pattern recognition methods are inefficient and inaccurate when dealing with the classification of part quality grades. Especially in the field of industrial manufacturing, the uncertainty of part size data affects the accuracy of the classification.
The fitness function of the improved genetic algorithm based on particle size and elastic network is adopted, and the optimal subset of features is screened through the feature selection genetic algorithm, and the part quality level classification is classified based on the evidence K nearest neighbor model, and the trust function is used to deal with uncertainty.
It significantly improves the efficiency and accuracy of part quality level classification in uncertain data environments, and achieves more accurate data classification by optimizing feature selection and processing of uncertain information.
Smart Images

Figure CN119760488B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and particularly to a method, device, equipment, medium and product for classifying part quality grades based on uncertain data pattern classification. Background Art
[0002] In the context of current artificial intelligence and data science, one of the main challenges is how to effectively process and analyze uncertain data. Uncertain data may come from various scenarios, such as sensor data acquisition, social media analysis, medical diagnosis, etc. They often contain noise, missing values or fuzzy information. Existing methods for processing uncertain data and pattern recognition may be inefficient or inaccurate when dealing with such uncertain data. Therefore, developing new algorithms and models to improve the ability to process uncertain data while maintaining high efficiency and accuracy has become an important research direction.
[0003] In the field of industrial manufacturing, part size data is affected by factors such as precision and wear, resulting in uncertainty, which is a typical type of uncertain data. If such uncertain data is directly used for classifying part quality grades, on the one hand, the processing efficiency will be low due to excessive invalid or redundant data, and on the other hand, the uncertainty of part size data will greatly affect the accuracy of part quality grade division. Summary of the Invention
[0004] The purpose of the present application is to provide a method, device, equipment, medium and product for classifying part quality grades based on uncertain data pattern classification, so as to improve the efficiency and accuracy of part quality grade classification in an uncertain data environment.
[0005] To achieve the above purpose, the present application provides the following solutions.
[0006] In the first aspect, the present application provides a method for classifying part quality grades based on uncertain data pattern classification, including:
[0007] Obtaining part size data with multiple dimensional features of different parts and different processing stages of a part monitored by a sensor;
[0008] Performing data preprocessing on the part size data to obtain preprocessed part size data;
[0009] Based on the fitness function of the genetic algorithm improved by granularity and elastic net, forming a feature selection genetic algorithm;
[0010] Using the feature selection genetic algorithm to screen multiple dimensional features of the preprocessed part size data to obtain an optimal feature subset;
[0011] Extract the feature data corresponding to the optimal feature subset from the preprocessed part size data, and construct a feature data set;
[0012] Based on the feature data set, use the evidence K-nearest neighbor model to classify the part quality level.
[0013] Optionally, the data preprocessing of the part size data to obtain the preprocessed part size data specifically includes:
[0014] Apply denoising, missing value filling, and normalization techniques to preprocess the part size data to obtain the preprocessed part size data.
[0015] Optionally, the fitness function of the genetic algorithm improved based on granularity and elastic net forms a feature selection genetic algorithm, specifically including:
[0016] Use the chromosome of each individual in the genetic algorithm to represent a set of feature selection schemes, and each gene on the chromosome represents whether the corresponding feature is selected; assume there are individuals in the population, and the gene coding length of each individual is , use to represent the chromosome of the rd individual, where represents 's th gene position; ; ; The entire population is represented as ;
[0017] Assume is the population of the current generation, contains all the genes of the current generation, and represents all the selected genes in the current generation population. According to the formula calculate the granularity of the individuals in the current generation; where is the selected feature subset in the current generation chromosome , , is the number of selected chromosome entries; is the number of selected features in the feature subset ; is the total number of genes in the current generation;
[0018] The elastic net calculates the prediction error of the folded data through cross-validation ; where is the true label of the test sample ; is the predicted label of the test sample ; is the number of test samples;
[0019] Based on the granularity of the current generation of individuals and the prediction error Construct an improved fitness function based on granularity and elastic net ; where is the weight factor; is the calculated fitness value;
[0020] Use the improved fitness function as the fitness function of the genetic algorithm to form a feature selection genetic algorithm.
[0021] Optionally, screening multiple dimensional features of the preprocessed part size data using the feature selection genetic algorithm to obtain an optimal feature subset, specifically including:
[0022] Randomly generate an initial population of the feature selection genetic algorithm. The chromosome of each individual in the initial population represents a set of feature selection schemes. Each gene on the chromosome represents whether the corresponding feature is selected, and the gene coding length of each individual is equal to the dimension of the features included in the preprocessed part size data;
[0023] Start iterative calculation of the genetic algorithm from the initial population. For each individual in each generation of the population, calculate the fitness value using the improved fitness function until a predetermined number of iterations is reached;
[0024] After reaching the predetermined number of iterations, output the feature subset represented by the chromosome of the individual with the highest fitness value as the optimal feature subset.
[0025] Optionally, extracting the feature data corresponding to the optimal feature subset from the preprocessed part size data to construct a feature dataset, specifically including:
[0026] Each feature has a corresponding index position in the optimal feature subset. In the preprocessed part size data, locate and extract the data of the corresponding feature column according to the index position corresponding to the feature as the corresponding feature data; the feature data extracted for all preprocessed part size data together constitutes a feature dataset.
[0027] Optionally, using the evidence K-nearest neighbor model for part quality grade classification based on the feature dataset, specifically including:
[0028] Divide the feature dataset into a training set and a test set, use the training set to train the evidence K-nearest neighbor model, and use the test set to test the prediction error of the evidence K-nearest neighbor model based on the optimal feature subset ; After training, use the trained evidence K-nearest neighbor model as the part quality grade classification model;
[0029] Input the feature vector of the part dimension data to be classified into the part quality grade classification model, and output the corresponding part quality grade; the part quality grades include qualified, slightly unqualified, and severely unqualified.
[0030] In a second aspect, the present application provides a part quality grade classification device based on uncertain data pattern classification, including:
[0031] An uncertain data acquisition module for acquiring part dimension data with multiple-dimensional features at different parts and different processing stages of a part monitored by a sensor;
[0032] A data preprocessing module for preprocessing the part dimension data to obtain preprocessed part dimension data;
[0033] A fitness function improvement module for improving the fitness function of a genetic algorithm based on granularity and elastic net to form a feature selection genetic algorithm;
[0034] A feature screening module for screening multiple-dimensional features of the preprocessed part dimension data using the feature selection genetic algorithm to obtain an optimal feature subset;
[0035] A feature data extraction module for extracting feature data corresponding to the optimal feature subset from the preprocessed part dimension data to construct a feature data set;
[0036] A part quality grade classification module for classifying part quality grades using an evidence K-nearest neighbor model based on the feature data set.
[0037] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement the part quality grade classification method based on uncertain data pattern classification.
[0038] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the part quality grade classification method based on uncertain data pattern classification.
[0039] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the part quality grade classification method based on uncertain data pattern classification.
[0040] According to the specific embodiments provided by the present application, the following technical effects are disclosed in the present application.
[0041] A method, device, equipment, medium and product for classifying part quality grades based on uncertain data pattern classification provided by the present application, which improves the fitness function of the genetic algorithm based on granularity and elastic net, optimizes and selects data features affecting the classification result through the genetic algorithm, and extracts feature data corresponding to the optimal feature subset; the selected feature data is input into the evidential K-nearest neighbor model, and the belief function is used to quantify the uncertainty, so as to perform more effective pattern classification on the uncertain data, and can significantly improve the efficiency and accuracy of part quality grade classification in an uncertain data environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0043] Figure 1 It is a schematic flowchart of a method for classifying part quality grades based on uncertain data pattern classification of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0045] In today's data-intensive application scenarios, the uncertainty and incompleteness of data have become an issue that cannot be ignored, and the existence of these issues significantly affects the accuracy and reliability of data analysis. To overcome this problem, the present application proposes a method, device, equipment, medium and product for classifying part quality grades based on uncertain data pattern classification, aiming to optimize feature selection and process uncertain information by combining an improved feature selection genetic algorithm and an evidential K-nearest neighbor model, so as to achieve more accurate data classification and improve the efficiency and accuracy of part quality grade classification in an uncertain data environment.
[0046] To make the above objects, features and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0047] In an exemplary embodiment, as Figure 1 shown, a method for classifying part quality grades based on uncertain data pattern classification is provided, including the following steps 1 to 6.
[0048] Step 1: Obtain the part size data with multiple dimensional features at different parts and different processing stages of the part monitored by sensors.
[0049] In the field of industrial manufacturing, the part size data is affected by factors such as accuracy and wear, resulting in uncertainty. Monitoring the part status using sensors or other data acquisition methods can obtain the part size data with multiple dimensional features at different parts and different processing stages of the part. The multiple dimensional features of the part size data can be, for example, the numerical range feature of the part size data, the change feature of different parts, the change feature of different processing stages and different times, and so on.
[0050] Step 2: Perform data preprocessing on the part size data to obtain the preprocessed part size data.
[0051] The original part size data obtained through sensors or other data acquisition methods may have noise, incompleteness or ambiguity, resulting in uncertainty. In the data preprocessing stage, techniques such as denoising, missing value filling, and normalization are applied to reduce the uncertainty in the data and obtain the preprocessed part size data. The preprocessed part size data is used in the subsequent feature selection, model training, and testing processes, and will not be specifically distinguished from the original part size data.
[0052] Step 3: Improve the fitness function of the genetic algorithm based on granularity and elastic net to form a feature selection genetic algorithm.
[0053] The genetic algorithm (GA) is a search algorithm that simulates natural selection and genetic mechanisms and solves optimization problems by simulating the principle of "survival of the fittest". The main steps of the genetic algorithm are as follows, S1 to S6.
[0054] S1 Population initialization: Randomly generate an initial population. Each individual's chromosome in the population represents a set of feature selection schemes, and each gene on each chromosome represents whether the corresponding feature is selected. In the genetic algorithm, each individual is represented by a string of binary codes, and this code is called a chromosome. Assume there are individuals in the population, and the gene coding length of each individual is . Use to represent the chromosome of the th individual, which can also be called individual . Among them, represents the th gene position of individual , and is a binary value, which can be 0 or 1. For example, being 0 means the corresponding feature is not selected, Let 1 indicate the selection of the corresponding feature. Then ; . This encoding process can be expressed as:
[0055] (1);
[0056] The entire population can be expressed as:
[0057] (2).
[0058] S2 Fitness calculation: For each individual in the population, its performance is evaluated through a fitness function. The fitness function is used to evaluate the performance of each individual in the current generation. Individuals with high fitness values are more likely to be selected as "excellent parents" and retained for mating to generate the next generation. The design of the fitness function should be closely related to the goal of feature selection. In this application, the fitness function of the genetic algorithm is improved based on granularity and elastic net, and the improved fitness function comprehensively considers the classification accuracy and the granularity information of feature selection.
[0059] S3 Selection operation: According to the calculated fitness values, a roulette wheel or other selection mechanism is used to select individuals with high fitness to be passed on to the next generation, ensuring the inheritance of excellent genes in the population.
[0060] S4 Crossover and mutation. Crossover: Randomly pair individuals and perform crossover operations according to the crossover probability to generate new offspring. Mutation: Randomly mutate the genes of individuals according to the mutation probability to increase the diversity of the population.
[0061] S5 Formation of the new generation population: After selection, crossover, and mutation operations, a new generation population is formed for the next round of iteration.
[0062] S6 Algorithm iteration: Repeat S2 to S5 until a predetermined number of iterations is reached. At this time, it is considered that the algorithm has converged to the optimal or approximate optimal solution, and the individual with the highest fitness value is output, called the optimal individual.
[0063] Finally, the genetic algorithm outputs the feature subset represented by the individual (chromosome) with the highest fitness value as the optimal feature subset for subsequent data analysis or pattern recognition tasks.
[0064] Step 3 specifically includes the following steps 3.1 to 3.5.
[0065] Step 3.1: Use the chromosome of each individual in the genetic algorithm to represent a set of feature selection schemes, and each gene on the chromosome represents whether the corresponding feature is selected; assume that there are Individuals, each with a gene encoding length of , using to represent the chromosome of the th individual, where represents the th gene position; ; ; The entire population is represented as .
[0066] Step 3.2: Assume is the population of the current generation, contains all the genes of the current generation, and represents all the selected genes in the current generation population. Calculate the granularity of the individuals in the current generation according to the formula ; where is the selected feature subset in the current generation chromosome , , is the number of selected chromosome entries; is the number of selected features in the feature subset ; is the total number of genes in the current generation.
[0067] Granularity is used to measure the coverage of the feature subset in the feature space and control the number and diversity of the selected features. Its main functions are as follows: ① Controlling the size of the feature subset: Granularity encourages the selection of relatively small feature subsets, which can avoid selecting too many features, thus reducing the risk of feature redundancy and overfitting. By introducing granularity, the genetic algorithm can tend to select more compact and efficient feature subsets. ② Promoting effective feature selection: Granularity guides the genetic algorithm to select subsets that can most effectively cover the feature space in each generation by measuring the "sparsity" or "density" of the feature subset. Therefore, granularity can help the genetic algorithm select features that can fully capture data changes, thereby improving the performance of the model. ③ Improving the convergence efficiency: In the genetic algorithm, granularity helps accelerate the convergence process and reduce the selection of invalid or redundant features. By gradually optimizing the granularity, the genetic algorithm can effectively find the optimal feature subset and accelerate the convergence to the optimal solution.
[0068] Specifically, assume is the population of the current generation, contains all the genes of the current generation, and represents all the selected genes in the current generation population, where is the selected feature subset in the current generation chromosome , and here , is the number of selected chromosome entries. Then the granularity of the individuals in the current generation is defined as:
[0069] (3);
[0070] where is the number of selected features in the feature subset ; is the total number of genes in the current generation. If , it means that the th chromosome selects the fewest features; if , it means that the th chromosome selects all features.
[0071] Step 3.3: The Elastic Net calculates the prediction error of the folded data through cross-validation ; where is the true label of the test sample ; is the predicted label of the test sample ; is the number of test samples.
[0072] The Elastic Net (Elastic Net, abbreviated as Elastic Net) regularization combines L1 (Lasso) and L2 (Ridge) regularizations and is widely used in regression problems, especially feature selection in high-dimensional data. Its main functions are as follows: 1) Improve the effectiveness of feature selection: The Elastic Net uses L1 regularization to force the coefficients of some features to become zero, thereby achieving feature selection. L2 regularization helps control the magnitude of the regression coefficients and prevent overfitting. The combination of the two enables the Elastic Net to find a balance between feature selection and regularization, removing unimportant features and retaining features useful for the model. 2) Adapt to high-dimensional data: In high-dimensional data, the number of features may be much larger than the number of samples. The Elastic Net can effectively perform feature selection and avoid overfitting problems caused by high data dimensionality. By improving the fitness function through the Elastic Net, the genetic algorithm can screen out the most valuable features in the high-dimensional feature space and improve the generalization ability of the model.
[0073] The definition of the Elastic Net is as follows. Suppose there is a linear regression polynomial model as , where is the response variable (target), is the design matrix containing all feature variables, is the coefficient vector to be estimated, is the error term. Suppose there are in total predictor variables, denoted as Based on linear regression, the response variable The estimate is expressed as The coefficient vector It is calculated by minimizing the sum of squared errors. The objective function of the elastic net can be written as:
[0074] (4);
[0075] in,
[0076] (5);
[0077] (6);
[0078] In the formula, Express request Minimum value of the intrinsic function. The first item It is used to minimize the sum of squared errors of the fitted data. represents the L1 norm, which promotes the sparsity of coefficients and encourages some coefficients to become zero, thereby achieving feature selection. Represents the L2 norm, which controls the sum of squares of coefficients to prevent overfitting.
[0079] Elastic Net calculates prediction error on fold data via cross validation , evaluate the performance of the algorithm.
[0080] (7);
[0081] in It is a test sample The true label of It is a test sample The predicted label of is the number of test samples.
[0082] Step 3.4: Granularity based on the current generation of individuals and prediction error Constructing an improved fitness function based on granularity and elastic network ;in is the weight factor; Calculated fitness value.
[0083] The improved fitness function based on granularity and elastic network is set as:
[0084] (8);
[0085] in, The model prediction error value obtained for the subset of features corresponding to the genetic code of an individual in the current generation; Represents the granularity of an individual in the current generation. Since the granularity is expected to increase over time and limit the impact on the fitness function, a weight factor is added in front. , usually = 0.15. This value is not arbitrary; rather, it is determined through multiple iterations and experimental verification to ensure that the granularity makes an appropriate contribution to the overall fitness assessment, balancing and the influence of the feature subset size on the performance of the genetic algorithm.
[0086] Feature selection in the genetic algorithm uses the improved fitness function (8) to identify the optimal individual in each iteration. The elastic net removes redundant features and retains the most valuable ones. Features that frequently appear in the optimal individuals are selected as the final features, forming the optimal feature subset. This process aims to reduce and increase the granularity of the iteration.
[0087] Step 3.5: Use the improved fitness function as the fitness function of the genetic algorithm to form a feature selection genetic algorithm.
[0088] Given the improved fitness function (8), the optimization problem of the genetic algorithm is to find the optimal chromosome, i.e., the optimal individual , which is the combination of the optimal features, called the optimal feature subset. This optimal feature subset satisfies:
[0089] (9);
[0090] where The function represents the variable value when the objective function reaches its minimum value.
[0091] Step 4: Use the feature selection genetic algorithm to screen multiple dimensional features of the preprocessed part size data to obtain the optimal feature subset.
[0092] For the part size data obtained from the industrial manufacturing field, the population initialization step in the genetic algorithm directly depends on this uncertain data. After the part size data is collected by sensors and preprocessed, it becomes the basic material for the genetic algorithm. The feature dimension corresponding to each part size data is like a gene locus on a chromosome. For example, different measurement dimensions of the part size data are encoded into binary values, and combined to form individuals representing feature selection schemes. Many such individuals constitute the initial population. This means that the uncertain data provides the original feature resources for the genetic algorithm, giving it an operable object and serving as the foundation for starting the subsequent algorithm process.
[0093] First, generate candidate feature subsets by initializing the population, and then use an improved fitness function to evaluate the contribution of each feature subset to the classification task. The improved fitness function not only considers the classification performance of the feature subset, but also introduces the Root Mean Square Error (RMSE) to measure the error of the model, and combines the elastic net for analysis. The elastic net can effectively reduce the weights of useless features by combining L1 and L2 regularization of the features, thereby reducing the error and improving the generalization ability of the model. On this basis, through genetic operations such as crossover and mutation, the feature subsets are iteratively optimized, and finally the most representative optimal feature subset is selected. The granularity benchmark mechanism quantifies the influence of each feature through the mathematical formula of the combined granularity, providing a quantitative evaluation criterion to guide the selection process in the genetic algorithm.
[0094] Step 4 specifically includes:
[0095] Step 4.1: Randomly generate the initial population of the feature selection genetic algorithm. The chromosome of each individual in the initial population represents a set of feature selection schemes. Each gene on the chromosome represents whether the corresponding feature is selected, and the gene encoding length of each individual is equal to the dimension of the features included in the preprocessed part size data;
[0096] Step 4.2: Start iterative calculation of the genetic algorithm from the initial population. For each individual in each generation of the population, calculate the fitness value using the improved fitness function (8) until the predetermined number of iterations is reached;
[0097] Step 4.3: After reaching the predetermined number of iterations, output the feature subset represented by the individual chromosome with the highest fitness value as the optimal feature subset.
[0098] This application adopts an improved binary genetic algorithm and a feature granulation strategy, and measures the candidate features through the granularity benchmark mechanism; the improved fitness function combines the classification accuracy and the feature granularity, greatly optimizing the feature selection process.
[0099] Step 5: Extract the feature data corresponding to the optimal feature subset from the preprocessed part size data to construct a feature dataset.
[0100] After screening out the optimal feature subset using the feature selection genetic algorithm, each feature has a corresponding index position in the optimal feature subset. In the preprocessed part size data, according to the corresponding index position of the feature, the data of the corresponding feature column can be accurately located and extracted as the corresponding feature data. All the feature data extracted for all the preprocessed part size data together constitute the feature dataset.
[0101] For example, all the part size data after preprocessing constitutes the original data set. If the optimal feature subset contains 5 features, corresponding to the features in the 3rd, 7th, 11th, 15th, and 19th columns of the original data set respectively, then from each row of data in the original data set, these 5 columns of data are extracted to form a new data set that only contains the feature data corresponding to the optimal feature subset, which is called the feature data set. This method is simple and direct, and can quickly complete data extraction according to established rules. It is especially suitable for structured data and can ensure the accuracy and efficiency of data extraction.
[0102] Subsequently, on the one hand, the feature data in the feature data set is divided into a training set and a test set according to a certain proportion. The feature subset represented by an individual is applied to the Evidential K-Nearest Neighbors (EKNN) model. After training with the training set, the classification accuracy of the model based on this feature subset is tested based on the test set, and is represented by This accuracy is a key component of the fitness function, measuring whether the feature subset can effectively handle the classification problem brought by data uncertainty. On the other hand, the granularity information refined from the uncertain data is incorporated into the fitness function. The granularity benchmark mechanism quantifies the granularity of the selected features in the chromosome through formula (3) according to data characteristics, such as the fluctuation granularity of part size data at different processing stages, and then quantifies the influence of each feature, providing an additional evaluation dimension for the fitness function to more accurately screen out the feature subset that can effectively handle uncertainty.
[0103] Step 6: Based on the feature data set, use the Evidential K-Nearest Neighbors model to classify the part quality level.
[0104] In this application, the K-Nearest Neighbors model based on evidence theory is used to handle uncertainty, abbreviated as the Evidential K-Nearest Neighbors (EKNN) model. This EKNN model is used to utilize the uncertain information of the data and perform pattern recognition and classification. If the class set of the uncertain data is composed of mutually exclusive elements, it is called the pattern recognition framework of the uncertain data, also known as the discernment framework. For the part size data of this application, its pattern recognition framework is .
[0105] Evidence theory provides the Dempster combination rule to orthogonally combine multiple pieces of information and realizes the fusion of multiple classification evidences. Among them, letters , etc. usually represent different focal elements, that is, subsets in the research problem domain. Taking the application of the EKNN model in the classification task as an example, assume that it is necessary to classify the part quality level in the industrial manufacturing field. Here, the universal set is the set of all possible part quality levels, such as {qualified, slightly unqualified, seriously unqualified}. Among them can represent the subset of "qualified", can represent the subset of "slightly unqualified", and so on.
[0106] and are respectively two basic probability assignment functions defined on the same identification framework . and represent the probabilities generated by different focal elements under the probability assignment function. The calculation formula of Dempster combination rule is as follows:
[0107] (10);
[0108] where is the new classification evidence obtained by fusing the evidence and ; represents the new focal element. It means that the probability assignment value for the empty set is 0. It can be set as:
[0109] (11);
[0110] is the conflict coefficient used to measure the conflict degree between evidences. The larger is, the greater the conflict between evidences. When = 1, the combination rule cannot be used. If is very large, it means there is a great conflict. The selection of the value will affect the classification accuracy. The optimal value can be determined by methods such as cross-validation.
[0111] In this application, the trust membership degree (abbreviation: belief degree) of EKNN is mapped using the Euclidean distance. In the algorithm for realizing belief degree assignment based on distance, calculate the distance between each scatter data and classification center points, and obtain distance values. Now assume a basic belief assignment function , and assume any two-dimensional scatter point is . Calculate the distance from this scatter point to cluster center points, and normalize all the distance values. After maps to obtain probability values, and each probability value is represented by , . needs to meet two conditions. One is that each Second, the sum of all allocated reliability results should be 1; second, the allocation function should be a monotonically decreasing function, that is, as the distance between the scatter point and the classification center point increases, the reliability allocation decreases.
[0112] Step 6 specifically includes the following steps 6.1 and 6.2.
[0113] Step 6.1: Divide the feature data set into a training set and a test set, use the training set to train the evidence K-nearest neighbor model, and use the test set to test the prediction error of the evidence K-nearest neighbor model based on the optimal feature subset ; after training, use the trained evidence K-nearest neighbor model as the part quality grade classification model.
[0114] Suppose a training set composed of feature data. Among them is a feature vector with a class label , is the dimensional space of the feature vector; . In this application, the obtained feature data is sorted out to conform to the input format of the EKNN model, which includes extracting and organizing each part size data sample according to the feature dimensions determined by the optimal feature subset to obtain feature data, and then forming it into a feature vector. For example, if the optimal feature subset contains three features, then each feature data is a feature vector composed of the values of these three features. The feature data is represented in the form of a concatenated vector, so it can also be called a feature vector. Let be the feature vector to be classified, coming from the test set, and here it is called the test vector. is in the set of nearest neighbors of . For each neighbor of , its label in constitutes a clear evidence set regarding
[0115] this evidence set can be calculated through the following distance mapping function:
[0116] (12);
[0117] (13);
[0118] where is a constant such that and is a positive parameter and is usually related to the class Usually fixed And for all , . Represents the probability assignment that a sample belongs to class under the premise of a known sample, Represents the probability assignment of all possible classes to which a sample belongs under the premise of a known sample.
[0119] If is "close" to , then it is believed that and belong to the same class; otherwise, there will be a lot of uncertainty about the class of . Regarding the class of , using Dempster's combination rule, all evidence items
[0120] can be combined as:
[0121] represents the fusion result for all evidence. The symbol represents taking the intersection of the evidence information involved.
[0122] Finally, a class decision is made through the fused evidence, and the class with the highest corresponding probability is the class to which the characteristic data of this sample belongs, that is, qualified, slightly unqualified, or severely unqualified.
[0123] Step 6.2: Input the feature vector of the part size data to be classified into the part quality grade classification model, and output the corresponding part quality grade; the part quality grade includes qualified, slightly unqualified, and severely unqualified.
[0124] In this application, feature selection is performed through a genetic algorithm combined with granularity and elastic net regularization. In multiple iterative processes, the genetic algorithm screens out the optimal feature subset according to the adaptive function improved based on granularity and elastic net. These optimal feature subsets can retain the key information of the data to the greatest extent and remove redundant features, thereby optimizing the subsequent classification task. Extract the feature data corresponding to the optimal feature subset from the part size data to be classified, and input it into the trained part quality grade classification model in the form of a feature vector for pattern recognition and classification, and the corresponding part quality grade can be output.
[0125] This application measures candidate features through a granularity benchmark mechanism, uses an improved fitness function that balances classification accuracy and granularity, and uses a genetic algorithm to optimize and select the features of the input data. By simulating the process of natural selection, it identifies the data features that have the most influence on pattern classification. The filtered feature data is input into the EKNN model, and the DS evidence theory is used to quantitatively process the conflicts between evidences. The innovative conflict quantification and allocation method of this application can effectively process and utilize uncertain information, has a wide range of applicability, can effectively process uncertain data from various sources, and improve the accuracy and efficiency of pattern recognition.
[0126] Therefore, the method of this application is not only applicable to the analysis and processing of part size data, but can also be extended to process uncertain data in other fields. For example, the temperature and humidity data in the production environment are affected by many factors in the workshop and are uncertain, and are collected by temperature and humidity sensors; in the field of transportation, the vehicle speed data is affected by road conditions, vehicle performance and external interference, etc., and uncertain data is obtained using in-vehicle sensors. The features corresponding to each physical quantity in the uncertain data (for example, the different dimensional features corresponding to temperature, wind speed, etc. in the meteorological field, such as the numerical range feature of temperature, the variation feature of wind speed at different times, etc.) will become the basic elements for constructing the candidate feature set. These original uncertain data features are combined and arranged to form many candidate feature sets in the initial population. That is to say, the various feature variables contained in the uncertain data become the specific objects available for selection and combination when initializing the population, so as to generate many different combinations of feature subsets as the starting point for the iterative optimization of the subsequent feature selection genetic algorithm.
[0127] Taking the industrial manufacturing field as an example, for part size data, a set of data containing multiple dimensional feature values measured for different parts and different processing stages of a part each time is an original data sample; for production environment temperature and humidity data, the combination of temperature and humidity values collected at a specific time period and a specific workshop location is also an original data sample. Similarly, in the field of transportation, the speed value of a vehicle at a certain moment and its associated road conditions, vehicle status and other feature values together constitute an original data sample. Applying the method of this application to the industrial manufacturing field can classify the part quality level as qualified, slightly unqualified, severely unqualified, etc. Applying it to the classification of the production environment status can determine whether the environment is suitable for production, whether the temperature and humidity are abnormal and need to be adjusted, etc. Applying it to the classification of vehicle safety conditions can evaluate whether the vehicle is currently safe or has potential safety hazards at different levels (mild, moderate, severe).
[0128] In an exemplary embodiment, this application also provides a part quality level classification device based on uncertain data pattern classification, including:
[0129] An uncertain data acquisition module for acquiring part dimension data with multiple dimensional features at different parts and different processing stages of a part monitored by sensors;
[0130] A data preprocessing module for preprocessing the part dimension data to obtain preprocessed part dimension data;
[0131] A fitness function improvement module for improving the fitness function of a genetic algorithm based on granularity and elastic net to form a feature selection genetic algorithm;
[0132] A feature screening module for screening multiple dimensional features of the preprocessed part dimension data using the feature selection genetic algorithm to obtain an optimal feature subset;
[0133] A feature data extraction module for extracting feature data corresponding to the optimal feature subset from the preprocessed part dimension data to construct a feature data set;
[0134] A part quality grade classification module for classifying the part quality grade using an evidence K-nearest neighbor model based on the feature data set.
[0135] This application uses a genetic algorithm to optimize the selection of features of input data. By simulating the process of natural selection, it identifies the data features that have the most influence on pattern classification. It measures candidate features through a granularity benchmark mechanism, defines a mathematical formula for granularity to quantify the granularity of the selected features in a chromosome, and optimally selects the data features that affect the classification result. By improving the fitness function that balances classification accuracy and granularity, it improves the accuracy and efficiency of pattern recognition. Then, the screened feature data is input into an EKNN model, and the belief function is used to quantify and process uncertainty to achieve more accurate pattern recognition and classification. The EKNN model uses the DS evidence theory to handle uncertainty and ambiguity, optimizes the application of the belief function in the classification process by quantifying and resolving conflicts between evidences. This method allows the model to make more effective decisions and classifications when facing incomplete or inconsistent information. Through an innovative conflict quantification and allocation method, the EKNN model can effectively process and utilize uncertain information without sacrificing classification accuracy, further improving the performance and reliability of the uncertain data pattern classification method based on the belief function theory. This innovation not only enhances the robustness of the model but also provides a new tool for pattern recognition in complex data environments.
[0136] In an exemplary embodiment, the present application further provides a computer device, which may be a server or a terminal. The computer device includes a processor, a memory, an input / output interface, and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. The computer program, when executed by the processor, implements the part quality grade classification method based on uncertain data pattern classification.
[0137] In an exemplary embodiment, the present application further provides a computer-readable storage medium, on which a computer program is stored, and the computer program, when executed by the processor, implements the part quality grade classification method based on uncertain data pattern classification.
[0138] In an exemplary embodiment, the present application further provides a computer program product, including a computer program, and the computer program, when executed by the processor, implements the part quality grade classification method based on uncertain data pattern classification.
[0139] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiment methods can be completed by hardware related to computer program instructions. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiment methods as described above. Among them, any reference to a memory or other medium provided in the embodiments of the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0140] It should be noted that the information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0141] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0142] In this article, specific examples are used to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for classifying the quality grades of parts based on the classification of uncertain data patterns, characterized in that, Including: Obtain the part size data including multiple dimensional features at different parts and different processing stages of the part monitored by the sensor; Perform data preprocessing on the part size data to obtain the preprocessed part size data; Improve the fitness function of the genetic algorithm based on granularity and elastic net to form a feature selection genetic algorithm; The fitness function of the genetic algorithm improved based on granularity and elastic net to form a feature selection genetic algorithm specifically includes: Use the chromosome of each individual in the genetic algorithm to represent a set of feature selection schemes, and each gene on the chromosome represents whether the corresponding feature is selected; assume that there are individuals in the population, and the gene encoding length of each individual is . Use to represent the chromosome of the -th individual, where represents the -th gene position; ; ; The entire population is represented as ; Hypothesis is the population of the current generation, which contains all the genes of the current generation, and represents all the selected genes in the population of the current generation. According to the formula calculate the granularity of the individuals in the current generation ; where is the chromosome of the current generation is the selected subset of features in , is the number of selected chromosome entries; is the number of selected features in the subset of features ; is the total number of genes in the current generation; The elastic net calculates the prediction error of the folded data through cross-validation ; where is the true label of the test sample ; is the predicted label of the test sample ; is the number of test samples; Based on the granularity of the current generation individuals and the prediction error Construct an improved fitness function based on granularity and elastic net ; where is the weight factor; is the fitness value calculated for the current generation; Use the improved fitness function as the fitness function of the genetic algorithm to form a feature selection genetic algorithm; Use the feature selection genetic algorithm to screen multiple dimensional features of the preprocessed part size data to obtain an optimal feature subset; Extract the feature data corresponding to the optimal feature subset from the preprocessed part size data to construct a feature dataset; Based on the feature dataset, use the evidence K-nearest neighbor model to classify the part quality grade.
2. The part quality grade classification method based on uncertain data pattern classification according to claim 1, characterized in that, The performing data preprocessing on the part size data to obtain the preprocessed part size data specifically includes: Apply denoising, missing value filling and normalization techniques to perform data preprocessing on the part size data to obtain the preprocessed part size data.
3. The method for classifying the part quality level based on the classification of uncertain data patterns according to claim 1, characterized in that, The using the feature selection genetic algorithm to screen multiple dimensional features of the preprocessed part size data to obtain an optimal feature subset specifically includes: Randomly generate the initial population of the feature selection genetic algorithm. The chromosome of each individual in the initial population represents a set of feature selection schemes. Each gene on the chromosome represents whether the corresponding feature is selected. The gene coding length of each individual is equal to the dimension of the features included in the preprocessed part size data; Start iterative calculation of the genetic algorithm from the initial population. For each individual in each generation of the population, calculate the fitness value using the improved fitness function until the predetermined number of iterations is reached; After reaching the predetermined number of iterations, output the feature subset represented by the individual chromosome with the highest fitness value as the optimal feature subset.
4. The method for classifying part quality grades based on uncertain data pattern classification according to claim 3, characterized in that The extracting the feature data corresponding to the optimal feature subset from the preprocessed part size data to construct a feature dataset specifically includes: Each feature has a corresponding index position in the optimal feature subset. In the preprocessed part size data, locate and extract the data of the corresponding feature column according to the index position corresponding to the feature as the corresponding feature data; The feature data extracted for all the preprocessed part size data together constitutes a feature dataset.
5. The method for classifying the part quality level based on the classification of uncertain data patterns according to claim 4, wherein The based on the feature dataset, using the evidence K-nearest neighbor model to classify the part quality grade specifically includes: Divide the feature dataset into a training set and a test set, use the training set to train the evidential K-nearest neighbor model, and use the test set to test the prediction error of the evidential K-nearest neighbor model based on the optimal feature subset. After training is completed, use the trained evidential K-nearest neighbor model as the part quality grade classification model. Input the feature vector of the part size data to be classified into the part quality grade classification model and output the corresponding part quality grade; The part quality grade includes qualified, slightly unqualified and seriously unqualified.
6. A part quality grade classification device based on uncertain data pattern classification, characterized in that, Including: An uncertain data acquisition module for obtaining the part size data including multiple dimensional features at different parts and different processing stages of the part monitored by the sensor; A data preprocessing module for performing data preprocessing on the part size data to obtain the preprocessed part size data; A fitness function improvement module for improving the fitness function of the genetic algorithm based on granularity and elastic net to form a feature selection genetic algorithm; The fitness function of the improved genetic algorithm based on granularity and elastic net forms a feature selection genetic algorithm, which specifically includes: Use the chromosome of each individual in the genetic algorithm to represent a set of feature selection schemes, and each gene on the chromosome represents whether the corresponding feature is selected; assume that there are individuals in the population, and the gene coding length of each individual is . Use to represent the chromosome of the -th individual, where represents the -th gene position; ; ; The entire population is represented as ; Hypothesis is the population of the current generation, which contains all genes of the current generation, and represents all selected genes in the population of the current generation. According to the formula calculate the granularity of individuals in the current generation ; where is the selected feature subset in the chromosome of the current generation, , is the number of selected chromosome entries; is the number of selected features in the feature subset ; is the total number of genes in the current generation; The elastic net calculates the prediction error of the folded data through cross-validation ; where is the true label of the test sample ; is the predicted label of the test sample ; is the number of test samples; Based on the granularity of the current generation individuals and the prediction error Construct an improved fitness function based on granularity and elastic net ; where is the weight factor; is the fitness value calculated for the current generation; Using the improved fitness function as the fitness function of the genetic algorithm to form a feature selection genetic algorithm; A feature screening module, which is used to use the feature selection genetic algorithm to screen multiple dimensional features of the preprocessed part size data to obtain an optimal feature subset; A feature data extraction module, which is used to extract feature data corresponding to the optimal feature subset from the preprocessed part size data to construct a feature dataset; A part quality grade classification module, which is used to classify the part quality grade based on the feature dataset using the evidence K-nearest neighbor model.
7. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the part quality grade classification method based on uncertain data pattern classification according to any one of claims 1-5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the part quality grade classification method based on uncertain data pattern classification according to any one of claims 1-5.
9. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the part quality grade classification method based on uncertain data pattern classification according to any one of claims 1-5.
Citation Information
Patent Citations
Internet financial fraud behavior detection method based on GA-SVM algorithm
CN112053223A
High-entropy alloy hardness prediction method based on machine learning and improved genetic algorithm feature screening
CN114464274A