Classification problem-oriented adaptive particle ball generation method
By selecting the sample with the smallest difference as the center in the pellet generation method and introducing the coverage threshold, the problems of unstable and complex calculation of pellet generation are solved, and more efficient and stable pellet division is achieved, and classification accuracy is improved.
Patent Information
- Application Number
- CN202510323926.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-20
AI Technical Summary
The existing pellet generation method adopts a random selection center strategy during the iteration process, resulting in high instability and high computational complexity of pellet division, affecting the accuracy and efficiency of classification results.
By calculating the median vector of each category sample set in the dataset, selecting the sample with the smallest difference as the center of the sphere division, and introducing the sphere coverage threshold and adaptive conditions to improve the sphere division process and avoid overlap detection and de-overlapping operations.
It improves the stability and efficiency of pellet generation, improves the accuracy and consistency of classification results, and reduces the sensitivity of the algorithm to outliers and outliers.
Smart Images

Figure CN120180255A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of multi-granularity cognitive computing and data mining, and particularly relates to an adaptive granule sphere generation method for classification problems. Background Art
[0002] With the advent of the big data era, efficient data analysis and processing have become the key tasks and challenges for mining data value and realizing knowledge discovery. Among numerous data processing paradigms, granular computing, as an effective and scalable method, simulates the multi-granularity cognitive mechanism of the human brain, effectively abstracts or generalizes relevant data from multiple levels and dimensions according to specific tasks or data characteristics, and then forms information granules with different sizes and hierarchical structures. Based on these information granules, granular computing can effectively conduct knowledge discovery and problem solving, significantly improving the efficiency of knowledge discovery, and becoming an effective data mining method for processing massive data.
[0003] Granule sphere computing, as an efficient learning method developed in the field of granular computing in recent years, has attracted extensive attention and research. Its core idea is to use granule spheres with different granularities to represent and cover the sample space, and use the granule spheres as information granules for subsequent learning. Usually, a single granule sphere contains multiple sample points, and the number of granule spheres is much lower than the number of sample points. Therefore, using coarse-grained granule spheres to replace fine-grained sample points as the model input can significantly reduce the scale of the input data, thereby improving the computing efficiency. Due to its high efficiency, robustness, and good interpretability, granule sphere computing has gradually been extended to various fields of artificial intelligence.
[0004] Granule sphere generation is the premise of granule sphere computing, and the quality of the generated granule spheres directly affects the effect of subsequent learning tasks. However, current granule sphere generation methods all adopt the strategy of randomly selecting the center during the iteration process. This randomness leads to the instability of granule sphere division, which in turn affects the accuracy and consistency of the classification results. In addition, the adaptive granule sphere generation method based on k-division significantly increases the computing complexity of the granule sphere iteration process due to the need to frequently perform granule sphere overlap detection and de-overlap operations, affecting the efficiency of granule sphere generation. Therefore, researching a more stable and efficient granule sphere generation method has important theoretical and practical significance. Summary of the Invention
[0005] To solve the above problems, the present invention provides an adaptive granule sphere generation method for classification problems. This method is based on the k-division adaptive granule sphere generation idea, effectively overcomes the instability and low efficiency problems of the existing methods by improving the granule sphere division center selection strategy and introducing a new adaptive granule sphere division condition.
[0006] The specific scheme includes the following steps:
[0007] S1. Calculate the median vector of each class sample set in the dataset, calculate the difference between each sample and the median vector corresponding to its class, and select k granule ball division centers according to the differences.
[0008] S2. Based on the k-division strategy, assign each sample to the nearest center and divide the dataset into k granule balls.
[0009] S3. For each granule ball, calculate the purity, coverage rate, and weighted purity sum of sub-granule balls.
[0010] S4. Calculate the granule ball coverage rate threshold and design an adaptive condition to achieve accelerated iterative division of granule balls.
[0011] S5. Repeat steps S1 - S3 until all granule balls meet the adaptive condition, and then execute step S6.
[0012] S6. Build a granule ball k-nearest neighbor classification model based on the generated granule balls and make a classification decision for the sample to be classified.
[0013] Furthermore, step S1 specifically includes:
[0014] S11. Obtain the labeled dataset D = {D1, D2, …, D k}, representing the q = 1, 2, …, k-th class sample set; the i-th sample in the class sample set D q where represents the sample the j = 1, 2, …, f-th eigenvalue of the sample, f represents the number of features, n represents the number of samples in the class sample set, represents the f-dimensional feature space;
[0015] S12. Calculate the median vector m q of the class sample set D q = (m1, m2,..., m f ), where m j represents the median of the j-th eigenvalue of all samples in the class sample set D q ;
[0016] S13. For each sample in the class sample set D q , calculate the difference between it and the median vector m q ;
[0017] S14. Select a sample corresponding to the minimum difference in each class sample set, and use the selected k samples as the granule ball division centers.
[0018] Further, the difference between each sample and the median vector corresponding to its category is expressed as
[0019]
[0020] where represents the set of category samples D q the sample in and the median vector m corresponding to its category q the difference between them.
[0021] Further, step S3 is for any one granule ball x i represents the i-th sample in the granule ball gb, and m represents the number of samples in the granule ball; where:
[0022] center
[0023] radius
[0024] purity
[0025] coverage rate
[0026] weighted sum of sub-granule ball purity
[0027] where, |·| represents the number of objects in set ·, d(x i , o) represents the Euclidean distance between the i-th sample x i in the granule ball gb and the center o, b j represents the j-th sub-granule ball of the granule ball gb, and k represents the number of categories of samples in the granule ball, gb * represents the set of majority-class samples in the granule ball gb; D represents the total set of all category samples, represents the number of majority-class samples in the j-th sub-granule ball of the granule ball gb, |gb * | represents the number of majority-class samples in the granule ball gb.
[0028] Further, the granule ball coverage rate threshold is expressed as
[0029]
[0030] where, CR T represents the granule ball coverage rate threshold, and D represents the data set.
[0031] Further, the adaptive conditions include:
[0032] The purity of the granule ball is equal to the weighted sum of the purity of its sub-granule balls;
[0033] The coverage rate of the granular balls is less than the granular ball coverage rate threshold.
[0034] Furthermore, a k-nearest neighbor classification model for granular balls is constructed based on the generated granular balls. The decision rule of the k-nearest neighbor classification model for granular balls is to predict the label of the sample as the label of the nearest granular ball, where the calculation formula for the distance between the sample and the granular ball is:
[0035]
[0036] where o and r respectively represent the center and radius of the granular ball, d(x, o) represents the Euclidean distance between the sample x and the ball center o, and dist(x, gb) represents the distance between the sample x and the granular ball gb.
[0037] Advantages of the present invention:
[0038] By selecting the sample with the smallest difference from the median vector in the dataset as the division center, the present invention effectively reduces the sensitivity of the algorithm to outliers and provides a more robust description of the central tendency, avoiding the inconsistency and instability caused by randomly selecting the center in the existing granular ball generation methods.
[0039] The present invention introduces the concept of granular ball coverage rate and proposes a new adaptive condition based on this, replacing the lower limit of the lowest quality and the overlap detection and de-overlap process in the traditional k-division adaptive granular ball generation method. This improvement significantly improves the efficiency of granular ball generation and further enhances the accuracy and efficiency of granular ball classification. Brief Description of the Drawings
[0040] Figure 1 is a flowchart of the present invention;
[0041] Figure 2 is a schematic diagram of the two-dimensional distribution of the synthetic dataset in a specific embodiment;
[0042] Figure 3 is a diagram of the division result of the synthetic dataset in the first iteration in a specific embodiment;
[0043] Figure 4 is a schematic diagram of the granular balls iteratively generated from the synthetic dataset in a specific embodiment. Detailed Embodiments
[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0045] The present invention provides an adaptive grain sphere generation method for classification problems, as Figure 1 shown, which includes the following steps:
[0046] S1. Calculate the median vector of each class sample set in the dataset, calculate the difference between each sample and the median vector corresponding to its class, and select k grain sphere division centers according to the difference.
[0047] Specifically, the sample data obtained by the present invention can be images, texts, etc., and subsequent classification processing is performed for image categories, text labels, etc.
[0048] Specifically, step S1 specifically includes:
[0049] S11. Obtain a dataset D = {D1, D2,..., D k} with labels, representing the q = 1, 2,..., k-th class sample set; the i-th sample in the class sample set D q where represents the j = 1, 2,..., f-th eigenvalue of the sample , f represents the number of features, n represents the number of samples in the class sample set, represents the f-dimensional feature space;
[0050] S12. Calculate the median vector m q = (m1, m2,..., m q ) of the class sample set D f , where m j represents the median of the j-th eigenvalue of all samples in the class sample set D q ;
[0051] S13. For each sample in the class sample set D q , calculate the difference between it and the median vector m q ;
[0052] Specifically, calculating the difference between each sample and the median vector corresponding to its class is expressed as
[0053]
[0054] where represents the difference between the sample q in the class sample set D and the median vector m q corresponding to its class.
[0055] S14. Select a sample corresponding to the minimum difference in each class sample set, and use the selected k samples as the grain sphere division centers.
[0056] Specifically, select the category sample set D q The granule division center of can be expressed as
[0057]
[0058] c = x index
[0059] where index represents the index of the sample corresponding to the minimum difference, x index represents the sample corresponding to the minimum difference, and c represents the granule division center.
[0060] S2. Based on the k-division strategy, assign each sample to the nearest center and divide the data set into k granules. S3. For each granule, calculate the purity, coverage rate, and weighted purity sum of sub-granules.
[0061] Specifically, step S3 is for any granule x i represents the i-th sample in the granule gb, and m represents the number of samples in the granule; where:[[]]
[0062] center
[0063] radius
[0064] purity
[0065] coverage rate
[0066] weighted purity sum of sub-granules
[0067] where |·| represents the number of objects in the set ·, d(x i , o) represents the Euclidean distance between the i-th sample in the granule gb and the center o, b j represents the j-th sub-granule of the granule gb, and k represents the number of categories of samples in the granule, gb * represents the set of majority-class samples in the granule gb; D represents the total set of all category samples, represents the number of majority-class samples in the j-th sub-granule of the granule gb, |gb * | represents the number of majority-class samples in the granule gb.
[0068] S4. Calculate the granule coverage rate threshold and design an adaptive condition to achieve accelerated iterative division of the granule.
[0069] Specifically, the granule coverage rate threshold is expressed as
[0070]
[0071] Among them, CR T represents the granule coverage threshold, and D represents the data set.
[0072] Specifically, based on the coverage rate of the granules, the weighted purity sum of the sub-granules, an adaptive condition for granule division is obtained:
[0073] (1) The purity of the granule is equal to the weighted purity sum of its sub-granules;
[0074] (2) The coverage rate of the granule is less than the coverage threshold.
[0075] Only when the granule satisfies both of the above two conditions, the division stops; otherwise, steps S1 to S3 need to be repeated to further divide the granule. When all granules satisfy the adaptive condition, the final granule set is obtained.
[0076] S5. Repeat steps S1 - S3 until all granules satisfy the adaptive condition, and then execute step S6.
[0077] Specifically, in the first iteration, after executing steps S1 - S3 for the collected data set, it is judged whether all granules satisfy the adaptive condition. If so, the iterative division stops; if not, for each granule that does not satisfy the adaptive condition, it is used as the data set to execute steps S1 - S3, and then it is judged again whether all the divided granules satisfy the adaptive condition; that is, in each iteration, only the granules that did not satisfy the adaptive condition in the previous iteration are processed, and the granules that satisfy the adaptive condition do not need to be split anymore until finally all granules satisfy the adaptive condition.
[0078] S6. Construct a granule k-nearest neighbor classification model based on the generated granules, and make a classification decision on the sample to be classified.
[0079] Specifically, construct a granule k-nearest neighbor classification model based on the generated granules. The decision rule of the model is to predict the label of the sample as the label of the nearest granule. The calculation formula for the distance between the sample and the granule is:
[0080]
[0081] where o and r respectively represent the center and radius of the granule, d(x, o) represents the Euclidean distance between the sample x and the center o of the sphere, and dist(x, gb) represents the distance between the sample x and the granule gb.
[0082] In one embodiment, obtain as Figure 2The synthetic binary classification dataset shown includes two classes, represented by "o" and "*" respectively. Among them, class "o" contains 50 samples, and class "*" contains 60 samples. Existing adaptive granule generation methods usually adopt the strategy of randomly selecting division centers, that is, randomly selecting a sample from class "x" and class "*" as the division center of the granule, thereby dividing the dataset into two granules, and then iteratively dividing the granules according to adaptive conditions. However, this method of randomly selecting division centers introduces a large degree of uncertainty, resulting in instability of the granule classification results. In addition, existing methods need to frequently perform granule overlap detection and de-overlap operations during the iterative process, thereby significantly increasing the computational complexity and reducing the efficiency of granule generation. To solve the above problems, the present invention proposes a method for determining the granule division center based on the median vector of the dataset, and introduces the concept of granule coverage rate, eliminating the process of granule overlap detection and de-overlap, thereby improving the stability and efficiency of granule generation.
[0083] For Figure 2 the dataset in, the present invention first calculates the median vector of each class, and obtains the median vectors of class "o" and class "*" as m1 = (0.399, 0.502) and m2 = (0.601, 0.551) respectively; then calculates the difference between each class sample and its median vector, and selects the sample with the smallest difference as the division center of this class, obtaining the division centers of class "o" and class "*" as o1 = (0.386, 0.502) and o2 = (0.605, 0.554) respectively; then assigns all samples in the initial dataset to the division center with the closest distance, obtaining two granules gb1 and gb2, as Figure 3 shown. For granules gb1 and gb2, their coverage rates are respectively And from the formula the granule coverage rate threshold CR T ≈0.085 is obtained, and neither gb1 nor gb2 simultaneously satisfies the following two adaptive conditions:
[0084] (1) The purity of the granule is equal to the weighted purity sum of its sub-granules;
[0085] (2) The coverage rate of the granule is less than the coverage rate threshold;
[0086] Therefore, gb1 and gb2 need to be continuously iteratively divided until the adaptive conditions are met. After continuous iteration, the initial dataset is divided into several granules of different sizes, as Figure 4 shown. For the sample to be classified x = (0.537, 0.574), the granule with the closest distance to it is gb * , and gb *Belonging to the category "*", the label of sample x is determined to be "*".
[0087] Since the samples of each category are fixed, the median vector of each category and the partitioning center selected based on this median vector are unique and determined, and the resulting grain balls are also determined. Therefore, the adaptive grain ball generation method proposed by the present invention has high stability. In addition, the present invention restricts the scale of grain balls by introducing the grain ball coverage rate condition, and abandons the grain ball overlap detection and de-overlap process in the traditional k-division based adaptive grain ball partitioning method, significantly improving the grain ball generation efficiency.
[0088] In the present invention, unless otherwise clearly defined and limited, terms such as "installation", "setting", "connection", "fixation", "rotation" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the internal communication of two components or the interaction relationship between two components. Unless otherwise clearly defined, for those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0089] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An adaptive sphere generation method for classification problems, characterized in that: The following steps are involved: S1. Calculate the median vector of each category of sample set in the data set, calculate the difference between each sample and the median vector corresponding to its category, and select k spheres to divide the centers according to the difference; S2. Based on the k-division strategy, each sample is assigned to the nearest center and the data set is divided into k spheres; S3. For each pellet, calculate the purity, coverage and weighted purity of the pellets; S4. Calculate the particle-sphere coverage threshold and design adaptive conditions to achieve accelerated iterative division of particles and spheres; S5. Repeat steps S1-S3 until all the balls meet the adaptive conditions, and then execute step S6; S6. Construct a sphere k-nearest neighbor classification model based on the generated spheres and make classification decisions for the samples to be classified.
2. The adaptive sphere generation method for classification problems according to claim 1, characterized in that: Step S1 specifically includes: S11. Get a labeled dataset D = {D1, D2, ..., D k }, represents the q=1, 2, ..., kth category sample set; category sample set D q The i-th sample in Representation sample The j-th eigenvalue of is 1, 2, ..., f, where f represents the number of features and n represents the number of samples in the class sample set. represents the f-dimensional feature space; S12. Calculate the category sample set D q The median vector m q =(m1,m2,...,m f ), where m j Denotes the category sample set D q The median of the jth eigenvalue of all samples in ; S13. For category sample set D q For each sample in, calculate its difference with the median vector m q The difference between S14. Select a sample corresponding to the minimum difference in each category of sample sets, and use the selected k samples as the particle sphere division centers.
3. The adaptive sphere generation method for classification problems according to claim 1, characterized in that: Calculate the difference between each sample and the median vector corresponding to its category as in, Denotes the category sample set D q Samples in The median vector m corresponding to the category to which it belongs q The difference between.
4. The adaptive sphere generation method for classification problems according to claim 1, characterized in that: Step S3: for any ball x i represents the i-th sample in the sphere gb, and m represents the number of samples in the sphere; where: center radius purity Coverage Grain weighted purity and Among them, |·| represents the number of objects in the set ·, d(x i ,o) represents the i-th sample x in the sphere gb i The Euclidean distance from the center o, b j represents the jth sub-ball of the ball gb, and k represents the number of categories of samples in the sphere, gb * represents the majority class sample set in the sphere gb; D represents the total set of samples of all categories, represents the number of majority class samples in the jth sub-sphere of sphere gb, |gb * | represents the number of majority class samples in the sphere gb.
5. The adaptive sphere generation method for classification problems according to claim 1, characterized in that: The particle coverage threshold is expressed as Among them, CR T represents the sphere coverage threshold, and D represents the total set of samples of all categories.
6. The adaptive sphere generation method for classification problems according to claim 1, characterized in that: Adaptive conditions include: The purity of a pellet is equal to the weighted purity of its sub-pellets; The coverage of the pellets is less than the pellet coverage threshold.
7. The adaptive sphere generation method for classification problems according to claim 1, characterized in that: A particle sphere k-nearest neighbor classification model is constructed based on the generated particles. The decision rule of the particle sphere k-nearest neighbor classification model is to predict the label of the sample as the label of the particle sphere closest to it, where the calculation formula of the distance between the sample and the particle sphere is: Where o and r represent the center and radius of the sphere respectively, d(x,o) represents the Euclidean distance between the sample x and the center o, and dist(x,gb) represents the distance between the sample x and the sphere gb.
Citation Information
Cited By
Centralized management method for data center cluster
CN120994145A