Semi-supervised clustering method based on gravitation equation

By employing a semi-supervised clustering method based on the gravitational equation, data weights and quality are calculated using the gravitational constant and a function to determine sub-centroids. This approach solves the problem of inconsistent classification across multiple domains and achieves high stability and effectiveness as data grows.

CN120849991AInactive Publication Date: 2025-10-28CHENGDU XIEEN SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511349473.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2025-10-28
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing semi-supervised clustering methods suffer from inconsistent classification when applied to multi-domain data, and the centroids need to be redefined as the data grows, which can render the model ineffective.

Method used

A semi-supervised clustering method based on the gravitational equation is adopted. The weights and masses of the data are calculated by using the gravitational constant, the spacing function, the weight function, and the mass function. The sub-centroids are determined by the gravitational function to achieve data clustering.

Benefits of technology

It maintains high stability and effectiveness when dealing with large-scale and diverse data, and can maintain the reliability of clustering tasks as data grows, making it particularly suitable for complex data environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849991A_ABST
    Figure CN120849991A_ABST
Patent Text Reader

Abstract

The invention discloses a semi-supervised clustering method based on a gravitational equation, and belongs to the technical field of machine learning, the method defines a gravitational constant and a spacing function, updates data in a training set and a test set through the spacing function, calculates the quality of each piece of data through the weight of the data, regards each piece of data in the training set as a sub-centroid, and obtains a semi-supervised clustering result. Putting the data of the training sets of the same category into one subset, marking category labels, updating gravitational constants, setting a gravitational function, calculating gravitational vectors of the data of each test set and each sub-centroid through the gravitational function and data quality, screening out the maximum gravitational vector of the data of each test set, and calculating the maximum gravitational vector of the data of each test set; according to the method and the device, a plurality of sub-centroids can be configured for the input data through the gravitation function, and the sub-centroids serve as key nodes and can effectively support quantization operation of mutual gravitation among the data; and a powerful technical support is provided for clustering analysis in a complex data environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of machine learning technology, specifically relating to a semi-supervised clustering method based on the gravitational equation. Background Technology

[0002] Clustering models aim to divide datasets into different categories, reflecting the affiliation of samples. Unsupervised or supervised clustering validity index models can be used to identify the category. This simple and effective process makes clustering models applicable to various fields, such as pattern recognition, applied statistics, machine learning, and information theory. However, traditional clustering models have various drawbacks, so recent research has combined clustering models with various semi-supervised learning (SSL) methods. SSL-based clustering improves the model's prediction process by clustering given data based on provided auxiliary information (constraints).

[0003] Recent studies have demonstrated that semi-supervised clustering (SSC) may be highly inconsistent with the true partitioning, and that SSC models may produce incorrect predictions when run multiple times on the same dataset.

[0004] When the given constraints are weak, the generated SSL constraints may have issues because they can affect the correct centroid determination by the SSC. Therefore, regardless of whether the centroid determination function depends on the generated constraints, the determined centroids influence the effectiveness of the generated constraints. However, data growth affects both constraints and centroids, potentially requiring retraining; correspondingly, specific data growth scenarios can render the model ineffective. The complex relationship between constraints, centroids, and data growth is a cause of classification inconsistencies and also affects the model's ability to perform well in multi-domain data environments (a single model may perform well in some domains but fail in others).

[0005] Therefore, while many existing SSC methods can improve prediction accuracy to some extent, they are generally not applicable to multi-domain data; that is, they work in some domains but fail to predict in others. Furthermore, when new data is added, it is necessary to redetermine the centroids and re-cluster all the data. Summary of the Invention

[0006] To address the problems mentioned in the background, this application provides a semi-supervised clustering method based on the gravitational equation, which solves the problem that existing clustering methods cannot be applied to data in multiple fields, and that when new data is added, it is necessary to redetermine the centroids and re-cluster all the data.

[0007] To achieve the above objectives, the present invention provides the following technical solution:

[0008] A semi-supervised clustering method based on the gravitational equation includes the following steps:

[0009] S1: Divide the given dataset into a training set and a test set according to a preset ratio;

[0010] S2: Define the gravitational constant and define the spacing function based on the gravitational constant, and update the data in the training set and the test set through the spacing function;

[0011] S3: Based on the updated data, define a weight function to calculate the weight of each data point in the training and test sets. Then, define a quality function to calculate the quality of each data point in the training and test sets.

[0012] S4: Treat each data point in the training set as a subcentroid, and group the data points of the same category in the training set into a subset and label them with the category.

[0013] S5: Update the gravitational constant and set the gravitational function. Calculate the gravitational vector between the data of each test set and each sub-mass using the gravitational function and the data quality calculated in S3.

[0014] S6: Filter out the maximum gravitational vector for each test set data. The category label of the sub-mass corresponding to the maximum gravitational vector is the category to which the test set data belongs, thus completing the classification.

[0015] Preferably, S1 specifically involves: randomly selecting a predetermined proportion of data from each category of the given dataset and combining them together as the training set data. The remaining data will be used as the test set. The number of features for all data is M.

[0016] Preferably, in S2, the center position of all training set data is found by averaging. Through the central position Define a gravitational constant G, set a spacing function using the gravitational constant G, and update the training and test set data using the spacing function.

[0017] Preferably, S2 is as follows:

[0018] S2.1: Find the center position of all training set data by averaging. ;

[0019] S2.2: Define the gravitational constant G, which is expressed as: ;

[0020] Where N is the number of data points in the training set. This indicates the position of the i-th data point in the training set. G represents the distance from each data point to the center, and G represents the average distance from the center of the entire training set. The larger G is, the more dispersed the data is; the smaller G is, the more concentrated the data is.

[0021] S2.3: Set Spacing Function Spacing function Expressed as: ;

[0022] in, Represent any data point in the training and test sets, and input each data point from the training and test sets into the distance function. Update.

[0023] Preferably, in S3, the weighting function Defined as:

[0024] ;

[0025] Using weight function Calculate the weight of each data point; the weights are an M-dimensional vector.

[0026] mass function Defined as:

[0027] ;

[0028] in, Representing data The j-th element, Representing data Corresponding weight vector The j-th element.

[0029] Preferably, S4 specifically involves: treating each data point in the training set as a subcentroid, grouping training data of the same category together to form a subset, and placing subsets of different categories into set C. This is specifically expressed as:

[0030] ;

[0031] in, This represents the training set data belonging to category label k. This represents a subset of all training data with category label k.

[0032] Preferably, S5 is as follows:

[0033] Update the gravitational constant G according to the formula in S2.2, set the gravitational function, and calculate each data point in the test set. With collection The gravitational force of each sub-center of mass in the matrix is ​​defined by the gravitational function shown below:

[0034] ;

[0035] in, This represents the test set data to be classified. Gravity vectors of all training set data, and These represent the quality and set of the test set data to be classified, respectively. The set of masses of each sub-center of mass. This indicates that each data point in the test set is computed. With collection Euclidean distance between them The vector elements representing the gravitational constant G are used to obtain the test set data through the gravitational function. With collection The gravitational force between each sub-center of mass, i.e. , where k is the number of training data.

[0036] Preferably, S6 specifically involves: for each test set data Searching for the gravitational vector Find the category label corresponding to the maximum gravitational force in the set. The subset to which the test set belongs, and the category represented by that subset, constitutes the test set data. Category: Complete the classification.

[0037] Compared with the prior art, the beneficial effects of the present invention are:

[0038] This invention innovatively introduces the core mechanism of the gravity function. Through the application of this mechanism, multiple sub-centroids can be configured for each input data. These sub-centroids serve as key nodes and can effectively support the quantitative calculation of the mutual attraction between data. This application demonstrates reliable effectiveness in clustering tasks when facing the increasing volume of data in specific scenarios. Of particular note is that this method can maintain extremely high operational stability when dealing with large-scale data and diverse data types, providing strong technical support for clustering analysis in complex data environments. Attached Figure Description

[0039] Figure 1 This is a schematic diagram illustrating the process of determining the sub-centroid in this application. Detailed Implementation

[0040] To facilitate understanding of the technical content of this invention by those skilled in the art, the invention will be further described in detail below with reference to the accompanying drawings and specific examples. It should be understood that the specific examples described herein are merely illustrative and not intended to limit the scope of the invention.

[0041] Example 1:

[0042] like Figure 1 As shown, a semi-supervised clustering method based on the gravitational equation includes the following steps:

[0043] S1: Randomly select 30% of the data from each category of the given dataset and combine them together as the training set data. The remaining 70% of the data will be used as the test set. The number of features for all data is M.

[0044] S2: Calculate the gravitational constant G and update the training and test set data using the spacing function, including the following steps:

[0045] S2.1: Find the center position of all training set data by averaging. ;

[0046] S2.2: Define the gravitational constant G:

[0047] ;

[0048] Where N is the number of data points in the training set. This indicates the position of the i-th data point in the training set. G represents the distance from each data point to the center, and G represents the average distance from the center of the entire training set. The larger G is, the more dispersed the data is; the smaller G is, the more concentrated the data is.

[0049] S2.3: Set the spacing function, define the spacing function as shown below:

[0050] ;

[0051] in This represents any training and test set data. Each data point in the training and test sets is input into this distance function for updating.

[0052] S3: Calculate the weight of each data point using a weighting function, and then calculate the quality of each data point using a quality function, including the following steps:

[0053] S3.1: Set the weight function, define the weight function as shown below:

[0054] ;

[0055] Calculate each data point using a weighting function The weights are M-dimensional vectors;

[0056] S3.2: Based on the weight obtained for each data point, calculate the mass of the object represented by that data point using the mass function, which is defined as follows:

[0057] ;

[0058] in, Representing data The j-th element, Representing data Corresponding weight vector The j-th element.

[0059] S4: Extract training data of the same category and place them in a subset. Arrange all extracted training data subsets of all categories into a set C. Calculate the gravitational force between each data point in the test set and set C using the gravity formula, obtaining a gravity vector. Determine the category of the test set data by finding the index corresponding to the maximum value in the vector, thus completing the data clustering. This includes the following steps:

[0060] S4.1: Treat each data point in the training set as a subcentroid, group training data of the same class together to form a subset, and place the subsets formed by different classes into set C as shown below:

[0061] ;

[0062] in, This represents the training set data belonging to category label k. This represents a subset of all training data with category label k.

[0063] S4.2: Update the gravitational constant G according to the formula in S2.2, set the gravitational function, and calculate each data point in the test set. With collection The gravitational force of each sub-center of mass in the matrix is ​​defined by the gravitational function shown below:

[0064] ;

[0065] in, This represents the test set data to be classified. Gravity vectors of all training set data, and These represent the quality and set of the test set data to be classified, respectively. The set of masses of each sub-center of mass. This indicates that each data point in the test set is computed. With collection Euclidean distance between them Let G be the vector element sum of the gravitational constant G. The test set data is obtained using the above gravitational function. With collection The gravitational force between each sub-center of mass, i.e. , where k is the number of training data.

[0066] S4.3: For each test set data Find the gravitational vector calculated by S4.2. Find the index corresponding to the maximum gravitational force in the set. The subset to which the test set belongs, and the category represented by that subset, constitutes the test set data. The data is classified by its category.

[0067] In this embodiment, as Figure 1 As shown, taking classes 1 and 2 as examples, both classes 1 and 2 have multiple sub-centroids. The green area on the left represents the data for class 1, where green circles represent sub-centroids and green stars represent data points. The red area on the right represents the data for class 2, where red circles represent sub-centroids and red stars represent data points. Compared to traditional clustering methods that only have one cluster center, this diagram illustrates that our designed clustering process involves a series of sub-centroids. The relationship between data points and sub-centroids is established using a gravity function, and the gravity vector is calculated to determine the category of each data point. In this method, the training data for each category is clustered using sub-centroids, while each test data point is determined to belong to a category based on the calculation of the maximum gravity vector. As shown in Table 1:

[0068] Table 1

[0069]

[0070] Table 1 shows the overall success rate of the semi-supervised gravitational clustering method of the present invention and existing methods under different growth percentages (data size gradually increases from 50% to 100%). The method proposed in this invention consistently maintains a high level of 93%-94%, while the density peak semi-supervised clustering method, which performs second best, only achieves 39%-43%, far lower than the method of the present invention. The method outperforms all the comparison methods throughout the entire data growth process, indicating that it has extremely high stability when dealing with large-scale and diverse data.

[0071] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A semi-supervised clustering method based on the gravitational equation, characterized in that, Includes the following steps: S1: Divide the given dataset into a training set and a test set according to a preset ratio; S2: Define the gravitational constant and define the spacing function based on the gravitational constant, and update the data in the training set and the test set through the spacing function; S3: Based on the updated data, define a weight function to calculate the weight of each data point in the training and test sets. Then, define a quality function to calculate the quality of each data point in the training and test sets. S4: Treat each data point in the training set as a subcentroid, and group the data points of the same category in the training set into a subset and label them with the category. S5: Update the gravitational constant and set the gravitational function. Calculate the gravitational vector between the data of each test set and each sub-mass using the gravitational function and the data quality calculated in S3. S6: Filter out the maximum gravitational vector for each test set data. The category label of the sub-mass corresponding to the maximum gravitational vector is the category to which the test set data belongs, thus completing the classification.

2. The semi-supervised clustering method based on the gravitational equation according to claim 1, characterized in that, S1 specifically involves randomly selecting a predetermined proportion of data from each category of the given dataset and combining them together as the training set data. The remaining data will be used as the test set. The number of features for all data is M.

3. The semi-supervised clustering method based on the gravitational equation according to claim 2, characterized in that, In S2, the center position of all training set data is found by averaging. Through the central position Define a gravitational constant G, set a spacing function using the gravitational constant G, and update the training and test set data using the spacing function.

4. The semi-supervised clustering method based on the gravitational equation according to claim 3, characterized in that, S2 specifically refers to: S2.1: Find the center position of all training set data by averaging. ; S2.2: Define the gravitational constant G, which is expressed as: ; Where N is the number of data points in the training set. This indicates the position of the i-th data point in the training set. G represents the distance from each data point to the center, and G represents the average distance from the center of the entire training set. The larger G is, the more dispersed the data is; the smaller G is, the more concentrated the data is. S2.3: Set Spacing Function Spacing function Expressed as: ; in, Represent any data point in the training and test sets, and input each data point from the training and test sets into the distance function. Update.

5. The semi-supervised clustering method based on the gravitational equation according to claim 4, characterized in that, In S3, the weight function Defined as: ; Using weight function Calculate the weight of each data point; the weights are an M-dimensional vector. mass function Defined as: ; in, Representing data The j-th element, Representing data Corresponding weight vector The j-th element.

6. The semi-supervised clustering method based on the gravitational equation according to claim 5, characterized in that, S4 specifically involves treating each data point in the training set as a subcentroid, grouping training data of the same class together to form a subset, and placing subsets of different classes into set C. This can be expressed as follows: ; in, This represents the training set data belonging to category label k. This represents a subset of all training data with category label k.

7. The semi-supervised clustering method based on the gravitational equation according to claim 6, characterized in that, S5 specifically refers to: Update the gravitational constant G according to the formula in S2.2, set the gravitational function, and calculate each data point in the test set. With collection The gravitational force of each sub-center of mass in the matrix is ​​defined by the gravitational function shown below: ; in, This represents the test set data to be classified. Gravity vectors of all training set data, and These represent the quality and set of the test set data to be classified, respectively. The set of masses of each sub-center of mass. This indicates that each data point in the test set is computed. With collection Euclidean distance between them The vector elements representing the gravitational constant G are used to obtain the test set data through the gravitational function. With collection The gravitational force between each sub-center of mass, i.e. , where k is the number of training data.

8. The semi-supervised clustering method based on the gravitational equation according to claim 7, characterized in that, S6 specifically refers to: for each test set data Searching for the gravitational vector Find the category label corresponding to the maximum gravitational force in the set. The subset to which the test set belongs, and the category represented by that subset, constitutes the test set data. Category: Complete the classification.