Method for processing multi-source data based on K-Means aggregation algorithm of large model
Through the K-Means aggregation algorithm based on the big model, the problems of initial clustering center randomness and outlier sensitivity are solved, and efficient and accurate clustering of electricity consumption data is achieved, which is suitable for multi-source data processing.
Patent Information
- Application Number
- CN202510601564.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-29
AI Technical Summary
The existing K-Means algorithm has high randomness in the initial cluster center selection, is sensitive to outliers, and has high computational complexity during large-scale high-dimensional data processing, resulting in unstable clustering results, low accuracy and slow convergence speed.
The K-Means aggregation algorithm based on the big model is used to determine the initial clustering center through the big model feature pre-extraction and the improved K-Means++ algorithm, and an outlier detection model is constructed, the cluster center update is optimized, and the double iteration termination conditions are set to improve cluster stability and accuracy.
It improves the stability and accuracy of power consumption data clustering, reduces the computational complexity, improves the efficiency of large-scale high-dimensional data processing, and better meets the requirements of real-time.
Smart Images

Figure CN120561797A_ABST
Abstract
Description
Technical Field
[0001] The invention discloses a method for processing multi-source data based on a large-model K-Means aggregation algorithm, and relates to the field of data processing. Background Art
[0002] The K-Means algorithm is a classic algorithm for processing data clustering. Its principle is to randomly select K initial cluster centers, then assign each sample in the data set to the cluster with the nearest cluster center. The cluster center is then updated based on the mean of the samples in the cluster, and the algorithm is iterated continuously until the cluster center no longer changes or the preset number of iterations is reached. However, the existing K-Means algorithm has some defects when processing data. For example, the random selection of initial cluster centers leads to unstable clustering results, and different initial selections may produce significantly different clustering effects. On the other hand, it is sensitive to outliers, and a small number of outliers may seriously affect the calculation of cluster centers, thereby reducing the accuracy of data clustering. Especially when faced with large-scale, high-dimensional data sets from different data sources, the existing K-Means algorithm has high computational complexity and slow convergence speed, which makes it difficult to meet the needs of fast and accurate data clustering in practical applications. Summary of the Invention
[0003] In response to the problems of the prior art, the present invention provides a method for processing multi-source data based on the K-Means aggregation algorithm of a large model, so as to solve the problems of the existing K-Means algorithm in that the initial clustering center selection is highly random, it is sensitive to outliers, and it has high computational complexity and slow convergence speed when processing large-scale high-dimensional data, thereby improving the stability, accuracy and efficiency of data clustering.
[0004] The specific scheme proposed by the present invention is:
[0005] The present invention provides a method for processing multi-source data using a K-Means aggregation algorithm based on a large model, comprising: step 1: performing feature pre-extraction on electricity consumption data based on the large model: performing feature extraction on an input data set using the large model, and converting raw data in the data set into a representative and discriminative feature vector, wherein the data set of the electricity consumption data includes a text data set, an image data set, a structured data set, and a geographic coordinate data set;
[0006] Step 2: Determine the initial electricity consumption cluster center: For each extracted feature vector, use the K-Means++ algorithm to determine the initial cluster center: randomly select a feature vector as the first initial cluster center, calculate the shortest distance from each feature vector to the selected initial cluster center, and repeat the selection process until K initial cluster centers are selected according to the principle that the greater the distance, the higher the probability of being selected as the next initial cluster center.
[0007] Step 3: Detect and process outliers: During the clustering process, the large model is used to build an outlier detection model. Each feature vector in different data sets is scored for abnormality. Samples with scores exceeding the set threshold are identified as outliers and temporarily removed from the data set.
[0008] Step 4: Update the cluster center of the optimized data set: Use the prediction results of the large model to adjust the cluster center update, where the large model is used to predict the distribution trend of samples in each cluster. When the new cluster center is calculated, the prediction information of the large model is combined to fine-tune the position of the cluster center, and the iteration is stopped according to the iteration termination condition to obtain the final cluster center, completing the electricity consumption clustering processing of the multi-source data set.
[0009] Step 5: Adjust the time-of-use electricity price and conduct grid dispatch based on the electricity consumption cluster center.
[0010] Furthermore, in step 1 of the method for processing multi-source data using a large-model K-Means aggregation algorithm, a language model based on a Transformer architecture, an image model based on a convolutional neural network, a model based on a multi-layer perceptron architecture, and a model based on an autoencoder architecture are selected for feature pre-extraction.
[0011] Furthermore, in step 1 of the method for processing multi-source data using a large model-based K-Means aggregation algorithm, feature extraction of the input data set using the large model includes:
[0012] When the dataset is a text data dataset, the text is preprocessed first, including word segmentation and stop word removal. Then the preprocessed text data is input into the corresponding large model to obtain the vector representation of the text.
[0013] When the data set is an image data set, the image data is input into the corresponding large model to extract the feature vector of the image, and the image data is preprocessed to unify the resolution before input.
[0014] When the data set is a structured data set, the structured data is organized into a suitable format and then input into the adapted large model, which outputs a feature vector that condenses the key information.
[0015] When the dataset is a geographic coordinate dataset, the longitude and latitude data are standardized and then input into the corresponding large model to output a low-dimensional feature vector that retains the key information of the spatial position relationship.
[0016] Furthermore, in step 4 of the method for processing multi-source data using a large-model K-Means aggregation algorithm, an iteration termination condition is set: a double iteration termination condition is set, wherein when the change in the cluster center in two consecutive iterations is less than a preset threshold, the iteration is stopped; a maximum number of iterations is set; when any one of the conditions is met, the iteration is terminated and the final clustering result is output.
[0017] The present invention also provides a device for processing multi-source data based on the K-Means aggregation algorithm of a large model, comprising a feature extraction module, an initial clustering module, a detection module, an update module and a scheduling module.
[0018] The feature extraction module performs feature pre-extraction on electricity consumption data based on a large model. The module uses the large model to extract features from the input dataset, converting the raw data in the dataset into representative and discriminative feature vectors. The electricity consumption data datasets include text data datasets, image data datasets, structured data datasets, and geographic coordinate datasets.
[0019] The initial clustering module determines the initial electricity consumption cluster center: the K-Means++ algorithm is used to determine the initial cluster center for each extracted feature vector: a feature vector is randomly selected as the first initial cluster center, and the shortest distance from each feature vector to the selected initial cluster center is calculated. The longer the distance, the higher the probability of being selected as the next initial cluster center. The selection process is repeated until K initial cluster centers are selected.
[0020] The detection module detects and processes outliers: During the clustering process, the large model is used to build an outlier detection model, which performs anomaly scores on each feature vector in different data sets. Samples with scores exceeding the set threshold are identified as outliers and temporarily removed from the data set.
[0021] The update module updates and optimizes the cluster center of the data set: the cluster center is updated and adjusted using the prediction results of the large model. The large model is used to predict the distribution trend of samples in each cluster. When the new cluster center is calculated, the position of the cluster center is fine-tuned in combination with the prediction information of the large model. The iteration is stopped according to the termination condition to obtain the final cluster center, completing the electricity consumption clustering processing of the multi-source data set.
[0022] The dispatching module adjusts the time-of-use electricity price and performs grid dispatching based on the electricity consumption cluster center.
[0023] Furthermore, the feature extraction module of the device for processing multi-source data based on the K-Means aggregation algorithm of a large model selects a language model based on a Transformer architecture, an image model based on a convolutional neural network, a model based on a multi-layer perceptron architecture, and a model based on an autoencoder architecture for feature pre-extraction.
[0024] Furthermore, the feature extraction module of the apparatus for processing multi-source data using the K-Means aggregation algorithm based on a large model extracts features from the input data set using the large model, including:
[0025] When the dataset is a text data dataset, the text is preprocessed first, including word segmentation and stop word removal. Then the preprocessed text data is input into the corresponding large model to obtain the vector representation of the text.
[0026] When the data set is an image data set, the image data is input into the corresponding large model to extract the feature vector of the image, and the image data is preprocessed to unify the resolution before input.
[0027] When the data set is a structured data set, the structured data is organized into a suitable format and then input into the adapted large model, which outputs a feature vector that condenses the key information.
[0028] When the dataset is a geographic coordinate dataset, the longitude and latitude data are standardized and then input into the corresponding large model to output a low-dimensional feature vector that retains the key information of the spatial position relationship.
[0029] Furthermore, the update module of the device for processing multi-source data using the K-Means aggregation algorithm based on a large model sets an iteration termination condition: a double iteration termination condition is set, wherein when the change in the cluster center in two consecutive iterations is less than a preset threshold, the iteration is stopped; a maximum number of iterations is set; when any one of the conditions is met, the iteration is terminated and the final clustering result is output.
[0030] The benefits of the present invention are:
[0031] Improve the stability of electricity consumption data clustering: By determining the initial cluster center based on large model feature pre-extraction and the improved K-Means++ algorithm, the selection of the initial cluster center is made more scientific and representative, reducing the differences in clustering results caused by random selection of the initial center and improving the stability of electricity consumption data clustering.
[0032] Enhanced clustering accuracy: The outlier detection model constructed by the large model effectively identifies and processes outliers, avoiding the interference of outliers on the cluster center. At the same time, the prediction results of the large model are used to optimize the update of the cluster center, making the cluster center more consistent with the actual distribution of the data, significantly improving the accuracy of electricity consumption data clustering.
[0033] Improving the computational efficiency of electricity consumption data: Pre-extraction of features from large models reduces data dimensionality, reducing the computational effort of the subsequent K-Means algorithm. When processing large-scale, high-dimensional data, the optimized algorithm reduces computational complexity and accelerates convergence, improving overall algorithm execution efficiency and better meeting application scenarios with high real-time requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a schematic flow chart of the method of the present invention. DETAILED DESCRIPTION
[0035] The present invention will be further described below with reference to the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it. However, the embodiments are not intended to limit the present invention.
[0036] Example 1
[0037] The present invention provides a method for processing multi-source data based on a K-Means aggregation algorithm of a large model, comprising: step 1: performing feature pre-extraction on electricity consumption data based on a large model: utilizing the large model to perform feature extraction on an input data set, and converting the original data in the data set into a representative and discriminative feature vector, wherein the data set of the electricity consumption data includes a text data set, an image data set, a structured data set, and a geographic coordinate data set.
[0038] In step 1, a language model based on the Transformer architecture, an image model based on a convolutional neural network, a model based on a multi-layer perceptron architecture, and a model based on an autoencoder architecture can be selected for feature pre-extraction.
[0039] Extract features from input datasets using large models, including:
[0040] When the dataset is a text data dataset, the text is preprocessed first, including word segmentation and stop word removal. Then the preprocessed text data is input into the corresponding large model to obtain the vector representation of the text.
[0041] When the data set is an image data set, the image data is input into the corresponding large model to extract the feature vector of the image, and the image data is preprocessed to unify the resolution before input.
[0042] When the data set is a structured data set, the structured data is organized into a suitable format and then input into the adapted large model, which outputs a feature vector that condenses the key information.
[0043] When the dataset is a geographic coordinate dataset, the latitude and longitude data is standardized and then fed into the corresponding large model, which outputs a low-dimensional feature vector that retains key information about spatial location relationships. For example, a dataset containing 10,000 coordinates is collected and filtered to exclude erroneous and non-compliant coordinates.
[0044] Large-Model Feature Pre-Extraction: A large model based on an autoencoder architecture is used to pre-extract features from geographic point data. By learning from a large amount of geographic point data, the autoencoder model can compress high-dimensional raw longitude and latitude data into low-dimensional, more representative feature vectors. The standardized geographic point data is input into the autoencoder, which outputs 32-dimensional feature vectors. These feature vectors not only preserve key information, such as the spatial relationship between the points, but also significantly reduce the data dimensionality, reducing the complexity of subsequent clustering calculations.
[0045] Step 2: Determine the initial electricity consumption cluster center: For each extracted feature vector, use the K-Means++ algorithm to determine the initial cluster center: randomly select a feature vector as the first initial cluster center, calculate the shortest distance from each feature vector to the selected initial cluster center, and repeat the selection process until K initial cluster centers are selected according to the principle that the greater the distance, the higher the probability of being selected as the next initial cluster center. For example, in the coordinate point vector space after the features are extracted by the autoencoder, use the improved K-Means++ algorithm to determine K = 15 initial cluster centers.
[0046] Step 3: Detect and process outliers: During the clustering process, the large model is used to build an outlier detection model. Each feature vector in different data sets is scored for abnormality. Samples with scores exceeding the set threshold are identified as outliers and temporarily removed from the data set.
[0047] Following the example in Step 2, we used the trained autoencoder model to construct an outlier detection model. This model reconstructs the input coordinate vectors and calculates the reconstruction error to assign an anomaly score to each coordinate vector. We set the anomaly score threshold to 0.88. Based on this threshold, we detected approximately 800 outlier coordinate data points and temporarily removed them from the dataset to prevent them from interfering with the cluster center calculation.
[0048] Step 4: Update the cluster centers of the optimized dataset: Use the prediction results of the large model to adjust the cluster center updates. The large model is used to predict the distribution trend of samples within each cluster. When the new cluster center is calculated, the location of the cluster center is fine-tuned in combination with the prediction information of the large model. The iteration is stopped according to the iterative termination condition to obtain the final cluster center, completing the electricity consumption clustering processing of the multi-source dataset. The iterative termination condition is set: a double iterative termination condition is set, in which the iteration is stopped when the change in the cluster center in two consecutive iterations is less than a preset threshold, such as when the Euclidean distance change is less than 0.01; the maximum number of iterations is set, such as 100 times; when any of these conditions are met, the iteration is terminated and the final clustering result is output.
[0049] Step 5: Adjust the time-of-use electricity price and conduct grid dispatch based on the electricity consumption cluster center.
[0050] Example 2
[0051] The present invention also provides a device for processing multi-source data based on the K-Means aggregation algorithm of a large model, comprising a feature extraction module, an initial clustering module, a detection module, an update module and a scheduling module.
[0052] The feature extraction module performs feature pre-extraction on electricity consumption data based on a large model. The module uses the large model to extract features from the input dataset, converting the raw data in the dataset into representative and discriminative feature vectors. The electricity consumption data datasets include text data datasets, image data datasets, structured data datasets, and geographic coordinate datasets.
[0053] The initial clustering module determines the initial electricity consumption cluster center: the K-Means++ algorithm is used to determine the initial cluster center for each extracted feature vector: a feature vector is randomly selected as the first initial cluster center, and the shortest distance from each feature vector to the selected initial cluster center is calculated. The longer the distance, the higher the probability of being selected as the next initial cluster center. The selection process is repeated until K initial cluster centers are selected.
[0054] The detection module detects and processes outliers: During the clustering process, the large model is used to build an outlier detection model, which performs anomaly scores on each feature vector in different data sets. Samples with scores exceeding the set threshold are identified as outliers and temporarily removed from the data set.
[0055] The update module updates and optimizes the cluster center of the data set: the cluster center is updated and adjusted using the prediction results of the large model. The large model is used to predict the distribution trend of samples in each cluster. When the new cluster center is calculated, the position of the cluster center is fine-tuned in combination with the prediction information of the large model. The iteration is stopped according to the termination condition to obtain the final cluster center, completing the electricity consumption clustering processing of the multi-source data set.
[0056] The dispatching module adjusts the time-of-use electricity price and performs grid dispatching based on the electricity consumption cluster center.
[0057] Since the information interaction, execution process and other contents between the modules in the above-mentioned device are based on the same concept as the embodiment of the method of the present invention, the specific contents can be found in the description of the embodiment of the method of the present invention and will not be repeated here.
[0058] Similarly, the device of the present invention improves the stability of electricity consumption data clustering: by determining the initial clustering center based on large model feature pre-extraction and the improved K-Means++ algorithm, the selection of the initial clustering center is made more scientific and representative, reducing the difference in clustering results caused by random selection of the initial center, and improving the stability of electricity consumption data clustering.
[0059] Enhanced clustering accuracy: The outlier detection model constructed by the large model effectively identifies and processes outliers, avoiding the interference of outliers on the cluster center. At the same time, the prediction results of the large model are used to optimize the update of the cluster center, making the cluster center more consistent with the actual distribution of the data, significantly improving the accuracy of electricity consumption data clustering.
[0060] Improving the computational efficiency of electricity consumption data: Pre-extraction of features from large models reduces data dimensionality, reducing the computational effort of the subsequent K-Means algorithm. When processing large-scale, high-dimensional data, the optimized algorithm reduces computational complexity and accelerates convergence, improving overall algorithm execution efficiency and better meeting application scenarios with high real-time requirements.
[0061] It should be noted that not all steps and modules in the above-mentioned processes and device structures are required, and certain steps or modules can be omitted according to actual needs. The execution order of each step is not fixed and can be adjusted as needed. The system structure described in the above-mentioned embodiments can be a physical structure or a logical structure, that is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities, or may be implemented by certain components in multiple independent devices.
[0062] The above embodiments are merely preferred embodiments for the purpose of fully illustrating the present invention, and the scope of protection of the present invention is not limited thereto. Equivalent substitutions or modifications made by those skilled in the art based on the present invention are within the scope of protection of the present invention. The scope of protection of the present invention shall be subject to the claims.
Claims
1. A method for processing multi-source data based on the K-Means aggregation algorithm of a large model, characterized by include: Step 1: Pre-extract features from electricity consumption data based on a large model: Use the large model to extract features from the input dataset, converting the raw data into representative and discriminative feature vectors. The electricity consumption data datasets include text data, image data, structured data, and geographic coordinate data. Step 2: Determine the initial electricity consumption cluster center: For each extracted feature vector, use the K-Means++ algorithm to determine the initial cluster center: randomly select a feature vector as the first initial cluster center, calculate the shortest distance from each feature vector to the selected initial cluster center, and repeat the selection process until K initial cluster centers are selected according to the principle that the greater the distance, the higher the probability of being selected as the next initial cluster center. Step 3: Detect and process outliers: During the clustering process, the large model is used to build an outlier detection model. Each feature vector in different data sets is scored for abnormality. Samples with scores exceeding the set threshold are identified as outliers and temporarily removed from the data set. Step 4: Update the cluster center of the optimized data set: Use the prediction results of the large model to adjust the cluster center update, where the large model is used to predict the distribution trend of samples in each cluster. When the new cluster center is calculated, the prediction information of the large model is combined to fine-tune the position of the cluster center, and the iteration is stopped according to the iteration termination condition to obtain the final cluster center, completing the electricity consumption clustering processing of the multi-source data set. Step 5: Adjust the time-of-use electricity price and conduct grid dispatch based on the electricity consumption cluster center.
2. The method for processing multi-source data using a large-model K-Means aggregation algorithm according to claim 1 is characterized in that in step 1, a language model based on a Transformer architecture, an image model based on a convolutional neural network, a model based on a multilayer perceptron architecture, and a model based on an autoencoder architecture are selected for feature pre-extraction.
3. The method for processing multi-source data using a K-Means aggregation algorithm based on a large model according to claim 1, characterized in that In step 1, the input data set is extracted using a large model, including: When the dataset is a text data dataset, the text is preprocessed first, including word segmentation and stop word removal. Then the preprocessed text data is input into the corresponding large model to obtain the vector representation of the text. When the data set is an image data set, the image data is input into the corresponding large model to extract the feature vector of the image, and the image data is preprocessed to unify the resolution before input. When the data set is a structured data set, the structured data is organized into a suitable format and then input into the adapted large model, which outputs a feature vector that condenses the key information. When the dataset is a geographic coordinate dataset, the longitude and latitude data are standardized and then input into the corresponding large model to output a low-dimensional feature vector that retains the key information of the spatial position relationship.
4. The method for processing multi-source data using a K-Means aggregation algorithm based on a large model according to claim 1, characterized in that In step 4, set the iteration termination condition: set the double iteration termination condition, where the iteration is stopped when the change of the cluster center in two consecutive iterations is less than the preset threshold; set the maximum number of iterations; when any one of the conditions is met, terminate the iteration and output the final clustering result.
5. A device for processing multi-source data based on a large-scale K-Means aggregation algorithm, characterized by Including feature extraction module, initial clustering module, detection module, update module and scheduling module, The feature extraction module performs feature pre-extraction on electricity consumption data based on a large model. The module uses the large model to extract features from the input dataset, converting the raw data in the dataset into representative and discriminative feature vectors. The electricity consumption data datasets include text data datasets, image data datasets, structured data datasets, and geographic coordinate datasets. The initial clustering module determines the initial electricity consumption cluster center: the K-Means++ algorithm is used to determine the initial cluster center for each extracted feature vector: a feature vector is randomly selected as the first initial cluster center, and the shortest distance from each feature vector to the selected initial cluster center is calculated. The longer the distance, the higher the probability of being selected as the next initial cluster center. The selection process is repeated until K initial cluster centers are selected. The detection module detects and processes outliers: During the clustering process, the large model is used to build an outlier detection model, which performs anomaly scores on each feature vector in different data sets. Samples with scores exceeding the set threshold are identified as outliers and temporarily removed from the data set. The update module updates and optimizes the cluster center of the data set: the cluster center is updated and adjusted using the prediction results of the large model. The large model is used to predict the distribution trend of samples in each cluster. When the new cluster center is calculated, the position of the cluster center is fine-tuned in combination with the prediction information of the large model. The iteration is stopped according to the termination condition to obtain the final cluster center, completing the electricity consumption clustering processing of the multi-source data set. The dispatching module adjusts the time-of-use electricity price and performs grid dispatching based on the electricity consumption cluster center.
6. The device for processing multi-source data based on the K-Means aggregation algorithm of a large model according to claim 5 is characterized in that feature extraction The module uses a language model based on the Transformer architecture, an image model based on a convolutional neural network, a model based on a multi-layer perceptron architecture, and a model based on an autoencoder architecture for feature pre-extraction.
7. The device for processing multi-source data based on the K-Means aggregation algorithm of a large model according to claim 5 is characterized in that feature extraction The module uses a large model to extract features from the input data set, including: When the dataset is a text data dataset, the text is preprocessed first, including word segmentation and stop word removal. Then the preprocessed text data is input into the corresponding large model to obtain the vector representation of the text. When the data set is an image data set, the image data is input into the corresponding large model to extract the feature vector of the image, and the image data is preprocessed to unify the resolution before input. When the data set is a structured data set, the structured data is organized into a suitable format and then input into the adapted large model, which outputs a feature vector that condenses the key information. When the dataset is a geographic coordinate dataset, the longitude and latitude data are standardized and then input into the corresponding large model to output a low-dimensional feature vector that retains the key information of the spatial position relationship.
8. The device for processing multi-source data using a large-model K-Means aggregation algorithm according to claim 5, characterized in that The update module sets the iteration termination condition: sets a double iteration termination condition, where the iteration stops when the change of the cluster center in two consecutive iterations is less than the preset threshold; sets the maximum number of iterations; when any one of the conditions is met, the iteration is terminated and the final clustering result is output.
Citation Information
Patent Citations
Discriminant text clustering method and system based on minimum normalized information distance
CN110955773A
Power grid user data analysis method based on improved k-means clustering algorithm
CN116894744A
Data identification method, system and equipment based on K-means clustering and storage medium
CN118503768A
Streaming speech recognition method, device and equipment based on adaptive AI large model
CN119252234A
Large model text auditing optimization method based on clustering preprocessing
CN119807427A