Multi-source heterogeneous data processing method based on K-means algorithm and computer equipment
Through the multi-source heterogeneous data processing method based on the K-means algorithm, the problems of high data processing complexity, insufficient real-timeness and low degree of false data removal in the prior art are solved, real-time processing and effectiveness of data are improved, and effective decision-making information is provided.
Patent Information
- Application Number
- CN202510677783.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing multi-source heterogeneous data processing algorithm has high computational complexity, insufficient real-timeness, low degree of false data removal, high redundancy of output data, and inability to give effective decision information.
Multi-source heterogeneous data processing method based on the K-means algorithm is adopted to achieve real-time processing and effectiveness improvement of data through data acquisition, cleaning, spatial and temporal registration, multi-dimensional feature extraction, dynamic adjustment of the number of clusters in the K-means algorithm, calculate the deviation between the sensor measurement value and the center of the cluster cluster, and use the weighted average algorithm to fusion the data to achieve real-time processing and efficiency improvement.
Real-time processing of multi-source heterogeneous data is realized, which improves the credibility and effectiveness of data, reduces data fuzziness and redundancy, and provides effective decision-making information.
Smart Images

Figure CN120196977A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multi-source heterogeneous data processing. Specifically, it relates to a multi-source heterogeneous data processing method and a computer device based on the K-means algorithm. Background Art
[0002] In the current era of rapid development of information technology, people are faced with data from different sources in various fields, including data measured by various sensors. These data are for the same target within the same time period. There are problems with inconsistent data sources and formats, and even incorrect and invalid data. Decision-makers cannot directly use these data for evaluation and analysis and give decisions. Therefore, data processing is required to make full use of the information provided by these data to meet the needs of in-depth data analysis and rapid decision-making. However, existing multi-source heterogeneous data processing algorithms have high computational complexity, insufficient real-time performance for processing large volumes of data, low degree of false data elimination, and high redundancy of output data, and cannot provide effective decision-making information. Summary of the Invention
[0003] The purpose of the present invention is to provide a multi-source heterogeneous data processing method, a computer device, a computer-readable storage medium, and a computer program product based on the K-means algorithm for the problems existing in the prior art, which can realize real-time processing of multi-source heterogeneous data, improve the credibility and effectiveness of data, and reduce data ambiguity and redundancy.
[0004] To achieve the above purpose, one aspect of the present invention provides a multi-source heterogeneous data processing method based on the K-means algorithm, including: Step S1, collecting and real-time receiving multi-source data from external data sources, and performing format standardization and consistency processing on the data through data parsing and recombination, where the multi-source data from external data sources are measurement data of multiple sensors; Step S2, cleaning, spatially registering, and temporally registering the data. Identifying and eliminating abnormal data through data cleaning, converting the data to a unified coordinate and performing time alignment through spatial registration, and unifying the data into a single time series through temporal registration; Step S3, extracting multi-dimensional features from the data, using the extracted features as the input of the K-means algorithm, and dynamically adjusting the number of clusters in the K-means algorithm until the sum of the squared distances from the cluster centers to all data objects is minimized, to realize the association of data belonging to the same target; Step S4, calculating the deviation between the associated measurement values of each sensor and the cluster center, calculating the standard deviation of the deviation using statistical theory, using the standard deviation as an approximation of the measurement accuracy of each sensor to calculate weights, and performing data fusion using the weighted average algorithm; Step S5: Integrate the fused data and output it in real time.
[0005] Another aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the above method.
[0006] Another aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.
[0007] Another aspect of the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the above method are implemented.
[0008] According to the multi-source heterogeneous data processing method, computer device, computer-readable storage medium, and computer program product based on the K-means algorithm in the above aspects of the present invention, real-time processing of multi-source heterogeneous data can be achieved, the credibility and effectiveness of data can be improved, and data ambiguity and redundancy can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts: Figure 1 is a flowchart of a multi-source heterogeneous data processing method based on the K-means algorithm according to an embodiment of the present invention; Figure 2 is a structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0010] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the drawings. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0011] An embodiment of the present invention provides a multi-source heterogeneous data processing method based on the K-means algorithm. As Figure 1 shown, the multi-source heterogeneous data processing method based on the K-means algorithm according to the embodiment of the present invention includes steps S1 to S5.
[0012] Step S1, data acquisition and reception. Multisource data from external data sources is acquired and received in real time through communication processing, and data format standardization and consistency processing are performed through data parsing and recombination to process the input data into data of the same format for the next step of processing. In this embodiment, the multisource data from external data sources is the measurement data of multiple sensors. In this step, data format standardization and consistency processing are performed through data parsing and recombination, and data with different measurement positions, measurement angles, and measurement accuracies are processed, which can improve the credibility and effectiveness of the data.
[0013] Step S2, data preprocessing. It includes cleaning, spatial registration, and temporal registration of the original data, aiming to improve the quality of the data and the efficiency of subsequent correlation and fusion.
[0014] The original multisource data usually contains noise, outliers, and redundant information, and these problems are solved through data preprocessing. First, data cleaning is used to identify and remove abnormal data, such as data coordinates outside the reasonable range, abnormal data timestamps (e.g., time is 0), two consecutive frames of data being the same, and the data rate being greater than the equipment data rate. Secondly, spatial registration is used to convert data points into a unified coordinate and perform temporal alignment to ensure data consistency. Then temporal registration is performed. Due to different initial sampling times of sensors, different temporal accuracies used, combined with different data transmission delays and different reporting frequencies of sensors, temporal registration technology is needed to process the data. Temporal registration unifies the local data of each sensor into a time series to achieve temporal unity. To meet the requirements of real-time and high-efficiency of the temporal registration method, an adaptive filtering method combined with Lagrange interpolation method is adopted. The adaptive filtering reduces the noise of the data and improves the performance of the Lagrange interpolation method in the temporal registration algorithm.
[0015] Step S3, data association. Data association is a crucial step before data fusion. Data association is to determine whether the collected data belongs to the same target object according to a certain algorithm. In this embodiment, it is judged whether the measurement results of different sensors belong to the same target according to the association criterion to ensure the association correct rate in cases such as dense targets, trajectory intersections, complex target maneuvers, and clutter interference. The basic method is to extract multi-dimensional features from the sampled measurement data, including information such as position, velocity, and acceleration, and exclude the obvious clutter or interference points to improve the processing accuracy of subsequent data fusion. Then, the K-means algorithm is used for association by utilizing the attribute information contained in the target set. This algorithm has simple parameter settings and low computational complexity, and is suitable for real-time data input and medium-dense target environments.
[0016] Data association is carried out using the K-means algorithm. The application of the K-means algorithm in multi-source data association is a complex but feasible process, involving multiple steps such as data preprocessing, clustering analysis, dynamic data processing, and post-processing. The principle is as follows: Given a data set X, which contains n data objects, and each data object is represented by d-dimensional attributes. At this time , where , the division of data clusters is represented by C , divide the n data into K data clusters, and each division is represented by . Select the Euclidean distance as the similarity measurement criterion. The sum of squares of the distances from the data objects in each cluster to the cluster center value can be calculated using the following formula:
[0017] where is the sum of squares of the distances from the data objects in each cluster to the cluster center value represents the i-th data object in the data set represents the k-th cluster center value represents the k-th data cluster
[0018] The goal of algorithm execution is to minimize the sum of squares of the distances between each data object and the corresponding cluster center value
[0019]
[0020] represents the sum of the sums of squares of the distances between each data object and the corresponding cluster center value in the K clusters
[0021] Step S3 specifically includes the following steps S31 to S37: Step S31, feature extraction: Extract the key features of the target, including position, speed, direction, and acceleration, which will be used as the input of the K-means algorithm
[0022] Step S32, normalization processing: Since the K-means algorithm is sensitive to data scales, it is necessary to normalize the features to eliminate the influence of different feature dimensions
[0023] Step S33, determine the initial K value and initialize the cluster center: K is the number of clusters. Use the silhouette coefficient method to determine the K value, and randomly select a data point in each cluster as the first clustering cluster center
[0024] Step S34, use the Euclidean distance to calculate the distance from each data object to the clustering cluster center, and assign each data object to the nearest data cluster
[0025] Step S35, clustering process: Recalculate the distance from each data point to each cluster center, and select the next cluster center according to this distance. The farther the point is, the higher the probability of being selected. In this way, the distribution between cluster centers will be more uniform, avoiding the over-concentration of initial cluster centers.
[0026] Step S36, combining with dynamic model: Since the target position changes over time, combine the K-means algorithm with dynamic models such as Kalman filter to update the target state and handle data association.
[0027] Step S37, dynamically adjust K, and repeat the processes of steps S34 and S36 until the sum of the squares of the distances from the cluster centers to all data objects is minimized.
[0028] Step S4, data fusion. Data fusion refers to making full use of the information obtained by multiple sensors at different spaces and times, performing complementary fusion advantages, ensuring the integrity of the original information of multiple targets to the greatest extent, and obtaining a more complete trajectory and a more accurate fusion result at the same time. In this embodiment, the measurement information from different sensors is integrated to improve the accuracy and robustness of the data, eliminate redundant information, reduce uncertainty, and thus obtain a more accurate target state estimate. When the sensor characteristics are unknown and the random errors are uncertain, a dynamic weighting method is used for fusion.
[0029] Since the true motion trajectory of the target is unknown, the clustering center points of the measurement data of each sensor are used as the reference points for the true position of the target, and the deviation between the measurement value of each sensor and the reference point is regarded as the measurement deviation. Use statistical theory to calculate the standard deviation of the deviation of the position information (distance, azimuth, pitch) of each sensor, and regard this standard deviation as an approximation of the measurement accuracy to obtain the weight values used in the weighted fusion calculation of multi-source data.
[0030] Step S4 specifically includes the following steps S41 to S46: Step S41, calculate the cluster center values of the measurement values of each sensor: , , ,
[0031] In the above formula, N is the total measurement time of the sensor, M is the number of measurement points after association, and m is the index of each associated point. , , are the distance value, azimuth angle value, and pitch angle value of the cluster center at the i-th measurement time point respectively, , , are the distance value, azimuth angle value, and pitch angle value of each associated measurement point at the i-th measurement time point respectively.
[0032] Step S42: Calculate the differences between the associated measurement values of each sensor and the cluster center value:
[0033] In the above formula, , , are respectively the distance difference, azimuth difference, and elevation angle difference between the associated measurement values of each sensor and the cluster center at the i-th measurement time point.
[0034] Step S43: Calculate the mean values of the differences between the associated measurement values of each sensor and the cluster center value:
[0035] In the above formula, , , are respectively the mean distance value, mean azimuth value, and mean elevation angle value of the differences.
[0036] Step S44: Calculate the standard deviations of the differences between the associated measurement values of each sensor and the cluster center value:
[0037] In the above formula, , , are respectively the standard deviations of the distance difference, azimuth difference, and elevation angle difference between the associated measurement value and the cluster center value.
[0038] Step S45: Take the standard deviation as an approximation of the accuracy of each sensor, and calculate the distance, azimuth, and elevation weights:
[0039] In the above formula, , and are respectively the distance weight, azimuth weight, and elevation weight of the m-th associated point.
[0040] Step S46: Perform data fusion using the weighted average algorithm:
[0041] In the above formula, , , are the fused distance value, azimuth angle value, and elevation angle value at the i-th measurement time point.
[0042] Step S5: Data integration and output. Integrate the fused data in a unified format and output it to the data communication interface in real time.
[0043] The multi-source heterogeneous data processing method based on the K-means algorithm in the embodiments of the present invention has the following beneficial effects: By processing the measurement data of various sensors with different measurement positions, measurement angles, and measurement accuracies, the credibility and effectiveness of the data can be improved. Through data preprocessing, data cleaning, spatial registration, and temporal registration are performed to improve the quality of the data and the efficiency of subsequent association and fusion. Through data association and fusion, the degree of control and reliability of the data are improved, data ambiguity and redundancy are reduced, and false data is eliminated. By using the detection data of sensors at different sites for the target, the spatio-temporal coverage range is expanded, and the integrity of the trajectory is improved. Through multi-source heterogeneous data processing, the measurement accuracy is improved, random errors are reduced, and it supports the realization of higher-level applications and services with intelligence and high efficiency.
[0044] An embodiment of the present invention further provides a computer device, which may be a server, and its internal structure diagram may be as Figure 2 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the operation parameter data of each framework. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the steps of the method in the embodiments of the present invention are implemented.
[0045] Those skilled in the art can understand that Figure 2 the structure shown in
[0046] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0046] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method in the embodiments of the present invention are implemented.
[0047] An embodiment of the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method in the embodiments of the present invention are implemented.
[0048] Only certain exemplary embodiments of the present invention have been described by way of illustration. Without doubt, for those of ordinary skill in the art, the described embodiments can be modified in various different ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. A multi-source heterogeneous data processing method based on the K-means algorithm, characterized in that Including: Step S1: Collect and receive in real time multi-source data from external data sources, and perform format standardization and consistency processing on the data through data parsing and recombination. The multi-source data from external data sources is the measurement data of multiple sensors; Step S2: Clean, spatially register, and temporally register the data. Identify and eliminate abnormal data through data cleaning, convert the data to a unified coordinate and perform time alignment through spatial registration, and unify the data into a single time series through temporal registration; Step S3: Extract multi-dimensional features from the data, use the extracted features as the input of the K-means algorithm, and dynamically adjust the number of clusters in the K-means algorithm until the sum of the squared distances from the cluster centers to all data objects is minimized, realizing the association of data belonging to the same target; Step S4: Calculate the deviation between the associated measurement values of each sensor and the cluster center, calculate the standard deviation of the deviation using statistical theory, use this standard deviation as an approximation of the measurement accuracy of each sensor to calculate the weights, and perform data fusion using the weighted average algorithm; Step S5: Integrate the fused data and output it in real time.
2. The method according to claim 1, wherein Step S3 includes: Step S31: Extract the key features of the target as the input of the K-means algorithm; Step S32: Normalize the features to eliminate the influence of different feature dimensions; Step S33: Use the silhouette coefficient method to determine the number of clusters, and randomly select a data point in each cluster as the first cluster center; Step S34: Calculate the distance from each data object to the cluster center using the Euclidean distance, and assign each data object to the nearest data cluster; Step S35: Recalculate the distance from each data point to the cluster centers, and select the next cluster center based on this distance; Step S36: Combine the K-means algorithm with the Kalman filter to update the target state; Step S37: Dynamically adjust the number of clusters, and repeat steps S34 - S36 until the sum of the squared distances from the cluster centers to all data objects is minimized.
3. The method according to claim 1 or 2, characterized in that Step S4 includes: Calculate the cluster center values of the measurement values of each sensor; Calculate the difference between the associated measurement values of each sensor and the cluster center values; Calculate the mean of the differences between the associated measurement values of each sensor and the cluster center values; Calculate the standard deviation of the differences between the associated measurement values of each sensor and the cluster center values; Use the standard deviation as an approximation of the accuracy of each sensor to calculate the distance, azimuth, and elevation weights; Perform data fusion using the weighted average algorithm.
4. The method according to claim 3, characterized in that The distance value, azimuth angle value, and elevation angle value after data fusion are respectively: where N is the total measurement time of the sensor, , , are the fused distance value, azimuth value, and elevation angle value at the i-th measurement time point respectively, M is the number of measurement points after association, , , are the distance value, azimuth value, and elevation angle value of each associated measurement point at the i-th measurement time point respectively, , and are the distance weight, azimuth weight, and elevation weight of the m-th associated measurement point respectively.
5. The method according to claim 4, wherein Calculate the distance, azimuth, and elevation weights as follows: Among them, , , are the standard deviations of the distance difference, azimuth difference, and elevation difference between the associated measurement value and the cluster center value, respectively.
6. The method according to claim 5, wherein Calculate the standard deviation of the differences between the associated measurement values and the cluster center values as follows: Among them, , , are respectively the mean distance, mean azimuth angle, and mean elevation angle of the difference between the associated measurement value and the cluster center value, , , are respectively the distance difference, azimuth angle difference, and elevation angle difference between the associated measurement value of each sensor at the i-th measurement time point and the cluster center value.
7. The method according to claim 6, wherein Calculate the differences between the associated measurement values of each sensor and the cluster center values as follows: Among them, , , are the distance value, azimuth value, and elevation angle value of each associated measurement point at the i-th measurement time point respectively, , , are the distance value, azimuth value, and elevation angle value of the cluster center at the i-th measurement time point respectively.
8. A computer device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 - 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 - 7.
10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 - 7.
Citation Information
Patent Citations
Double-station radar cross positioning method and system based on distance weighted fusion, and medium
CN111366921A
Multi-sensor multi-target cooperative detection information fusion method and system
CN111860589A
Heterogeneous sensor information fusion method and device
CN116340736A
Passive sound positioning fusion method, system and device for underwater small platform detection and medium
CN118962590A
Virtual intelligence and optimization through multi-source, real-time, and context-aware real-world data
US20210056459A1
Cited By
Bayesian optimization-based multi-source traffic block data fusion weight dynamic optimization method
CN121459584A
Multi-source traffic block data fusion weight dynamic optimization method based on bayesian optimization
CN121459584B