A distribution network data comparison method and terminal
By performing cluster analysis of distribution network data and the application of Heming distance algorithm, the problem of low matching accuracy in distribution network data comparison is solved, and higher data matching accuracy and business applicability are achieved.
Patent Information
- Application Number
- CN202210696758.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-20
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-06-20
AI Technical Summary
The existing data comparison algorithms have low matching accuracy in distribution network data, especially when the data characteristics are obvious and the amount of data is large.
By clustering and analyzing the sample data of the distribution network, a feature data set is determined, and the data to be compared are marked and spliced based on the feature data set to form a long string, and then the similarity between the data is calculated using the Heming distance algorithm.
Through the combination of clustering analysis and Heming distance algorithm, the characteristic classification and similarity of the data can be accurately determined, and similar data can be matched to the maximum extent, improving the matching accuracy of distribution network data.
Smart Images

Figure CN115186138B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data comparison, and in particular to a distribution network data comparison method and terminal. Background Art
[0002] Since the archive data and operating data of the distribution network data may exist in multiple systems, when conducting business analysis and auxiliary decision-making based on the archive data and operating data, we often encounter the problem of inconsistent data calibers, but cannot use unified codes and names to strongly associate them. For example, the archive data needs to be based on system A, and the operating data needs to be based on system B, but there is no unique association between the archive data of system A and the operating data of system B. In this business scenario, it is necessary to perform similarity matching on the data of the two systems A and B and take the data intersection, which involves the comparison between different data.
[0003] For data comparison solutions, the most commonly used technology is to format and standardize the data, then form a unified file format or database model, and then perform associative fuzzy matching on fixed columns of the file or data model. The applied algorithms mainly include text fuzzy matching algorithm, similarity algorithm and distance algorithm.
[0004] Text fuzzy matching algorithm takes SequenceMatcher as an example. The SequenceMatcher class can be used to compare two data of any type as long as it can be hashed. It uses an algorithm to calculate the longest continuous subsequence of a sequence and ignores meaningless "useless data". The idea is to find the longest continuous matching subsequence that does not contain "junk" elements. These "junk" elements are uninteresting in some sense, such as blank lines or blanks (junk information processing is an extension of the Ratcliff and Obershelp algorithm). Then, the same idea is recursively applied to the left and right subsequences of the matching subsequence. This does not produce the smallest edit sequence, but it produces matches that people "look right". SequenceMatcher supports a heuristic method that automatically treats certain sequence items as junk. The heuristic counts the number of times each individual item appears in the sequence. If the repetitions of an item (after the first) account for more than 1% of the sequence and the sequence is at least 200 items long, the item will be marked as "popular" and treated as junk for sequence matching. This heuristic can be turned off by setting the autojunk parameter to False when creating a SequenceMatcher.
[0005] Similarity algorithms, such as cosine similarity, use the cosine value of the angle between two vectors in a vector space as a measure of the difference between two individuals. When the cosine value is close to 1 and the angle is close to 0, the two vectors are more similar. When the cosine value is close to 0 and the angle is close to 90 degrees, the two vectors are less similar.
[0006] Distance algorithms, such as the Hamming distance, calculate the number of different characters in corresponding positions between two equal-length strings by performing an exclusive-OR (xor) operation on two bit strings. The shorter the Hamming distance, the higher the similarity.
[0007] However, the above algorithms have their own shortcomings. For example, in certain business scenarios, such as distribution network data with obvious data features and large data volume, the performance of the text fuzzy matching algorithm is not ideal; and in the case of small text content comparison, the calculation error of the Hamming distance is large. Therefore, if the existing comparison algorithm is used to compare the distribution network data, the matching accuracy is not high. Summary of the invention
[0008] The technical problem to be solved by the present invention is to provide a distribution network data comparison method and terminal, which can improve the matching accuracy of the distribution network data.
[0009] In order to solve the above technical problems, a technical solution adopted by the present invention is:
[0010] A method for comparing distribution network data comprises the following steps:
[0011] S1. Perform cluster analysis on the distribution network sample data to determine a feature data set corresponding to the distribution network sample data;
[0012] S2. Mark the first data and the second data to be compared respectively according to the feature data set to determine a first feature classification corresponding to the first data and a second feature classification corresponding to the second data;
[0013] S3, according to the feature classification result, concatenate the data with the same feature classification in the first data and the second data by taking the same fields, respectively, to obtain a first comparison string corresponding to the first data and a second comparison string corresponding to the second data;
[0014] S4. Calculate the Hamming distance between the first comparison string and the second comparison string, and determine the similarity between the first data and the second data according to the Hamming distance.
[0015] In order to solve the above technical problems, another technical solution adopted by the present invention is:
[0016] A distribution network data comparison terminal comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program:
[0017] S1. Perform cluster analysis on the distribution network sample data to determine a feature data set corresponding to the distribution network sample data;
[0018] S2. Mark the first data and the second data to be compared respectively according to the feature data set to determine a first feature classification corresponding to the first data and a second feature classification corresponding to the second data;
[0019] S3, according to the feature classification result, concatenate the data with the same feature classification in the first data and the second data by taking the same fields, respectively, to obtain a first comparison string corresponding to the first data and a second comparison string corresponding to the second data;
[0020] S4. Calculate the Hamming distance between the first comparison string and the second comparison string, and determine the similarity between the first data and the second data according to the Hamming distance.
[0021] The beneficial effects of the present invention are as follows: when comparing distribution network data, firstly cluster analysis is performed on distribution network sample data to obtain a feature data set, and then the first data and the second data to be compared are marked based on the feature data set to determine the feature classifications corresponding to the first data and the second data respectively, and then the data with the same feature classification and the same field in the first data and the second data are spliced to form a long character string, and finally the Hamming distance algorithm is used to perform distance calculation on the long character string of the first data and the second data to determine the similarity of the first data and the second data, firstly the feature classification corresponding to each data can be accurately determined by cluster analysis, so that the long character string belonging to the same category can be accurately spliced, and then the Hamming distance algorithm is used to perform distance calculation on the long character string to realize data comparison, and by combining the feature classification result with the Hamming distance algorithm, similar data can be matched to the maximum extent, and the data intersection of different data systems can be determined, thereby greatly improving the matching accuracy of distribution network data. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 A flow chart of the steps of a method for comparing distribution network data according to an embodiment of the present invention;
[0023] Figure 2 A schematic diagram of the structure of a distribution network data comparison terminal according to an embodiment of the present invention;
[0024] Figure 3 1 is a daily load curve diagram of a distribution transformer under different power consumption modes according to an embodiment of the present invention. DETAILED DESCRIPTION
[0025] In order to explain the technical content, achieved objectives and effects of the present invention in detail, the following is an explanation in combination with the implementation modes and the accompanying drawings.
[0026] Please refer to Figure 1 , a method for comparing distribution network data, comprising the steps of:
[0027] S1. Perform cluster analysis on the distribution network sample data to determine a feature data set corresponding to the distribution network sample data;
[0028] S2. Mark the first data and the second data to be compared respectively according to the feature data set to determine a first feature classification corresponding to the first data and a second feature classification corresponding to the second data;
[0029] S3, according to the feature classification result, concatenate the data with the same feature classification in the first data and the second data by taking the same fields, respectively, to obtain a first comparison string corresponding to the first data and a second comparison string corresponding to the second data;
[0030] S4. Calculate the Hamming distance between the first comparison string and the second comparison string, and determine the similarity between the first data and the second data according to the Hamming distance.
[0031] From the above description, it can be seen that the beneficial effect of the present invention is that: when comparing the distribution network data, the distribution network sample data is first clustered and analyzed to obtain a feature data set, and then the first data and the second data to be compared are marked based on the feature data set to determine the feature classifications corresponding to the first data and the second data respectively, and then the data with the same feature classification and the same field in the first data and the second data are spliced to form a long character string, and finally the Hamming distance algorithm is used to calculate the distance between the long character string of the first data and the second data to determine the similarity between the first data and the second data, firstly, the feature classification corresponding to each data can be accurately determined through cluster analysis, so that the long character string belonging to the same category can be accurately spliced, and then the Hamming distance algorithm is used to calculate the distance of the long character string to realize data comparison, and by combining the feature classification results with the Hamming distance algorithm, similar data can be matched to the maximum extent, and the data intersection of different data systems can be determined, which greatly improves the matching accuracy of the distribution network data.
[0032] Furthermore, the step S1 comprises:
[0033] Normalize the distribution network sample data to obtain the feature vector set for cluster analysis;
[0034] Using K-means clustering algorithm to perform cluster analysis on the feature vector set to obtain a clustering result;
[0035] According to the clustering result, a feature vector set corresponding to the clustering result is selected to obtain feature data corresponding to each cluster, and a feature data set corresponding to the distribution network sample data is determined according to the feature data corresponding to each cluster.
[0036] From the above description, we can see that the K-means algorithm is an unsupervised machine algorithm. When the business data sample is large enough, it can calculate sufficiently accurate feature classification results. On the basis of determining the clustering results, further selection is performed to determine the feature data set, which further improves the accuracy of the feature classification results.
[0037] Furthermore, the step S2 comprises:
[0038] Compare the first data and the second data to be compared with the feature data set respectively, and determine a first similarity set between the first data and the feature data set and a second similarity set between the second data and the feature data set respectively;
[0039] A first feature classification of the first data is determined according to feature data corresponding to the highest similarity in the first similarity set, and a second feature classification of the second data is determined according to feature data corresponding to the highest similarity in the second similarity set.
[0040] From the above description, it can be seen that by comparing the data to be compared with the feature data set, the corresponding similarity set is determined, and the feature data corresponding to the highest similarity in the similarity set is determined as the feature classification of the data to be compared, thereby ensuring the accuracy of the determined feature classification of the data to be compared.
[0041] Further, the calculating the Hamming distance between the first comparison string and the second comparison string includes:
[0042] Performing word segmentation operations on the first comparison string and the second comparison string respectively to obtain corresponding first keyword sets and second keyword sets;
[0043] Steps S31-S34 are performed on the first keyword set and the second keyword set respectively to obtain corresponding first dimensionality reduction sequence strings and second dimensionality reduction sequence strings:
[0044] S31, mapping each keyword in the keyword set to a corresponding hash code according to the sample library;
[0045] S32, weighting the hash code corresponding to each keyword according to the weight of the keyword;
[0046] S33, accumulating and merging each weighted hash sequence in the keyword set to form a sequence string corresponding to the keyword set;
[0047] S34, performing a dimensionality reduction operation on the sequence string to obtain a dimensionality reduction sequence string corresponding to the keyword set;
[0048] The Hamming distance between the first comparison character string and the second comparison character string is calculated according to the first reduced dimension sequence string and the second reduced dimension sequence string.
[0049] From the above description, it can be seen that before calculating the Hamming distance, word segmentation, mapping, weighting, merging accumulation and dimension reduction operations are performed in sequence, which further ensures the accuracy of data matching.
[0050] Further, the determining the similarity between the first data and the second data includes:
[0051] It is determined whether the Hamming distance is less than a preset value. If so, the first data is similar to the second data; otherwise, the first data is not similar to the second data.
[0052] It can be seen from the above description that by comparing the Hamming distance with a preset value, it is convenient and quick to determine whether the first data and the second data are similar based on the comparison result.
[0053] Please refer to Figure 2 , a distribution network data comparison terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program:
[0054] S1. Perform cluster analysis on the distribution network sample data to determine a feature data set corresponding to the distribution network sample data;
[0055] S2. Mark the first data and the second data to be compared respectively according to the feature data set to determine a first feature classification corresponding to the first data and a second feature classification corresponding to the second data;
[0056] S3, according to the feature classification result, concatenate the data with the same feature classification in the first data and the second data by taking the same fields, respectively, to obtain a first comparison string corresponding to the first data and a second comparison string corresponding to the second data;
[0057] S4. Calculate the Hamming distance between the first comparison string and the second comparison string, and determine the similarity between the first data and the second data according to the Hamming distance.
[0058] From the above description, it can be seen that the beneficial effect of the present invention is that: when comparing the distribution network data, the distribution network sample data is first clustered and analyzed to obtain a feature data set, and then the first data and the second data to be compared are marked based on the feature data set to determine the feature classifications corresponding to the first data and the second data respectively, and then the data with the same feature classification and the same field in the first data and the second data are spliced to form a long character string, and finally the Hamming distance algorithm is used to calculate the distance between the long character string of the first data and the second data to determine the similarity between the first data and the second data, firstly, the feature classification corresponding to each data can be accurately determined through cluster analysis, so that the long character string belonging to the same category can be accurately spliced, and then the Hamming distance algorithm is used to calculate the distance of the long character string to realize data comparison, and by combining the feature classification results with the Hamming distance algorithm, similar data can be matched to the maximum extent, and the data intersection of different data systems can be determined, which greatly improves the matching accuracy of the distribution network data.
[0059] Furthermore, the step S1 comprises:
[0060] Normalize the distribution network sample data to obtain the feature vector set for cluster analysis;
[0061] Using K-means clustering algorithm to perform cluster analysis on the feature vector set to obtain a clustering result;
[0062] According to the clustering result, a feature vector set corresponding to the clustering result is selected to obtain feature data corresponding to each cluster, and a feature data set corresponding to the distribution network sample data is determined according to the feature data corresponding to each cluster.
[0063] From the above description, we can see that the K-means algorithm is an unsupervised machine algorithm. When the business data sample is large enough, it can calculate sufficiently accurate feature classification results. On the basis of determining the clustering results, further selection is performed to determine the feature data set, which further improves the accuracy of the feature classification results.
[0064] Furthermore, the step S2 comprises:
[0065] Compare the first data and the second data to be compared with the feature data set respectively, and determine a first similarity set between the first data and the feature data set and a second similarity set between the second data and the feature data set respectively;
[0066] A first feature classification of the first data is determined according to feature data corresponding to the highest similarity in the first similarity set, and a second feature classification of the second data is determined according to feature data corresponding to the highest similarity in the second similarity set.
[0067] From the above description, it can be seen that by comparing the data to be compared with the feature data set, the corresponding similarity set is determined, and the feature data corresponding to the highest similarity in the similarity set is determined as the feature classification of the data to be compared, thereby ensuring the accuracy of the determined feature classification of the data to be compared.
[0068] Further, the calculating the Hamming distance between the first comparison string and the second comparison string includes:
[0069] Performing word segmentation operations on the first comparison string and the second comparison string respectively to obtain corresponding first keyword sets and second keyword sets;
[0070] Steps S31-S34 are performed on the first keyword set and the second keyword set respectively to obtain corresponding first dimensionality reduction sequence strings and second dimensionality reduction sequence strings:
[0071] S31, mapping each keyword in the keyword set to a corresponding hash code according to the sample library;
[0072] S32, weighting the hash code corresponding to each keyword according to the weight of the keyword;
[0073] S33, accumulating and merging each weighted hash sequence in the keyword set to form a sequence string corresponding to the keyword set;
[0074] S34, performing a dimensionality reduction operation on the sequence string to obtain a dimensionality reduction sequence string corresponding to the keyword set;
[0075] The Hamming distance between the first comparison character string and the second comparison character string is calculated according to the first reduced dimension sequence string and the second reduced dimension sequence string.
[0076] From the above description, it can be seen that before calculating the Hamming distance, word segmentation, mapping, weighting, merging accumulation and dimension reduction operations are performed in sequence, which further ensures the accuracy of data matching.
[0077] Further, the determining the similarity between the first data and the second data includes:
[0078] It is determined whether the Hamming distance is less than a preset value. If so, the first data is similar to the second data; otherwise, the first data is not similar to the second data.
[0079] It can be seen from the above description that by comparing the Hamming distance with a preset value, it is convenient and quick to determine whether the first data and the second data are similar based on the comparison result.
[0080] Embodiment 1
[0081] Please refer to Figure 1 , a method for comparing distribution network data, comprising the steps of:
[0082] S1. Perform cluster analysis on the distribution network sample data to determine a feature data set corresponding to the distribution network sample data;
[0083] Among them, actual business sample data in the distribution network is prepared, such as sample data related to load conditions;
[0084] S2. Mark the first data and the second data to be compared respectively according to the feature data set to determine a first feature classification corresponding to the first data and a second feature classification corresponding to the second data;
[0085] Specifically, the first data and the second data to be compared are respectively compared with the feature data set, and a first similarity set between the first data and the feature data set and a second similarity set between the second data and the feature data set are respectively determined;
[0086] Determine a first feature classification of the first data according to the feature data corresponding to the highest similarity in the first similarity set, and determine a second feature classification of the second data according to the feature data corresponding to the highest similarity in the second similarity set;
[0087] When performing the comparison, the first data and the second data are first converted into data with the same format as the feature data in the feature data set, and then the comparison is continued. For example, the first data contains data X1, X2, X3, ..., Xi; the second data contains data Y1, Y2, Y3, ..., Yj; the feature data set contains feature data A1, A2, A3, ..., Am; then X1 in the first data is compared with A1, A2, A3, ..., Am respectively to obtain the corresponding similarity results B1, B2, ..., Bm, and the feature data corresponding to the similarity result with the smallest value is selected from B1, B2, ..., Bm, and the feature classification corresponding to A1 is determined based on the feature data. By analogy, the feature classifications corresponding to X2, X3, ..., Xi and Y1, Y2, Y3, ..., Yj can be calculated in sequence.
[0088] S3, according to the feature classification result, concatenate the data with the same feature classification in the first data and the second data by taking the same fields, respectively, to obtain a first comparison string corresponding to the first data and a second comparison string corresponding to the second data;
[0089] For example, if both the first data and the second data have a feature classification corresponding to the feature data Ak, and the data of the feature classification have the fields: device name, asset type, city name, and district / county name, then the data of these fields are extracted from the first data and the second data respectively, and concatenated: device name + asset type + city name + district / county name; and special characters are removed from the concatenated string to finally obtain a comparison string;
[0090] S4, calculating a Hamming distance between the first comparison string and the second comparison string, and determining a similarity between the first data and the second data according to the Hamming distance;
[0091] Specifically, determining the similarity between the first data and the second data includes:
[0092] Determine whether the Hamming distance is less than a preset value. If so, the first data is similar to the second data. Otherwise, the first data is not similar to the second data. For example, the preset value can be set to 3. A distance less than 3 indicates that the first data and the second data are similar, that is, they are associated. In this way, the data intersection of the same equipment of two different systems in the distribution network can be matched to meet actual business needs.
[0093] Embodiment 2
[0094] This embodiment further limits the use of K-means clustering algorithm to perform cluster analysis on the distribution network sample data, and finally obtains a data feature set, specifically:
[0095] Normalize the distribution network sample data to obtain the feature vector set for cluster analysis;
[0096] In this embodiment, 24-point load data of the distribution transformer are selected to form the characteristic vector of cluster analysis. The load power at each time point reflects the electricity consumption of users in different periods of time. Users in the same industry have similar load characteristics, so the daily load curves of users in different industries have strong distinguishability. Therefore, classification can be carried out based on different industries to achieve cluster analysis.
[0097] For users in the same industry, in order to avoid inaccurate classification when the load levels vary greatly, it is necessary to normalize the load power at each test time point:
[0098] Let P i =[p i1 , p i2 , p i3 ,…,p in ] is the power value of the i-th distribution transformer at point n, then P iThe corresponding standard value P′ can be obtained by normalizing according to the following formula: i :
[0099]
[0100] Where j = 1, 2, ..., n is the number of the power sampling point of the distribution transformer, p imax and p imin are the maximum and minimum values of the n-point power values of the i-th distribution transformer respectively;
[0101] Using K-means clustering algorithm to perform cluster analysis on the feature vector set to obtain a clustering result;
[0102] Selecting a feature vector set corresponding to the clustering result according to the clustering result to obtain feature data corresponding to each cluster, and determining a feature data set corresponding to the distribution network sample data according to the feature data corresponding to each cluster;
[0103] After cluster analysis, the clustering results are obtained, that is, the sample data is divided into a certain number of categories. In this embodiment, the sample data is the 24-point load data of the distribution transformer. After clustering, the distribution transformer industry can be clustered. After the clustering is completed, the clustering results can be further refined;
[0104] In this embodiment, selection can be performed in the following manner:
[0105] Determine the sample data corresponding to each category, count the number of sample data under each category, and eliminate low-probability events according to the quantitative difference of the result set. For example, you can eliminate the category whose sample number is greater than the first sample threshold and the category whose sample number is less than the second sample threshold. For example, if the clustered result set has 7 categories, and the number of samples in each category is 1, 2, 3, 4, 5, 6, and 7, then you can eliminate the categories with sample numbers of 1 and 7. The first sample threshold and the second sample threshold can be determined by statistical analysis of the clustering results. For example, you can count the average number of sample data for each category, and then determine the value corresponding to the first preset value less than the average number of sample data as the second sample threshold, and determine the value corresponding to the second preset value greater than the average number of sample data as the first sample threshold.
[0106] After cluster analysis, the data to be compared needs to be marked. In this embodiment, the daily load curve of a typical industry is obtained by cluster analysis of the load type of the distribution transformer (the 24-point daily load characteristic data set of the distribution transformer can be displayed in the form of a curve, so it can be called a daily load curve, such as Figure 3 As shown), load type identification is performed on distribution transformers with unknown industry attributes:
[0107] First, the daily load data of the distribution transformer to be marked is normalized. This normalization process is the same as the normalization process of sample data in cluster analysis.
[0108] Calculate the square of the spatial distance between the normalized typical daily load curve of the distribution transformer and the typical daily load curve of each industry. The smaller the distance, the higher the similarity between the distribution transformer and the industry. Select the industry with the highest similarity as the industry affiliation of the unknown type of distribution transformer. The calculation formula for the square of the spatial distance is as follows:
[0109]
[0110] where k = 1, 2, ..., n is the number of the power sampling point of the distribution transformer; Xj = [x j1 , x j2 , …, x jn ] is the power value (normalized) of the typical industry j at point n; Xi = [x i1 , x i2 , …, x in ] is the n-point power value (normalized) of the i-th distribution transformer.
[0111] Embodiment 3
[0112] This embodiment further defines how to calculate the Hamming distance, specifically:
[0113] The calculating the Hamming distance between the first comparison string and the second comparison string comprises:
[0114] Performing word segmentation operations on the first comparison string and the second comparison string respectively to obtain corresponding first keyword sets and second keyword sets;
[0115] The word segmentation server can be used to compare and segment the string and extract all keywords;
[0116] Steps S31-S34 are performed on the first keyword set and the second keyword set respectively to obtain corresponding first dimensionality reduction sequence strings and second dimensionality reduction sequence strings:
[0117] S31, mapping each keyword in the keyword set to a corresponding hash code according to the sample library;
[0118] The sample library stores various keywords and their corresponding hash codes. For each keyword to be mapped, the corresponding keyword is retrieved in the sample library through searching, and then matched to its corresponding hash code. For example, it can be mapped to a six-bit hash code 1 0 0 1 0 0, 1 0 000 1, etc.
[0119] S32, weighting the hash code corresponding to each keyword according to the weight of the keyword;
[0120] A weight may be attached to each keyword in the sample library according to the distribution of the keywords, and then the corresponding hash code may be weighted based on the weight corresponding to the keyword to form a character string. In an optional implementation, the hash code and 1 may be subjected to a bitwise operation. If the bit is 1, the bit is weighted according to the weight of the corresponding keyword. If the bit is not 1, the bit is downgraded according to the weight of the corresponding keyword. For example, for the hash code in the above example, the weight of the first corresponding keyword is 2, and the weight of the second corresponding keyword is 4. After weighting, the weights are: 2 -2 -2 2 -2 -2, 4 -4 -4 -4 -4 -4 4;
[0121] S33, accumulating and merging each weighted hash sequence in the keyword set to form a sequence string corresponding to the keyword set;
[0122] After weighting, the weighted hash codes corresponding to all keywords are accumulated and merged to form a sequence string. For example, after the comparison string is segmented, a total of 20 hash codes are obtained. Then, these 20 weighted hash codes are accumulated and merged, and finally the result is: -26 35 28 -31 22 19;
[0123] S34, performing a dimensionality reduction operation on the sequence string to obtain a dimensionality reduction sequence string corresponding to the keyword set;
[0124] Traverse the merged result and do the same bit comparison. If the bit is greater than 0, record it as 1; if the bit is less than 0, record it as 0, such as: 0 1 1 0 1 1;
[0125] Calculating the Hamming distance between the first comparison string and the second comparison string according to the first reduced dimension sequence string and the second reduced dimension sequence string;
[0126] The first reduced-dimensional sequence string and the second reduced-dimensional sequence string can be XOR-ed bitwise compared to obtain the Hamming distance;
[0127] In an optional implementation, considering the comprehensive performance of time and space, the 64-bit hash code of the sample library text can be split into 4 segments. The hashcode is 64 bits and divided into 4 segments in order. Each segment has 16 bits and is combined and stored. An exact match can be performed before calculating the Hamming distance, which can greatly improve the calculation efficiency.
[0128] Embodiment 4
[0129] Please refer to Figure 2A distribution network data comparison terminal includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, each step of the distribution network data comparison method of any one of the above-mentioned embodiments 1 to 3 is implemented.
[0130] In summary, the present invention provides a method and terminal for comparing distribution network data. When comparing distribution network data, the distribution network sample data is first clustered and selected by the K-means algorithm to obtain a feature data set, and then the first data and the second data to be compared are marked based on the feature data set to determine the feature classifications corresponding to the first data and the second data respectively. Then, the data with the same feature classification and the same field in the first data and the second data are spliced to form a long character string. Finally, the Hamming distance algorithm is used to perform distance calculation on the long character string of the first data and the second data to determine the similarity of the first data and the second data. First, the feature classification corresponding to each data can be accurately determined by cluster analysis, so that the long character string belonging to the same category can be accurately spliced. Then, the Hamming distance algorithm is used to perform distance calculation on the long character string to realize data comparison. By combining the feature classification result with the Hamming distance algorithm, similar data can be matched to the maximum extent, and the data intersection of different data systems can be determined, which greatly improves the matching accuracy of distribution network data, can solve many business pain points caused by different sources of business data, and plays an important technical support for business planning improvement and auxiliary decision-making, with high accuracy and business applicability.
[0131] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent transformations made using the contents of the present invention's specification and drawings, or directly or indirectly applied in related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for comparing distribution network data. It is characterized in that Includes steps: S1. Perform cluster analysis on the distribution network sample data to determine a feature data set corresponding to the distribution network sample data; S2. Mark the first data and the second data to be compared respectively according to the feature data set to determine a first feature classification corresponding to the first data and a second feature classification corresponding to the second data: Comparing the first data and the second data to be compared with the feature data set respectively, and determining a first similarity set between the first data and the feature data set and a second similarity set between the second data and the feature data set respectively; determining a first feature classification of the first data according to the feature data corresponding to the highest similarity in the first similarity set, and determining a second feature classification of the second data according to the feature data corresponding to the highest similarity in the second similarity set; S3, according to the feature classification result, concatenate the data with the same feature classification in the first data and the second data by taking the same fields, respectively, to obtain a first comparison string corresponding to the first data and a second comparison string corresponding to the second data; S4, calculating a Hamming distance between the first comparison string and the second comparison string, and determining a similarity between the first data and the second data according to the Hamming distance; The calculating the Hamming distance between the first comparison string and the second comparison string comprises: Performing word segmentation operations on the first comparison string and the second comparison string respectively to obtain corresponding first keyword set and second keyword set; Steps S31-S34 are performed on the first keyword set and the second keyword set respectively to obtain corresponding first dimensionality reduction sequence strings and second dimensionality reduction sequence strings: S31, mapping each keyword in the keyword set to a corresponding hash code according to the sample library; S32, weighting the hash code corresponding to each keyword according to the weight of the keyword; S33, accumulating and merging each weighted hash sequence in the keyword set to form a sequence string corresponding to the keyword set; S34, performing a dimensionality reduction operation on the sequence string to obtain a dimensionality reduction sequence string corresponding to the keyword set; The Hamming distance between the first comparison character string and the second comparison character string is calculated according to the first reduced dimension sequence string and the second reduced dimension sequence string.
2. A method for comparing distribution network data according to claim 1, It is characterized in that The step S1 comprises: Normalize the distribution network sample data to obtain the feature vector set for cluster analysis; Using K-means clustering algorithm to perform cluster analysis on the feature vector set to obtain a clustering result; According to the clustering result, a feature vector set corresponding to the clustering result is selected to obtain feature data corresponding to each cluster, and a feature data set corresponding to the distribution network sample data is determined according to the feature data corresponding to each cluster.
3. A method for comparing distribution network data according to claim 1 or 2, It is characterized in that Determining the similarity between the first data and the second data includes: It is determined whether the Hamming distance is less than a preset value. If so, the first data is similar to the second data. Otherwise, the first data is not similar to the second data.
4. A distribution network data comparison terminal, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, It is characterized in that When the processor executes the computer program, the following steps are implemented: S1. Perform cluster analysis on the distribution network sample data to determine a feature data set corresponding to the distribution network sample data; S2. Mark the first data and the second data to be compared respectively according to the feature data set to determine a first feature classification corresponding to the first data and a second feature classification corresponding to the second data: Comparing the first data and the second data to be compared with the feature data set respectively, and determining a first similarity set between the first data and the feature data set and a second similarity set between the second data and the feature data set respectively; determining a first feature classification of the first data according to the feature data corresponding to the highest similarity in the first similarity set, and determining a second feature classification of the second data according to the feature data corresponding to the highest similarity in the second similarity set; S3, according to the feature classification result, concatenate the data with the same feature classification in the first data and the second data by taking the same fields, respectively, to obtain a first comparison string corresponding to the first data and a second comparison string corresponding to the second data; S4, calculating a Hamming distance between the first comparison string and the second comparison string, and determining a similarity between the first data and the second data according to the Hamming distance; The calculating the Hamming distance between the first comparison string and the second comparison string comprises: Performing word segmentation operations on the first comparison string and the second comparison string respectively to obtain corresponding first keyword set and second keyword set; Steps S31-S34 are performed on the first keyword set and the second keyword set respectively to obtain corresponding first dimensionality reduction sequence strings and second dimensionality reduction sequence strings: S31, mapping each keyword in the keyword set to a corresponding hash code according to the sample library; S32, weighting the hash code corresponding to each keyword according to the weight of the keyword; S33, accumulating and merging each weighted hash sequence in the keyword set to form a sequence string corresponding to the keyword set; S34, performing a dimensionality reduction operation on the sequence string to obtain a dimensionality reduction sequence string corresponding to the keyword set; The Hamming distance between the first comparison character string and the second comparison character string is calculated according to the first reduced dimension sequence string and the second reduced dimension sequence string.
5. A distribution network data comparison terminal according to claim 4, It is characterized in that The step S1 comprises: Normalize the distribution network sample data to obtain the feature vector set for cluster analysis; Using K-means clustering algorithm to perform cluster analysis on the feature vector set to obtain a clustering result; According to the clustering result, a feature vector set corresponding to the clustering result is selected to obtain feature data corresponding to each cluster, and a feature data set corresponding to the distribution network sample data is determined according to the feature data corresponding to each cluster.
6. A distribution network data comparison terminal according to claim 4 or 5, It is characterized in that Determining the similarity between the first data and the second data includes: It is determined whether the Hamming distance is less than a preset value. If so, the first data is similar to the second data. Otherwise, the first data is not similar to the second data.
Citation Information
Patent Citations
Text similarity comparison method and device for all-media review
CN113407693A
Power supply and distribution method and system for data center
CN113839389A