An IDMapping method based on factor analysis and graph clustering
Through the IDMapping method of factor analysis and graph clustering, the problem of insufficient computing performance of massive fragmented data is solved, and more efficient and accurate user portrait formation is achieved.
Patent Information
- Application Number
- CN202111269662.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2041-10-29
AI Technical Summary
When the existing IDMapping technology faces massive fragmented data, its computing performance is insufficient, which makes it difficult for data accuracy and computing power to meet the needs, and multiple user data are mixed together due to inaccurate data sources.
Using a method based on factor analysis and graph clustering, data preprocessing, weight calculation and graph clustering are used to perform distributed calculations using the SparkGraphX framework, and a user portrait is formed in combination with the connected sub-graph algorithm.
It improves the accuracy and computing efficiency of data merging, ensures that each user's data portrait is more accurate, and through the combination of weight calculation and graph clustering algorithm, more efficient data association and merging are achieved.
Smart Images

Figure CN114077865B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data fusion, and particularly relates to an IDMapping method based on factor analysis and graph clustering. Background Art
[0002] With the development of technologies such as mobile Internet, Internet of Things, and cloud computing, various APPs have emerged as the times require. Along with the improvement of the quality of life, people can use these APPs on terminals such as mobile phones, pads, and laptops at the same time, generating huge amounts of data while satisfying their spiritual life. These data include numerous registration information, buried point data, associated information, etc. However, the data generated by each terminal and each APP exists in a fragmented manner, and there is an association between them. By using the IDMapping technology, they can be associated through keys to form a complete data, and finally form a user portrait.
[0003] There are some implementation methods of IDMapping in the prior art:
[0004] 1. Adopting a kv database to store data through kv key-value pairs. If the key already exists, the existing unified identifier is used for integration.
[0005] 2. Aggregating each key through a distributed computing platform (such as MapReduce) and gradually merging them into one piece of data.
[0006] Nowadays, with the continuous expansion of data sources and the rapid growth of data volume, the previous computing performance can no longer meet the daily needs. Various data sources, the time when data is generated, etc. will all affect the accuracy of data. There are many problems with the data obtained by most IDMapping technologies, such as multiple users being mixed together due to inaccurate certain data sources and many other problems. At this time, a more efficient data merging method that combines the use of weights is needed to improve the computing power and data accuracy. Summary of the Invention
[0007] Aiming at the deficiency that the existing computing performance is difficult to meet due to the rapid growth of data volume, the present invention provides an IDMapping method based on factor analysis and graph clustering, which merges massive fragmented data from various sources, improves the data quality, and finally forms a piece of user portrait data.
[0008] Aiming at the deficiencies of the traditional fusion method in terms of accuracy and computing power, the present invention adopts the following technical solutions: an IDMapping method based on factor analysis and graph clustering, including the following steps:
[0009] S1, data preprocessing:
[0010] (1)Based on the data obtained from each data source, extract the pairwise relationships of the data, sort and number the data with different attributes in each pair of relationship data according to the relationship starting point and the relationship ending point, and obtain the data collection times count, data collection time ctime, data source domain, and data source reliability rel of each pair of relationship data;
[0011] (2)Select the time span Tsapn, collection time T, collection times N, data source reliability REL, and the number of data source types TYPE as the feature dimensions of the data, and perform normalization processing on each feature dimension according to Equations 1-5;
[0012]
[0013] In the formula, x is the difference between the earliest collection time and the latest collection time in a set of the same relationship data, ctime j is the data collection time of the j-th relationship in this set of the same relationship data, m is the total number of data in this set of the same relationship data; x i is the difference between the earliest collection time and the latest collection time in the i-th set of the same relationship data, n is the total number of sets of the same relationship data; Tsapn i is the normalized value of the difference between the earliest collection time and the latest collection time of the same relationship data; j represents the serial number of a certain collected relationship data in a set of the same relationship data, m is the total number of relationship data in this set of the same relationship data, i represents the serial number of a certain set of the same relationship data, and n is the total number of sets of the same relationship data;
[0014]
[0015] In the formula, T is the normalized value of the number of days from the latest collection time of the same relationship data to the current time, now is the current time, day is the number of days from the latest collection time of the same relationship data to the current time, day() is the function of converting time to days, day i is the value of the collection time of the i-th relationship data;
[0016]
[0017] In the formula, N is the normalized value of the total number of collections of the same relationship data from different sources and at different times, c is the number of the same data generated by collecting from different sources and at different times for a certain same relationship, c i is the collection times of the i-th relationship data;
[0018]
[0019] In the formula, REL is the normalized value of the data source reliability, r is the reliability score of a certain relationship data, relj is the credibility of the jth relational data source (judged and evaluated by domain experts), rel∈{0.1,0.5,1}, k is the number of source credibility scores (e.g., {0.1,0.5,1} is 3), C l is the number of credibility scores of the lth source, r is the reliability score of the same relationship,
[0020]
[0021] In the formula, TYPE is the normalized value of the number of data source types, y is the number of different sources of the same relational data, and y i is the number of times the i-th relation data is collected;
[0022] (3) Remove abnormal data nodes;
[0023] (4) Sort the starting point of the relationship data by NO Star and relationship endpoint sort number NO End , and the five feature dimensions of time span Tsapn, collection time T, collection times N, data source reliability REL and data source type TYPE are obtained after preprocessing and normalization for data output. The output data format is as follows: {NO Star NO End Tspan TN RELTYPE};
[0024] Step S2, weight calculation:
[0025] (5) Based on the KMO test statistic, the reliability weight scores of the above five feature dimensions are calculated. The KMO calculation formula is as follows:
[0026]
[0027] Where X and Y are the vectors of the five feature dimensions mentioned above, r XY is the Pearson correlation coefficient between X and Y, α XY is the partial correlation coefficient between X and Y;
[0028] (6) After the factor analysis passes the test, the contribution rate of all eigenvalues of each pairwise relationship data is calculated. First, the 5×5 covariance matrix cov of the sample is calculated:
[0029]
[0030] Where X is the vector of the five characteristic dimensions of the relational data, T is the transpose of the X vector, and D is the number of dimensions;
[0031] Then, the eigenvalue λ and eigenvalue contribution rate f are calculated by the following formula: i
[0032]
[0033] Wherein, A is the matrix of the result of formula (8), E is the identity matrix, λ is the eigenvalue matrix, the subscript d is the number of dimensions, and di is the number of summation dimensions;
[0034] (7) Finally, the weight w of each feature dimension in the final pairwise relationship data is calculated through the following formula as the output data,
[0035]
[0036] Wherein, f d is the contribution rate value of the d-th feature dimension, and y d is the value of the d-th dimension;
[0037] Step S3, perform graph clustering processing on the data:
[0038] (8) Use SparkGraphX to create point objects (EdgeRDD) and edge objects (VertexRDD) for the output data in step (7), thereby generating a graph structure object (Graph object graph);
[0039] (9) Use the connected subgraph algorithm to split the generated Graph object graph to obtain several interconnected subgraphs. The obtained subgraphs are the portraits of each user, and the smallest ID value among all nodes in the subgraph is set as the unique key (OneID) of the subgraph,
[0040] where each of the subgraphs is the data of one user or several users with conflicts.
[0041] Further, when the subgraph in step (9) is the data of several users with conflicts (referring to the situation where multiple unique single-user attributes are related to a non-single-user attribute at the same time, such as telephone, QQ, etc.), it is processed in the following manner:
[0042] a) When multiple single-user attributes are related to one attribute, retain the relationship with the largest weight and discard other relationships;
[0043] b) When one attribute is associated with more than 2 identical attributes, retain the first N associated data with larger weight results;
[0044] c) When multiple single-user attributes appear in a subgraph and are indirectly associated through other attributes, take the shortest path from these attributes to each single-user attribute's indirect association, and use the relationship with the largest weight of the shortest path as the attribute of the user;
[0045] d) In the above three cases, if the weights are the same, no processing is performed temporarily. Processing will be carried out when subsequent data incremental updates cause changes in the weights.
[0046] Further, the unique single-user attributes include ID card number, user ID, and IMSI.
[0047] Further, in step (1), for each pair of relationship data, different attribute data are numbered in the order of the first letters of the attribute names. The attribute with the first letter earlier is set as the relationship starting point, and the attribute with the first letter later is set as the relationship ending point.
[0048] Further, in step (3), the abnormal data nodes include the situation where the number of attributes in the subgraph exceeds the warning quantity N, and the warning quantity N is set artificially.
[0049] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0050] 1. Analyze some features affecting data accuracy for the associated data, and through factor analysis, determine whether these features meet the requirements, and finally calculate the weights of the data.
[0051] 2. For fragmented data from various sources, use the distributed graph computing framework SparkGraphX for data association, and use the connected subgraph algorithm to form IDMapping. In this way, through the graph clustering algorithm in cooperation with the distributed computing framework, IDMapping can be carried out more conveniently and quickly.
[0052] 3. For the data of each subgraph, further segment it using its own business to make each piece of data more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is the flow chart of the IDMapping method based on factor analysis and graph clustering according to the present invention;
[0054] Figure 2 is the schematic diagram of the connected subgraph according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] The following further clarifies the present invention with reference to the drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent modifications made by those skilled in the art fall within the scope defined by the appended claims of this application.
[0056] The IDMapping method based on factor analysis and graph clustering according to the present invention mainly includes data preprocessing, weight calculation, and graph clustering in specific steps. The step process is as follows:
[0057] Step S1. Data preprocessing:
[0058] Based on various data sources, we extract pairwise relationships from the data, such as the relationship between four codes (TEL, IMEI, IMSI, and MAC), and the relationship between TEL and accounts like QQ and ID cards. We assign a distinguishing number to each attribute, such as "01" for phone and "02" for QQ. We then arrange the relationships alphabetically. For example, in the relationship between TEL and QQ, the first letter of QQ comes before the first letter of TEL, so QQ is placed first, marking the starting point of the relationship, and TEL is placed second, marking the end point. We also obtain the number of data collections, the collection time, and the data source for each pair of relationships.
[0059] Identify the characteristic dimensions that influence your data, merge data with the same relationship, calculate the value of each characteristic dimension, and normalize each characteristic dimension. In our solution, the characteristic dimensions include time span (the difference between the earliest and latest collection times), collection time (the number of days between the latest collection time and the current time), collection times (the number of identical data items generated from different sources or at different times), data source reliability (judgment by domain experts), and data source type.
[0060] For the above five features, different functions are selected to normalize the original value of each feature to between [0, 1].
[0061] F1: The normalized formula of time span Tsapn is:
[0062]
[0063] In the formula, x is the difference between the earliest collection time and the latest collection time in a set of the same relational data, ctime j is the data collection time of the jth relation in the set of identical relation data, and m is the total number of data in the set of identical relation data; i Tsapn is the difference between the earliest collection time and the latest collection time in the i-th group of identical relationship data, and n is the total number of identical relationship data groups; i It is the normalized value of the difference between the earliest collection time and the latest collection time of the same relational data;
[0064] j represents the sequence number of a collected relational data in a group of identical relational data, m represents the total number of relational data in the group of identical relational data, i represents the sequence number of a group of identical relational data, and n represents the total number of identical relational data groups;
[0065] F2: Acquisition time T, the step-by-step normalization formula divided by period is:
[0066]
[0067] In the formula, T is the normalized value of the number of days from the latest collection time of the same relationship data to the current time, now is the current time, day is the number of days from the latest collection time of the same relationship data to the current time, day() is the function to convert time to days, and day i is the value of the collection time of the i-th relationship data.
[0068] The data updated within 7 days is defaulted to have a very good update degree, which is set to 1; for the data that is more than 7 days and less than 90 days, a decreasing function is used for calculation. When the update time is longer, its value will become smaller and smaller; for the data that has not been updated for more than 90 days, its update degree is very poor and is defaulted to 0.1.
[0069] F3: The number of collections N, and the normalization formula is:
[0070]
[0071] In the formula, N is the normalized value of the total number of collections of the same relationship data from different sources and at different times, c is the number of the same data generated by different sources and at different times of a certain same relationship, and c i is the number of collections of the i-th relationship data.
[0072] F4: The reliability of the data source REL i , and the normalization formula is:
[0073]
[0074] In the formula, REL is the normalized value of the reliability of the data source, r is the reliability score of a certain relationship data, and rel j is the credibility of the j-th relationship data source (judged and evaluated by domain experts), rel ∈ {0.1, 0.5, 1}, k is the number of credibility score values (such as {0.1, 0.5, 1} is 3), and C l is the number of credibility score values of the l-th source, and r is the reliability score of the same relationship.
[0075] F5: The type of data source TYPE, and the normalization formula is:
[0076]
[0077] In the formula, TYPE is the normalized value of the number of types of data sources, y is the number of different sources of the same relationship data, and y i is the number of collections of the i-th relationship data.
[0078] Next, remove abnormal data nodes. For example, if a data node has N (which can be set, default is 50) associated relationships and there are too many associated data, it is regarded as an abnormal node, and this node and its associated relationships need to be deleted.
[0079] Finally, generate the preprocessed output data. The output data format is as follows: {NO Star NO End Tspan T NREL TYPE}, where the sorting number NO of the relationship starting point Star and the sorting number NO of the relationship ending point End , and the preprocessed time span Tsapn, collection time T, collection times N, data source reliability REL, and data source type TYPE are separated by spaces.
[0080] Step S2. Weight calculation:
[0081] a) In the first step, calculate the KMO of each type of relationship for the several eigenvalue obtained from data preprocessing.
[0082] b) Determine the influence of these 5 eigenvalue dimensions on the reliability of the data through the KMO test, and select the eigenvalue with a high KMO score for weight calculation. If the KMO test score is very low, reselect according to the business.
[0083] The KMO (Kaiser - Meyer - Olkin) test statistic test is a sampling suitability test. This test is to test the relative magnitudes of the simple correlation coefficients and partial correlation coefficients between the original variables. The calculation formula is:
[0084]
[0085] In the formula, X and Y are vectors of the above five eigenvalue dimensions, r XY is the Pearson correlation coefficient between X and Y, and α XY is the partial correlation coefficient between X and Y.
[0086] c) Through factor analysis, obtain the contribution rate of all eigenvalues for each pairwise relationship. First, calculate the covariance matrix of the sample, which is a 5×5 matrix cov:
[0087]
[0088] In the formula, X is the vector of 5 eigenvalue dimensions of the relationship data, T is the transpose of the X vector, and D is the number of dimensions;
[0089] Then, calculate the eigenvalue λ and the eigenvalue contribution rate f through the following formula i :
[0090]
[0091] In the formula, A is the matrix of the result of formula (8), E is the identity matrix, λ is the eigenvalue matrix, the subscript d is the number of dimensions, and di is the number of summation dimensions.
[0092] d) Calculate the weights of the final pairwise relationships
[0093] Finally, the weights w of each feature dimension in the final pairwise relationship data are calculated and obtained through the following formula as the output data.
[0094]
[0095] In the formula, f d is the contribution rate value of the d-th feature dimension, and y d is the value of the d-th dimension. Generate the weight calculation output data, and the data format is: relationship starting point value relationship ending point value weight value (separated by spaces). To ensure the freshness of the accuracy result, the weight calculation process is set to be updated once a month.
[0096] Step S3. Graph clustering
[0097] 3.1 Clustering
[0098] 1) Use SparkGraphX to create EdgeRDD and VertexRDD objects for the output result of the second step, and finally generate a Graph object.
[0099] 2) Use the connected subgraph algorithm to cut the large graph into several interconnected subgraphs. Each subgraph is the data of one user or several users with conflicts.
[0100] Subgraph conflict means that in the attribute nodes of a connected subgraph, due to reasons such as dirty data in the source data or associated relationship data, there are multiple single-user attributes (unique attributes that only appear in one person's data) in terms of business, or there are multiple associations in the one-to-one corresponding relationship.
[0101] For example, if there are multiple single-user attributes with uniqueness such as ID card and user ID in a subgraph, it is considered a conflict; if a mobile phone number belongs to multiple user ID numbers, it is considered a conflict; if a mobile phone number should only correspond to one IMSI within a certain period of time, if there is a one-to-many or many-to-one relationship, it is considered a conflict.
[0102] 3.2 Business segmentation
[0103] a) For the conflict data, we need to perform business segmentation on the subgraph. Since the scale of each subgraph is very small, the adjacency list structure can be used to handle it. With the help of Spark's concurrent processing, the computing power can be greatly improved.
[0104] b) Determine whether there is a sub-graph conflict situation. If there is no conflict, directly go to step h.
[0105] c) Determine whether there are multiple single-user attributes in the sub-graph connecting to another attribute at the same time. A single-user attribute can only exist in one user, such as an ID card, user ID, etc.
[0106] d) Identify the relationships that should be one-to-one in the actual situation but are shown as one-to-many. According to the weight results, retain the relationship with the largest weight in the one-to-many relationships. For example, if a mobile phone number has relationships with multiple IMSIs, select the relationship with the largest weight and remove the other relationships.
[0107] e) If an attribute is associated with more than N identical attributes, retain the top N associated data with larger weight results. For example, a person can have at most 3 to 4 mobile phone numbers within a certain period.
[0108] f) If the attributes representing single users are indirectly connected in the sub-graph, such as an ID card or user ID appearing multiple times in a sub-graph and being indirectly connected through other attributes, at this time, perform the shortest path algorithm on these attributes and the single-user attributes, and use the largest weight in the results as the attribute of this user.
[0109] g) If there is a conflict and the weights are the same, do not process it temporarily. Wait until the weights change during subsequent data incremental updates and then process it.
[0110] The finally obtained sub-graph is the portrait of each user. And set the smallest value among the unique ID values of all nodes in the sub-graph as the OneID of this sub-graph.
[0111] The following is a detailed description through a specific case:
[0112] For data with multiple data sources, these data are fragmented. Through IDMapping, they can be associated, and through factor analysis, weight values can be obtained, the relationship weights can be calculated, and an accurate user portrait can be segmented. Its overall process is as Figure 1 shown. This process is divided into 4 steps:
[0113] 1. Data preprocessing: According to multiple data sources, determine the types of pairwise relationships obtained, obtain the data collection time, and data sources. Aggregate and calculate the data with the same relationship type value to obtain the characteristic values of each piece of data: time span, collection time, collection times, reliability of data sources, types of data sources, and perform normalization. Number the corresponding relationships, as shown in Table 1. A sample of a preprocessed data for the TEL-QQ relationship is shown in Table 2:
[0114] Table 1 shows the correspondence between numbers and attributes
[0115]
[0116]
[0117] Table 2 shows the data after preprocessing
[0118]
[0119] 2. Weight calculation: Factor analysis is performed on all data of each relationship to obtain the KMO test value, and the eigenvalue and contribution rate are calculated, and finally the weight is obtained. Table 3 shows five sample feature information of the sample data of the TEL-QQ relationship
[0120] Table 3
[0121] Tspan T N REL TYPE 1.000000 0.279257 0.000000 0.302794 0.000000 0.904265 0.299435 0.083333 0.331307 0.000000 0.824645 0.112994 0.166667 0.256937 0.125000 1.000000 0.207022 0.000000 0.315402 0.000000 0.919431 0.279257 1.000000 0.217488 0.375000 1.000000 0.234463 0.000000 0.331307 0.000000 1.000000 0.260291 0.000000 0.331307 0.000000 0.640284 0.104923 0.166667 0.234745 0.250000 1.000000 0.207022 0.000000 0.315402 0.000000 1.000000 0.110169 0.000000 0.297392 0.000000 … … … … … ,
[0122] KMO measure: 0.8096926811363232. The standard for suitable factor analysis is to select the KMO score greater than 0.8. Therefore, the KMO measure of the TEL-QQ relationship meets the standard. The characteristic contribution rate is shown in Table 4
[0123] Table 4
[0124]
[0125]
[0126] Then, referring to Table 1, the weight is calculated as follows
[0127] 0.461538×0.840492 + 1.0×0.113064 + 0.000493×0.044652 + 0.908020×0.000300 + 0.0×0.001492 = 0.501278
[0128] The final output is shown in Table 5: Weight result
[0129] Table 5
[0130]
[0131] 3. Graph clustering
[0132] The input relationship weight data is shown in Table 6
[0133] Table 6 Relationship weight data
[0134]
[0135] Aggregate the relationships into a graph through SparkGraphX, and obtain each cluster through connected subgraphs, such as the connected subgraph shown Figure 2 as
[0136] 0212***56 conflicts with 033201***111 and 033201***112. Since the weight of 0212***56 and 033201***112 is 0.4, which is smaller than the weight of 0212***56 and 033201***111, the association between 0212***56 and 033201***112 is removed.
[0137] 04userid1 and 04userid2 have an indirect conflict. By comparing the weights of each point between them to their shortest paths, when reaching 05460***906, the indirect weight on the left of 068697***596 is smaller, so the association between 068697***596 and 05460***906 is removed. Set the IDMapping value for each subgraph, and the result is shown in Table 7:
[0138] Table 7 IDMapping Results
[0139]
[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An IDMapping method based on factor analysis and graph clustering, characterized by The following steps are involved: S1, data preprocessing: (1) Based on the data obtained from each data source, pairwise relationship extraction is performed on the data, and different attribute data in each pair of relationship data are sorted and numbered according to the relationship starting point and relationship end point, and the data collection count, data collection time ctime, data source domain, and data source reliability rel of each pair of relationship data are obtained; the pairwise relationship extracted data includes pairwise relationship data between four codes, or pairwise relationship data between a phone number and a QQ number and an ID card number respectively, and the four codes include phone number, IMEI, IMSI, and MAC; (2) Select the time span Tsapn, the collection time T, the number of collections N, the data source reliability REL, and the number of data source types TYPE as the characteristic dimensions of the data, and perform normalization processing on each characteristic dimension according to formulas 1 to 5; In the formula, x is the difference between the earliest collection time and the latest collection time in a set of the same relational data, ctime j is the data collection time of the jth relation in the set of identical relation data, and m is the total number of data in the set of identical relation data; i Tsapn is the difference between the earliest collection time and the latest collection time in the i-th group of identical relationship data, and n is the total number of identical relationship data groups; i It is the normalized value of the difference between the earliest collection time and the latest collection time of the same relationship data; j represents the sequence number of a collected relationship data in a group of the same relationship data, m is the total number of relationship data in the group of the same relationship data, i represents the sequence number of a group of the same relationship data, and n is the total number of the same relationship data groups; In the formula, T is the normalized value of the number of days between the latest collection time of the same relational data and the current time, now is the current time, day is the number of days between the latest collection time of the same relational data and the current time, Day() is the function of converting time to days, and day i is the value of the i-th relation data collection time; Where N is the normalized value of the total number of times the same relational data is collected from different sources at different times, c is the number of identical data items generated from different sources and collected at different times for the same relation, and c is the number of identical data items generated from different sources and collected at different times for the same relation. i is the number of times the i-th relation data is collected; In the formula, REL is the normalized value of data source reliability, r is the reliability score of a certain relational data, and rel j is the source credibility of the jth relation data, rel∈{0.1,0.5,1}, k is the number of source credibility scores, C l is the number of credibility scores of the lth source, r is the reliability score of the same relationship, e r is the exponential function of the reliability score r In the formula, TYPE is the normalized value of the number of data source types, y is the number of different sources of the same relational data, and y i is the number of times the i-th relation data is collected; (3) Remove abnormal data nodes; (4) Sort the starting point of the relationship data by NO Star and relationship endpoint sort number NO End , and the five feature dimensions of time span Tsapn, collection time T, collection times N, data source reliability REL and data source type TYPE are obtained after preprocessing and normalization for data output. The output data format is as follows: {NO Star NO End Tspan TN RELTYPE}; Step S2, weight calculation: (5) Based on the KMO test statistic, the reliability weight scores of the above five feature dimensions are calculated. The KMO calculation formula is as follows: Where X and Y are the vectors of the five feature dimensions mentioned above, r XY is the Pearson correlation coefficient between X and Y, α XY is the partial correlation coefficient between X and Y; (6) After the factor analysis passes the test, the contribution rate of all eigenvalues of each pairwise relationship data is calculated. First, the 5×5 covariance matrix cov of the sample is calculated: Where X is the vector of the five characteristic dimensions of the relational data, T is the transpose of the X vector, and D is the number of dimensions; Then, the eigenvalue λ and eigenvalue contribution rate f are calculated by the following formula: i Where A is the matrix of the result of formula (8), E is the identity matrix, λ is the eigenvalue matrix, the subscript d is the number of dimensions, and di is the number of summation dimensions; (7) Finally, the weight w of each feature dimension in the final pairwise relationship data is calculated as the output data by the following formula: Where, f d is the contribution rate value of the d-th feature dimension, y d is the value of the dth dimension; Step S3: Data is clustered: (8) Using SparkGraphX to create the point object EdgeRDD and edge object VertexRDD for the output data in step (7), thereby generating a graph structure object Graph object graph; (9) The generated Graph object graph is divided into several interconnected subgraphs by the connected subgraph algorithm. The obtained subgraph is the portrait of each user, and the smallest ID value among all nodes in the subgraph is set as the unique key OneID of the subgraph. Each of the subgraphs is data of one user or several conflicting users.
2. The IDMapping method based on factor analysis and graph clustering according to claim 1, characterized in that: When the subgraph in step (9) contains data of several users with conflicts, it is processed in the following manner: a) When multiple single-user attributes are related to one attribute, the relationship with the largest weight is retained and the other relationships are discarded; b) When an attribute is associated with more than two identical attributes, the first N associated data with the larger weight result are retained; c) When multiple single-user attributes appear in a subgraph and are indirectly related through other attributes, the shortest path from these attributes to each single-user attribute is taken, and the relationship with the largest weight on the shortest path is taken as the attribute of the user; d) If the weights are the same in the above three situations, they will not be processed for the time being and will be processed when the weights change due to subsequent incremental data updates.
3. The IDMapping method based on factor analysis and graph clustering according to claim 2, characterized in that: The single user attributes include identity card number, user ID or IMSI.
4. The IDMapping method based on factor analysis and graph clustering according to claim 1, characterized in that: In step (1), the different attribute data in each pair of relationship data are numbered in alphabetical order according to the first letters of the attribute names, the attribute with the first letter is set as the relationship starting point, and the attribute with the second letter is set as the relationship end point.
5. The IDMapping method based on factor analysis and graph clustering according to claim 1, characterized in that: The abnormal data node in step (3) includes a situation where the attributes in the subgraph have an association relationship exceeding the warning number N, and the warning number N is set manually.
Citation Information
Patent Citations
Method and system for organizing manufacturing process information
CN108133338A
Aggregating data from a plurality of data sources
US9105000B1