Data matching method and device based on federated learning
By calculating the matching degree of labels and data content between the first data set and the multiple second data sets, and determining the target data set as federated matching data data, the problem of poor data matching effect in vertical federated learning is solved, and the comprehensiveness and accuracy of the matching results are improved.
Patent Information
- Application Number
- CN202210108127.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-28
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-01-28
AI Technical Summary
In the process of vertical federated learning, the data matching effect in the prior art is poor, resulting in limited data volume, poor data quality and serious data homogeneity.
By calculating the label matching degree and data content matching degree between the first data set and the plurality of second data sets, a comprehensive matching degree is generated, thereby determining the target data set as the federated matching data of the first data set.
It significantly improves the comprehensiveness, accuracy and accuracy of vertical federal matching results, broadens the scope of data matching, enriches the feature dimensions of data, and improves the subsequent model effect.
Smart Images

Figure CN114492852B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of federated learning, and in particular, to a data matching method and device based on federated learning. Background Art
[0002] When conducting vertical federated learning training, the initiator and the data provider first need to align the data, and then complete the subsequent model training with the participation of the coordinator. Before that, it is necessary to first perform the matching of federated data. In the related matching technologies, the main matching strategies include local search, three-party recommendation, and fixed participants, etc. However, the data obtained by these methods has problems such as limited data volume, poor data quality, and serious data homogenization. Summary of the Invention
[0003] The present invention provides a data matching method and device based on federated learning, which is used to solve the defect of poor data matching effect in the process of vertical federated learning in the prior art and achieve high-quality data matching.
[0004] The present invention provides a data matching method based on federated learning, including:
[0005] Calculating the similarity between the first data label corresponding to the first data set and the multiple second data labels corresponding to the multiple second data sets respectively, and generating multiple label matching degrees;
[0006] Calculating the similarity between the first data feature set corresponding to the first data set and the multiple second data feature sets corresponding to the multiple second data sets respectively, and generating multiple data content matching degrees;
[0007] Based on the label matching degree and the data content matching degree between the first data set and the same second data set, determining a target data set from the multiple second data sets as the federated matching data of the first data set.
[0008] According to the data matching method based on federated learning provided by the present invention, the step of determining a target data set from the multiple second data sets as the federated matching data of the first data set based on the label matching degree and the data content matching degree between the first data set and the same second data set includes:
[0009] Generating a comprehensive matching degree based on the label matching degree and the data content matching degree between the first data set and the same second data set;
[0010] Sorting the comprehensive matching degrees in descending order, and determining the second data sets corresponding to the target number of the comprehensive matching degrees with the top rankings as the target data sets;
[0011] Alternatively, sort the comprehensive matching degrees in ascending order, and determine the target data set as the target number of the second data sets corresponding to the comprehensive matching degrees at the end of the sorting.
[0012] According to a data matching method based on federated learning provided by the present invention, generating a comprehensive matching degree based on the label matching degree and the data content matching degree between the first data set and the same second data set includes:
[0013] Obtain the target first weight value corresponding to the label matching degree, the target second weight value corresponding to the data content matching degree, and the target evaluation score;
[0014] Perform normalization processing on the label matching degree and the data content matching degree respectively to generate a normalized label matching degree and a normalized data content matching degree;
[0015] Generate the comprehensive matching degree based on the normalized label matching degree, the normalized data content matching degree, the target first weight value, the target second weight value, and the target evaluation score.
[0016] According to a data matching method based on federated learning provided by the present invention, calculating the similarity between the first data label corresponding to the first data set and the multiple second data labels corresponding to the multiple second data sets respectively to generate multiple label matching degrees includes:
[0017] Perform cosine similarity calculation on the first data label and the second data label to generate the label matching degree between the first data set and the second data set.
[0018] According to a data matching method based on federated learning provided by the present invention, calculating the similarity between the first data feature set corresponding to the first data set and the multiple second data feature sets corresponding to the multiple second data sets respectively to generate multiple data content matching degrees includes:
[0019] Perform private set intersection on the first data feature set and the second data feature set, and obtain the first ratio of the same data in the first data feature set and the second data feature set to the first data feature set;
[0020] Generate the data content matching degree between the first data feature set and the second data feature set based on the first ratio and the data quality information of the second data feature set.
[0021] A data matching method based on federated learning provided by the present invention, wherein the data quality information of the second data feature set includes at least one of: a second ratio of missing data in the second data feature set, a third ratio of abnormal data in the second data feature set, and a fourth ratio of incorrect data in the second data feature set.
[0022] The present invention also provides a data matching device based on federated learning, including:
[0023] A first processing module, configured to calculate the similarity between the first data label corresponding to the first data set and the multiple second data labels corresponding to the multiple second data sets respectively, and generate multiple label matching degrees;
[0024] A second processing module, configured to calculate the similarity between the first data feature set corresponding to the first data set and the multiple second data feature sets corresponding to the multiple second data sets respectively, and generate multiple data content matching degrees;
[0025] A third processing module, configured to determine a target data set from the multiple second data sets as the federated matching data of the first data set based on the label matching degree and the data content matching degree between the first data set and the same second data set.
[0026] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the data matching method based on federated learning as described in any one of the above.
[0027] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the data matching method based on federated learning as described in any one of the above.
[0028] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the data matching method based on federated learning as described in any one of the above.
[0029] The data matching method and device based on federated learning provided by the present invention aggregate the second data sets provided by multiple data providers, and calculate the matching degree between the second data sets provided by the data providers and the first data set of the initiator from multiple dimensions, so as to determine the target data set with the highest matching degree from the multiple second data sets as the federated matching data of the first data set, greatly broadening the scope of data matching, enriching the feature dimensions of the data, significantly improving the comprehensiveness, accuracy and precision of the vertical federated matching result, and helping to improve the subsequent model effect. Description of the Drawings
[0030] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0031] Figure 1 It is one of the schematic flowcharts of the data matching method based on federated learning provided by the present invention;
[0032] Figure 2 It is another schematic flowchart of the data matching method based on federated learning provided by the present invention;
[0033] Figure 3 It is the third schematic flowchart of the data matching method based on federated learning provided by the present invention;
[0034] Figure 4 It is the fourth schematic flowchart of the data matching method based on federated learning provided by the present invention;
[0035] Figure 5 It is the fifth schematic flowchart of the data matching method based on federated learning provided by the present invention;
[0036] Figure 6 It is the schematic structural diagram of the data matching device based on federated learning provided by the present invention;
[0037] Figure 7 It is the schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0038] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0039] The inventors found during the R & D process that in the related art, when performing vertical federated multi-party data matching, the following several methods are mainly used:
[0040] (1) Local search. This method refers to the initiator of the federated data selecting data providing objects from the existing organizations with business interactions. Under this matching method, the scale of the federation is limited by the number of docking services. Especially when the number of its docking business parties is limited, the shortage of data providers will lead to the inability to carry out joint training.
[0041] (2) Third-party recommendation. This method is to select a suitable data provider through a trusted third party based on the initiator's requirements and expert experience. However, affected by the subjectivity of expert experience and the credibility of the third-party organization, it will affect the data quality and ultimately the model effect.
[0042] (3) Fixed participants. This method refers to specifying the participants of the federation through a superior organization to achieve modeling and joint training in specific scenarios. The federation built under this mode is limited to fixed scenarios and the mode is not general. Moreover, the participants in this mode of federation are very likely to be similar or overlapping in business, and their data is highly homogenized. In the vertical federation scenario, more feature expansion cannot be done, which is insufficient to support model training.
[0043] Most of the above methods are to qualitatively find suitable data providers through oneself or a third-party organization, resulting in high communication costs and low efficiency in the matching process, and fewer data providers that can be found, and poor dataset matching. Secondly, the subjective and qualitative methods cannot accurately describe the data quality and matching situation of the participants, resulting in large fluctuations in the data quality of the federated joint modeling, thus affecting the final model effect.
[0044] The following combines Figures 1 to 5 to describe the data matching method based on federated learning of the present invention.
[0045] It should be noted that the execution subject of the data matching method based on federated learning can be a data matching device based on federated learning, or a data matching system for vertical federation, or a server, or a user's terminal, where the user's terminal includes non-mobile terminals such as desktop computers, and mobile terminals such as mobile phones, tablets, watches, and learning machines.
[0046] As Figure 1 shown, the data matching method based on federated learning includes: Step 110, Step 120, and Step 130.
[0047] Step 110: Calculate the similarity between the first data label corresponding to the first dataset and the multiple second data labels corresponding to the multiple second datasets respectively, and generate multiple label matching degrees.
[0048] In this step, it should be noted that in vertical federated learning, there are generally three roles: initiator, data provider, and coordinator. Among them, the initiator refers to the creator of the federated learning task and the participant who starts and executes the task; the data provider refers to the participant who provides the private data required for the federated learning task; the coordinator refers to the participant who manages and coordinates the configuration and execution of the federated learning task.
[0049] In the actual execution process, each institution participating in the federation needs to report its dataset metadata information to the central node. The dataset metadata information includes: dataset name, data labels, the missing value ratio of each feature value, and feature values, etc. The central node will save and record this metadata information to provide a reference for the basic information for subsequent federated node pairing.
[0050] Among them, the federated node is the node corresponding to the data provider, and each data provider corresponds to a node.
[0051] For example, in the actual execution process, the method provided by the embodiments of the present invention can be executed by setting up a vertical federated data matching system for federated data matching and federated training.
[0052] The vertical federated data matching system can be an integrated platform that aggregates multiple participants, and this platform can provide corresponding data evaluation and matching algorithms.
[0053] Such as Figure 2 As shown, among the four institutions A, B, C, and D, institution A is the initiator, and institutions B, C, and D are data providers. Among them, institutions B, C, and D have been pre-registered in this system. Institution B corresponds to the first node, institution C corresponds to the second node, and institution D corresponds to the third node.
[0054] The dataset of institution A includes: the first sub-dataset A1, the first sub-dataset A2, etc. Among them, the first sub-dataset A1 includes data labels L1, L2, and the first data feature set used to characterize the data features of the first sub-dataset A1 itself. The first sub-dataset A2 includes data labels L3, L4, and the first data feature set used to characterize the data features of the first sub-dataset A2 itself.
[0055] The dataset of institution B includes: the second sub-dataset B1, the second sub-dataset B2, etc. Among them, the second sub-dataset B1 includes data labels L1, L3, and the second data feature set used to characterize the data features of the second sub-dataset B1 itself. The second sub-dataset B2 includes data label L3 and the second data feature set used to characterize the data features of the second sub-dataset B2 itself.
[0056] The dataset of institution C includes: the second sub-dataset C1, the second sub-dataset C2. Among them, the second sub-dataset C1 includes data labels L2, L3, and the second data feature set used to characterize the data features of the second sub-dataset C1 itself. The second sub-dataset C2 includes data label L3 and the second data feature set used to characterize the data features of the second sub-dataset C2 itself.
[0057] The data set of institution D includes: a second sub - data set D1 and a second sub - data set D2; among them, the second sub - data set D1 includes a data label L1, a data label L5, and a second data feature set for characterizing the data features of the second sub - data set D1 itself, and the second sub - data set D2 includes a data label L2 and a second data feature set for characterizing the data features of the second sub - data set D2 itself.
[0058] In this embodiment, the first data set is the data set provided by the initiator that needs to be matched, such as the first data set corresponding to institution A.
[0059] The first data set may include one or more first sub - data sets, such as the first sub - data set A1 and the first sub - data set A2, etc.
[0060] The first sub - data set includes: at least one first data label and a first data feature set. The first data label is a data label that the expected data provider can provide, and the first data feature set is used to characterize the data features in the first data set.
[0061] The second data set is the data set provided by the data provider for matching, such as the second data set corresponding to institution B, the second data set corresponding to institution C, and the second data set corresponding to institution D.
[0062] It can be understood that the data provider can be one or more. Each data provider corresponds to a second data set, and each second data set can include multiple second sub - data sets. For example, the data set of institution B includes: the second sub - data set B1 and the second sub - data set B2, etc.
[0063] The second sub - data set includes: at least one second data label and a second data feature set. The second data label is a data label provided by the data provider corresponding to the second sub - data set, and the second data feature set is used to characterize the data features in the second sub - data set.
[0064] For example, the first data label in the first data set can be the user identification of institution A, and the first data feature set can be the set of user features corresponding to the users of institution A, such as user historical purchase data, etc.; the second data label in the second data set can be the user identification of institution B, and the second data feature set can be the set of user features corresponding to the users of institution B, such as user exercise data, etc.
[0065] The label matching degree is used to characterize the similarity between the first data label corresponding to the first sub - data set and the second data label corresponding to the second data set.
[0066] For example, continue to refer to Figure 2, Institution A and Institution B are two institutions in the same region. The user groups of the two have a large intersection. That is, the overlap degree of the first data tags corresponding to Institution A and the second data tags corresponding to Institution B is relatively high. Then it is considered that the similarity degree between the first data tags of the first data set corresponding to Institution A and the second data tags of the second data set corresponding to Institution B is relatively high.
[0067] For another example, Institution A and Institution D are two institutions at home and abroad. Due to geographical restrictions, the user groups of Institution A and Institution D have a very small intersection. That is, the overlap degree of the first data tags corresponding to Institution A and the second data tags corresponding to Institution D is relatively low. Then it is considered that the similarity degree between the first data tags of the first data set corresponding to Institution A and the second data tags of the second data set corresponding to Institution D is relatively low.
[0068] In this step, calculate the similarity between the first data tag and each second data tag corresponding to the second data set respectively, and multiple tag matching degrees can be generated.
[0069] In the actual execution process, as Figure 3 shown, the initiator (such as Institution A) submits a data matching application to the central node of the system. The application information includes the first data set. Among them, the first data set includes information such as the first data tag of the expected data provider and the local first data feature set. The central node receives the matching application and processes it.
[0070] Among them, the first data set may include multiple first sub-data sets, such as the first sub-data set A1 and the first sub-data set A2, and multiple first data tags corresponding to each first sub-data set.
[0071] The central node calculates the matching degree according to the registered node situation in the current system. Among them, each node corresponds to a data provider, such as the nodes corresponding to Institution B, Institution C, and Institution D;
[0072] Each data provider corresponds to a second data set. Each second data set may include multiple second sub-data sets and multiple second data tags corresponding to the second sub-data sets; for example, the second data set corresponding to the node corresponding to Institution B includes: the second sub-data set B1 and the second sub-data set B2.
[0073] Then the central node needs to calculate the similarity between each second sub-data set at each node and the first sub-data set.
[0074] For example, similarity calculation methods such as the cosine similarity algorithm, the Euclidean distance similarity algorithm, and the Pearson correlation coefficient similarity algorithm can be used to calculate the similarity, so as to generate the label matching degree between the first data label corresponding to each first sub-dataset and the second data label corresponding to each second sub-dataset at each node.
[0075] By the same method, the label matching degree between the first data label corresponding to each first sub-dataset and the second data label corresponding to each second sub-dataset at each node can be calculated, so as to obtain multiple label matching degrees.
[0076] For example, generate the label matching degree between the first sub-dataset A1 and the second sub-dataset B1, generate the label matching degree between the first sub-dataset A1 and the second sub-dataset B2, generate the label matching degree between the first sub-dataset A1 and the second sub-dataset C1, etc.
[0077] Taking the cosine similarity algorithm as an example below, the specific implementation manner of step 110 will be described.
[0078] In some embodiments, step 110 may include: calculating the cosine similarity between the first data label and the second data label to generate the label matching degree between the first dataset and the second dataset.
[0079] In this embodiment, continuing with the above-mentioned data matching system for vertical federation as an example for illustration.
[0080] For example, the system has m built-in industries, and each industry has n labels. When the initiator submits a matching application, it needs to check several labels of several industries. The system sets the checked labels to 1 and the unchecked labels to 0 to generate an m*n-dimensional matrix, which is the matrix of the first data label corresponding to the first sub-dataset in the first dataset.
[0081] For example, the matrix of the first data label can be shown as:
[0082]
[0083] Among them, X is the matrix of the first data label, 1 represents the data label included in the first sub-dataset, and 0 represents the data label not included in the first sub-dataset.
[0084] Using the same method, the second dataset already registered in the system can be encoded, and each registered second sub-dataset can obtain an m*n-dimensional matrix, which is the matrix of the second data label corresponding to the second sub-dataset in the second dataset.
[0085] Similarly, the matrix of the second data label can be expressed as
[0086]
[0087] where Y is the matrix of the second data label, 1 indicates the data label included in the second subset of data, and 0 indicates the data label not included in the second subset of data.
[0088] It should be noted that the matrix Y of the second data label can be pre-generated.
[0089] After obtaining the matrix X of the first data label and the matrix Y of the second data label, by processing the matrix X of the first data label and the matrix Y of the second data label based on the cosine similarity algorithm, the label matching degree between the first data label corresponding to the first subset of data and the second data label corresponding to the second subset of data can be obtained.
[0090] For example, the formula:
[0091]
[0092] can be used to generate the label matching degree, where ls is the label matching degree between the first data label corresponding to the first subset of data and the second data label corresponding to the second subset of data, X i is the i-th value in matrix X, Y i is the i-th value in matrix Y, m is the number of industries, and n is the number of labels corresponding to each industry.
[0093] It can be understood that the closer the value of ls is to 1, the higher the similarity between the first data label corresponding to the first subset of data and the second data label corresponding to the second subset of data.
[0094] The cosine similarity calculation is respectively performed on the matrix X corresponding to the first subset of data and the matrix Y corresponding to each registered second subset of data to obtain the label matching degree between the first subset of data and each second subset of data at this node.
[0095] By adopting a similar algorithm for each first subset of data in the first data set, multiple label matching degrees between the first data label corresponding to the first data set and the multiple second data labels corresponding to the multiple second data sets can be generated.
[0096] According to the data matching method based on federated learning provided by the embodiments of the present invention, by calculating the label matching degree between the first data set and the second data set, the second data set corresponding to the appropriate node is screened for the initiator from the data label dimension, shortening the matching time and improving the matching efficiency.
[0097] Step 120: Calculate the similarity between the first data feature set corresponding to the first data set and the multiple second data feature sets corresponding to the multiple second data sets respectively, and generate multiple data content matching degrees;
[0098] In this step, the first data feature set is a set of data features of the first sub - data set in the first data set, and is used to characterize the data content of the first sub - data set.
[0099] It can be understood that each first sub - data set corresponds to a first data feature set.
[0100] The second data feature set is a set of data features of the second sub - data set in the second data set, and is used to characterize the data content of the second sub - data set.
[0101] It can be understood that each second sub - data set corresponds to a second data feature set.
[0102] The data content matching degree is used to characterize the similarity degree of the data features in the first sub - data set and the second sub - data set, that is, it is used to characterize the similarity of the data content of the first sub - data set and the second sub - data set.
[0103] For example, continuing to refer to Figure 2 , Institution A and Institution B are two institutions in the same region, and there is a large overlap in the user characteristics of Institution A and Institution B. Then it is considered that the similarity degree of the data features of the first data set corresponding to Institution A and the second data set corresponding to Institution B is relatively high.
[0104] Another example is that Institution A and Institution D are two institutions at home and abroad. Due to geographical restrictions, the intersection of the user groups of Institution A and Institution D is very small, and only a small part of the user characteristics overlap. Then it is considered that the similarity degree of the data features of the first data set corresponding to Institution A and the second data set corresponding to Institution D is relatively low.
[0105] By calculating the similarity between the first data feature set and the second data feature set, the data content matching degree can be generated.
[0106] The following describes the specific implementation manner of step 120.
[0107] In some embodiments, step 120 may include:
[0108] Perform a private set intersection on the first data feature set and the second data feature set, and obtain the first ratio of the same data in the first data feature set and the second data feature set to the first data feature set;
[0109] Based on the first ratio and the data quality information of the second data feature set, generate the data content matching degree between the first data feature set and the second data feature set.
[0110] In this embodiment, the first ratio is the ratio of the same data in the first data feature set corresponding to the first sub-dataset and the second data feature set corresponding to the second sub-dataset at the target node to all the data in the first data feature set.
[0111] The target node can be any node in the system.
[0112] This first ratio is used to characterize the feature overlap degree between the first sub-dataset and the second sub-dataset.
[0113] During the actual execution process, as Figure 4 shown, in the scenario of vertical federated learning, the system can use the Private Set Intersection (PSI) algorithm to achieve the intersection between the first data feature set corresponding to the initiator (such as institution A) and the second data feature set corresponding to the data provider (such as institution X) without exchanging the original data.
[0114] After the intersection is completed, the initiator reports the intersection result, that is, the first ratio, to the central node.
[0115] After obtaining the first ratio, based on the first ratio and the data quality information of the second data feature set, the data content matching degree between the first data feature set and the second data feature set can be generated.
[0116] Among them, the data quality information of the second data feature set is used to characterize the quality level of the data in the second data feature set.
[0117] In some embodiments, the data quality information of the second dataset may include at least one of: the second ratio of missing data in the second dataset, the third ratio of abnormal data in the second dataset, and the fourth ratio of incorrect data in the second dataset.
[0118] In this embodiment, the second ratio is the ratio of the missing data in the second data feature set to all the data in the second data feature set; the second ratio is used to characterize the level of completeness of the second data feature set.
[0119] The third ratio is the ratio of the abnormal data in the second data feature set to all the data in the second data feature set. Among them, the abnormal data is the data that does not conform to the data standard included in the second data feature set. For example, when the second data feature set is used to characterize the gender of users, where 0 represents male and 1 represents female, if a non-0 or 1 number appears in the second data feature set, then the data is determined to be abnormal data.
[0120] The fourth ratio is the ratio of the incorrect data in the second data feature set to all the data in the second data feature set. Among them, the incorrect data is the data that does not conform to the actual situation included in the second data feature set. For example, when the second data feature set is used to represent the gender of a user, where males are represented by the number 0 and females are represented by the number 1, if the number corresponding to a certain data in the second data feature set is 0, and in fact the gender of the user corresponding to this data is female, then this data is determined to be incorrect data.
[0121] The third ratio and the fourth ratio are used to characterize the correctness level of the second data feature set.
[0122] In the actual execution process, the formula:
[0123]
[0124] can be used to generate the data content matching degree between the first data feature set and the second data feature set. Among them, ds is the data content matching degree, ir is the first ratio, α i is the second ratio, β i is the third ratio, γ i is the fourth ratio, and n is the number of columns of the second data feature set (that is, the number of features in the second data feature set).
[0125] By processing the first ratio corresponding to the first sub-dataset and the data quality information corresponding to the second sub-dataset at each registered node in the system respectively, the data content matching degree between the first sub-dataset and each second sub-dataset at that node can be obtained.
[0126] Adopting a similar algorithm for each first sub-dataset in the first dataset, multiple data content matching degrees between the first data feature set corresponding to the first dataset and multiple second data feature sets corresponding to multiple second datasets can be generated.
[0127] It should be noted that the execution order of step 110 and step 120 can be a parallel order, or it can also be a sequential order, which is not limited in the present invention.
[0128] Of course, in some other embodiments, other similarity calculation methods can also be used to obtain the data content matching degree, such as calculation methods like cosine similarity or Jaccard similarity, which are not limited in the present invention.
[0129] According to the data matching method based on federated learning provided by the embodiments of the present invention, by calculating the data content matching degree between the first dataset and the second dataset, the second dataset corresponding to a suitable node is screened for the initiator from the data content dimension, shortening the matching time and improving the matching efficiency.
[0130] Step 130: Based on the label matching degree and data content matching degree between the first data set and the same second data set, determine the target data set from multiple second data sets as the federated matching data of the first data set.
[0131] In this step, the federated matching data is the data that can be used for federated training obtained from the data provider.
[0132] The target data set is a second sub - data set with a relatively high matching degree with the first data set, which is screened from multiple second data sets.
[0133] The target data set may include one or more second sub - data sets.
[0134] In the actual execution process, the label matching degree and data content matching degree between each first sub - data set and each second sub - data set at each node in the system can be fused, such as through weighted fusion calculation or calculating the average value, etc., and based on the fusion result, determine the second sub - data set corresponding to the node with a relatively high matching degree with the first sub - data set from multiple nodes in the system, so as to be able to screen out the target data set with a relatively high matching degree with the first data set from all the second sub - data sets as the federated matching data corresponding to this first sub - data set.
[0135] In the embodiments of the present invention, on the one hand, by aggregating multiple participants to achieve data matching and joint training between the initiator and the data provider in the federation, by supporting the access of data from multiple enterprises in multiple industries (i.e., the access of multiple second data sets), the scope of data matching is greatly broadened, and the feature dimension of the data is enriched; by providing a larger range of data providers for screening and more accurate data quality for screening, the comprehensiveness, accuracy, and precision of the matching result are significantly improved, thereby improving the data quality to improve the subsequent model effect.
[0136] On the other hand, by calculating the matching degree between the data provider and the initiator from multiple dimensions such as data labels and data content, to determine the target data set with the highest matching degree from multiple second data sets as the federated matching data of the first data set, the comprehensiveness and precision of the evaluation are effectively improved, thereby improving the quality of the matched data; in addition, it can also effectively shorten the matching time to improve the matching efficiency.
[0137] According to the data matching method based on federated learning provided by an embodiment of the present invention, by aggregating the second data sets provided by multiple data providers and calculating the matching degree between the second data sets provided by the data providers and the first data set of the initiator from multiple dimensions, the target data set with the highest matching degree is determined from multiple second data sets as the federated matching data of the first data set, greatly broadening the scope of data matching, enriching the feature dimensions of the data, significantly improving the comprehensiveness, accuracy, and precision of the vertical federated matching results, and helping to improve the subsequent model effect.
[0138] In some embodiments, step 130 may include:
[0139] Generate a comprehensive matching degree based on the label matching degree and content matching degree between the first data set and the same second data set;
[0140] Sort the comprehensive matching degrees in descending order, and determine the second data sets corresponding to the top target number of comprehensive matching degrees as the target data sets.
[0141] In this embodiment, the comprehensive matching degree is the final matching degree after fusing the label matching degree and content matching degree;
[0142] The comprehensive matching degree is used to represent the overall similarity degree between the first sub-data set and the second sub-data set.
[0143] It can be understood that the larger the value of the comprehensive matching degree, the higher the overall similarity degree between the first sub-data set and the second sub-data set.
[0144] The comprehensive matching degree can be generated by means such as average calculation or weighted calculation. Its specific generation method will be described in subsequent embodiments and will not be elaborated here.
[0145] The target number can be user-defined. For example, the target number can be set to 10 or 50, etc., and the present invention does not make any limitations.
[0146] For example, after sorting the comprehensive matching degrees in descending order, select the second data sets corresponding to the top 10 or top 50 comprehensive matching degrees as the target data sets.
[0147] Such as Figure 5 As shown, in the actual execution process, the label matching degree and content matching degree between each first sub-data set and each second sub-data set at this node are respectively fused to obtain the comprehensive matching degree between this first sub-data set and this second sub-data set.
[0148] Using a similar method, the comprehensive matching degrees between each first sub-data set in the first data set and each corresponding second sub-data set at each node can be obtained, thereby obtaining multiple comprehensive matching degrees.
[0149] It is understandable that the comprehensive matching degrees corresponding to two different second sub - data sets may be the same or different.
[0150] After obtaining all the comprehensive matching degrees, sort the comprehensive matching degrees in descending order, and use the second sub - data sets corresponding to the TopN comprehensive matching degrees as the target data sets and return them to the initiator.
[0151] Of course, in some other embodiments, step 130 may further include:
[0152] Generate a comprehensive matching degree based on the label matching degree and the content matching degree between the first data set and the same second data set;
[0153] Sort the comprehensive matching degrees in ascending order, and determine the second data sets corresponding to the target number of comprehensive matching degrees with lower rankings as the target data sets.
[0154] In this embodiment, the target number can also be user - defined. For example, the target number can be set to 10 or 50, etc., and the present invention does not make any limitations.
[0155] For example, after obtaining all the comprehensive matching degrees, the comprehensive matching degrees can also be sorted in ascending order, and the second sub - data sets corresponding to the N comprehensive matching degrees with lower rankings are used as the target data sets and returned to the initiator.
[0156] According to the data matching method based on federated learning provided by the embodiments of the present invention, by fusing the matching degrees of multiple dimensions such as data labels and data content to generate the final comprehensive matching degree, and based on the comprehensive matching degree, determining the second data sets corresponding to the nodes with higher comprehensive matching degrees among multiple nodes as the federated matching data, it effectively improves the comprehensiveness and accuracy of the evaluation, thereby improving the quality of the matched data; in addition, it can also effectively shorten the matching time to improve the matching efficiency.
[0157] Next, the generation method of the comprehensive matching degree will be described through specific embodiments.
[0158] In some embodiments, generating a comprehensive matching degree based on the label matching degree and the content matching degree between the first data set and the same second data set may include:
[0159] Obtain the target first weight value corresponding to the label matching degree, the target second weight value corresponding to the data content matching degree, and the target evaluation score;
[0160] Respectively perform normalization processing on the label matching degree and the content matching degree to generate a normalized label matching degree and a normalized data content matching degree;
[0161] Generate a comprehensive matching degree based on the normalized label matching degree, the normalized data content matching degree, the target first weight value, the target second weight value, and the target evaluation score.
[0162] In this embodiment, the target first weight value and the target second weight value are the weight values corresponding to each parameter value in the process of fusing the label matching degree and the content matching degree.
[0163] Among them, the target first weight value and the target second weight value can be customized by the initiator.
[0164] It can be understood that in the case where the initiator attaches importance to the data label matching degree, the target first weight value can be set higher. On the contrary, in the case where the initiator attaches importance to the data content matching degree, the target second weight value can be set higher.
[0165] Of course, in some other embodiments, the target first weight value and the target second weight value can also be automatically generated by the system, and the present invention does not make a limitation.
[0166] The target evaluation score is the evaluation score corresponding to the target node among the registered nodes in the system, that is, the evaluation score corresponding to the data provider. Among them, the target node can be any node in the system.
[0167] The evaluation score is used to characterize the credibility of the data provider.
[0168] Each node corresponds to an evaluation score.
[0169] The target evaluation score can be generated by the system based on historical federation records.
[0170] For example, in the actual execution process, after each federation is completed, the initiator can evaluate and score the data provider, so as to form the evaluation score corresponding to this node.
[0171] The system processes multiple evaluation scores at this node to generate the target evaluation score corresponding to this node.
[0172] In addition, in the actual execution process, the label matching degree and the data content matching degree can also be normalized to scale the values of the label matching degree and the data content matching degree to between [0, 1] respectively, so as to improve the accuracy of subsequent calculation results.
[0173] For example, through the formula:
[0174]
[0175] Perform normalization processing on the label matching degree, where ls’ is the normalized label matching degree and ls is the label matching degree.
[0176] Through the formula:
[0177]
[0178] Normalize the data content matching degree, where ds’ is the normalized data content matching degree and ds is the data content matching degree.
[0179] After obtaining the target first weight value, the target second weight value, the target evaluation score at this node corresponding to the second data set, the normalized label matching degree, and the normalized data content matching degree, through the formula:
[0180] score=gs×(w 1 ×ls’+w 2 ×ds’)
[0181] The comprehensive matching degree can be generated, where score is the comprehensive matching degree, gs is the target evaluation score corresponding to the node, w 1 is the target first weight value, ls’ is the normalized label matching degree, w 2 is the target second weight value, ds’ is the normalized data content matching degree, and w 1 +w 2 =1.
[0182] Of course, in other embodiments, the comprehensive matching degree can also be generated by setting other weight values, which is not limited in the present invention.
[0183] According to the data matching method based on federated learning provided by the embodiments of the present invention, by introducing the target evaluation score, the credibility of the data provider can be incorporated into the reference factors for screening, so that the second sub-data set provided by the data provider with a higher credibility can be screened to improve the accuracy of the data matching result; by introducing the weight value, the user can adjust the weight value based on actual needs to screen the required matching data, which has high flexibility and universality in use.
[0184] Next, the data matching device based on federated learning provided by the present invention will be described. The data matching device based on federated learning described below can be correspondingly referred to the data matching method based on federated learning described above.
[0185] As Figure 6 shown, the data matching device based on federated learning includes: a first processing module 610, a second processing module 620, and a third processing module 630.
[0186] The first processing module 610 is used to calculate the similarity between the first data label corresponding to the first data set and the multiple second data labels corresponding to the multiple second data sets, and generate multiple label matching degrees;
[0187] A second processing module 620, configured to calculate the similarity between the first data feature set corresponding to the first data set and the multiple second data feature sets corresponding to the multiple second data sets respectively, and generate multiple data content matching degrees.
[0188] A third processing module 630, configured to determine a target data set from the multiple second data sets as the federated matching data of the first data set based on the label matching degree and the data content matching degree between the first data set and the same second data set.
[0189] According to the data matching device based on federated learning provided by the embodiments of the present invention, by aggregating the second data sets provided by multiple data providers and calculating the matching degree between the second data sets provided by the data providers and the first data set of the initiator from multiple dimensions, to determine the target data set with the highest matching degree from the multiple second data sets as the federated matching data of the first data set, which greatly broadens the scope of data matching, enriches the feature dimensions of the data, significantly improves the comprehensiveness, accuracy and precision of the vertical federated matching result, and helps to improve the subsequent model effect.
[0190] In some embodiments, the third processing module 630 may further be configured to:
[0191] Generate a comprehensive matching degree based on the label matching degree and the data content matching degree between the first data set and the same second data set;
[0192] Perform a descending order sorting on the comprehensive matching degree, and determine the second data sets corresponding to the target number of comprehensive matching degrees ranked at the front as the target data sets;
[0193] Alternatively, perform an ascending order sorting on the comprehensive matching degree, and determine the second data sets corresponding to the target number of comprehensive matching degrees ranked at the back as the target data sets.
[0194] In some embodiments, the third processing module 630 may further be configured to:
[0195] Obtain a target first weight value corresponding to the label matching degree, a target second weight value corresponding to the data content matching degree, and a target evaluation score;
[0196] Perform a normalization process on the label matching degree and the data content matching degree respectively to generate a normalized label matching degree and a normalized data content matching degree;
[0197] Generate a comprehensive matching degree based on the normalized label matching degree, the normalized data content matching degree, the target first weight value, the target second weight value, and the target evaluation score.
[0198] In some embodiments, the first processing module 610 may further be configured to: calculate the cosine similarity between the first data tag and the second data tag to generate a tag matching degree between the first data set and the second data set.
[0199] In some embodiments, the second processing module 620 may further be configured to:
[0200] Perform a private set intersection on the first data feature set and the second data feature set to obtain a first ratio of the same data in the first data feature set and the second data feature set to the first data feature set;
[0201] Generate a data content matching degree between the first data feature set and the second data feature set based on the first ratio and the data quality information of the second data feature set.
[0202] In some embodiments, the data quality information of the second data feature set includes at least one of: a second ratio of missing data in the second data feature set, a third ratio of abnormal data in the second data feature set, and a fourth ratio of incorrect data in the second data feature set.
[0203] Figure 7 An example of a schematic physical structure diagram of an electronic device is shown as Figure 7 shown. The electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communications interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 may call logic instructions in the memory 730 to execute a data matching method based on federated learning. The method includes: calculating the similarity between the first data tag corresponding to the first data set and multiple second data tags corresponding to multiple second data sets respectively to generate multiple tag matching degrees; calculating the similarity between the first data feature set corresponding to the first data set and multiple second data feature sets corresponding to multiple second data sets respectively to generate multiple data content matching degrees; determining a target data set from multiple second data sets as the federated matching data of the first data set based on the tag matching degree and the data content matching degree between the first data set and the same second data set.
[0204] In addition, when the logical instructions in the above-mentioned memory 730 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0205] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the data matching method based on federated learning provided by the above-mentioned various methods. The method includes: respectively calculating the similarity between the first data labels corresponding to the first data set and the multiple second data labels corresponding to the multiple second data sets to generate multiple label matching degrees; respectively calculating the similarity between the first data feature set corresponding to the first data set and the multiple second data feature sets corresponding to the multiple second data sets to generate multiple data content matching degrees; based on the label matching degree and the data content matching degree between the first data set and the same second data set, determining a target data set from the multiple second data sets as the federated matching data of the first data set.
[0206] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the data matching method based on federated learning provided by the above-mentioned various methods. The method includes: respectively calculating the similarity between the first data labels corresponding to the first data set and the multiple second data labels corresponding to the multiple second data sets to generate multiple label matching degrees; respectively calculating the similarity between the first data feature set corresponding to the first data set and the multiple second data feature sets corresponding to the multiple second data sets to generate multiple data content matching degrees; based on the label matching degree and the data content matching degree between the first data set and the same second data set, determining a target data set from the multiple second data sets as the federated matching data of the first data set.
[0207] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0208] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0209] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data matching method based on federated learning, characterized in that, it includes: Calculate the similarity between the first data labels corresponding to the first data set and the multiple second data labels corresponding to the multiple second data sets respectively, and generate multiple label matching degrees; Calculate the similarity between the first data feature set corresponding to the first data set and the multiple second data feature sets corresponding to the multiple second data sets respectively, and generate multiple data content matching degrees; Based on the label matching degree and the data content matching degree between the first data set and the same second data set, determine a target data set from the multiple second data sets as the federated matching data of the first data set; Among them, vertical federated learning includes three roles: the initiator, the data provider, and the collaborator; the first data set is the data set to be matched provided by the initiator, and the second data set is the data set provided by the data provider for matching; The step of calculating the similarity between the first data feature set corresponding to the first data set and the multiple second data feature sets corresponding to the multiple second data sets respectively, and generating multiple data content matching degrees includes: Perform private set intersection on the first data feature set and the second data feature set, and obtain a first ratio of the same data in the first data feature set and the second data feature set to the first data feature set; Based on the first ratio and the data quality information of the second data feature set, generate the data content matching degree between the first data feature set and the second data feature set.
2. The data matching method based on federated learning according to claim 1, characterized in that, The step of determining a target data set from the multiple second data sets as the federated matching data of the first data set based on the label matching degree and the data content matching degree between the first data set and the same second data set includes: Generate a comprehensive matching degree based on the label matching degree and the data content matching degree between the first data set and the same second data set; Sort the comprehensive matching degrees in descending order, and determine the second data sets corresponding to the target number of the comprehensive matching degrees with the highest rankings as the target data sets; Or, sort the comprehensive matching degrees in ascending order, and determine the second data sets corresponding to the target number of the comprehensive matching degrees with the lowest rankings as the target data sets.
3. The data matching method based on federated learning according to claim 2, characterized in that, The step of generating a comprehensive matching degree based on the label matching degree and the data content matching degree between the first data set and the same second data set includes: Obtain a target first weight value corresponding to the label matching degree, a target second weight value corresponding to the data content matching degree, and a target evaluation score; the target evaluation score is the evaluation score corresponding to the data provider; Perform normalization processing on the label matching degree and the data content matching degree respectively, and generate a normalized label matching degree and a normalized data content matching degree; Generate the comprehensive matching degree based on the normalized label matching degree, the normalized data content matching degree, the target first weight value, the target second weight value, and the target evaluation score.
4. The data matching method based on federated learning according to any one of claims 1-3, wherein, the calculating the similarity between the first data label corresponding to the first data set and the multiple second data labels corresponding to the multiple second data sets respectively to generate multiple label matching degrees includes: calculating the cosine similarity between the first data label and the second data label to generate the label matching degree between the first data set and the second data set.
5. The data matching method based on federated learning according to claim 1, wherein, the data quality information of the second data feature set includes at least one of: a second ratio of missing data in the second data feature set, a third ratio of abnormal data in the second data feature set, and a fourth ratio of incorrect data in the second data feature set.
6. A data matching device based on federated learning, wherein, it includes: a first processing module, configured to calculate the similarity between the first data label corresponding to the first data set and the multiple second data labels corresponding to the multiple second data sets respectively to generate multiple label matching degrees; a second processing module, configured to calculate the similarity between the first data feature set corresponding to the first data set and the multiple second data feature sets corresponding to the multiple second data sets respectively to generate multiple data content matching degrees; a third processing module, configured to determine a target data set as the federated matching data of the first data set from the multiple second data sets based on the label matching degree and the data content matching degree between the first data set and the same second data set; wherein, vertical federated learning includes three roles: an initiator, a data provider, and a collaborator; the first data set is the data set to be matched provided by the initiator, and the second data set is the data set for matching provided by the data provider; the second processing module is specifically configured to: perform a private set intersection on the first data feature set and the second data feature set to obtain a first ratio of the same data in the first data feature set and the second data feature set to the first data feature set; generate the data content matching degree between the first data feature set and the second data feature set based on the first ratio and the data quality information of the second data feature set.
7. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the program, it implements the data matching method based on federated learning according to any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium, on which a computer program is stored, wherein, when the computer program is executed by a processor, it implements the data matching method based on federated learning according to any one of claims 1 to 5.
9. A computer program product, comprising a computer program, wherein, when the computer program is executed by a processor, it implements the data matching method based on federated learning according to any one of claims 1 to 5.
Citation Information
Patent Citations
Federation learning method, device and equipment and storage medium
CN111768008A
Model training method and device based on longitudinal federated learning system and storage medium
CN112001500A