ID Mapping method, device, electronic device, and storage medium

By acquiring and analyzing relationship data from multiple channels and using a bipartite graph model to perform optimal matching calculations, we solved the problem of ID Mapping accuracy in complex scenarios, achieved detailed relationship strength characterization and association calculations, and improved the accuracy of ID Mapping.

CN115827794BActive Publication Date: 2025-09-30RUN TECH CO LTD BEIJING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211379449.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-04
Publication Date
2025-09-30
Estimated Expiration
2042-11-04

AI Technical Summary

Technical Problem

Existing ID Mapping solutions have difficulty achieving accurate calculation goals in scenarios with complex ID types, large differences in data channels, and data acquisition continuity.

Method used

By obtaining the relationship data between identifiers within the same and different entity types, using the bipartite graph model to perform the best matching calculation, combining the data channel weight and relationship type weight, detailed relationship strength characterization and association calculation are performed to determine the ID Mapping result.

Benefits of technology

It improves the calculation accuracy of ID Mapping in complex scenarios, realizes the association calculation between various IDs, avoids data conflicts, and improves the consistency and accuracy of calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115827794B_ABST
    Figure CN115827794B_ABST
Patent Text Reader

Abstract

The present invention provides an ID Mapping method, device, electronic device, and storage medium. First, ID Mapping is performed within an entity type based on relational data corresponding to multiple data channels to obtain an ID Mapping library for each entity object within the same entity type. A comprehensive score is then calculated between pairwise identifiers of different entity types. Combined with the ID Mapping library for each entity object within the same entity type, the relationship scores between different entity objects are determined. A best match calculation is then performed using a bipartite graph model to obtain the association relationship between different entity objects, ultimately determining the ID Mapping results between entity types. This approach takes into account relational data corresponding to multiple data channels in complex scenarios, achieving a detailed characterization of relationship strengths and weaknesses, and implementing association calculations through a bipartite graph model, thereby improving the calculation accuracy of ID Mapping in complex scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data computing technology, and in particular to an ID Mapping method, device, electronic device and storage medium. Background Art

[0002] ID Mapping involves using various technical means to identify data from multiple sources as belonging to the same object or entity. ID Mapping is an essential task for all types of big data companies. For example, internet giants encompass numerous applications within their product ecosystems, each with its own account system. Individuals can register accounts across multiple applications. To fully understand a data subject's preferences while ensuring privacy, it's necessary to link their online behavior across different applications for big data analysis. This requires designing an algorithm to calculate the correlation between different application account IDs (identifiers) and identify the underlying data entity. This is ID Mapping.

[0003] In the aforementioned internet companies' ID Mapping calculation scenarios, the types of IDs involved are relatively simple, primarily network account IDs and IDs associated with devices used to log in to these accounts. Calculation data is generally collected through a stable and continuous collection mechanism, and the relationship data between various account IDs is relatively objective. Existing ID Mapping solutions can achieve these calculation goals. However, when ID types become more complex and data channels and the degree of continuity of data acquisition vary significantly, existing ID Mapping solutions struggle to achieve these goals. Summary of the Invention

[0004] The object of the present invention is to provide an ID Mapping method, device, electronic device and storage medium to improve the calculation accuracy of ID Mapping in complex scenarios.

[0005] In a first aspect, an embodiment of the present invention provides an ID Mapping method, including:

[0006] Obtaining first basic data under corresponding relationship types between two identifiers of the same entity type; wherein the relationship type is defined by identifier types of the two corresponding identifiers, and the first basic data includes relationship data corresponding to multiple data channels;

[0007] Determine, based on the first basic data, an ID Mapping library for each entity object within the same entity type;

[0008] Obtain the second basic data under each relationship type between two identifiers of different entity types;

[0009] Calculate, based on the second basic data, a comprehensive score between two identifiers of different entity types;

[0010] Determine the relationship score between different entity objects based on the comprehensive score between the two identifiers of the different entity types and the ID Mapping library of each entity object within the same entity type;

[0011] According to the relationship scores between the different entity objects, the best matching calculation is performed through a bipartite graph model to obtain the association relationship between the different entity objects;

[0012] According to the association relationship between the different entity objects, the ID Mapping results between the entity types are determined.

[0013] Furthermore, the step of determining the IDMapping library of each entity object within the same entity type based on the first basic data includes:

[0014] Calculating, based on the first basic data, a first score for a corresponding relationship type between two identifiers of the same entity type;

[0015] According to the first scores under the corresponding relationship types between the two identifiers in the same entity type, the best matching calculation is performed through the bipartite graph model to obtain the association relationship between the identifiers in the same entity type;

[0016] According to the association relationship between the identifiers within the same entity type, the entity objects are split according to the priority of the identifier type to obtain an ID Mapping library for each entity object within the same entity type.

[0017] Furthermore, the step of calculating, based on the first basic data, a first score for a corresponding relationship type between two identifiers within the same entity type includes:

[0018] For each group of identifiers within the same entity type, the basic score corresponding to each data channel under the corresponding relationship type of the group identifier is calculated based on the relationship data corresponding to each data channel under the corresponding relationship type of the group identifier;

[0019] According to the channel weight corresponding to each data channel, the basic scores corresponding to each data channel under the corresponding relationship type of the group identifier are weightedly calculated to obtain the first score under the corresponding relationship type of the group identifier.

[0020] Furthermore, the step of performing a best match calculation using a bipartite graph model based on the first scores under the corresponding relationship types between the two identifiers within the same entity type to obtain the association relationship between the identifiers within the same entity type includes:

[0021] All identifiers corresponding to two identifier types with a relationship within the same entity type are divided into two types of nodes in the bipartite graph model according to the identifier type;

[0022] Connecting nodes having relationships in the bipartite graph model into edges, where the weights of the edges are the corresponding first scores;

[0023] According to the saturation corresponding to each type of node and the weight of the edge between nodes, excess edges are deleted to obtain the association relationship between the identifiers corresponding to the bipartite graph model.

[0024] Furthermore, the step of calculating, based on the second basic data, a comprehensive score between two identifiers of different entity types includes:

[0025] Calculating, based on the second basic data, a second score for each relationship type between two identifiers of different entity types;

[0026] According to the relationship weight corresponding to each relationship type, the second scores under each relationship type between the two identifiers of the different entity types are weightedly calculated to obtain a comprehensive score between the two identifiers of the different entity types.

[0027] Furthermore, the step of determining the IDMapping results between entity types based on the association relationships between the different entity objects includes:

[0028] generating at least one association network based on the association relationships between the different entity objects, wherein the nodes of the association network are entity objects, and the weights of the paths connecting the nodes are corresponding relationship scores;

[0029] For each of the associated networks, determining whether there are multiple target nodes corresponding to the entity type with the highest priority in the associated network;

[0030] If so, the paths connected to the target nodes in the associated network are split in ascending order of path weights, and the multiple sub-networks obtained by splitting are used as ID Mapping results between entity types; wherein each sub-network contains one target node.

[0031] Furthermore, the relationship data includes one or more of the latest discovery time corresponding to the general data channel, the number of days between the earliest and latest discovery times, the number of discovery days, the number of discoveries, and / or the number of days since the data acquisition time corresponding to the specific data channel.

[0032] In a second aspect, an embodiment of the present invention further provides an ID Mapping device, including:

[0033] A first acquisition module is configured to acquire first basic data under a corresponding relationship type between two identifiers within the same entity type; wherein the relationship type is defined by the identifier types of the two corresponding identifiers, and the first basic data includes relationship data corresponding to multiple data channels;

[0034] A first determining module, configured to determine an ID Mapping library for each entity object of the same entity type according to the first basic data;

[0035] A second acquisition module is used to acquire second basic data under each relationship type between two identifiers of different entity types;

[0036] A first calculation module is configured to calculate, based on the second basic data, a comprehensive score between two identifiers of different entity types;

[0037] A score determination module, configured to determine a relationship score between different entity objects based on a comprehensive score between two identifiers of the different entity types and an ID Mapping library of each entity object within the same entity type;

[0038] A second calculation module is configured to perform a best matching calculation using a bipartite graph model based on the relationship scores between the different entity objects to obtain the association relationships between the different entity objects;

[0039] The second determining module is used to determine the IDMapping results between entity types according to the association relationship between the different entity objects.

[0040] In a third aspect, an embodiment of the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the ID Mapping method of the first aspect is implemented.

[0041] In a fourth aspect, an embodiment of the present invention further provides a storage medium having a computer program stored thereon, and when the computer program is run by a processor, the ID Mapping method of the first aspect is executed.

[0042] The ID Mapping method, device, electronic device, and storage medium provided by the embodiments of the present invention first perform ID Mapping within an entity type based on the relational data corresponding to multiple data channels, obtain an ID Mapping library for each entity object within the same entity type, then calculate a comprehensive score between pairwise identifiers of different entity types, and determine the relationship score between different entity objects by combining the ID Mapping library of each entity object within the same entity type. The best matching calculation is then performed using a bipartite graph model to obtain the association relationship between different entity objects, and finally determine the ID Mapping result between entity types. This ID Mapping method based on an entityized data system takes into account the relational data corresponding to multiple data channels in complex scenarios, achieves a detailed characterization of the strength of relationships, and implements association calculations through a bipartite graph model, thereby improving the calculation accuracy of ID Mapping in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0044] Figure 1 A flowchart of an ID Mapping method provided by an embodiment of the present invention;

[0045] Figure 2 A flowchart of another ID Mapping method provided by an embodiment of the present invention;

[0046] Figure 3 A schematic diagram of a logic for determining relationship scores between different entity objects provided by an embodiment of the present invention;

[0047] Figure 4 A schematic diagram of association matching logic based on a bipartite graph model provided by an embodiment of the present invention;

[0048] Figure 5 A schematic diagram of splitting an associated network provided by an embodiment of the present invention;

[0049] Figure 6 A schematic diagram of the structure of an ID Mapping device provided in an embodiment of the present invention;

[0050] Figure 7 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0051] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0052] For the ID Mapping calculation scenario of internet companies, the idea of ​​an existing ID Mapping solution is to classify and organize the raw data collected by various ecological applications, extract the object's unique identity ID, non-unique identity ID, and attribute information, and construct a two-dimensional ID relationship pair. Then, based on the input ID relationship pair, a machine learning algorithm is used to output an ID relationship pair with a stable relationship. Then, a connectivity graph algorithm is used to generate a UID as an identification code that uniquely identifies the object, achieving the goal of ID Mapping calculation.

[0053] Another existing ID Mapping solution uses terminal devices as a bridge, combines the relationship pairs between various accounts and device models, and user data such as device usage patterns, adopts rules to characterize the strength of relationships, and calculates the ID Mapping results based on connectivity graph partitioning and community discovery algorithms.

[0054] However, in complex scenarios, ID types are more complex, and data channels and the degree of continuity of data acquisition vary significantly, making it difficult to achieve the calculation goals of ID Mapping using the above solution. The embodiments of the present invention design a data system and calculation method for this typical scenario. Based on this, the embodiments of the present invention provide an ID Mapping method, device, electronic device, and storage medium. These methods organize the relationship data between various ID types based on a materialized data system design and provide a corresponding ID Mapping calculation method. These methods can implement association calculations between various ID types, improving the accuracy of ID Mapping calculations in complex scenarios.

[0055] To facilitate understanding of this embodiment, an ID Mapping method disclosed in an embodiment of the present invention is first introduced in detail.

[0056] An embodiment of the present invention provides an ID Mapping method that organizes the relationships between various IDs based on a materialized data system and designs relationship strength scoring rules and a relationship matching algorithm to calculate the degree of association between the IDs. The ID Mapping method can be executed by an electronic device with data processing capabilities.

[0057] In this embodiment, based on the idea of ​​entity-based data system design, various types of IDs commonly involved in ID Mapping are divided into different entity types. Various relationship types are defined between various identification types within entity types and between entity types. For example, personnel, vehicles, and mobile phones are three different entity types. The ID card number and passport number are two identification types for personnel, the frame number and license plate number are two identification types for vehicles, and the mobile phone number, IMEI code, and MEID are three identification types for mobile phones. The relationship type between the ID card number and the passport number is the same subject relationship, the relationship type between the frame number and the license plate number is the same subject relationship, and the relationship type between the ID card number and the license plate number is the owner relationship, etc. It should be noted that there is one relationship type between various identification types within an entity type, and there may be multiple relationship types between various identification types between entity types. For example, between the ID card number and the mobile phone number, a card registration relationship is generated when applying for a card, and a usage relationship is generated when checking in at a hotel.

[0058] Under this data system design, ID Mapping mainly includes the following two scenarios:

[0059] One is the association between various unique IDs within an entity type. This is essentially the connection between various parallel ID systems. Once discovered, this type of association will not change easily, and its time attribute is relatively weak.

[0060] One is the association between entities, which essentially reflects a stable binding relationship with strong time attributes.

[0061] The ID Mapping method provided in the embodiment of the present invention is mainly designed for the above two scenarios.

[0062] See also Figure 1 The flowchart of an ID Mapping method shown in FIG. 1 mainly includes the following steps S101 to S107:

[0063] Step S101: Acquire first basic data under corresponding relationship types between two identifiers in the same entity type; wherein the relationship type is defined by the identifier types of two corresponding identifiers, and the first basic data includes relationship data corresponding to multiple data channels.

[0064] The above-mentioned data channels can be divided into general data channels and specific data channels; among them, general data channels can include registration channels, card application channels, registration information channels, etc. The relationship data corresponding to general data channels can include one or more of the latest discovery time since today, the number of days between the earliest and latest discovery times, the number of discovery days, and the number of discoveries; the data of specific data channels can be high-quality data provided by customers, etc., and can be marked with a special "data channel". The relationship data corresponding to specific data channels can include the number of days since data acquisition time since today, etc.

[0065] An example of relational data corresponding to a possible common data channel is shown in Table 1 below:

[0066] Table 1

[0067]

[0068]

[0069] It should be noted that the same identification value can correspond to different identification types. For example, the identification value is a mobile phone number, but its identification type can be a payment APP account or a communication APP account.

[0070] Step S102: Determine an ID Mapping library for each entity object within the same entity type according to the first basic data.

[0071] In some possible embodiments, the existing ID Mapping scheme can be used to determine the ID Mapping library of each entity object within the same entity type. For example, the strength of the relationship between each identifier within the entity type is first characterized based on the first basic data, and then the ID Mapping result is calculated based on the connectivity graph partitioning and community discovery algorithm.

[0072] In other possible embodiments, the strength of the relationship between each pair of identifiers within the entity type can be first characterized based on the first basic data, and then the best matching calculation can be performed on the relationship between the identifiers within the entity type according to the bipartite graph model and the strength of the relationship. Finally, the entity objects within the entity type can be split to obtain the ID Mapping result within the entity type.

[0073] Based on this, the above step S102 can be implemented as follows:

[0074] 1.1 Based on the first basic data, a first score under a corresponding relationship type between two identifiers in the same entity type is calculated.

[0075] 1.2 Based on the first score of the corresponding relationship type between two identifiers in the same entity type, the best matching calculation is performed through the bipartite graph model to obtain the association relationship between the identifiers in the same entity type.

[0076] 1.3 Based on the association relationship between each identifier within the same entity type, the entity objects are split according to the priority of the identifier type to obtain the ID Mapping library of each entity object within the same entity type.

[0077] The relationship between each identifier within an entity type can be scored based on the data channel, number of discoveries, number of days discovered, earliest and latest discovery time, and their combined dimensions. Based on this, the above step 1.1 can be implemented through the following process: for each group of identifiers within the same entity type (a group of identifiers includes two related identifiers), the basic score corresponding to the data channel under the corresponding relationship type of the group identifier is calculated based on the relationship data corresponding to each data channel under the corresponding relationship type of the group identifier; according to the channel weight corresponding to each data channel, the basic score corresponding to each data channel under the corresponding relationship type of the group identifier is weighted and calculated to obtain the first score under the corresponding relationship type of the group identifier. Among them, the relationship data includes one or more of the latest discovery time corresponding to the general data channel, the number of days between the earliest and latest discovery times, the number of days discovered, the number of discoveries, and / or the number of days from today when the data was acquired corresponding to the specific data channel.

[0078] Due to the existence of dirty data or accompanying relationship data in the relational data, there may be multiple single-user attributes (unique attributes that only appear in one character data), or multiple associations may appear in a one-to-one correspondence, that is, there are conflicting association relationships. For example, if a unique single-user attribute such as an ID card or user ID has multiple related identifiers, it is considered that a conflict has occurred; if a mobile phone number belongs to multiple user ID numbers, it is considered that a conflict has occurred; if a mobile phone number only corresponds to one IMSI for a period of time, if a one-to-many or many-to-one relationship occurs, it is considered that a conflict has occurred. In response to this situation, this embodiment proposes an association matching algorithm based on a bipartite graph model, that is, the above step 1.2 can be implemented by the following process: all identifiers corresponding to two types of identifiers with relationships in the same entity type are divided into two types of nodes of the bipartite graph model according to the identifier type, and one type of node represents an identifier of one type of identifier; the nodes with relationships in the bipartite graph model are connected into edges, and the weight of the edge is the corresponding first score; according to the saturation corresponding to each type of node and the weight of the edge between nodes, excess edges are deleted to obtain the association relationship between each identifier corresponding to the bipartite graph model. This will delete the conflicting relationships.

[0079] The association relationships between identifiers within the same entity type (IDMapping association pairs within the entity type) calculated in step 1.2 will form connected association networks. Each association network may involve multiple entities. In order to minimize the association network to the corresponding data entity, the ID Mapping association network needs to be split. Based on this, the above step 1.3 can be implemented through the following process: based on the association relationships between identifiers within the same entity type, at least one association network is generated, the nodes of the association network are identifiers, and the weights of the paths connecting the nodes are the corresponding first scores; for each association network, determine whether there are multiple target nodes corresponding to the identifier type with the highest priority in the association network; if so, split the paths connected to the target nodes in the association network in order of the path weight from small to large, and use the multiple sub-networks obtained by splitting as the ID Mapping results within the entity type; wherein each sub-network contains one target node; if not, directly use the association network as the ID Mapping result within the entity type.

[0080] Step S103: Acquire second basic data of each relationship type between two identifiers of different entity types.

[0081] The second basic data here is similar to the first basic data mentioned above and will not be described again here.

[0082] Step S104: Calculate the comprehensive scores between each pair of identifiers of different entity types based on the second basic data.

[0083] Taking into account that there may be multiple relationship types between various identification types between entity types, such as multiple relationship types under the same "A entity type-B entity type", if there are the same "A identification type-A identification value-B identification type-B identification value", then after calculating the second score under each relationship type, according to the relationship weight corresponding to each relationship type, the second score under each relationship type is weighted averaged to obtain a comprehensive score under "A identification type-A identification value-B identification type-B identification value". Based on this, the above step S104 can be implemented through the following process: according to the second basic data, the second score under each relationship type between the two identifications of different entity types is calculated; according to the relationship weight corresponding to each relationship type, the second score under each relationship type between the two identifications of different entity types is weightedly calculated to obtain a comprehensive score between the two identifications of different entity types.

[0084] The specific process of calculating the second score for each relationship type between two identifiers of different entity types based on the second basic data can be referred to the specific process of calculating the first score for the corresponding relationship type between two identifiers of the same entity type based on the first basic data, and will not be repeated here.

[0085] Step S105 : determining the relationship scores between different entity objects based on the comprehensive scores between the identifiers of each of the different entity types and the ID Mapping library of each entity object within the same entity type.

[0086] The comprehensive scores between the pairwise identifiers of different entity types are combined with the results of ID Mapping within the entity type. The comprehensive scores between the pairwise identifiers of the above different entity types are merged according to the entity objects to obtain multiple comprehensive scores between the entity objects. The maximum value of the comprehensive score can be taken (other functions can also be designed) as the relationship score between the entity objects, which is used for subsequent ID Mapping calculations between entity types.

[0087] Step S106 , performing best matching calculations through a bipartite graph model based on the relationship scores between different entity objects to obtain the association relationships between different entity objects.

[0088] Considering the potential for conflicting relationships between different entity objects, for example, if a person is associated with at most five mobile phones, but person X is associated with six, a conflict is considered to exist. To address this situation, the aforementioned bipartite graph-based association matching algorithm can be used to perform a best match calculation on the relationship scores between different entity objects, thereby determining the relationships between them. The specific implementation process is similar to the bipartite graph-based association matching algorithm used to determine the relationships between identifiers within the same entity type. Specifically, the entity objects with relationships are divided into two types of nodes in the bipartite graph model based on entity type, with one type of node representing an entity object of each entity type. The nodes with relationships in the bipartite graph model are connected into edges, with the edge weights representing the corresponding relationship scores. Based on the saturation of each node type and the edge weights between nodes, excess edges are deleted to determine the relationships between different entity objects corresponding to the bipartite graph model. This method can remove conflicting relationships.

[0089] Step S107: Determine the ID Mapping result between entity types according to the association relationship between different entity objects.

[0090] The association relationships between different entity objects (ID Mapping association pairs between entity types) calculated in step S106 will form connected association networks. Each association network may involve multiple entities. In order to narrow the association network to the corresponding data entity as much as possible, the ID Mapping association network needs to be split. Based on this, the above step S107 can be implemented through the following process: based on the association relationships between different entity objects, at least one association network is generated, the nodes of the association network are entity objects, and the weights of the paths connecting the nodes are the corresponding relationship scores; for each association network, it is determined whether there are multiple target nodes corresponding to the entity type with the highest priority in the association network; if so, the paths connected to each target node in the association network are split in order of the path weight from small to large, and the multiple sub-networks obtained by the split are used as the ID Mapping results between entity types; wherein each sub-network contains one target node; if not, the association network is directly used as the ID Mapping result between entity types.

[0091] The ID Mapping method provided in an embodiment of the present invention first performs ID Mapping within an entity type based on the relational data corresponding to multiple data channels, obtains an ID Mapping library for each entity object within the same entity type, then calculates a comprehensive score between pairwise identifiers of different entity types, and combines the ID Mapping library of each entity object within the same entity type to determine the relationship score between different entity objects. A best match calculation is then performed using a bipartite graph model to obtain the association relationship between different entity objects, ultimately determining the ID Mapping result between entity types. This ID Mapping method based on an entityized data system takes into account the relational data corresponding to multiple data channels in complex scenarios, achieves a detailed characterization of the strength of relationships, and implements association calculations through a bipartite graph model, thereby improving the calculation accuracy of ID Mapping in complex scenarios.

[0092] For ease of understanding, the embodiment of the present invention also provides an ID Mapping overall data flow, such as Figure 2 As shown:

[0093] 1. Data Preparation

[0094] That is, extract the relational data between various types of identifiers.

[0095] 2. Association Calculation

[0096] (1) Association calculation of ID Mapping within entity type

[0097] 1) Obtain the relationship data between each identifier within the entity type;

[0098] 2) Calculate the strength of the relationship between each pair of identifiers within the entity type;

[0099] 3) Based on the association matching algorithm of the bipartite graph model, the best matching of the relationship between each pair of identifiers within the entity type is calculated;

[0100] 4) Split entity objects according to the priority of identification type.

[0101] (2) Associations between various entity objects within an entity type

[0102] Get the ID Mapping library of each entity object, such as the ID Mapping library of entity A, the ID Mapping library of entity B, and the ID Mapping library of entity M.

[0103] (3) Association calculation of ID Mapping between entity types

[0104] 1) Obtain the relationship data between each identifier between entity types;

[0105] 2) Calculate the strength of the relationship between each pair of identifiers between entity types;

[0106] 3) Combine the association results of various entity objects within the entity type and calculate the strength of the relationship between the two entity objects of different entity types;

[0107] 4) Based on the association matching algorithm of the bipartite graph model, the optimal matching of pairwise relationships between different entity objects is calculated.

[0108] (4) Associations between entity types

[0109] Get the ID Mapping library between different entity objects of different entity types, such as the ID Mapping library of entity A-entity B, the ID Mapping library of entity A-entity M...the ID Mapping library of entity B-entity M.

[0110] (5) Association network between entity types

[0111] For example, A entity object 1 is associated with B entity object 1, B entity object 1 is also associated with M entity object 1, A entity object 2 is associated with B entity object 2, B entity object 2 is also associated with M entity object 1, and so on.

[0112] (6) Splitting the association network of ID Mapping between entity types

[0113] 1) Split the association network according to the priority of entity type;

[0114] 2) Get the ID Mapping subnetwork.

[0115] For ease of understanding, the following describes in detail the calculation of the first score under the corresponding relationship type between two identifiers within the same entity type in step 1.1 based on the first basic data.

[0116] (1) Basic scoring

[0117] 1) Basic dimension scoring

[0118] Basic dimensions refer to the various dimensions corresponding to common data channels, including the number of days from the latest discovery time to today, the number of days between the earliest and latest discovery times, the saturation of discovery days, and the average number of discoveries per day.

[0119] ①The number of days since the latest discovery time

[0120] The score s1 under this dimension can be determined by the following formula:

[0121]

[0122] Among them, x1 is the number of days from the latest discovery time to today.

[0123] This dimension primarily considers the freshness of the data. The more recent the discovery time, the higher the score. After 15 days, the score begins to decay exponentially, rapidly decreasing to a lower value. The scoring range is (1, 10).

[0124] ② Number of days between the earliest and latest discovery times

[0125] The score s2 under this dimension can be determined by the following formula:

[0126] s2=s1+s1*min(x2 / 180,1)

[0127] Where x2 is the number of days between the earliest and latest discovery times (including the earliest and latest discovery times).

[0128] This dimension primarily considers the score gain from adding the time interval between the earliest and latest discovery times to the score of the latest discovery time. If the latest discovery time is the same, the further the earliest discovery time is from the latest discovery time, the longer the relationship is, and the higher the score (up to double). The score range is (1, 20).

[0129] ③Discover the saturation of days

[0130] The score s3 under this dimension can be determined by the following formula:

[0131] s3=s2+s2*(x3 / x2)

[0132] Among them, x3 is the number of days to discovery.

[0133] This dimension primarily considers the saturation of the number of days between the earliest and latest discovery times, based on the aforementioned scoring. When the interval between the earliest and latest discovery times is the same, a greater number of discovery days indicates a more stable relationship, and the corresponding score is higher (up to doubled). The score range is (1, 40).

[0134] ④Average number of discoveries per day

[0135] The score under this dimension can be determined by the following formula:

[0136]

[0137] Where x4 is the number of discoveries, and score is the basic score corresponding to the common data channel.

[0138] This dimension combines the number of discoveries and the number of days since discovery. The more discoveries per day, the more stable the relationship is considered, and the higher the corresponding score (up to a maximum of 0.25 times). The score range is (1, 50).

[0139] 2) Special dimension scoring

[0140] Special dimensions refer to dimensions corresponding to specific data channels, and can include the number of days since the data was acquired. For data from specific data channels, scoring follows a separate logic, and the score for that dimension, kh_score, can be determined using the following formula:

[0141] hk_score=base_score+(100-base_score)*2 (-x / T)

[0142] Where x is the number of days since the data was acquired; base_score guarantees the lower limit of the score for a specific channel relationship data; different values ​​of T correspond to different rates of score decay; and kh_score is the base score for a specific data channel.

[0143] The data from a specific data channel is scored 100 when it is first stored, and then the score decays exponentially over time. If base_score = 1 and T = 180, the score is halved approximately every six months.

[0144] (2) Data channel weighting

[0145] The first score s under a certain relationship type i can be calculated by the following formula i :

[0146]

[0147] Among them, s ijis the basic score corresponding to a data channel j under a certain relationship type i; w ij is the channel weight corresponding to a data channel j under a relationship type i; m is the number of data channels with basic scores under a relationship type i; n is the number of all data channels under a relationship type i (including both data channels with basic scores and historical data channels not involved in this calculation).

[0148] The basic score is weighted and averaged according to the channel weight to ensure that the "dimension" of the score is at the basic score level; when weighted averaging, the denominator is the sum of the weights of all data channels under all relationship types to ensure the "benefits" of data under data channels with heavy weights; the logarithmic part ensures that when the scores in the weighted average part are consistent, multiple data channels will ultimately score more.

[0149] Refer to the following Figure 3 The logic for determining the relationship scores between the above different entity objects is introduced as an example. Figure 3 As shown, entity types A and B are involved. Each entity type A and B includes multiple entity objects, each corresponding to a dashed box within the entity type. Each entity object includes one or more identification values ​​of an identification type. For example, the third entity object from left to right in entity type A includes identification type A 3_IDm and identification type A 2_IDp; the first entity object from left to right in entity type B includes identification type B 2_ID1 and identification type B 1_ID1. A relationship score p is set between entity objects of different entity types. For example, the relationship score between the third entity object in entity type A and the first entity object in entity type B is p3. Based on the comprehensive scores between the pairwise identifications of different entity types (e.g., w1, w2, ... wi), the comprehensive relationship strength (i.e., relationship score) between the entity objects can be calculated: p = f(w1, w2, ... wi). For example, the relationship score between the third entity object in entity type A and the first entity object in entity type B is p3, which can be the maximum value among w1, w2, and w3.

[0150] For ease of understanding, refer to Figure 4 This paper introduces the association matching logic based on the bipartite graph model. The association matching algorithm based on the bipartite graph model implements the mVn parameter matching of a bipartite graph, that is, a left-side ID can match at most n right-side IDs according to the relationship strength score from high to low, and conversely, a right-side ID can match at most m left-side IDs. Figure 4 As shown, the matching parameters m = 2, n = 2. The logic of the association matching algorithm is as follows:

[0151] Step 1: The left ID expresses its willingness to pair with all connected IDs on the right (initializing the edge state), i.e., the left ID selects all connected right IDs;

[0152] Step 2: The right-side ID selects the left-side ID that it likes the most (i.e., each right-side ID selects the left-side ID with the highest relationship strength score), achieving a preliminary pairing (updating the edge state). In other words, the right-side ID reversely selects the strongest left-side ID.

[0153] Step 3: After completing a pairing process, according to the requirements of mVn, determine whether the nodes on both sides have reached saturation. For excessive matches on the nodes, it is necessary to remove the edges with weaker relationships (determine node saturation and delete excess edges). For example, if there is excess deletion in Step 3, delete the edge between C and e;

[0154] Repeat steps 1-3 until all nodes on either side are saturated, and the calculation ends. For example, in step 4, the left-side ID positively selects all remaining connected right-side IDs; in step 5, the right-side ID negatively selects the strongest remaining left-side ID; in step 6, during over-deletion, delete the edge between B and a, even if there are still over-matched nodes; in step 7, perform over-deletion again, deleting the edge between B and c. At this point, all nodes on either side are saturated, and the calculation ends.

[0155] Note: When the ID on the right is reversed to the optimal matching ID on the left, if there are multiple optimal choices, they will all be preliminarily matched and then eliminated when the node is judged to be saturated. When eliminating, if there are multiple elimination options, the elimination method that achieves more matches will be selected.

[0156] For ease of understanding, refer to Figure 5 This section provides an example of how to split an association network. The ID Mapping association pairs within and between entity types, calculated sequentially using the aforementioned technology and the ID Mapping calculation process, form association networks. For example, ID Mapping between entity types may involve multiple A entity objects, B entity objects, C entity objects, and so on. In order to minimize the association network to the corresponding data entities, the ID Mapping association network needs to be split. The logic for network splitting is as follows:

[0157] Step 1: Set the priority of entity types, such as entity type A > entity type B > entity type C > entity type D, etc. Split a certain association network according to the priority of entity type.

[0158] Step 2: First check whether there is an A entity object in the network. If there is one A entity object, the splitting is complete. If there are multiple A entity objects, check all the associated paths for two A entity objects and cut the edges from small to large according to the strength of the association on the associated path until the two A entity objects can be separated. Split in sequence until all A entity objects in the network are split, and the splitting is complete.

[0159] Step 3: If there is no entity object A in the network, check whether there is entity object B in the network. The splitting logic is the same as above.

[0160] Step 4: If there is no B entity object in the network, check whether there is a C entity object in the network. The splitting logic is the same as above.

[0161] Step 5: Split entity objects at each level in turn.

[0162] The sub-network results obtained by the final split are the final ID Mapping results.

[0163] like Figure 5 As shown, there are two A entity objects (A identifier ID1 and A identifier ID2) in the original network, so these two A entity objects are split; the association path (path1 and path2) between the two A entity objects in the A entity network is determined; the strength of the association on all association paths is checked, and it can be seen that the association between A identifier ID2 and C identifier ID1 in path1 (0.82) is the weakest, and the association between B identifier ID2 and C identifier ID2 in path2 (1.59) is the weakest. Therefore, these two edges are cut to obtain result network 1 and result network 2, which are the final ID Mapping results.

[0164] In summary, the ID Mapping calculation process corresponding to the two scenarios mainly includes:

[0165] ID Mapping within an entity type: The relationship between each identifier within an entity type is scored and calculated based on the data channel, number of discoveries, number of days of discovery, earliest and latest discovery time, and their combination. The relationship between identifiers within an entity type is calculated based on the bipartite graph model and the strength of the relationship for optimal matching. Entity objects within an entity type are split, etc.

[0166] ID Mapping between Entity Types: The relationships between identifiers between entity types are scored and calculated based on the relationship type, data channel, number of discoveries, days of discovery, earliest and latest discovery times, and their combined dimensions. Identifiers between entity types are merged by entity objects. The relationships between entity objects are calculated for optimal matching based on the bipartite graph model and the strength of the relationship. The association network between entity types is split, etc.

[0167] The embodiment of the present invention takes into account the multi-dimensional relationship strength calculation logic; based on the detailed characterization of relationship strength, it implements preliminary association calculation based on the association matching algorithm of the bipartite graph model; designs the object splitting within the entity type and the association network splitting logic between entity types to avoid obvious data conflicts in the same association network and increase data consistency.

[0168] Corresponding to the above-mentioned ID Mapping method, the embodiment of the present invention also provides an ID Mapping device, see Figure 6 The structure diagram of an ID Mapping device shown in FIG. 1 includes:

[0169] A first acquisition module 601 is configured to acquire first basic data corresponding to a relationship type between two identifiers of the same entity type; wherein the relationship type is defined by the identifier types of the two corresponding identifiers, and the first basic data includes relationship data corresponding to multiple data channels;

[0170] A first determining module 602 is configured to determine an ID Mapping library for each entity object of the same entity type based on the first basic data;

[0171] The second acquisition module 603 is used to acquire second basic data of each relationship type between two identifiers of different entity types;

[0172] A first calculation module 604 is configured to calculate, based on the second basic data, comprehensive scores between two identifiers of different entity types;

[0173] A score determination module 605 is configured to determine the relationship score between different entity objects based on the comprehensive score between the two identifiers of different entity types and the ID Mapping library of each entity object within the same entity type;

[0174] A second calculation module 606 is configured to perform a best matching calculation using a bipartite graph model based on the relationship scores between different entity objects to obtain the association relationships between different entity objects;

[0175] The second determining module 607 is configured to determine the IDMapping results between entity types according to the association relationships between different entity objects.

[0176] The ID Mapping device provided in an embodiment of the present invention first performs ID Mapping within an entity type based on the relational data corresponding to multiple data channels, obtains an ID Mapping library for each entity object within the same entity type, then calculates a comprehensive score between pairwise identifiers of different entity types, and combines the ID Mapping library of each entity object within the same entity type to determine the relationship score between different entity objects. It then performs the best matching calculation using a bipartite graph model to obtain the association relationship between different entity objects, and finally determines the ID Mapping result between entity types. This ID Mapping method based on an entityized data system takes into account the relational data corresponding to multiple data channels in complex scenarios, achieves a detailed characterization of the strength of relationships, and implements association calculations through a bipartite graph model, thereby improving the calculation accuracy of ID Mapping in complex scenarios.

[0177] Furthermore, the above-mentioned first determination module 602 is specifically used to: calculate the first score under the corresponding relationship type between each pair of identifiers within the same entity type based on the first basic data; perform the best matching calculation through the bipartite graph model based on the first score under the corresponding relationship type between each pair of identifiers within the same entity type to obtain the association relationship between each identifier within the same entity type; based on the association relationship between each identifier within the same entity type, split the entity objects according to the priority of the identifier type to obtain the ID Mapping library of each entity object within the same entity type.

[0178] Furthermore, the above-mentioned first determination module 602 is also used to: for each group of identifiers within the same entity type, calculate the basic score corresponding to the data channel under the corresponding relationship type of the group identifier based on the relationship data corresponding to each data channel under the corresponding relationship type of the group identifier; and perform weighted calculation on the basic scores corresponding to each data channel under the corresponding relationship type of the group identifier according to the channel weight corresponding to each data channel to obtain the first score under the corresponding relationship type of the group identifier.

[0179] Furthermore, the above-mentioned first determination module 602 is also used to: divide all identifiers corresponding to two identifier types with relationships within the same entity type into two types of nodes of the bipartite graph model according to the identifier type; connect the nodes with relationships in the bipartite graph model into edges, and the weight of the edge is the corresponding first score; delete excess edges according to the saturation corresponding to each type of node and the weight of the edge between nodes, and obtain the association relationship between each identifier corresponding to the bipartite graph model.

[0180] Furthermore, the above-mentioned first calculation module 604 is specifically used to: calculate the second score under each relationship type between pairwise identifiers of different entity types based on the second basic data; perform weighted calculation on the second score under each relationship type between pairwise identifiers of different entity types according to the relationship weight corresponding to each relationship type, and obtain a comprehensive score between pairwise identifiers of different entity types.

[0181] Furthermore, the second determination module 607 is specifically configured to: generate at least one association network based on the association relationships between different entity objects, wherein the nodes of the association network are entity objects, and the weights of the paths connecting the nodes are the corresponding relationship scores; for each association network, determine whether there are multiple target nodes corresponding to the entity type with the highest priority in the association network; if so, split the paths connected to the target nodes in the association network in ascending order of the path weights, and use the multiple sub-networks obtained by the splitting as the ID Mapping results between the entity types; wherein each sub-network contains one target node.

[0182] Furthermore, the above-mentioned relationship data includes one or more of the latest discovery time corresponding to the general data channel, the number of days between the earliest and latest discovery times, the number of discovery days, the number of discoveries, and / or the number of data acquisition time corresponding to the specific data channel from today.

[0183] The ID Mapping device provided in this embodiment has the same implementation principle and technical effects as those of the aforementioned ID Mapping method embodiment. For the sake of brief description, for matters not mentioned in the ID Mapping device embodiment, reference may be made to the corresponding contents in the aforementioned ID Mapping method embodiment.

[0184] like Figure 7 As shown, an embodiment of the present invention provides an electronic device 700, including: a processor 701, a memory 702 and a bus, the memory 702 stores a computer program that can be run on the processor 701, when the electronic device 700 is running, the processor 701 and the memory 702 communicate through the bus, and the processor 701 executes the computer program to implement the above-mentioned IDMapping method.

[0185] Specifically, the memory 702 and processor 701 can be general-purpose memories and processors, which are not specifically limited here.

[0186] An embodiment of the present invention further provides a storage medium storing a computer program that, when executed by a processor, executes the ID Mapping method described in the preceding method embodiments. The storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), RAM, a magnetic disk, or an optical disk.

[0187] In all examples shown and described herein, any specific values ​​should be interpreted as merely exemplary and not limiting, and thus other examples of the exemplary embodiments may have different values.

[0188] The flowcharts and block diagrams in the accompanying drawings show the possible implementation architecture, functions and operations of the methods and computer program products according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the module, program segment or part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0189] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, the indirect coupling or communication connection of the device or unit can be electrical, mechanical or other forms.

[0190] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0191] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0192] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An ID Mapping method, characterized in that: include: Obtaining first basic data under corresponding relationship types between two identifiers of the same entity type; wherein the relationship type is defined by identifier types of the two corresponding identifiers, and the first basic data includes relationship data corresponding to multiple data channels; Determine, based on the first basic data, an IDMapping library for each entity object within the same entity type; Obtain the second basic data under each relationship type between two identifiers of different entity types; Calculate, based on the second basic data, a comprehensive score between two identifiers of different entity types; Determine the relationship score between different entity objects based on the comprehensive score between the two identifiers of the different entity types and the ID Mapping library of each entity object within the same entity type; According to the relationship scores between the different entity objects, the best matching calculation is performed through a bipartite graph model to obtain the association relationship between the different entity objects; According to the association relationship between the different entity objects, the ID Mapping results between the entity types are determined.

2. The ID Mapping method according to claim 1, characterized in that: The step of determining, based on the first basic data, an ID Mapping library for each entity object within the same entity type includes: Calculating, based on the first basic data, a first score for a corresponding relationship type between two identifiers of the same entity type; According to the first scores under the corresponding relationship types between the two identifiers in the same entity type, the best matching calculation is performed through the bipartite graph model to obtain the association relationship between the identifiers in the same entity type; According to the association relationship between the identifiers within the same entity type, the entity objects are split according to the priority of the identifier type to obtain an ID Mapping library for each entity object within the same entity type.

3. The ID Mapping method according to claim 2, characterized in that: The step of calculating, based on the first basic data, a first score for a corresponding relationship type between two identifiers within the same entity type includes: For each group of identifiers within the same entity type, the basic score corresponding to each data channel under the corresponding relationship type of the group identifier is calculated based on the relationship data corresponding to each data channel under the corresponding relationship type of the group identifier; According to the channel weight corresponding to each data channel, the basic scores corresponding to each data channel under the corresponding relationship type of the group identifier are weightedly calculated to obtain the first score under the corresponding relationship type of the group identifier.

4. The ID Mapping method according to claim 2, characterized in that: The step of performing a best match calculation using a bipartite graph model based on the first scores under the corresponding relationship types between the two identifiers within the same entity type to obtain the association relationship between the identifiers within the same entity type includes: All identifiers corresponding to two identifier types with a relationship within the same entity type are divided into two types of nodes in the bipartite graph model according to the identifier type; Connecting nodes having relationships in the bipartite graph model into edges, where the weights of the edges are the corresponding first scores; According to the saturation corresponding to each type of node and the weight of the edge between nodes, excess edges are deleted to obtain the association relationship between the identifiers corresponding to the bipartite graph model.

5. The ID Mapping method according to claim 1, characterized in that: The step of calculating, based on the second basic data, a comprehensive score between two identifiers of different entity types includes: Calculating, based on the second basic data, a second score for each relationship type between two identifiers of different entity types; According to the relationship weight corresponding to each relationship type, the second scores under each relationship type between the two identifiers of the different entity types are weightedly calculated to obtain a comprehensive score between the two identifiers of the different entity types.

6. The ID Mapping method according to claim 1, characterized in that: The step of determining the ID Mapping results between entity types based on the association relationships between the different entity objects includes: generating at least one association network based on the association relationships between the different entity objects, wherein the nodes of the association network are entity objects, and the weights of the paths connecting the nodes are corresponding relationship scores; For each of the associated networks, determining whether there are multiple target nodes corresponding to the entity type with the highest priority in the associated network; If so, the paths connected to the target nodes in the associated network are split in ascending order of path weights, and the multiple sub-networks obtained by splitting are used as ID Mapping results between entity types; wherein each sub-network contains one target node.

7. The ID Mapping method according to claim 1, characterized in that: The relationship data includes one or more of the latest discovery time corresponding to the general data channel, the number of days between the earliest and latest discovery times, the number of discovery days, the number of discoveries, and / or the number of days since the data acquisition time corresponding to the specific data channel.

8. An ID Mapping device, characterized in that: include: A first acquisition module is configured to acquire first basic data under a corresponding relationship type between two identifiers within the same entity type; wherein the relationship type is defined by the identifier types of the two corresponding identifiers, and the first basic data includes relationship data corresponding to multiple data channels; A first determining module, configured to determine an IDMapping library for each entity object within the same entity type according to the first basic data; A second acquisition module is used to acquire second basic data under each relationship type between two identifiers of different entity types; A first calculation module is configured to calculate, based on the second basic data, a comprehensive score between two identifiers of different entity types; A score determination module, configured to determine a relationship score between different entity objects based on a comprehensive score between two identifiers of the different entity types and an ID Mapping library of each entity object within the same entity type; A second calculation module is configured to perform a best matching calculation using a bipartite graph model based on the relationship scores between the different entity objects to obtain the association relationships between the different entity objects; The second determining module is used to determine the IDMapping results between entity types according to the association relationship between the different entity objects.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, wherein: When the processor executes the computer program, the ID Mapping method according to any one of claims 1 to 7 is implemented.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the IDMapping method according to any one of claims 1 to 7 is executed.

Citation Information

Patent Citations

  • Multivariate relation extraction method and extraction system based on multi-model fusion

    CN114925693A

  • Systems and methods for near-real or real-time contact tracing

    US20170024531A1