Entity alignment method and system for dispersed storage data of large-scale mechanism
By calculating attribute information entropy and using the attention mechanism of a single-layer graph neural network for entity alignment, the problem of data redundancy and inconsistency in large organizations is solved, and the accuracy of data alignment and overall data quality are improved.
Patent Information
- Application Number
- CN202510250231.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Due to the decentralization of geographical and logical structures, large organizations have led to insufficient data storage dispersion and correlation between entities, resulting in data redundancy, inconsistency, incompleteness and information loss, affecting the identification and management of important data.
A solid alignment method for dispersed storage data for large-scale institutions is adopted, and the initial weight is assigned by calculating the attribute information entropy, and a single-layer graph neural network is used to embed attributes, and the cosine similarity of the entity feature vector is calculated in combination with the attention mechanism to achieve solid alignment.
Improves the accuracy of data alignment, reduces data redundancy, enhances data consistency and integrity, and supports the identification and management of important data.
Smart Images

Figure CN120105115A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data fusion processing, and in particular to an entity alignment method and system for decentralized storage data of large-scale organizations. Background Art
[0002] As the country pays more and more attention to data security, identifying and managing important data has become a top priority, especially for large organizations with huge data accumulation. However, since large organizations usually have subordinate organizations dispersed in geography and logical structure, each organization has an independent management system, resulting in the decentralization of data storage and the lack of correlation between entities. The complexity of data entity attributes further exacerbates the problem. Different subsidiaries and departments have redundancy, inconsistency, incompleteness and missing information in the data content of each organization due to business independence, geographical independence, and logical structure independence. This has brought serious obstacles to the identification of important data. In the identification process, the following problems will occur: redundant data is classified and graded multiple times, there are obstacles in the classification and grading of important data, and the lack of data information leads to classification and grading errors, which seriously affect the management and protection of important data.
[0003] Raoufi et al. compared the performance of different models on real-world datasets and found that many existing models performed poorly on real datasets. One of the important reasons is the difference in the distribution of structural information and attribute information between real data and benchmark data. Public datasets such as DBpedia-YAGO and DBpedia-Freebase select entities with rich structural relationships during the construction process, which means that the model can use a lot of structural information, such as the connection mode, path, relationship type (directed relationship, undirected relationship, etc.), in-degree, out-degree, adjacent subgraphs and shared neighbors between entities, and perform entity alignment tasks by comparing the structural relationships of entities. This is also a current research hotspot for entity alignment. However, the data accumulated in actual production and life may have only a small amount of structural information, or even no structural information, forming data orphans, but it has a large amount of attribute information. Entities have a large number of attributes and attribute values, which is significantly different from public datasets.
[0004] Most existing entity alignment methods rely on the structural information of entities and perform alignment through subgraph matching. However, the lack of structural information in decentralized storage data leads to poor alignment results. Although decentralized storage data lacks structural information, it has rich attribute information, so using attribute information for entity alignment becomes a better choice. Summary of the invention
[0005] In view of the above-mentioned deficiencies in the prior art, the present invention provides an entity alignment method and system for decentralized storage of data in large-scale organizations, which solves the problems of insufficient entity alignment accuracy and excessive data redundancy in the prior art.
[0006] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is: an entity alignment method for distributed storage data of large-scale organizations, comprising the following steps: S1. Assign an initial weight to each entity attribute in the known entity seed, and obtain the initial weight matrix composed of the initial weights of all attributes of the entity; the entity attribute is represented as a triple (E, A, V ), E is the entity, A is the attribute, V is the attribute value; S2, input the data sources KG1 and KG2 to be aligned into a single-layer graph neural network for attribute embedding, and obtain the entity feature vectors of KG1 and KG2; S3. Calculate the cosine similarity between the entity feature vectors of KG1 and KG2 to obtain the alignment result.
[0007] Further: Step S1 specifically includes the following sub-steps: S11. Calculate the probability distribution of attribute values , whose expression is:
[0008] in, Indicates that the value of attribute A is V The number of times N Take the total number of all attribute values for attribute A; S12. Calculate the information entropy of attribute A , whose expression is:
[0009] in, n Represents the total number of values of attribute A, The first i A value, Indicates that attribute A takes the i Value probability; S13. Information entropy Perform normalization processing to obtain normalized information entropy; S14. Initialize each attribute weight of entity E to the normalized information entropy, and obtain an initial weight matrix composed of initial weights of all attributes of entity E.
[0010] Further: the attribute embedding in step S2 specifically includes the following steps: S21. Use the pre-trained BERT model to embed the attribute value vector of the data source to obtain the attribute-attribute value vector ;in Represents attribute characteristics, Represents the attribute value characteristics, Indicates splicing; S22. Calculate attention score , whose expression is:
[0011] in, Represents the comprehensive feature vector formed by the concatenation of the initial entity features and attribute features. represents a linear mapping, For variable parameters, is the bias vector; is the activation function; is the intermediate parameter; S23. Calculate attention weight , whose expression is:
[0012] in, is the initial weight matrix obtained in step S1, is the scaling factor; S24. Fusion of attention score and attention weight to get total attention , whose expression is:
[0013] in It means fusion; S25. Use total attention to weight the attribute-attribute value vector to obtain the entity feature vector , whose expression is:
[0014] in, represents the initial entity features; Indicates splicing.
[0015] The present invention also provides a system based on an entity alignment method for large-scale organization distributed storage data, comprising: Initialization weighting module, used to assign initial weights to each entity attribute in the known entity seed, and obtain an initial weight matrix consisting of initial weights of all attributes of the entity; An input module, used for receiving and transmitting the data sources KG1 and KG2 to be aligned; A single-layer graph neural network module is used to embed the attributes of KG1 and KG2 to obtain the entity feature vectors of KG1 and KG2; The output module is used to calculate the cosine similarity between the entity feature vectors of KG1 and KG2 to obtain the alignment result.
[0016] Further: The single-layer graph neural network module includes: BERT model, used to embed attribute value vectors of data sources to obtain attribute-attribute value vectors ; The fused attention mechanism encoder is used to calculate and fuse the attention score and attention weight to obtain the total attention; The attention weighting unit is used to weight the attribute-attribute value vector using the total attention to obtain the entity feature vector.
[0017] Further: The fusion attention mechanism encoder calculates the alignment loss between the embedding of the source entity and the embedding of the target entity when calculating the attention weight, and then uses the back-propagation gradient descent method to update the attention weight matrix to minimize the loss value.
[0018] The beneficial effects of the present invention are: 1. The present invention calculates the attribute information entropy and then assigns initial weights to quantify the contribution of each attribute to the entity. It can quickly distinguish the importance of attributes in the initial stage and assign higher weights to attributes with higher information content, thereby better allocating the algorithm's attention to the attributes and further improving the accuracy of data alignment.
[0019] 2. The present invention integrates attention scores and attention weights, comprehensively considers local and global information, fully reflects the important relationship between attributes and entities, and combines local and global perspectives to characterize the importance of different attributes in the data alignment process, which can more comprehensively characterize entities.
[0020] 3. The single-layer graph neural network used in the present invention can avoid information interference from secondary neighbor nodes, focus on capturing attribute relationships directly connected to entity nodes, and improve the performance and generalization ability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 A flow chart of the method proposed by the present invention; Figure 2 It is a schematic diagram of the structure of the system proposed by the present invention. DETAILED DESCRIPTION
[0022] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.
[0023] like Figure 1 As shown, the present invention provides an entity alignment method for distributed storage data of large-scale institutions, comprising the following steps: S1. Assign an initial weight to each entity attribute in the known entity seed, and obtain the initial weight matrix composed of the initial weights of all attributes of the entity; the entity attribute is represented as a triple (E, A, V ), E is the entity, A is the attribute, V is the attribute value.
[0024] Step S1 specifically includes the following sub-steps: S11. Calculate the probability distribution of attribute values , whose expression is:
[0025] in, Indicates that the value of attribute A is V The number of times N Take the total number of all attribute values for attribute A; S12. Calculate the information entropy of attribute A In information theory, information entropy is a measure of the uncertainty of a random variable. The size of information entropy reflects the amount of information contained in the random variable. The larger the information entropy, the higher the uncertainty of the random variable and the greater the amount of information contained. According to the definition of information entropy, the information entropy of attribute A is It is expressed as:
[0026] in, n Represents the total number of values of attribute A, The first i A value, Indicates that attribute A takes the i Value probability.
[0027] In data alignment, information entropy is used to preliminarily measure the contribution of attribute A in characterizing entity E. The smaller the The larger the value of attribute A is, the more diverse the value of attribute A is, and the greater the contribution of attribute A to entity E. On the contrary, The larger the The smaller the value of attribute A is, the smaller the contribution of attribute A to entity E is. For example, for a hotel entity, the contribution of attribute "address" to a hotel is greater than that of attribute "star" (the value range is from 1 to 5, indicating the star rating of the hotel).
[0028] S13, through After measuring the amount of information of attribute A, the information entropy Normalization is performed to obtain the normalized information entropy.
[0029] S14. Initialize each attribute weight of entity E to the normalized information entropy, and obtain the initial weight matrix composed of the initial weights of all attributes of entity E .
[0030] S2. Input the data sources KG1 and KG2 to be aligned into a single-layer graph neural network for attribute embedding respectively to obtain the entity feature vectors of KG1 and KG2.
[0031] The attribute embedding in step S2 specifically includes the following steps: S21. Use the pre-trained BERT model to embed the attribute value vector of the data source to obtain the attribute-attribute value vector ;in Represents attribute characteristics, Represents the attribute value characteristics, Indicates splicing.
[0032] S22. Calculate attention score First, use the linear layer to concatenate the initial entity features and attribute features to form a comprehensive feature vector Perform linear mapping to obtain the initial attention score of attribute A . Linear layers contain variable parameters and the bias vector , where the variable parameter is used to learn the importance of attribute features to entity features. Then use the activation function Perform nonlinear transformation to get the final attention score , so that it can learn more complex relationships and have stronger fitting capabilities. and The expression is:
[0033] in, Represents the comprehensive feature vector formed by the concatenation of the initial entity features and attribute features. represents a linear mapping, For variable parameters, is the bias vector; is the activation function; is the intermediate parameter.
[0034] In the entity encoding process, the attention score reflects the relationship between the attribute and a specific entity. This relationship is local and is only used for the encoding of the entity. In the entity alignment task, it is also necessary to fully consider the common importance of the attribute to all entities, so the attention weight is introduced. , further increase the weight of public attributes from a global perspective and weaken the weight of private attributes to enhance the comparability between entity vectors. Assume that there is an entity E1 and its potential matching object entity E2, whose attribute sets are A1 and A2 respectively, the intersection of the attribute sets of E1 and E2 is the public attribute A, the private attribute set of E1 is A1, and the private attribute set of E2 is A2. In the attribute-based entity alignment model, the weight of public attribute A in entity encoding should be higher, and its comparability is stronger, while the weight of private attributes A1 and private attributes A2 should be lower, and their comparability is lower. If there is no intersection between A1 and A2, it will be difficult to align E1 and E2. For example, if more entities have the "location" attribute and fewer entities have the "fare" attribute, the "location" attribute is more public. Therefore, the attention weight of the "location" attribute from a global perspective will be higher, and more attention will be paid to the "location" attribute shared by the entities during encoding, thereby enhancing the comparability between different entities and further improving the alignment accuracy.
[0035] S23. Calculate attention weight , whose expression is:
[0036] in, is the initial weight matrix obtained in step S1, is the scaling factor; the attention weight matrix The shape of is attr_num×1, where attr_num represents the number of attribute types in the dataset. When calculating the attention weight using the fusion attention mechanism encoder, the alignment loss between the embedding of the source entity and the embedding of the target entity is calculated, and then the attention weight matrix is updated using the gradient descent method of backpropagation to minimize the loss value, so that the attention weight is gradually stabilized.
[0037] S24. Fusion of attention score and attention weight to get total attention , whose expression is:
[0038] in Representation fusion. Total attention integrates the importance of attributes to specific entities and the entire dataset, taking into account both global and local information, and comprehensively reflects the important relationship between attributes and entities.
[0039] S25. Use total attention to weight the attribute-attribute value vector to obtain the entity feature vector , whose expression is:
[0040] in, represents the initial entity features; Indicates splicing.
[0041] In the entity encoding process, the attention score is responsible for measuring the importance of the attribute to a single entity, focusing on the local features of each entity; the attention weight is responsible for measuring the general importance of the attribute in the entire dataset, taking into account the correlation between all entities and attributes. The fusion of the two can more comprehensively and accurately reflect the attributes and attribute information contained in the entity, thereby better encoding the entity features, achieving more accurate entity encoding, and improving the robustness and accuracy of entity alignment. At the same time, this method enables the system to better handle the differences between different entities and attributes in the dataset, so as to better meet the needs of practical applications. Therefore, the encoding method that combines attention scores and attention weights not only improves the performance of the entity alignment task, but also enhances the generalization ability of the system, enabling it to achieve better results in various practical scenarios.
[0042] S3. Calculate the cosine similarity between the entity feature vectors of KG1 and KG2 to obtain the alignment result.
[0043] like Figure 2 As shown, the present invention also provides a system for implementing the above method, comprising: Initialization weighting module, used to assign initial weights to each entity attribute in the known entity seed, and obtain an initial weight matrix consisting of initial weights of all attributes of the entity; An input module, used for receiving and transmitting the data sources KG1 and KG2 to be aligned; The single-layer graph neural network module is used to embed the attributes of KG1 and KG2 to obtain the entity feature vectors of KG1 and KG2; the single-layer graph neural network module includes: BERT model, which is used to embed the attribute value vector of the data source to obtain the attribute-attribute value vector ; The fused attention mechanism encoder is used to calculate and fuse the attention score and attention weight to obtain the total attention; the attention weighting unit is used to weight the attribute-attribute value vector using the total attention to obtain the entity feature vector. Among them, the fused attention mechanism encoder calculates the alignment loss between the embedding of the source entity and the embedding of the target entity when calculating the attention weight, and then uses the gradient descent method of backpropagation to update the attention weight matrix to minimize the loss value.
[0044] The single-layer graph neural network used by the system has a relatively simple structure and low complexity. It is not prone to overfitting during training and is easy to generalize to unseen data. At the same time, the single-layer graph neural network focuses more on capturing the relationship between directly connected nodes without over-deepening consideration of secondary neighbor nodes. In attribute-based entity alignment tasks, the important information of a node is usually concentrated on its directly connected attribute nodes. Since the single-layer network focuses more on these direct relationships, it can capture the local features of the node more accurately. This effective capture of local information enables the single-layer network to better handle the correlation between nodes, improving the performance and generalization ability of the system.
[0045] The output module is used to calculate the cosine similarity between the entity feature vectors of KG1 and KG2 to obtain the alignment result.
[0046] In order to verify the beneficial effects of the present invention, the following comparative experiments were conducted.
[0047] Comparative experiment 1: The two data sources DataSource1 and DataSource2 in the dataset CTD are aligned using the proposed method, BERT-INT and AttrGNN respectively. The two data sources (DataSource1, DataSource2) are two subordinate institutions of a large-scale organization. There are problems such as data redundancy, data islands and incomplete data between the dispersedly stored data. The dataset CTD covers entities such as hotels, homestays, farmhouses, museums, memorial halls, natural scenic spots, and historical sites in Yibin City; it contains multiple attributes such as ratings, traffic conditions, ticket prices, number of reviews, addresses, and contact numbers; according to the association of entities, relationship information such as chain, proximity, and the same type is added. The details of the CTD dataset are shown in Table 1.
[0048] Table 1
[0049] The experimental results are shown in Table 2.
[0050] Table 2
[0051] According to the above experimental results, in the CTD dataset, the proposed method performs best in all indicators, with the highest Hits@1, Hits@10 and MRR, and the lowest MR, indicating that the present invention has very strong performance in entity alignment tasks, can accurately find matching entity pairs, and is closer to the top in the ranking. In terms of Hits@1 and Hits@10 indicators, the present invention is ahead of other models, indicating that in the entity alignment task, the present invention can find a match in the first potential matching entity with an accuracy of 90.45%, which is 10.21% higher than the current best performing model, and find a match in the first 10 potential matching entities with an accuracy of 99.60%, which is 5.04% higher than the current best performing model. The MR indicator of the present invention is 1.28, which is much lower than other models, indicating that the correct entity pair can be placed in a more advanced position.
[0052] Comparative experiment 2: In order to verify the effect of assigning initial weights to entity attributes in the present invention, three groups of ablation experiments are carried out on attribute initialization weighting. The experimental results are shown in Table 3.
[0053] Table 3
[0054] According to the comparison of experimental results, the method of using information entropy to weight attributes can significantly improve the performance of the system in terms of Hits@1, MR and MRR indicators, which is better than uninitialized weighting in all aspects, and Hits@10 is the same as or slightly better than uninitialized weighting. This shows that after using information entropy to initialize attribute weights, it helps to improve the accuracy of entity alignment, the average ranking (MR) of correct results in the sorting results is lower, and matching entities can be searched more quickly.
[0055] In summary, the present invention can complete the entity alignment task more effectively, and provides a powerful data fusion solution for large-scale organizations, thereby reducing data redundancy, enhancing data consistency and integrity, and greatly supporting important data identification and management.
Claims
1. An entity alignment method for distributed storage data of large-scale organizations, characterized in that: The following steps are involved: S1. Assign an initial weight to each entity attribute in the known entity seed, and obtain the initial weight matrix composed of the initial weights of all attributes of the entity; the entity attribute is represented as a triple (E, A, V ), E is the entity, A is the attribute, V is the attribute value; S2, input the data sources KG1 and KG2 to be aligned into a single-layer graph neural network for attribute embedding, and obtain the entity feature vectors of KG1 and KG2; S3. Calculate the cosine similarity between the entity feature vectors of KG1 and KG2 to obtain the alignment result.
2. According to claim 1, a method for aligning entities for decentralized storage of data in large-scale organizations, characterized in that: Step S1 specifically includes the following sub-steps: S11. Calculate the probability distribution of attribute values , whose expression is: in, Indicates that the value of attribute A is V The number of times N Take the total number of all attribute values for attribute A; S12. Calculate the information entropy of attribute A , whose expression is: in, n Represents the total number of values of attribute A, The first i A value, Indicates that attribute A takes the i Value probability; S13. Information entropy Perform normalization processing to obtain normalized information entropy; S14. Initialize each attribute weight of entity E to the normalized information entropy, and obtain an initial weight matrix composed of initial weights of all attributes of entity E.
3. The entity alignment method for distributed storage data of large-scale institutions according to claim 1 is characterized in that: The attribute embedding in step S2 specifically includes the following steps: S21. Use the pre-trained BERT model to embed the attribute value vector of the data source to obtain the attribute-attribute value vector ;in Represents attribute characteristics, Represents the attribute value characteristics, Indicates splicing; S22. Calculate attention score , whose expression is: in, Represents the comprehensive feature vector formed by the concatenation of the initial entity features and attribute features. represents a linear mapping, For variable parameters, is the bias vector; is the activation function; is the intermediate parameter; S23. Calculate attention weight , whose expression is: in, is the initial weight matrix obtained in step S1, is the scaling factor; S24. Fusion of attention score and attention weight to get total attention , whose expression is: in It means fusion; S25. Use total attention to weight the attribute-attribute value vector to obtain the entity feature vector , whose expression is: in, represents the initial entity features; Indicates splicing.
4. A system based on the entity alignment method for distributed storage data of large-scale organizations according to any one of claims 1 to 3, characterized in that: include: Initialization weighting module, used to assign initial weights to each entity attribute in the known entity seed, and obtain an initial weight matrix consisting of initial weights of all attributes of the entity; An input module, used for receiving and transmitting the data sources KG1 and KG2 to be aligned; A single-layer graph neural network module is used to embed the attributes of KG1 and KG2 to obtain the entity feature vectors of KG1 and KG2; The output module is used to calculate the cosine similarity between the entity feature vectors of KG1 and KG2 to obtain the alignment result.
5. The system according to claim 4, characterized in that The single-layer graph neural network module includes: BERT model, used to embed attribute value vectors of data sources to obtain attribute-attribute value vectors ; The fused attention mechanism encoder is used to calculate and fuse the attention score and attention weight to obtain the total attention; The attention weighting unit is used to weight the attribute-attribute value vector using the total attention to obtain the entity feature vector.
6. The system according to claim 5, characterized in that The fused attention mechanism encoder calculates the alignment loss between the embedding of the source entity and the embedding of the target entity when calculating the attention weights, and then uses the gradient descent method of backpropagation to update the attention weight matrix to minimize the loss value.