An entity alignment method and device based on semantic and structural sampling strategy
Patent Information
- Application Number
- CN202311596284.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-11-27
AI Technical Summary
现有方式一是采用端到端实体匹配的神经网络模型实现实体对齐的模型,但是需要依赖大量种子对齐数据作为训练数据,而这些种子对齐数据的标注成本非常高;现有方式二是专注于具有文字属性的表格数据,其提出相似性度量或深度学习模型来比较文字属性,并生成主动学习的特征向量
[0016]本申请提供的基于语义与结构采样策略的实体对齐方法,包括提取未标记数据池中的所有未标记实体,将未标记实体的上一次迭代得到的边界不确定性数值、未标记实体链接的其他实体的上一次迭代的边界不确定性数值,以及控制未标记实体的不确定性和所链接的其他实体不确定性的比重值,进行迭代计算得到未标记实体的边界不确定性的数值,选取预设数量的未标记实体进行标注更新标记数据集,利用更新后的标记数据集实体对齐模型进行训练,重复上述步骤,直到实体对齐模型满足预设训练结果。本申请利用语义表征模型以及实体对齐模型,对标注数据进行采样,优先标注对知识图谱融合更有价值的数据。在标注完一个批次数据后,更新语义表征模型和实体对齐模型,提升采样策略的效果,再进行下一次采样。不断迭代上述过程,在有限的预算下,可以实现更好的实体对齐效果。
Smart Images

Figure QLYQS_1 
Figure QLYQS_2 
Figure QLYQS_3
Abstract
Description
Technical Field
[0001] This application relates to the field of text processing technology, and in particular to an entity alignment method and apparatus based on semantic and structural sampling strategies. Background Technology
[0002] Currently, identifying equivalent entities from different knowledge graphs for entity alignment in knowledge graph fusion is a key technology. Existing approaches include: 1) using end-to-end entity matching neural network models for entity alignment, but this relies heavily on seed alignment data for training, which is extremely costly to annotate; and 2) focusing on tabular data with textual attributes, proposing similarity measures or deep learning models to compare textual attributes and generate actively learned feature vectors. However, entities in knowledge graphs differ significantly from those in databases, and different knowledge graphs are often represented by heterogeneous schemas. Therefore, how to generate entity alignment models with lower annotation costs and higher efficiency is a pressing technical problem that needs to be solved. Summary of the Invention
[0003] In order to generate entity alignment models with less annotation cost and higher efficiency, this application provides an entity alignment method and apparatus based on semantic and structural sampling strategies.
[0004] Firstly, this application provides an entity alignment method based on a semantic and structural sampling strategy, the method comprising:
[0005] Extract all unlabeled entities from the unlabeled data pool;
[0006] The boundary uncertainty value obtained from the previous iteration of the unlabeled entity, the boundary uncertainty value of the other entities linked to the unlabeled entity from the previous iteration, and the weighting values of the uncertainty of the unlabeled entity and the uncertainty of the other linked entities are input into the iterative algorithm for calculation until the preset iteration result is met, and the boundary uncertainty value of the unlabeled entity is obtained.
[0007] Based on the boundary uncertainty values of all the unlabeled entities, a preset number of the unlabeled entities are selected as entities to be labeled, and the labeled data is updated to the labeled dataset.
[0008] After training the entity alignment model to be trained using the updated labeled dataset, the unlabeled data pool is updated, and the above steps are repeated until the entity alignment model meets the preset training results.
[0009] Secondly, this application also provides an entity alignment device based on a semantic and structural sampling strategy, the device comprising:
[0010] The selection module is used to extract all unlabeled entities from the unlabeled data pool;
[0011] The iteration module is used to input the boundary uncertainty value obtained from the previous iteration of the unlabeled entity, the boundary uncertainty value of the previous iteration of other entities linked to the unlabeled entity, and the weighting value of the uncertainty of the unlabeled entity and the uncertainty of the linked other entities into the iteration algorithm for calculation until the preset iteration result is met, so as to obtain the boundary uncertainty value of the unlabeled entity.
[0012] The annotation module is used to select a preset number of unlabeled entities as entities to be annotated based on the boundary uncertainty values of all the unlabeled entities, and update the labeled data to the labeled dataset.
[0013] The training module is used to train the entity alignment model to be trained using the updated labeled dataset, then update the unlabeled data pool, and repeat the above steps until the entity alignment model meets the preset training results.
[0014] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the entity alignment method based on the semantic and structural sampling strategy described in any one of the first aspects.
[0015] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the entity alignment method based on the semantic and structural sampling strategy described in any one of the first aspects.
[0016] The entity alignment method based on semantic and structural sampling strategies provided in this application includes extracting all unlabeled entities from an unlabeled data pool; iteratively calculating the boundary uncertainty value of the unlabeled entity based on the boundary uncertainty values of other entities linked to it from the previous iteration, as well as the weighting values controlling the uncertainty of the unlabeled entity and the uncertainty of other linked entities; selecting a preset number of unlabeled entities for labeling and updating the labeled dataset; training the entity alignment model using the updated labeled dataset; and repeating the above steps until the entity alignment model meets the preset training results. This application utilizes a semantic representation model and an entity alignment model to sample labeled data, prioritizing the labeling of data more valuable for knowledge graph fusion. After labeling a batch of data, the semantic representation model and entity alignment model are updated to improve the effectiveness of the sampling strategy before the next sampling. By continuously iterating the above process, better entity alignment results can be achieved within a limited budget. Attached Figure Description
[0017] The accompanying drawings, which form part of this specification, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:
[0018] Figure 1 This is a flowchart illustrating the entity alignment method based on semantic and structural sampling strategies provided in this application embodiment;
[0019] Figure 2 This is a schematic diagram of the graph architecture to be fused in an entity alignment method based on semantic and structural sampling strategies provided in another embodiment of this application;
[0020] Figure 3 This is a schematic diagram of the training framework for an entity alignment method based on semantic and structural sampling strategies provided in another embodiment of this application;
[0021] Figure 4 This is a flowchart illustrating an entity alignment method based on semantic and structural sampling strategies provided in another embodiment of this application;
[0022] Figure 5 This is a schematic diagram of a module of an entity alignment device based on a semantic and structural sampling strategy provided in another embodiment of this application. Detailed Implementation
[0023] The present invention will now be described in detail with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0024] The following detailed description is exemplary and intended to provide further detailed explanation of the invention. Unless otherwise specified, all technical terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this invention is for describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention.
[0025] Example 1:
[0026] The following will be combined with the appendix Figure 1 This application provides a detailed description of the entity alignment method based on semantic and structural sampling strategies provided in its embodiments, including the following steps:
[0027] S1. Extract all unlabeled entities from the unlabeled data pool. S2. Input the boundary uncertainty values of the unlabeled entity from the previous iteration, the boundary uncertainty values of other entities linked to the unlabeled entity from the previous iteration, and the weighting values controlling the uncertainty of the unlabeled entity and the uncertainty of other linked entities into the iterative algorithm for calculation until the preset iteration result is met, thus obtaining the boundary uncertainty value of the unlabeled entity.
[0028] S3. Based on the boundary uncertainty values of all unlabeled entities, select a preset number of unlabeled entities as entities to be labeled and update the labeled data to the labeled dataset.
[0029] S4. After training the entity alignment model to be trained using the updated labeled dataset, update the unlabeled data pool and return to step S1 until the entity alignment model to be trained meets the preset training results.
[0030] Based on the above embodiments, specifically, step S2 includes:
[0031] S21. Determine the unmarked entity e. i The boundary uncertainty value f obtained in the t-th iteration t (e i The boundary uncertainty value f obtained from the (t-1)th iteration and the (t-1)th iteration t-1 (e i Does the difference satisfy the preset iteration result?
[0032] If so, proceed to step S22;
[0033] Otherwise, proceed to step S23 and return to step S21;
[0034] S22, the i-th unmarked entity e i The numerical value of the boundary uncertainty obtained in the t-th iteration, f t (e i ) is set to a numerical value representing the boundary uncertainty of the unmarked entity;
[0035] S23, move the i-th unlabeled entity e i The boundary uncertainty value f obtained in the t-th iteration t (e i The unmarked entity e i Other entities linked to e j The boundary uncertainty value f obtained in the t-th iteration t (e j ), and control of the unmarked entity e i The weighting α of the uncertainty of the entity and the uncertainty of other linked entities is entered into the formula. In the process, the i-th unlabeled entity e is obtained. i The numerical value of the boundary uncertainty f obtained in the (t+1)th iteration t+1 (e i );
[0036] in, It is the unmarked entity e i Other linked entity sets, t≥1, i≥1, j≥1.
[0037] S24. Determine whether all unlabeled entities have been iterated through to obtain the value of boundary uncertainty;
[0038] If not, then the next unlabeled entity whose boundary uncertainty value has not been obtained will be taken as the unlabeled entity e. i Then return to step S21. Based on the above embodiments, specifically, step S2 further includes:
[0039] Calculate unlabeled entity e i and each entity e in the graph to be merged m similarity F(e) i ,e m ).
[0040] For all similarities F(e) i ,e m After sorting, a preset number of similarities from high to low are selected as the selected similarities to calculate the mean similarity.
[0041] The mean similarity and the similarity F(e) i ,e m Input the variance formula to obtain the unlabeled entity e. i The initial boundary uncertainty value f0(e i ).
[0042] Based on the above embodiments, specifically, the unlabeled entity e is calculated. i and each entity e in the graph to be merged m similarity F(e) i ,e m Specifically, this includes:
[0043] Unmarked entity e i and entity e m Input entity alignment model to obtain matching score F EA (e i ,e m ).
[0044] Unmarked entity e i and entity e mThe semantic similarity F obtained from the input semantic representation model S (e i ,e m ).
[0045] Based on the matching score F EA (e i ,e m semantic similarity F S (e i ,e m ) and preset weight values, to obtain the unlabeled entity e i and each entity e in the graph to be merged m similarity F(e) i ,e m ).
[0046] Based on the above embodiments, specifically, the unlabeled entity e... i and entity e m The semantic similarity F obtained from the input semantic representation model S (e i ,e m Specifically, this includes:
[0047] Unmarked entity e i Input the semantic representation model Sbert model and obtain the unlabeled entity e. i The representation vector.
[0048] Entity e m Input the semantic representation model Sbert model to obtain entity e m The representation vector.
[0049] Calculate unlabeled entity e i The representation vector and entity e m The representation vector is used to obtain the unlabeled entity e. i and each entity e in the graph to be merged m similarity F(e) i ,e m ).
[0050] Based on the above embodiments, specifically, step S3 includes sorting the boundary uncertainty values of all unlabeled entities and labeling a preset number of unlabeled entities at the top of the sort as entities to be labeled.
[0051] The entity alignment method based on semantic and structural sampling strategies provided in Embodiment 1 of this application includes...
[0052] This paper extracts all unlabeled entities from the unlabeled data pool. It iteratively calculates the boundary uncertainty value of each unlabeled entity by taking the boundary uncertainty values of the unlabeled entity from the previous iteration, the boundary uncertainty values of other entities linked to the unlabeled entity from the previous iteration, and the weighting values controlling the uncertainty of the unlabeled entity and the uncertainty of other linked entities. A predetermined number of unlabeled entities are selected for labeling and updating the labeled dataset. The entity alignment model is then trained using the updated labeled dataset. This process is repeated until the entity alignment model meets the predetermined training results. This application utilizes a semantic representation model and an entity alignment model to sample labeled data, prioritizing the labeling of data more valuable for knowledge graph fusion. After labeling a batch of data, the semantic representation model and entity alignment model are updated to improve the effectiveness of the sampling strategy before the next sampling. By continuously iterating this process, better entity alignment results can be achieved within a limited budget.
[0053] Example 2:
[0054] The following will be combined with the appendix Figures 2 to 4 This paper provides a detailed description of the entity alignment method based on semantic and structural sampling strategies provided in this application, and its application in a real-world environment, including the following steps:
[0055] 110. Initialize the graph database environment and prepare the required standard data.
[0056] Specifically, the database environment can be Neo 4j or other database environments, which will not be elaborated in this embodiment. The standard data can be, for example... Figure 2 The standard knowledge graph architecture shown.
[0057] 120. Select the entity alignment model.
[0058] Specifically, in this implementation, BootEA is selected, and entity e is defined. i and entity e j The matching score F returned by the entity alignment model EA (e i ,e j ).
[0059] 130. The query system extracts all unlabeled entities from the unlabeled data pool, performs iterative calculations on each unlabeled entity, and obtains the numerical value of the boundary uncertainty of the unlabeled entity.
[0060] Based on the above embodiments, step 130 specifically includes:
[0061] 131. Determine if entity e is unlabeled. i The boundary uncertainty value f obtained in the t-th iteration t (ei The boundary uncertainty value f obtained from the (t-1)th iteration and the (t-1)th iteration t-1 (e i Does the difference satisfy the preset iteration result?
[0062] If so, proceed to step 132;
[0063] Otherwise, proceed to step 133 and return to step 131;
[0064] 132. The i-th unlabeled entity e i The numerical value of the boundary uncertainty obtained in the t-th iteration, f t (e i ) is set to the numerical value of the boundary uncertainty of unlabeled entities;
[0065] 133. The i-th unlabeled entity e i The boundary uncertainty value f obtained in the t-th iteration t (e i ), unmarked entity e i Other entities linked to e j The boundary uncertainty value f obtained in the t-th iteration t (e j ), and control of unlabeled entity e i The weighting α of the uncertainty of the entity and the uncertainty of other linked entities is entered into the formula. In the process, the i-th unlabeled entity e is obtained. i The numerical value of the boundary uncertainty f obtained in the (t+1)th iteration t+1 (e i );
[0066] in, It is an unmarked entity e i Other linked entity sets, t≥1, i≥1, j≥1.
[0067] 134. Determine whether all unlabeled entities have completed the iteration to obtain the value of the boundary uncertainty;
[0068] If not, then the next unlabeled entity whose boundary uncertainty value has not been obtained will be taken as the unlabeled entity e. i Return to step 131. It should be understood that defining an entity's influence on its context as the degree to which it can help its neighbors eliminate uncertainty is based on the formula... In the equation, function f is entity e. i Based on the uncertainty of the boundaries, It is entity e i The linked set of other entities, α, is used to control the ratio of its own uncertainty to the uncertainty of its neighbors, and entity e is obtained through iteration.i The boundary uncertainty value f obtained in the t-th iteration t (e i Iterate continuously until entity e i The boundary uncertainty value f obtained in the t-th iteration t (e i The boundary uncertainty value f obtained from the (t-1)th iteration and the (t-1)th iteration t-1 (e i The difference f) t (e i )-f t-1 (e i <0.1.
[0069] Based on the above embodiments, specifically, the unlabeled entity e i The initial boundary uncertainty value f0(e i The value of is obtained through the following method:
[0070] Calculate unlabeled entity e i and each standard entity e in the graph to be merged m similarity F(e) i ,e m ).
[0071] For all similarities F(e) i ,e m After sorting, a preset number of similarities from high to low are selected as the selected similarities to calculate the mean similarity.
[0072] The mean similarity and the similarity F(e) i ,e m Input the variance formula to obtain the unlabeled entity e. i The initial boundary uncertainty value f0(e i ).
[0073] For example, after ranking similarity, the variance of the top k=100 similarity scores is calculated. The larger the variance, the greater the information content. The formula for calculating the variance is as follows:
[0074]
[0075]
[0076] Based on the above embodiments, specifically, the unlabeled entity e is calculated. i and each standard entity e in the graph to be merged m similarity F(e) i ,e m Specifically, this includes:
[0077] Unmarked entity ei and standard entity e m Input entity alignment model to obtain matching score F EA (e i ,e m ).
[0078] Unmarked entity e i and standard entity e m The semantic similarity F obtained from the input semantic representation model S (e i ,e m ).
[0079] Based on the matching score F EA (e i ,e m semantic similarity F S (e i ,e m ) and preset weight values, to obtain the unlabeled entity e i and each standard entity e in the graph to be merged m similarity F(e) i ,e m ).
[0080] It should be understood that when calculating the similarity F between two entities, a large-scale pre-trained language model is introduced to vectorize the entity names and descriptions, and to calculate the semantic matching score between the entities. Based on these two uncertainties, F... EA (e i ,e j The final uncertainty of the entity can be obtained by weighting the following formula:
[0081] F(e i ,e m )=(1-β)F EA (e i ,e m )+βF S (e i ,e m )
[0082] Among them, F EA (e i ,e m F is the matching score returned by the entity alignment model. S (e i ,e j ) is the semantic similarity score returned by the semantic representation model, and β is a weight value of 0-1, which was set to 0.2 in the experiment.
[0083] Based on the above embodiments, specifically, the unlabeled entity e... i and standard entity em The semantic similarity F obtained from the input semantic representation model S (e i ,e m Specifically, this includes:
[0084] Unmarked entity e i Input the semantic representation model Sbert model and obtain the unlabeled entity e. i The representation vector.
[0085] Standard entity e m Input the semantic representation model Sbert model to obtain the standard entity e. m The representation vector.
[0086] Calculate unlabeled entity e i The representation vector and standard entity e m The representation vector is used to obtain the unlabeled entity e. i and each standard entity e in the graph to be merged m similarity F(e) i ,e m ).
[0087] It should be understood that the Sentence-BERT model is used to represent the textual description information of entities as vectors, and the similarity between two entities is measured by calculating the distance between the vectors. Sentence-BERT borrows the framework of the Siamese network model, inputting different sentences into two BERT models to obtain the sentence representation vector of each sentence. The sentence representation vectors obtained by training can be used for semantic similarity calculation.
[0088] 140. The 100 entities with the highest value selected in step 130 are sent to the annotation system for annotation of the actual corresponding entities, and the new annotation data is added to the label dataset.
[0089] 150. Apply the updated labeled dataset to the entity alignment model F. EA (e i ,e j Train and update the query system.
[0090] 160. Repeat steps 130 to 150 above until the entity alignment model reaches the preset training result.
[0091] The entity alignment method based on semantic and structural sampling strategies provided in Embodiment 2 of this application samples labeled data using a semantic representation model and an entity alignment model, prioritizing the labeling of data more valuable for knowledge graph fusion. After labeling a batch of data, the semantic representation model and entity alignment model are updated to improve the effectiveness of the sampling strategy before the next sampling is performed. By continuously iterating the above process, better entity alignment results can be achieved within a limited budget.
[0092] Example 3:
[0093] The following will be combined with the appendix Figure 5 This application provides a detailed description of the entity alignment device based on semantic and structural sampling strategies provided in its embodiments, specifically including:
[0094] The selection module is used to extract all unlabeled entities from the unlabeled data pool;
[0095] The iteration module is used to input the boundary uncertainty value obtained from the previous iteration of the unlabeled entity, the boundary uncertainty value of the previous iteration of other entities linked to the unlabeled entity, and the weight value controlling the uncertainty of the unlabeled entity and the uncertainty of the linked other entities into the iteration algorithm for calculation until the preset iteration result is met, and obtain the boundary uncertainty value of the unlabeled entity.
[0096] The annotation module is used to select a preset number of unlabeled entities as entities to be annotated based on the boundary uncertainty values of all unlabeled entities, and then update the labeled data to the labeled dataset.
[0097] The training module is used to train the entity alignment model to be trained using the updated labeled dataset, then update the unlabeled data pool, and repeat the iteration module to the labeling module until the entity alignment model to be trained meets the preset training results.
[0098] Based on the above embodiments, the iteration module further includes a first iteration module, a second iteration module, a third iteration module, and a fourth iteration module;
[0099] The first iteration module is specifically used to determine the unlabeled entity e. i The boundary uncertainty value f obtained in the t-th iteration t (e i The boundary uncertainty value f obtained from the (t-1)th iteration and the (t-1)th iteration t-1 (e i Does the difference satisfy the preset iteration result?
[0100] If so, execute the second iteration module;
[0101] Otherwise, execute the third iteration module and return to the first iteration module;
[0102] The second iteration module is specifically used to process the i-th unlabeled entity e i The numerical value of the boundary uncertainty obtained in the t-th iteration, f t (e i ) is set to a numerical value representing the boundary uncertainty of the unmarked entity;
[0103] The third iteration module is specifically used to process the i-th unlabeled entity e i The boundary uncertainty value f obtained in the t-th iteration t (e i The unmarked entity e i Other entities linked to e j The boundary uncertainty value f obtained in the t-th iteration t (e j ), and control of the unmarked entity e i The weighting α of the uncertainty of the entity and the uncertainty of other linked entities is entered into the formula. In the process, the i-th unlabeled entity e is obtained. i The numerical value of the boundary uncertainty f obtained in the (t+1)th iteration t+1 (e i );
[0104] in, It is the unmarked entity e i Other linked entity sets, t≥1, i≥1, j≥1.
[0105] The fourth iteration module is used to determine whether all the unlabeled entities have been iterated to obtain the value of the boundary uncertainty;
[0106] If not, then the next unlabeled entity whose boundary uncertainty value has not been obtained will be taken as the unlabeled entity e. i Return to the first iteration module.
[0107] Based on the above embodiments, furthermore, the iteration module is also used to calculate the unlabeled entity e. i and each entity e in the graph to be merged m similarity F(e) i ,e m );
[0108] For all similarities F(e) i ,e m After sorting, a preset number of similarities from high to low are selected as the selected similarities to calculate the average similarity.
[0109] The mean similarity and the similarity F(e) i ,e m Input the variance formula to obtain the unlabeled entity e. i The initial boundary uncertainty value f0(e i ).
[0110] Based on the above embodiments, the iteration module is further configured to process unlabeled entity e. i and entity e m Input entity alignment model to obtain matching score F EA (e i ,e m );
[0111] Unmarked entity e i and entity e m The semantic similarity F obtained from the input semantic representation model S (e i ,e m );
[0112] Based on the matching score F EA (e i ,e m Semantic similarity F S (e i ,e m ) and preset weight values, to obtain the unlabeled entity e i and each entity e in the graph to be merged m similarity F(e) i ,e m ).
[0113] Based on the above embodiments, the iteration module is further configured to process unlabeled entity e. i Input the semantic representation model Sbert model and obtain the unlabeled entity e. i The representation vector;
[0114] Entity e m Input the semantic representation model Sbert model to obtain entity e m The representation vector;
[0115] Calculate unlabeled entity e i The representation vector and entity e m The representation vector is used to obtain the unlabeled entity e. i and each entity e in the graph to be merged m similarity F(e) i ,e m ).
[0116] Based on the above embodiments, the annotation module is further configured to sort the boundary uncertainty values of all unlabeled entities and annotate a preset number of unlabeled entities at the top of the sort as entities to be annotated.
[0117] The entity alignment device based on semantic and structural sampling strategies provided in Embodiment 3 of this application includes a selection module that extracts all unlabeled entities from an unlabeled data pool; an iteration module that iteratively calculates the boundary uncertainty value of the unlabeled entity by taking the boundary uncertainty value of the unlabeled entity obtained in the previous iteration, the boundary uncertainty value of other entities linked to the unlabeled entity in the previous iteration, and the weighting of the uncertainty of the unlabeled entity and the uncertainty of other linked entities; a labeling module that selects a preset number of unlabeled entities to label and update the labeled dataset; and a training module that trains the entity alignment model using the updated labeled dataset. The above steps are repeated until the entity alignment model meets the preset training results. This application utilizes a semantic representation model and an entity alignment model to sample labeled data, prioritizing the labeling of data more valuable for knowledge graph fusion. After labeling a batch of data, the semantic representation model and entity alignment model are updated to improve the effectiveness of the sampling strategy before the next sampling. By continuously iterating the above process, better entity alignment results can be achieved within a limited budget.
[0118] Furthermore, embodiments of this application include a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the entity alignment method based on semantic and structural sampling strategies as described in any of the above technical solutions.
[0119] This application also includes a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the entity alignment method based on semantic and structural sampling strategy as described in any of the above technical solutions.
[0120] As is known from common technical knowledge, this invention can be implemented through other embodiments that do not depart from its spirit or essential characteristics. Therefore, the disclosed embodiments described above are merely illustrative in all respects and are not the only ones. All modifications within the scope of this invention or equivalent to the scope of this invention are included in this invention.
[0121] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0122] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0123] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0124] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. An entity alignment method based on semantic and structural sampling strategies, characterized in that, The method includes: Extract all unlabeled entities from the unlabeled data pool; The boundary uncertainty value obtained from the previous iteration of the unlabeled entity, the boundary uncertainty value of the other entities linked to the unlabeled entity from the previous iteration, and the weighting values of the uncertainty of the unlabeled entity and the uncertainty of the other linked entities are input into the iterative algorithm for calculation until the preset iteration result is met, and the boundary uncertainty value of the unlabeled entity is obtained. Based on the boundary uncertainty values of all the unlabeled entities, a preset number of the unlabeled entities are selected as entities to be labeled, and the labeled data is updated to the labeled dataset. After training the entity alignment model to be trained using the updated labeled dataset, update the unlabeled data pool and repeat the above steps until the entity alignment model meets the preset training results. The step of inputting the boundary uncertainty value obtained from the previous iteration of the unlabeled entity, the boundary uncertainty values from the previous iteration of other entities linked to the unlabeled entity, and the weighting values controlling the uncertainty of the unlabeled entity and the uncertainty of the linked other entities into the iterative algorithm for calculation until a preset iteration result is met, to obtain the boundary uncertainty value of the unlabeled entity, specifically includes: S1. Determine the unmarked entity. The The boundary uncertainty value obtained in the second iteration and the The boundary uncertainty value obtained in the second iteration Does the difference satisfy the preset iteration result? If so, proceed to step S2; Otherwise, proceed to step S3 and return to step S1; S2, the first Unlabeled entities The The numerical value of the boundary uncertainty obtained in the second iteration Set to a numerical value representing the boundary uncertainty of the unmarked entity; S3, the first Unlabeled entities The The boundary uncertainty value obtained in the second iteration The unmarked entity Other entities linked The The boundary uncertainty value obtained in the second iteration and control of the unmarked entity The proportion of uncertainty of the entity and the uncertainty of other linked entities Enter the formula In the middle, the first one is obtained Unlabeled entities The The numerical value of the boundary uncertainty obtained in the second iteration ; in, It is the unmarked entity Other sets of entities linked to, , , ; S4. Determine whether all the unlabeled entities have been iterated to obtain the value of the boundary uncertainty; If not, then the next unlabeled entity whose boundary uncertainty value has not been obtained will be taken as the unlabeled entity. Return to step S1; The method further includes: Calculate the unlabeled entity and entities in the graph to be merged similarity ; For all the aforementioned similarities After sorting, a preset number of similarities from high to low are selected as the selected similarities to calculate the average similarity. The mean similarity and the similarity Input the variance formula to obtain the unlabeled entity. Initial boundary uncertainty numerical ; The calculation of the unlabeled entity and entities in the graph to be merged similarity Specifically, it includes: The unmarked entity and the entity The matching score is obtained by inputting the entity alignment model. ; The unmarked entity and the entity The textual description information is represented as a vector, and input into the semantic representation model to obtain semantic similarity. ; Based on the matching score The semantic similarity The unlabeled entity is obtained by combining the preset weight value. and entities in the graph to be merged similarity .
2. The method according to claim 1, characterized in that, The unmarked entity and the entity The textual description information is represented as a vector, and the semantic similarity is obtained by inputting it into the semantic representation model. Specifically, it includes: The unmarked entity Inputting the first semantic representation model, the Sbert model, yields the unlabeled entity. The representation vector; The entity Input the second semantic representation model, the Sbert model, to obtain the entity. The representation vector; Calculate the unlabeled entity The representation vector and the entity The representation vector is used to obtain the unlabeled entity. and entities in the graph to be merged semantic similarity .
3. The method according to any one of claims 1-2, characterized in that, The step of selecting a preset number of unlabeled entities as entities to be labeled based on the boundary uncertainty values of all the unlabeled entities specifically includes: Sort the boundary uncertainty values of all the unlabeled entities from largest to smallest, and label the first preset number of unlabeled entities as entities to be labeled.
4. An entity alignment device based on a semantic and structural sampling strategy, characterized in that, The device includes: The selection module is used to extract all unlabeled entities from the unlabeled data pool; The iteration module is used to input the boundary uncertainty value obtained from the previous iteration of the unlabeled entity, the boundary uncertainty value of the previous iteration of other entities linked to the unlabeled entity, and the weighting value of the uncertainty of the unlabeled entity and the uncertainty of the linked other entities into the iteration algorithm for calculation until the preset iteration result is met, so as to obtain the boundary uncertainty value of the unlabeled entity. The annotation module is used to select a preset number of unlabeled entities as entities to be annotated based on the boundary uncertainty values of all the unlabeled entities, and update the labeled data to the labeled dataset. The training module is used to train the entity alignment model to be trained using the updated labeled dataset, then update the unlabeled data pool, and repeat the iteration module to the labeling module until the entity alignment model meets the preset training result. The iterative module includes a first iterative module, a second iterative module, a third iterative module, and a fourth iterative module; The first iteration module is specifically used to determine the unlabeled entity. The The boundary uncertainty value obtained in the second iteration and the The boundary uncertainty value obtained in the second iteration Does the difference satisfy the preset iteration result? If so, execute the second iteration module; Otherwise, execute the third iteration module and return to the first iteration module; The second iteration module is specifically used to process the first... Unlabeled entities The The numerical value of the boundary uncertainty obtained in the second iteration Set to a numerical value representing the boundary uncertainty of the unmarked entity; The third iteration module is specifically used to... Unlabeled entities The The boundary uncertainty value obtained in the second iteration The unmarked entity Other entities linked The The boundary uncertainty value obtained in the second iteration and control of the unmarked entity The proportion of uncertainty of the entity and the uncertainty of other linked entities Enter the formula In the middle, the first one is obtained Unlabeled entities The The numerical value of the boundary uncertainty obtained in the second iteration ; in, It is the unmarked entity Other sets of entities linked to, , , ; The fourth iteration module is used to determine whether all the unlabeled entities have been iterated to obtain the value of the boundary uncertainty; If not, then the next unlabeled entity whose boundary uncertainty value has not been obtained will be taken as the unlabeled entity. Return to the first iteration module; The iteration module is also used to calculate the unlabeled entities. and entities in the graph to be merged similarity ; For all the aforementioned similarities After sorting, a preset number of similarities from high to low are selected as the selected similarities to calculate the average similarity. The mean similarity and the similarity Input the variance formula to obtain the unlabeled entity. Initial boundary uncertainty numerical ; The iteration module is also used to process the unlabeled entities. and the entity The matching score is obtained by inputting the entity alignment model. ; The unmarked entity and the entity The textual description information is represented as a vector, and input into the semantic representation model to obtain semantic similarity. ; Based on the matching score The semantic similarity The unlabeled entity is obtained by combining the preset weight value. and entities in the graph to be merged similarity .
5. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the entity alignment method based on the semantic and structural sampling strategy as described in any one of claims 1 to 3.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the entity alignment method based on the semantic and structural sampling strategy as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Knowledge graph alignment model training method, alignment method, device and equipment
CN112966124A