Training data generation method, model training method, knowledge graph completion method and electronic equipment

By dividing entities in the knowledge graph into sparse and dense sets, generating high-quality positive triples using dense entities and optimizing negative sampling, the problem of distorted representation of sparse entities is solved, improving the model training effect and the accuracy of knowledge graph completion, especially showing excellent performance in the black and gray industry field.

CN120975263APending Publication Date: 2025-11-18WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511002815.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-11-18

Smart Images

  • Figure CN120975263A_ABST
    Figure CN120975263A_ABST
Patent Text Reader

Abstract

The invention provides a training data generation method, a model training method, a knowledge graph completion method and electronic equipment. According to the method, the problems of sparse entity training distortion and negative sampling noise interference are effectively solved by introducing a positive sampling mechanism based on entity similarity and a negative sampling optimization method of multistage verification. Specifically, firstly, entities in an original knowledge graph are divided into a sparse entity set and a dense entity set, and a high-confidence positive triple is generated for the sparse entities by using semantic information of the dense entities, so that the embedding representation quality of the sparse entities is improved. And secondly, through an importance-based negative sampling algorithm, a negative triple with a high training value is generated, and interference of false negative examples is avoided, so that the effect and accuracy of model training are improved.
Need to check novelty before this filing date? Find Prior Art