Structured and unstructured fact knowledge fusion corpus construction method and device

By transforming fact retrieval queries from structured knowledge bases and combining them with entity recognition and linking, a high-quality corpus is constructed, which solves the problem of poor corpus quality in existing technologies and improves the performance of neural network models.

CN118296157BActive Publication Date: 2026-06-02HARBIN INST OF TECH AT WEIHAI

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HARBIN INST OF TECH AT WEIHAI
Filing Date
2024-04-07
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

The corpus constructed by existing technologies is of poor quality, resulting in poor performance of the trained neural network models.

Method used

Structured factual knowledge is selected from a pre-built structured knowledge base and transformed into factual retrieval queries. Entity recognition and entity linking are used to retrieve unstructured factual knowledge text candidates from an unstructured text library. The relevance of facts is judged by matching sets of triples, and candidates with strong relevance are saved to the corpus.

Benefits of technology

A high-quality factual knowledge corpus was constructed, which improved the modeling ability of neural network models, especially the performance of large-scale neural network models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118296157B_ABST
    Figure CN118296157B_ABST
Patent Text Reader

Abstract

The application discloses a structured and unstructured fact knowledge fusion corpus construction method and device. Structured fact knowledge is selected from a pre-constructed structured knowledge base and is converted into a fact retrieval query. A plurality of unstructured fact knowledge text candidates are retrieved from a pre-constructed unstructured text base according to the fact retrieval query. Entities in the unstructured fact knowledge text candidates are obtained by using entity recognition and entity linking. The fact knowledge text candidates are matched with the structured fact knowledge in the knowledge base. The fact relevance of the fact knowledge text candidates is judged based on the matching result. Text candidates with strong relevance and the structured fact knowledge matched therewith are reserved as a piece of structured and unstructured fact knowledge matched corpus and are saved into a corpus. The above process is repeatedly performed, and finally a high-quality corpus containing a plurality of matched corpora is formed.
Need to check novelty before this filing date? Find Prior Art