一种文档体裁的识别方法、装置、设备及存储介质

By acquiring summary information and chapter content of documents, and using a pre-trained model to identify document genres, the accuracy and long-tail problems of document genre identification are solved, achieving more accurate document classification.

CN117271769BActive Publication Date: 2026-07-17BEIJING ZITIAO NETWORK TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2023-09-19
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In existing technologies, document genre recognition is greatly affected by subjective factors, has low accuracy, and suffers from the problem of long-tail genre labeling, resulting in insufficient or excessive genre labeling.

Method used

By acquiring summary information and chapter content of the target document, and using a pre-trained genre recognition model, combined with the title, summary, and a set number of chapter contents, the document genre is identified. The mapping relationship between the title, summary, and chapter contents is considered to increase the number of genres that can be identified.

Benefits of technology

It improves the accuracy of document genre recognition, solves the long-tail problem of genres, increases the number of genres identified, and enhances the effect of document classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117271769B_ABST
    Figure CN117271769B_ABST
Patent Text Reader

Abstract

本公开实施例提供了一种文档体裁的识别方法、装置、设备及存储介质。该方法包括获取目标文档的第一信息和第二信息,其中,所述第一信息表征目标文档内容的概括信息,所述第二信息表征所述目标文档的章节内容;按照文档章节划分所述第二信息,得到至少两个第二信息片段,根据所述第一信息和至少两个第二信息片段确定设定数量的模型输入信息;将所述设定数量的模型输入信息分别输入预训练的体裁识别模型,通过所述体裁识别模型基于所述模型输入信息识别所述目标文档的体裁。本公开的技术方案,实现通过体裁识别模型基于目标文档的概括信息和章节内容识别目标文档的体裁,解决了人工体裁识别方式存在的准确度及体裁长尾问题。
Need to check novelty before this filing date? Find Prior Art