一种PDF合同文件的识别方法、系统、介质及程序产品

By combining deep learning and natural language understanding models, the problem of identifying key fields in PDF contract documents with complex layouts has been solved, achieving efficient and accurate automated processing and electronic signature authentication, thus improving the processing efficiency and recognition accuracy of contract documents.

CN121189323BActive Publication Date: 2026-07-17BEIJING QIANRUNHE TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING QIANRUNHE TECH CO LTD
Filing Date
2025-09-17
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing PDF contract document information extraction technologies struggle to accurately identify key fields when faced with complex formats, leading to identification errors and requiring manual verification. This severely restricts contract processing efficiency and increases operating costs.

Method used

Deep learning models are used for image parsing and structured data generation. Adaptive semantic analysis and natural language understanding models are combined for multi-dimensional feature matching and entity relationship verification. Dependency parsing and semantic role labeling are used to correct extraction errors and generate a list of contract elements that meet electronic signature authentication requirements.

Benefits of technology

It achieves high-precision automated processing of PDF contract documents, improves the accuracy of key field identification, reduces manual intervention, and enhances system adaptability and contract data flow efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121189323B_ABST
    Figure CN121189323B_ABST
Patent Text Reader

Abstract

本申请公开了一种PDF合同文件的识别方法、系统、介质及程序产品,涉及信息识别技术领域,方法包括:采用深度学习模型对获取的PDF合同文件进行图像解析,生成可编辑文本并同步识别文档页面布局,输出包含页码、文本段落及坐标信息的结构化数据;依据自适应语义分析算法、预设的合同模板库以及动态关键字库,对文本段落进行多维度特征匹配,定位合同核心要素信息;对合同核心要素信息的内容进行实体关系校验,通过依存句法分析和语义角色标注修正提取误差,建立包含置信度权重的合同要素数据组;将校验的合同要素数据组进行序列化封装,生成符合电子签名认证的合同要素清单。本申请提升对PDF合同文件的处理效率和识别精度。
Need to check novelty before this filing date? Find Prior Art