Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

71 results about "Document layout" patented technology

File arrangement processing method and system

The invention relates to the technical field of intelligent document processing, and particularly discloses a file arrangement processing method and system. The method comprises the following steps: acquiring an input document, analyzing the structure and the style of the input document, and generating standardized format data; extracting a semantic relationship of the document through a pre-trained text understanding model, constructing a semantic relationship graph and optimizing a chapter sequence; loading a document processing plug-in to execute cross-format arrangement, and generating intermediate format data; executing processing nodes based on a task scheduling mechanism to generate optimized content; calling a multi-format output module to generate a target document, and adjusting metadata in combination with the global variable; and finally evaluating the document quality and updating the model. According to the method, automatic analysis, intelligent optimization and format unification of the multi-source heterogeneous document are realized, and the intelligent degree and the output efficiency of document processing are improved.
Owner:SHENZHEN HAIGUI NETWORK TECH CO LTD

Machine Learning-Based Techniques for Document Layout Identification and Data Extraction

Techniques for identifying a document layout are disclosed. In one embodiment, attribute data associated with an electronic document is accessed. A machine learning model is then applied to the attribute data. The machine learning model is configured to classify the electronic document based on the attribute data and feature sets of a plurality of document classes. Based on the document class predicted by the machine learning model, the system identifies a layout associated with the document class. The layout specifies layout elements and content types associated with the layout elements. The system extracts and stores information from the electronic document according to its content type.
Owner:ORACLE INT CORP

Teaching document anti-copy protection method and system based on OCR interference resistance

The invention provides a teaching document anti-copy protection method and system based on OCR interference resistance, and belongs to the technical field of information security and digital copyright protection. According to the method, anti-OCR font processing is applied to document characters, invisible disturbance symbols are inserted between characters or paragraphs, watermark information is embedded in document typesetting or font strokes, and finally a standard PDF or DOCX file is compiled. Watermark information is embedded through stroke thickness change, paragraph row spacing indentation perturbation and other modes. The system comprises a font disturbance module, a character disturbance module, a watermark embedding module, a compiling output module and a verification module. According to the method, the human eyes can normally read the document, but messy codes or disorder can occur when the OCR tool extracts the document, and the document cannot be directly reused; meanwhile, watermark information cannot be stripped, so that leakage tracking is facilitated; and the output file keeps a standard format and is compatible with the existing office teaching process. The method is suitable for anti-copying protection of teaching documents, scientific research data, test papers, contract files and the like.
Owner:SICHUAN HEALTH REHABILITATION VOCATIONAL COLLEGE

Document auditing method and device, electronic equipment and storage medium

The invention provides a document auditing method and device, electronic equipment and a storage medium. The method comprises the steps of obtaining a first sample data set and a second sample data set; performing model training on the first original model according to the first sample data set to obtain a first target model for identifying the document layout; performing model training on the second original model according to the second sample data set to obtain a second target model for identifying the document content; obtaining a to-be-audited document, and inputting the to-be-audited document into the first target model to obtain a first document with target layout labeling information; inputting the first document into a second target model to obtain target text content information corresponding to the to-be-audited document; and auditing the to-be-audited document according to the target layout labeling information and the target text content information. According to the method, the accuracy of document auditing can be improved.
Owner:WUHAN BIG PULP IND DEV CO LTD

Method and framework for optimizing scanning copy content recognition quality by using large model

The invention relates to the technical field of artificial intelligence, in particular to a method and a framework for optimizing scanning copy content recognition quality by using a large model. The method comprises the steps of document image analysis and processing, character extraction and formatting and context-based OCR correction by using a large model. The invention aims to provide the method and the framework for optimizing the identification quality of the scanned copy content by using the large model, the powerful functions of large language models such as a visual model and a text model are combined, the deep understanding of the document content and layout is realized, the document layout is accurately analyzed, different elements such as text blocks, tables and images are identified, and the identification quality of the scanned copy content is improved. And in combination with an analysis result of the visual model, converting the document content into a graceful and smooth Markdown format, and retaining an original layout of the document.
Owner:FUJIAN YIRONG INFORMATION TECH

Systems and methods for extracting tables from documents

Embodiments of computer-implemented systems and methods for extracting tables from documents are described. Document layout data is generated based on a plurality of glyphs associated with document lines, comprising data identifying: a plurality of text segments within each line; a plurality of text segment links; a plurality of text blocks of one or more of the text segments; and a plurality of text block links. A document is generated including at least one editable document table corresponding to at least one document table identified based on the plurality of glyphs associated with document lines.
Owner:CANVA PTY LTD

File data content accurate and deep analysis and interpretation method based on AI

The invention belongs to the technical field of artificial intelligence, and particularly relates to an AI-based file data content accurate and deep analysis and interpretation method, which comprises the following steps: acquiring multi-format file data and file meta-information, constructing an AI analysis network, extracting text semantic vectors, image visual features, table structure information and document layout features, and constructing a multi-dimensional semantic map. According to a user query intention, semantic extension is performed in combination with a domain knowledge base, an enhanced semantic description vector is generated through a graph attention mechanism, semantic reasoning and relation mining are performed by adopting an improved knowledge distillation Transform model, and a deep analysis conclusion is generated through a multi-hop reasoning model in combination with file data complexity, information density and user requirements. And generating a personalized interpretation report in combination with a user role and a task scene, and outputting an analysis result through a visual interface. Therefore, the problems of poor understanding ability, poor file adaptability and the like in the prior art are solved.
Owner:WUHAN CHANGYUAN HONGTIAN DATA INFORMATION TECHNOLOGY CO LTD

Graph-based document layout detection

A document layout system for determining a layout of a document. The document layout system is configured to apply an OCR technique to identify the text in the document. The document layout system is further configured to generate a graph representation of the document, wherein the graph representation comprises a plurality of nodes and a plurality of edges that connect different ones of the plurality of nodes, wherein individual ones of the nodes correspond to different portions of the text. The document layout system is also configured to apply a graph cluster network machine learning model to the graph representation to identify a layout of different sections of the document according to respective merge inferences determined for individual ones of the plurality of edges. The document layout system is also configured to provide the layout of different sections of the document.
Owner:AMAZON TECH INC

Structuring-based key information extraction in multimodal models for enhancing document understanding

A system and method for extracting structured key information from diverse document types using large multimodal models (LMMs) is disclosed. The invention employs a zero-shot analysis to identify candidate keys within an input document, then selects a document schema from a document schema database based on the identified keys. The LMM is prompted with the selected document schema to generate structured key-value pairs, with field constraints enforced by the document schema. Relationships among extracted keys are mapped to a graph representation, enabling robust handling of complex document layouts. The system supports nested structures, tabular data, and alias definitions for fields, and can update document schemas based on ground truth feedback. The resulting structured output is provided in a machine-readable format, enabling reliable and scalable document understanding across varied domains such as invoices, health cards, and driving licenses.
Owner:ORACLE INT CORP

Method for identifying PDF (Portable Document Format) file layout based on mask processing

The invention discloses a PDF (Portable Document Format) file layout identification method based on mask processing, and aims to realize accurate identification of various layout elements such as titles, texts, page headers, page footers, tables and pictures in PDF files. According to the method, characters and coordinates in a PDF file are analyzed through MuPDF, different color masks are adopted to replace texts and punctuation marks respectively, original content of a non-text area is reserved, a page picture subjected to mask processing is generated, a training data set is constructed based on the page picture, training is conducted through a target detection model (such as YOLO), and a high-precision layout recognition model is obtained. According to the method, the accuracy and the automation level of PDF layout recognition are effectively improved, and the method has a wide application prospect.
Owner:GOKE HUANYU (NANJING) ELECTRONIC TECH CO LTD

Layout detection based on individual document compression with compression dictionaries

The disclosure generally describes methods, software, and systems for assigning incoming documents to a pre-defined layout class. A digitalized document corresponding to an original document is obtained. The digitalized document can be compressed, using a compression algorithm and a plurality of compression dictionaries, to generate a plurality of compressed documents. A respective compression ratio for each compressed document can be generated. A matching compression ratio associated with a first compressed document can be identified. The matching compression ratio can be identified as matching a selection criterion to identify a document layout matching the digitalized document. A first document layout associated with the compression dictionary used to generate the first compressed document can be assigned to the digitalized document. The assigned layout can be used to extract one or more data entries from the digitalized document to generate a record.
Owner:SAP SE

Document streaming layout rendering method and device based on configuration driving

The invention relates to the technical field of document typesetting, in particular to a document streaming layout rendering method and device based on configuration driving. The method comprises the following steps: defining a four-dimensional generation key value pair configuration file according to document attributes, page layout, component geometry and styles, and verifying the key value pair configuration file; converting the configuration node into a document object model through an analysis engine; receiving external business data, positioning a dynamic slot position in the model, filling data, and generating a to-be-rendered instance tree; constructing a component rendering factory based on the strategy mode, and dynamically instantiating a corresponding rendering strategy class; and calculating the geometric position of the component by adopting a streaming typesetting algorithm, paging according to requirements, generating and executing a drawing instruction, and outputting a PDF (Portable Document Format) document. Aiming at the problems of code and style coupling, out-of-control typesetting and high operation threshold in the existing scheme, layout and code decoupling, high-precision typesetting and low-threshold configuration are realized, the method is adaptive to the requirements of multiple scenes such as finance and medical treatment, and the expansibility is high.
Owner:SILIDI SEMICON (SUZHOU) CO LTD

Format conversion method and system based on multi-modal fusion and generative adversarial network

The invention relates to the technical field of document and image processing, and provides a format conversion method and system based on multi-modal fusion and generative adversarial networks, and the method comprises the steps: obtaining document layout information of a document format original file through a deep learning model according to the extracted document multi-modal information of the original file, obtaining document conversion style information by utilizing a document generative adversarial network model according to the document layout information and the first target style requirement, and converting the original file into a PDF format file according to the document conversion style information; for the image format original file, obtaining image understanding information by using an image analysis model according to the extracted image multi-modal information of the original file, and obtaining image conversion style information by using an image generative adversarial network model according to the image understanding information, a target image type requirement and a second target format requirement, and converting the original file into a PDF format file according to the image conversion style information. And PDF files meeting specific requirements are efficiently and accurately formed.
Owner:CHINA CITIC BANK CO LTD

Document layout detection method, training method and device for text processing model

The present disclosure provides a method for detecting document layout, a method and device for training a text processing model, which relate to the field of artificial intelligence technology, and particularly relate to technologies such as natural language processing, computer vision, and deep learning. Document layout detection includes: determining multiple texts in the document and the position information of the multiple texts; constructing a target input based on the multiple texts and the position information of the multiple texts; and using the text processing model to process the target input to obtain a document layout detection result, where the document layout detection result indicates at least one text corresponding to a preset document component among the multiple texts and the order of the at least one text in the document.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

System and method for semantic parsing of digital documents using visual and textual features

PCT designated stageWO2026143433A1DocumentationUser interface
A system for semantic parsing of an input digital document using visual and textual features is provided. The system includes a user interface, a document layout classification module, a semantic recovery module, and a text structuring module. The user interface enables users to upload the input digital document. The document layout classification module processes the document to categorize its elements based on page images and textual data, outputting layout information with tags and locations. The semantic recovery module uses this layout information to derive content, including tables, lists, and charts, and generates a hierarchical structure. The text structuring module organizes tokens based on the layout tags, groups text into sections by topic relevance, and handles page boundaries, producing another hierarchical structure.
Owner:HONG KONG APPLIED SCI & TECH RES INST

Document information entry method, device and storage medium

The present application discloses a document information entry method, device and storage medium, and the present application relates to the field of electronic digital data processing technology. The document information entry method includes: extracting structured document data of a document image through a priori model, and the priori model is pre-trained through document layout data; dividing the structured document data into windows according to fields to generate a detection task; verifying the target document data in the detection task through a rule engine to obtain a grammatical result, and verifying the target document data in the detection task through a knowledge graph to obtain a semantic result; if the grammatical result and the semantic result are correct, storing the target document data in a document database. It solves the technical problem that the related technology focuses on text extraction and cannot understand the context logic, which leads to a low error detection rate when entering documents, such as inconsistencies between the previous and next content, and achieves the technical effect of improving the error detection rate.
Owner:GUANGZHOU PINGYUN CRAFTSMAN TECH CO LTD

Document typesetting difference detection method and device, electronic equipment and storage medium

The embodiment of the invention relates to a difference detection method and device for document typesetting, electronic equipment and a storage medium. The method comprises the following steps: opening a first document and a second document of which the typesetting difference is to be detected; performing screenshot on the first document and the second document to obtain a first screenshot and a second screenshot respectively; and generating typesetting difference information between the first document and the second document based on the first screenshot and the second screenshot. Therefore, the accuracy of generating the typesetting difference information between the first document and the second document can be improved.
Owner:ZHUHAI KINGSOFT OFFICE SOFTWARE +2

Method and system for exporting PDF (Portable Document Format) of operator report

PendingCN120930614AText processingData setLine wrap and word wrap
The invention discloses an operator report PDF exporting method and system, and belongs to the technical field of report generation and document typeset.The method comprises the following steps that a data set and a field name list of a report to be exported are received; dynamically generating a header according to the field name list, and calling a PdfPTable to create a table; for each row of data in the table, estimating the height of the row and comparing the height with the remaining available height of the current page, if the height is not enough, calling document.newPage () for paging, and reconstructing a header on a new page; and for the cells, the Phrase is used for measuring and calculating the text size, and setNoWrap (false) and the minimum height are set so as to realize column width self-adaption and automatic line feed. According to the method, the report export flexibility, attractiveness and readability are remarkably improved, and the multi-scene and multi-format customization requirements of operators are met.
Owner:INSPUR TIANYUAN COMM INFORMATION SYST CO LTD

A knowledge question and answer method based on a multi-modal document understanding large model of a graph structure

The application relates to the technical field of knowledge question answering, in particular to a knowledge question answering method based on a multi-modal document understanding large model of a graph structure, which comprises the following steps: obtaining a target task and a target document graph corresponding to the target task; generating a target heterogeneous document graph corresponding to the target document graph according to the target document graph; inputting the target task and the target heterogeneous document graph into a preset large model to obtain a target question and answer text corresponding to the target task; the method can achieve higher accuracy and robustness on multiple public benchmark tasks (such as document question answering, graph table question answering and table understanding); the method has excellent generalization ability for novel and complex document layouts; and the method realizes more simple and efficient system deployment and application through an end-to-end unified architecture.
Owner:北京中科闻歌科技股份有限公司

A low-cost entity annotation method and system based on user behavior analysis

The application relates to a low-cost entity labeling method and system based on user behavior analysis, which comprises the following steps: S1, data collection: using a state machine to provide a document layout service, collecting user historical documents, entity recognition results and user revision records; S2, data labeling: generating a labeling data set according to the entity recognition results, finding out suspicious incorrect labeling by using the user revision records and reconfirming to optimize the labeling data set; S3, model updating: training an NER model by using the labeling data set, and replacing the state machine in the step S1 with the NER model when the accuracy of the NER model exceeds that of the state machine. The method and system can obtain a higher labeling accuracy under the premise of less labeling workload.
Owner:FUZHOU UNIV ZHICHENG COLLEGE

A document page identification method and device, electronic equipment and storage medium

Embodiments of the present application provide a document layout recognition method and device, electronic equipment and storage medium, the method comprises: obtaining a to-be-recognized document, extracting visual features and semantic features of the to-be-recognized document, wherein the visual features identify the visual characteristics on the overall layout of the image corresponding to the to-be-recognized document, and the semantic features at least include character-level features and text line-level features; fusing the image features and the semantic features to obtain multi-modal document features; and based on the multi-modal document features, recognizing the element positions and categories of each element in the to-be-recognized document. The character-level semantic features can extract text-level elements such as formulas embedded in the text, and the visual features can recognize visual elements such as images, and then the multi-modal document features can obtain elements including visual features such as images and character-level elements within text lines such as formulas, so that the recognition result of the document layout is more comprehensive, and the accuracy of the document layout recognition result is greatly improved.
Owner:HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD

Document typesetting method and device

The invention relates to a document typesetting method and device, and the method comprises the steps: obtaining a to-be-typeset document, and recognizing an element in a document page of the to-be-typeset document, a first position of the element, and a category label of the element; grouping the elements according to the category labels to obtain element groups; according to the category labels corresponding to the element groups, performing target processing on the elements in the element groups to obtain target elements, and creating a target document page corresponding to the document page; and typesetting the target element on the target document page according to the first position, and forming a target document by the target document page subjected to the typesetting of the target element. Therefore, the suitability is improved, the labor cost is reduced, and the efficient and accurate document typesetting requirement is met.
Owner:ZHUHAI KINGSOFT OFFICE SOFTWARE +2

A method and system for automatic identification, generation, and decision-making of financial documents

This application relates to the field of financial document management technology, and discloses a method and system for automatic identification, generation, and decision-making of financial documents. The method first receives and preprocesses the original document image, then uses computer vision technology to perform layout analysis and classification on the enhanced image, distinguishing between known standard layouts and unknown layouts. For known layouts, a template-based OCR module is used to accurately extract key text blocks; for unknown layouts, a general document understanding model is used to obtain preliminary key-value pairs of key information. Subsequently, the extracted information is transformed into structured data through a key information structure extraction module, and verified by a data validation and business rule engine. If the validation passes, the structured document information is output; otherwise, it is sent to a manual intervention queue for further review. This method effectively improves the efficiency and accuracy of automated financial document processing, while enhancing the system's flexibility in adapting to different document layouts.
Owner:ANHUI SCI & TECH UNIV

Construction and evaluation method of manufacturing industry multi-mode question and answer data set

The invention discloses a construction and evaluation method of a manufacturing industry multi-modal question and answer data set. The construction and evaluation method comprises the following steps: S100: constructing a corpus of manufacturing industry standard documents; s200, respectively carrying out metadata annotation on the standard documents in the corpus; s300, multi-modal element decoupling separation is carried out, multi-modal heterogeneous elements are extracted, text recognition is carried out, and recognition results are rearranged according to an original standard document layout; s400, post-processing the extracted text, inputting the post-processed text into a language model, carrying out term extraction according to a cue word template, and constructing a global manufacturing industry term dictionary by adopting the extracted terms; s500, constructing an initial question and answer pair seed set; and S600, performing batch question and answer pair generation and answer reasoning step generation by adopting a multi-modal model, and constructing a multi-modal manufacturing industry standard question and answer data set. According to the method, efficient and accurate construction of the manufacturing industry multi-modal question and answer data set can be realized, and a reliable data basis can be provided for training and evaluation of a subsequent multi-modal large model in a manufacturing industry scene.
Owner:BEIJING INFORMATION SCI & TECH UNIV

A document layout element detection method, device, storage medium and equipment

The application discloses a document layout element detection method and device, a storage medium and equipment. The method comprises the following steps: firstly, obtaining a target image in which a target document to be detected is located; then, constructing a preset coding vector corresponding to a preset layout element type; next, inputting the target image and the coding vector into a pre-constructed document layout element detection model to predict a layout element detection result corresponding to the target document; wherein, the document layout element detection model is trained according to a preset document mixed element by using a pre-training mode of contrast learning and mask prediction. It can be seen that, since the document layout element detection model trained according to the preset document mixed element is used to detect the layout element of the target document, the detection efficiency and accuracy of the layout element can be effectively improved, and the self-defined detection can be performed according to the preset layout element type on demand in the detection process, thereby improving the user experience.
Owner:IFLYTEK CO LTD

Document processing method and device and terminal equipment

The invention discloses a document processing method and device and terminal equipment. The document processing method comprises the steps that an image of a document is acquired; the image of the document is recognized, text information is obtained, and the text information comprises content and position information of a text block; based on the position information of the text block, the content of the text block is mapped into a configuration area, and one or more symbol numbers are arranged at the position, where the content of the text block is not mapped, in the configuration area. According to the method, the document information representation mode based on the language model is provided, the document layout can be implicitly coded through the document information representation mode, the layout or format of the document can be accurately or completely conveyed, and high-quality information extraction is supported.
Owner:HANGZHOU RUIZHEN TECH CO LTD

Concurrent conversion method for converting multi-type documents into target documents

The invention discloses a concurrent conversion method for converting multi-type documents into target documents, and belongs to the technical field of document processing. The conversion method comprises the following steps: inputting and detecting a document, inputting the detected document to be converted into a message queue, and sending a processing request to the message queue; the message queue allocates tasks and initializes the tasks; performing OCR processing on the document to generate an OCR full-element result packet and a unique identifier of the result packet; generating a target document file, performing document layout detection, and determining a complete reading sequence; and generating and outputting a complete target document file. According to the method, the message queue is adopted as a concurrency basis, the OCR technology, the page layout checking technology and the Markdown text formatting technology are combined, format conversion from multiple types of documents to the Markdown document is achieved, and the document conversion efficiency is high.
Owner:NO 30 INST OF CHINA ELECTRONIC TECH GRP CORP +1

Document segmentation methods, apparatus, computer equipment and storage media

ActiveCN121902816BAccurately identify structural boundariesImprove Segmentation AccuracyDocument structuringEngineering
This application discloses a document segmentation method, apparatus, computer device, and storage medium. In response to a segmentation command, a document to be segmented is acquired; the document is parsed to obtain multiple text units; visual features related to the document layout are determined based on the text units; a segmentation score is calculated based on the visual features; and segmentation is performed based on the segmentation score. In this application, the visual layout information of the document is referenced from a visual layout perspective, avoiding the text extraction quality defects of treating the document as a plain text stream without relying on OCR. Instead, segmentation is performed by combining the visual geometric layout characteristics of the document when the user browses the text, conforming to the browsing patterns of users reading documents, accurately identifying document structural boundaries, and improving the accuracy of text segmentation.
Owner:HANGZHOU YOUZAN TECH CO LTD

System and method for semantic parsing of digital documents using visual and textual features

PendingUS20260187338A1DocumentationUser interface
A system for semantic parsing of an input digital document using visual and textual features is provided. The system includes a user interface, a document layout classification module, a semantic recovery module, and a text structuring module. The user interface enables users to upload the input digital document. The document layout classification module processes the document to categorize its elements based on page images and textual data, outputting layout information with tags and locations. The semantic recovery module uses this layout information to derive content, including tables, lists, and charts, and generates a hierarchical structure. The text structuring module organizes tokens based on the layout tags, groups text into sections by topic relevance, and handles page boundaries, producing another hierarchical structure.
Owner:HONG KONG APPLIED SCI & TECH RES INST