Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

251 results about "Structured document" patented technology

A structured document is an electronic document where some method such as markup or embedded coding, is used to identify the whole and parts of the document as having various meanings beyond their formatting. For example, a structured document might identify a certain portion as a "chapter title" (or "code sample" or "quatrain") rather than as "Helvetica bold 24" or "indented Courier". Such portions in general are commonly called "components" or "elements" of a document.

Generating structured documents with traceable source lineage

Systems and methods disclosed herein are enabled to dynamically generate structured documents using one or more artificial intelligence models. A computing device receives an output generation request and uses a first AI model to retrieve data chunks from source documents and applicable templates. A second AI model ranks the retrieved chunks based on one or more metrics, such as vector similarity, keyword density, and temporal relevance. A third AI model subsequently generates a response using the ranked chunks, templates, and predefined operational boundaries for each chunk. The generated response is tagged with source identifiers to enable the traceability of the response back to corresponding chunks. The system transmits, via the computing device, the response, the retrieved chunks, and / or the source identifiers.
Owner:CITIBANK N A

Enterprise-level large model agent application system supporting multi-modal collaboration

The invention relates to the technical field of artificial intelligence, and discloses an enterprise-level large-model agent application system supporting multi-modal collaboration, and the system comprises a user interaction terminal, an agent engine server, a knowledge engine server, a plug-in integration center, and a distributed storage unit. The agent engine server is responsible for intention recognition and task arrangement of a multi-modal input signal, and dynamically loads a differential reasoning strategy based on an environment isolation mechanism. And the knowledge engine server constructs a cross-modal semantic anchor point, analyzes an unstructured document into a tetrad knowledge unit, and realizes accurate recall of images and texts by using a hybrid retrieval algorithm. And the plug-in integration center executes outbound replacement and inbound restoration of the sensitive data through the context-aware dynamic desensitization gateway. According to the method, through a multi-modal semantic association and closed-loop verification mechanism, the problems of low complex document retrieval precision and leakage of external calling data are solved, and the service processing capacity and safety of the system are improved.
Owner:LINGRUIDA (XIAMEN) TECHNOLOGY CO LTD

Dynamic rule generation and self-adaptive auditing system and method for material management

The invention relates to the technical field of material management, and discloses a dynamic rule generation and self-adaptive auditing system and method for material management, and the system comprises a rule intelligent extraction module, a rule management knowledge base, an enhanced auditing engine, a man-machine cooperation calibration module and a self-adaptive execution module. The method comprises the steps of automatic rule extraction, rule storage and management, enhanced auditing and reasoning, man-machine collaborative calibration, knowledge base real-time optimization and adaptive routing execution. According to the method, the rule is automatically extracted from the unstructured document, the problem that a traditional system rule depends on manpower and is lagged in updating is solved, dynamic optimization of the rule and confidence is achieved by introducing a man-machine collaborative feedback closed loop, the accuracy and transparency of an audit decision are improved by enhancing reasoning and explainable decision technologies, and the audit efficiency is improved. And the optimal balance between auditing efficiency and risk control is realized through self-adaptive routing execution based on credibility.
Owner:PANGU CLOUD CHAIN (TIANJIN) DIGITAL TECH CO LTD

Document interpretation and report generation method and device, equipment and medium

The invention relates to the technical field of natural language processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a document interpretation and report generation method, device, equipment and medium, which comprises the following steps: receiving an original document set to generate a structured document object, executing optical character recognition on an image content set to generate a recognition text set, the recognition text set and the text content set are combined into a unified text sequence, element item extraction is executed based on the interpretation template parameter set to generate an interpretation element set, a retrieval enhancement context is retrieved and generated from the domain knowledge base, and the unified text sequence, the interpretation template parameter set and the retrieval enhancement context are input into a language model to generate an interpretation result. And generating report content based on the historical report template set. According to the method, automatic closed loop of document interpretation and report generation is realized through multi-modal unified processing and semantic enhanced reasoning, the efficiency is improved, and the manual dependence and compliance risk are reduced.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Bidding document compliance auditing method based on multi-modal knowledge graph

PendingCN121504484ASemantic analysisKnowledge representationEngineeringAutomated reasoning
The invention relates to the technical field of bidding document auditing, and particularly discloses a bidding document compliance auditing method based on a multi-modal knowledge graph, which comprises the following steps: extracting a standard reference file from an authoritative data source, and constructing the multi-modal knowledge graph in combination with natural language processing and image recognition technologies; through modal separation and structured processing, texts, tables and images in the bidding document file are converted into structural features, and standardized expression of an unstructured document is achieved; through semantic mapping and node matching, a one-to-one correspondence relationship is established between text modal features of a bidding document file and standard term nodes in a knowledge graph, and text difference features are extracted, so that the accuracy and pertinence of an auditing result are improved; based on an automatic reasoning mechanism of a knowledge graph logic rule, automatic compliance judgment is completed, and subjectivity and omission risks of manual auditing are reduced; the bidding document file is subjected to hierarchical labeling through cross-modal consistency verification, a compliance evaluation result is formed, and the intelligent level of auditing is improved.
Owner:江苏省设备成套股份有限公司

Bidding knowledge graph construction method and system based on OCR and NLP

The invention discloses a bidding and tendering knowledge graph construction method and system based on OCR and NLP. The method comprises the following steps: collecting a multi-source bidding document, and converting an unstructured document into text data by utilizing OCR (Optical Character Recognition); key entities, relations and attributes are extracted through the NLP technology; constructing a knowledge graph with entities as nodes and relationships as edges, and generating entity node vectors; analyzing the bidding and tendering file, extracting text demands such as qualification requirements and scoring standards, and converting the text demands into demand vectors; performing semantic matching on the demand vector and the entity node vector, and calculating the similarity; and extracting high-similarity entities and attribute relationships thereof from the knowledge graph based on a matching result, and automatically filling and generating bidding document core chapters according to a preset template. The system can realize structured storage and intelligent reuse of bidding and tendering knowledge, and improves the bidding document compiling efficiency and quality.
Owner:JINHUA BADA GRP CO LTD

Test case library construction method and system based on large model

The invention provides a test case library construction method and system based on a large model, and relates to the technical field of intelligent automobile tests.The method comprises the steps that semantic analysis of a tested function document is separated from test scene recognition, and an initial test case is generated in combination with a rule-guided large language model; systematization, intellectualization and standardization of a test case library construction process are realized; formats and content specifications output by the large model are effectively restrained through preset professional rules, and the professionality and structural consistency of the generated use cases are guaranteed; meanwhile, problems are found and fed back in time through a multi-dimensional verification and quality evaluation mechanism, self-adaptive optimization of a construction strategy is driven, and the quality and generation efficiency of the test cases are improved; finally, qualified cases are integrated into standardized records, a high-quality test case library with both manageability and technicality is formed, and efficient and automatic conversion from unstructured documents to reusable and maintainable test assets is achieved on the whole.
Owner:CHINA FAW CO LTD

Carbon border adjustment mechanism intelligent auxiliary customs declaration method and system based on natural language processing

The invention relates to the field of intelligent information processing, in particular to a carbon border adjustment mechanism intelligent auxiliary customs declaration method and system based on natural language processing, and the method comprises the steps: carrying out the real-time translation and key information extraction of CBAM related rule files and unstructured documents; a dynamic CBAM knowledge graph is constructed based on related rule texts, rule updating is monitored in real time through a text classification and event extraction technology, and calculation rules and declaration logic are dynamically adjusted; performing standardization processing on supplier data in different formats, integrating an industry emission factor library, automatically matching a calculation model according to a product type, generating a carbon emission result meeting a CBAM requirement, and mapping the data to a corresponding position of a CBAM declaration form; based on historical declaration data and related rule texts, analysis and compliance risk prediction are carried out by using a natural language processing technology, so that the declaration data are ensured to accord with latest rules, and violation risks caused by rule changes are avoided.
Owner:BEIJING SHU INTELLIGENT CARBON TECHNOLOGY CO LTD

Logic rule-based relative support and confidence for semi-structured document content extraction

One method includes extracting word-elements, each corresponding to a respective element of a ground truth cell-item array from an annotated document, applying logic rules to the extracted word-elements so that the applicability, or not, of each logic rule to each element of the ground truth cell-item array is determined. Based on the applying of the logic rules, metrics are obtained that indicate, for each word-element of the annotated document, the applicability of the logic rules, and the frequency with which applicable logic rules is satisfied. A first aggregation process is performed that aggregates the metrics across a group of unstructured, and annotated, documents, and a second aggregation process is performed that aggregates the metrics regarding a model-generated cell item array that was created based on the group of annotated documents. Finally, respective outcomes of the first and second aggregation processes are compared so as to identify logic rules of interest.
Owner:DELL PROD LP

Information extraction from unstructured documents with hybrid retrieval augmentation using multi-modal language models

A system for extracting a number of data elements from one or more data sources. Image-based documents are indexed using optical character recognition and a text embedding model to convert the document text to vector embeddings. Relevant portions of the document are identified by comparing the vector embedding of the documents to a vector embedding of a prompt or a request to extract information. The relevant text is mapped to a corresponding page of the documents. The page may be provided to a multi-modal language model for information extraction. The multi-modal language model can process contextual information included in the layout, figures, markings, etc. of the document to extract the information. The system populates an ontological data store based on the response from the language model. Extraction accuracy is improved without significant increases in computations performed by the system.
Owner:AMERICAN INTERNATIONAL GROUP INC

Product scheme PPT automatic generation method, system, device and medium

The invention provides a product scheme PPT automatic generation method, system and device and a medium, and belongs to the technical field of artificial intelligence. The method comprises the steps that product demand information is collected, an NLP tool is called for entity recognition and relation extraction, and a demand feature vector is generated; based on a MySQL database, retrieval is performed according to the demand feature vector, a structured record set is returned, and an associated document ID is obtained; based on a Milvus vector database, retrieval is carried out according to the demand feature vectors, and an unstructured document fragment set is returned; filtering the unstructured document fragment set according to the document ID, retrieving the unstructured document fragment set by adopting a mixed strategy of a BM25 algorithm and semantic retrieval, comprehensively sequencing retrieval results through a weighted scoring mechanism, and generating a final content fragment set; and calling a PPT template according to user requirements, and filling the final content fragment set into the PPT template to generate a PPT draft file.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Multi-modal document analysis method, electronic equipment and storage medium

The invention provides a multi-modal document analysis method, electronic equipment and a storage medium, and the method comprises the steps: preprocessing an original document, and recognizing key elements (at least including formulas, tables, texts and images) by using a target detection model to obtain an element positioning labeling table; roughly identifying the task type based on the label table to obtain an identifier, and generating a field adaptation parameter in combination with an original document metadata identification field; determining an element range according to a labeling table, and performing layered detection and targeted repair on obstacles in combination with task identification to obtain a barrier-free element document; carrying out collaborative coding on barrier-free elements, splitting sub-tasks and carrying out parallel processing on the basis of task identifiers and coding results, and calling adaptive parameters to adjust precision; and checking a correction processing result, and integrating into a final structured document report according to an adaptive parameter format. According to the method and the device, the multi-modal document analysis efficiency and precision can be improved.
Owner:北京中科闻歌科技股份有限公司

Intelligent document understanding method and system combining natural language processing and deep learning

The invention provides an intelligent document understanding method and system combining natural language processing and deep learning, and relates to the technical field of natural language processing and information management. Fusing the semantic pre-annotation result and the original text feature of the document, inputting the fused semantic pre-annotation result and the original text feature of the document into a cross-modal semantic enhancement model to obtain a document enhanced semantic representation, performing bidirectional semantic interaction with a dynamic business knowledge network based on the document enhanced semantic representation, generating an associated interaction result, and constructing a structured semantic asset; finally, the structured semantic assets are input into an intelligent document application engine, semantic feedback data are collected to optimize model parameters, document understanding accuracy and practicability can be improved, and diversified business requirements are met.
Owner:NANTONG INST OF TECH +1

Method and system of converting unstructured digital documents to a structure format using a secure API

In one aspect, a computerized method for document extraction workflow for unstructured documents includes the steps of implementing a text mining operation on a set of digital documents the incoming documents. This is done by defining a document type of each digital document. Based on the document type, the method defines a set of data dictionaries to extract any data from each digital document. The method uses the defined set of data dictionaries to extract any data from each digital document.
Owner:YERRAMSETTY VENKATA SAI RAMAN +1

Embedded data synthesis method and device integrating retrieval and large model distillation and medium

The invention provides an embedded data synthesis method and device fusing retrieval and large model distillation and a medium. The method comprises the following steps of: preprocessing an unstructured document in a vertical field, and dividing the unstructured document into multi-granularity text blocks with a hierarchical association relationship; forming a context based on the combination of the multi-granularity text blocks, injecting disturbance information corresponding to the priori knowledge in the vertical field into the context, calling a generative model to generate a retrieval query according to the context, and determining a target text block corresponding to the retrieval query as an initial positive sample; false negative sample text blocks are filtered according to the incidence relation between the text blocks, and a positive sample set and a negative sample set are formed; and constructing a comparative learning training sample, and training the semantic representation model by using the comparative learning training sample to generate an embedded vector for the retrieval task. According to the method, the retrieval task construction efficiency and authenticity can be improved, the positive sample coverage integrity is improved, and the contrast learning training stability and retrieval precision are enhanced.
Owner:北京衔远有限公司

Document revision method, document revision device and storage medium

The invention discloses a document revision method, a document revision device and a storage medium, and relates to the technical field of data processing. The method comprises the following steps: firstly, acquiring a reference document and a to-be-revised document in a revision mode; and analyzing the reference document, and extracting structured revision information corresponding to the reference revision operation trace. And analyzing the to-be-revised document to generate structured document information containing target content and target format information. And then, generating a cue word based on the semantic intention of the structured revision information, and inputting the cue word and the structured document information into the target model to obtain target revision information. And finally, mapping the target revision information to the to-be-revised document, and generating a target document with a target revision operation trace. According to the document revision method, the normalization and accuracy of the revision result can be ensured, the manual intervention cost is reduced, and format errors and semantic errors are reduced.
Owner:CHENGDU HONGRUI TECH

Intelligent technical supervision analysis decision method and system

The invention relates to the technical field of data processing, in particular to an intelligent technical supervision analysis decision method and system, and the method comprises the steps: obtaining a natural language supervision demand of a user, extracting a supervision target, a supervision dimension and a constraint condition through semantic analysis, and generating a structured supervision demand expression; the method comprises the following steps: generating a multi-source data query instruction set by adopting a model for converting a natural language to a database query statement on the basis of a structured supervision demand expression, obtaining a structured query result, an unstructured document fragment and external dynamic data in combination with a retrieval enhancement architecture, and generating an enhanced decision data set after fusion filtering; constructing a supervision analysis model, inputting the enhanced decision data set, and outputting a preliminary supervision analysis result; and verifying the preliminary supervision analysis result, if yes, outputting a final supervision decision result, and if not, returning to adjust the multi-source data query instruction set, and repeating the subsequent steps. The method can improve the efficiency and reliability of supervision decision, is adaptive to multi-source heterogeneous data, and achieves the dynamic optimization of a whole process.
Owner:LONGTAN HYDROPOWER DEV +1

Intelligent processing method for standard documents in ship industry

The embodiment of the invention provides a ship industry standard document intelligent processing method, which comprises the following steps: S1, document analysis: carrying out structured analysis and restoration on a ship industry standard document to generate a structured text supporting vectorization storage and semantic retrieval; s2, database construction: constructing a document vector index database based on the structured text to form a semantic retrieval task-oriented efficient data structure; and S3, content retrieval: performing user-oriented query, constructing a semantic retrieval mechanism, and completing high-precision matching from a natural language problem to a structured document content and result return. According to the embodiment of the invention, the document structure can be accurately restored, the content segmentation is more reasonable, the retrieval matching is more accurate, the retrieval result has more content depth, the technical path is clear, and the project landing is easy.
Owner:SHANGHAI WAIGAOQIAO SHIP BUILDING CO LTD

Large model information extraction and structure restoration system for long text document

The invention belongs to the technical field of intelligent document processing, and particularly relates to a large model information extraction and structure restoration system for a long text document. The system comprises an extraction point defining and text preprocessing module which is used for carrying out deep preprocessing and analysis on an input unstructured document, extracting text and coordinate information and identifying and protecting a special structure; meanwhile, an intelligent dynamic blocking strategy based on semantic boundaries is adopted, a dynamic overlapping mechanism is combined, and an input document is divided into text blocks keeping semantic integrity; the high-concurrency processing and scheduling control module is used for processing the text blocks in batches, calling a large language model to carry out distributed reasoning and outputting a dispersion result; and the extraction reasoning and result fusion module is used for performing semantic deduplication, entity alignment and confidence fusion on the dispersion result returned by the large language model to generate globally consistent structured output. The method has the advantages of being high in precision, high in speed, controllable in cost and high in robustness.
Owner:浙江实在智能科技有限公司

Reservoir hydropower station carbon accounting key information extraction method and system based on large language model

The invention relates to a reservoir hydropower station carbon accounting key information extraction method and system based on a large language model, and belongs to the technical field of artificial intelligence and environmental science. The method mainly comprises the following steps that firstly, domain entities, attributes and relations are extracted through a large language model, and a reservoir hydropower station carbon accounting domain knowledge graph is constructed; secondly, core concept nodes are recognized based on the knowledge graph, retrieval keywords are generated, and target literatures are accurately obtained from the multi-source heterogeneous data; thirdly, performing multi-modal analysis on the literature, extracting unstructured text and table data, and converting the unstructured text and table data into a structured document; secondly, performing semantic partitioning and vectorization processing on the document content, and constructing a vector knowledge base; and finally, responding to user query, realizing enhanced retrieval through multi-path recall and reordering, and accurately obtaining carbon accounting information meeting multiple limitation conditions. According to the method provided by the invention, efficient and accurate extraction of the carbon accounting core parameters can be realized.
Owner:CHONGQING INST OF GREEN & INTELLIGENT TECH CHINESE ACAD OF SCI

Enhanced OCR data processing through data enrichment and contextual tagging for llms

A method and system are disclosed for improving the accuracy, efficiency, and scalability of data interpretation and extraction of structured data from unstructured documents utilizing Large Language Models (LLMs). Applicable in finance, healthcare, legal, and government contexts, the disclosed invention addresses limitations of conventional Optical Character Recognition (OCR), machine learning, and LLM-based methods. In particular, the system and method incorporate feedback loops for continuous learning and leverage pre-processing, contextual tagging, customized prompt engineering, and post-processing to achieve robust data extraction. By integrating data enrichment techniques, the invention manages the inherent complexities of multilingual documents and evolving content standards.
Owner:ETON SOLUTIONS LP

Supply chain data real-time integration processing method and system based on multi-source heterogeneous data

The invention relates to the technical field of industrial data management, in particular to a supply chain data real-time integration processing method and system based on multi-source heterogeneous data, and the method comprises the steps: obtaining an external compliance data document, calling a document layout analysis model based on a multi-modal converter, extracting a text semantic vector, fusing the text semantic vector with a two-dimensional space coordinate feature to generate a composite index key, and carrying out the real-time integration processing of the supply chain data. Executing clustering analysis to recognize the text blocks and generating structured metadata; carrying out discretization processing on the business transaction data flow, and constructing an associated database with production environment and logistics data as attributes; executing cross-modal entity alignment by using the structured metadata, and establishing a bidirectional pointer link between the unstructured document index and the associated database; receiving a query vector, executing a recursive traversal algorithm based on a bidirectional pointer link, calculating a cascade influence probability as a correlation sorting score, and outputting a retrieval result; through multi-modal feature fusion and cross-modal entity alignment, deep integration and risk linkage retrieval of supply chain heterogeneous data are realized.
Owner:SUZHOU JINZHIYUAN TECHNOLOGY CO LTD

Lightweight multi-modal fusion document information structured extraction method and system

The invention provides a lightweight multi-modal fusion document information structured extraction method and system, and relates to the technical field of data processing, and the method comprises the steps: obtaining a document image; preprocessing the document image to obtain an optimized image; through the MobileNetV3, text features of the optimized image are extracted; performing multi-scale feature fusion on the multiple text features to obtain a fused feature map; through a double-branch Tokenized MLP module, extracting a context relation feature between a position feature of the fusion feature map and a spatial position; detecting a text region of the optimized image according to the position features of the fused feature map and the context relationship features between the spatial positions; performing text recognition on the text region of the optimized image through LPRNet; through SLANetplus, a table area of the optimized image is detected; and extracting structured document information by combining the text region and the table region through a multi-modal encoder.
Owner:WUXI YIMAIDE TECH CO LTD

Unstructured document multi-modal analysis and semantic metadata deep extraction method based on multi-agent collaboration

The invention discloses an unstructured document multi-modal analysis and semantic metadata deep extraction method and system based on multi-agent collaboration. The method comprises the steps that multi-modal elements such as texts, images and tables are extracted in parallel through a multi-format unified analysis engine; a four-agent (table structure analysis, multi-modal image analysis, semantic text understanding and metadata extraction) collaborative architecture is adopted, and task decomposition, asynchronous scheduling and result fusion are realized based on a multi-agent collaborative scheduling algorithm; the performance is improved by combining dynamic concurrency control, memory optimization and an intelligent cache mechanism; and finally generating structured output. The system integrates the modules and supports micro-service architecture and stability guarantee. According to the method, multi-modal deep semantic understanding is achieved, the processing efficiency is improved by 3-5 times compared with that of a traditional method, the text / table analysis precision reaches 98% and 95% respectively, localized efficient processing and incremental updating are supported, and a high-quality data source is provided for enterprise knowledge management and content retrieval.
Owner:CHINA ELECTRONICS CLOUD DIGITAL INTELLIGENCE TECH CO LTD

Structured document generation using different embedding space regions

A method and related system for generating a document using different portions of an embedding space includes obtaining a related document based on a first text, generating first vectors in an embedding space based on the first text and second vectors in the embedding space based on the related document, and determining a first region in the embedding space based on the first vectors and a second region in the embedding space based on the second vectors. The method further includes generating a first portion of a structured document based on the first vectors and third vectors in a third region within the first region but not within the second region. The method further includes generating a second portion of the structured document based on the first and second vectors and the first portion of the structured document.
Owner:CAPITAL ONE SERVICES LLC

Method and system for extracting information from documents with varying formats

Certain aspects of the disclosure provide a method for extracting attributes from documents with varying formats, layouts and complexities. The method displays a user interface (UI) that enables a user to obtain an unstructured document from a knowledge base. The method converts the unstructured document into a text document using a text recognition. The method obtains, as output from a large language model (LLM), an extracted page attribute from the text document. The extracted page attribute contains a first type of information recorded in text on a single page of the text document. The extracted document attribute contains a second type of information recorded in text on more than one page of the text document. The method obtains, as output from the LLM, an extracted document attribute from the text document. The extracted page attribute and the extracted document attribute are displayed in the UI.
Owner:SCHLUMBERGER TECH CORP

Data augmentation and feature selection for table decomposition

A method comprises obtaining an unstructured document and font information for the document, wherein the unstructured document includes a table; generating location information for an element of the table based on the font information; and generating a structured representation of the table based on the location information.
Owner:ADOBE INC

General document structured analysis method

The invention discloses a general document structured analysis method, and relates to the field of document intelligent analysis, and the method comprises the following steps: carrying out global layout analysis based on a thumbnail to obtain each layout area position; performing local content identification on each layout area based on the native resolution to obtain corresponding element contents; and sorting and splicing the element contents to obtain the structured document. Through decoupling global layout analysis and local content recognition, computing resources are accurately put into an information area, and unification of efficiency and precision is achieved; and multiple tasks are processed through a unified model and special elements are processed through a special algorithm, so that the integrity and accuracy of analysis are ensured.
Owner:SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT

Method for automatically generating complex-structure document based on large-model multi-agent

PendingCN121835641AResolving structural impedance mismatch issuesprecise alignmentDigital data information retrievalSemantic analysisExecution planAlgorithm
The invention is applicable to the technical field of artificial intelligence, and provides a large-model multi-agent-based method for automatically generating a complex-structure document, which comprises the following steps of: mapping an original structured document template into a compressed structure mark sequence with a reserved layout, and performing zero-sample semantic understanding and planning through a structure planning agent to obtain a complex-structure document; generating an atomic operation execution plan; according to the semantic logic relationship among the atomic operations in the structured execution plan, constructing a field-dependent directed acyclic graph, carrying out topological sorting on the field-dependent directed acyclic graph, and determining an execution sequence of content generation; through collaborative iteration of the content generation agent and the symbol verification agent, each field content is generated under the condition that a hard constraint condition is met; and finally synthesizing a complete document. According to the method, accurate understanding, reliable generation and format guarantee of the complex document under the zero sample condition are realized.
Owner:GUANGXI UNIV OF FOREIGN LANGUAGES

Unstructured document OCR error correction method and system based on multi-modal large model

The invention discloses an unstructured document OCR (Optical Character Recognition) error correction method based on a multi-modal large model, which comprises the following steps: step S101, acquiring an unstructured document to be processed, and decomposing the document into a page image unit and a preliminary optical character recognition text unit corresponding to the page image unit; step S102, selecting an error correction page, combining a page image unit with a preliminary optical character recognition text unit corresponding to the page image unit, and constructing a multi-modal input containing visual information and text information; step S103, providing the multi-modal input and a preset composite cue word to at least one multi-modal large model; step S104, the multi-modal large model generates a structured error correction suggestion according to the multi-modal input and the composite cue word; s105, displaying the original page image and the text attached with the error correction suggestions to the user on the interactive interface, receiving the processing operation of the user on each error correction suggestion, and recording the operation of the user as feedback data; and S106, optimizing an error correction system by using the feedback data.
Owner:SHANGHAI MITSUBISHI ELEVATOR CO LTD