Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

371 results about "Structured document" patented technology

A structured document is an electronic document where some method such as markup or embedded coding, is used to identify the whole and parts of the document as having various meanings beyond their formatting. For example, a structured document might identify a certain portion as a "chapter title" (or "code sample" or "quatrain") rather than as "Helvetica bold 24" or "indented Courier". Such portions in general are commonly called "components" or "elements" of a document.

System and method for adaptive semantic parsing and structured data transformation of digitized documents

A computing system is disclosed for transforming document data into schema-conformant structured outputs. The system obtains document data comprising multi-format structured documents and classifies each document by type and class using vector-based modeling and structural feature analysis. An extraction configuration is selected for each document, the configuration comprising machine-executable instructions for parsing based on semantic and layout characteristics. The system extracts semantic data using structured inference, transforms the semantic data into schema-conformant outputs, and validates the outputs using temporal and domain-specific constraints. Validated structured data may be used for downstream processing, visualizations, or optimization based on performance metrics.
Owner:ALTHQ INC

Document processing method and system based on text content extraction

The invention relates to a document processing method and system based on text content extraction. The method comprises the steps that an original document containing text, image and format information is received, the encoding format of the document is automatically detected, character set conversion is executed, and hierarchical indexes including page numbers, paragraphs and tables are established for an unstructured document; the method comprises the following steps: synchronously processing text content and visual layout through a pre-trained visual-language model, extracting word-level and sentence-level semantic features by a text stream embedding layer, analyzing spatial distribution features of document elements by a visual encoder, and fusing text and visual features through a cross-modal attention mechanism; and loading the domain knowledge graph matched with the document type, and executing entity linking to associate the text mentions to the knowledge nodes. According to the document processing method and system based on text content extraction, through the synergistic effect of vision-text joint coding and knowledge enhancement, the accuracy of financial contract key clause recognition tasks is improved, the error rate is lower than that of industry benchmark products, and the semantic understanding precision is remarkably improved.
Owner:WIN THE BID HUIKANG TECH CO LTD

Knowledge base intelligent deep auditing method and system based on AI large model

The invention provides a knowledge base intelligent deep auditing method and system based on an AI large model, and the method comprises the steps: obtaining an unstructured file uploaded by a user, analyzing the unstructured file, extracting a semantic auditing unit, and constructing a semantic auditing unit set; constructing a graph structure based on context dependence according to the semantic auditing unit set, marking continuous, neighbor, reference and reverse logic edge types, and fusing a compliance attention mechanism to update node representation to obtain a node semantic representation set; identifying candidate problem items through a double-layer risk discrimination function by combining the graph structure and the node semantic representation set design according to the introduced law and regulation knowledge embedding library; and executing causal chain backtracking on the candidate problem items to generate a structured auditing conclusion. The auditing accuracy and efficiency are remarkably improved, and the manual rechecking burden is reduced.
Owner:GUANGZHOU TAIXIN INFORMATION TECH CO LTD

Chatbot System For Structured And Unstructured Data

Techniques for operating a chatbot system for enterprise-level conversational agents are disclosed. These techniques are performed by an application or cloud service executing on one or more computing devices. An enterprise system can deploy conversational agents onto user devices to run as chat interfaces for logging analytics question-answering. One example application or cloud service may be a multi-model chat mechanism configured to support these chat interfaces with backend functionality. In response to an incoming question, the chat mechanism first consolidates the question with any conversation history and then, classifies the user's question as either a question regarding unstructured document data, a question regarding structured log data, or a hybrid question. Based on the classification, the chat mechanism can generate a proper large language model (LLM) response.
Owner:ORACLE INT CORP

Information extraction system for unstructured documents using retrieval augmentation providing source traceability and error control

A system for extracting a number of data elements from one or more unstructured data sources. The system may separate the text from the tables in a document, such that only the table data may be sent to the large language model (LLM), when the LLM only needs to review the table data. The system generates chunks from the document. The system associates unique identifiers with each chunk to provide traceability. The system identifies relevant chunks from the documents and includes the relevant chunks with a request to extract the data elements in a prompt to the LLM. The system also includes a request for the LLM to report the chunks used during extraction of the data elements. The reported chunks are stored with the extracted data for verification, auditing, and error control.
Owner:AMERICAN INTERNATIONAL GROUP INC

Multi-modal index knowledge base, construction method thereof and question and answer processing method

The invention discloses a multi-modal index knowledge base and a construction method thereof. The construction method comprises the following steps: processing a heterogeneous document to obtain a semi-structured document; identifying a title hierarchical relationship of the document to construct a document logic structure; carrying out minimum chapter blocking on the text of the semi-structured document to obtain logic blocks; performing semantic segmentation on each logic block to obtain text blocks; the method comprises the following steps: constructing text block nodes by meta-information of text blocks, constructing non-text block nodes by meta-information of non-text elements, extracting document nodes, chapter nodes and chapter-chapter inclusion relationships according to a document logic structure, respectively extracting semantic information from the text blocks and the non-text elements, and storing the semantic information in a database; recording the corresponding relationship between the text block nodes and the semantic information and between the non-text block nodes and the semantic information; constructing a knowledge graph based on each node and relationship and storing the knowledge graph into a graph database; and constructing semantic knowledge based on the semantic information and storing the semantic knowledge into a vector database. According to the scheme, lossless retention of multi-modal information and structured organization of document logic are realized, and efficient indexing and accurate recall are facilitated.
Owner:浙江泰隆商业银行股份有限公司

Document analysis method and device, equipment and storage medium

The invention provides a document analysis method and device, equipment and a storage medium, and relates to the technical field of computers, in particular to the technical field of deep learning, data processing and document analysis. According to the specific implementation scheme, at least one layout area divided by an article to which the document image belongs is determined according to the layout of the document image; identifying element contents of a plurality of layout elements in the document image; sorting the reading sequence of the layout elements in the same layout area; and obtaining structured document information according to the sorting result and the corresponding element content. According to the technical scheme, different article areas on the same page can be accurately identified and separated, the respective reading sequence is correctly reconstructed on the basis, and the content in the document image is converted into structured information.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Intelligent analysis method and device for unstructured PDF document, equipment and medium

The invention discloses an intelligent analysis method and device for an unstructured PDF document, equipment and a medium, and relates to the field of document analysis, the method comprises the steps of obtaining a to-be-analyzed PDF document, analyzing page elements in the PDF document, and generating a document metadata dictionary; if the PDF document does not contain the extractable text, converting the PDF document into an image and performing optical character recognition to generate first structured data; if the PDF document contains the extractable text, judging whether the PDF document contains a table or not; if the PDF document does not contain the table, a PDFMiner is adopted to extract the text, and second structured data is generated; if the PDF document contains the table, performing multi-modal feature extraction and feature fusion on the PDF document according to the document metadata dictionary to obtain multi-modal fusion features, and generating third structured data according to the multi-modal fusion features; according to the method and the device, the analysis precision and efficiency of the PDF document are improved.
Owner:LU ZE TECH CO LTD

Generating structured documents with traceable source lineage

Systems and methods disclosed herein are enabled to dynamically generate structured documents using one or more artificial intelligence models. A computing device receives an output generation request and uses a first AI model to retrieve data chunks from source documents and applicable templates. A second AI model ranks the retrieved chunks based on one or more metrics, such as vector similarity, keyword density, and temporal relevance. A third AI model subsequently generates a response using the ranked chunks, templates, and predefined operational boundaries for each chunk. The generated response is tagged with source identifiers to enable the traceability of the response back to corresponding chunks. The system transmits, via the computing device, the response, the retrieved chunks, and / or the source identifiers.
Owner:CITIBANK N A

Enterprise-level large model agent application system supporting multi-modal collaboration

The invention relates to the technical field of artificial intelligence, and discloses an enterprise-level large-model agent application system supporting multi-modal collaboration, and the system comprises a user interaction terminal, an agent engine server, a knowledge engine server, a plug-in integration center, and a distributed storage unit. The agent engine server is responsible for intention recognition and task arrangement of a multi-modal input signal, and dynamically loads a differential reasoning strategy based on an environment isolation mechanism. And the knowledge engine server constructs a cross-modal semantic anchor point, analyzes an unstructured document into a tetrad knowledge unit, and realizes accurate recall of images and texts by using a hybrid retrieval algorithm. And the plug-in integration center executes outbound replacement and inbound restoration of the sensitive data through the context-aware dynamic desensitization gateway. According to the method, through a multi-modal semantic association and closed-loop verification mechanism, the problems of low complex document retrieval precision and leakage of external calling data are solved, and the service processing capacity and safety of the system are improved.
Owner:LINGRUIDA (XIAMEN) TECHNOLOGY CO LTD

Dynamic rule generation and self-adaptive auditing system and method for material management

The invention relates to the technical field of material management, and discloses a dynamic rule generation and self-adaptive auditing system and method for material management, and the system comprises a rule intelligent extraction module, a rule management knowledge base, an enhanced auditing engine, a man-machine cooperation calibration module and a self-adaptive execution module. The method comprises the steps of automatic rule extraction, rule storage and management, enhanced auditing and reasoning, man-machine collaborative calibration, knowledge base real-time optimization and adaptive routing execution. According to the method, the rule is automatically extracted from the unstructured document, the problem that a traditional system rule depends on manpower and is lagged in updating is solved, dynamic optimization of the rule and confidence is achieved by introducing a man-machine collaborative feedback closed loop, the accuracy and transparency of an audit decision are improved by enhancing reasoning and explainable decision technologies, and the audit efficiency is improved. And the optimal balance between auditing efficiency and risk control is realized through self-adaptive routing execution based on credibility.
Owner:PANGU CLOUD CHAIN (TIANJIN) DIGITAL TECH CO LTD

Multi-agent power grid project intelligent monitoring, control and evaluation system and method

The invention relates to the technical field of project intelligent management and control, and discloses a multi-agent power grid project intelligent monitoring, management and control and evaluation system and method, and the system comprises a mixed data resource library, a multi-agent system, an index calculation module and a process control module. The mixed data resource library stores structured business data and unstructured documents; the multi-agent system comprises an information extraction agent, an index design agent, a code generation and treatment agent and the like, works cooperatively, and converts a high-level natural language management and control rule into an executable code; the index calculation module is responsible for executing codes according to a scheduling strategy and calculating quantitative indexes; and the process control module automatically identifies risks and executes management and control operations such as process locking according to the index result and a preset threshold value. According to the invention, the complex business logic can be automated and coded, the accuracy and timeliness of risk identification are improved, and refined and prospective intelligent management and control of the power grid project are realized.
Owner:BEIJING JINGHANG TIANLI TECH CO LTD

Document interpretation and report generation method and device, equipment and medium

The invention relates to the technical field of natural language processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a document interpretation and report generation method, device, equipment and medium, which comprises the following steps: receiving an original document set to generate a structured document object, executing optical character recognition on an image content set to generate a recognition text set, the recognition text set and the text content set are combined into a unified text sequence, element item extraction is executed based on the interpretation template parameter set to generate an interpretation element set, a retrieval enhancement context is retrieved and generated from the domain knowledge base, and the unified text sequence, the interpretation template parameter set and the retrieval enhancement context are input into a language model to generate an interpretation result. And generating report content based on the historical report template set. According to the method, automatic closed loop of document interpretation and report generation is realized through multi-modal unified processing and semantic enhanced reasoning, the efficiency is improved, and the manual dependence and compliance risk are reduced.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Document generation method based on agent collaboration

The invention discloses a document generation method based on agent collaboration, which comprises the following steps: acquiring initial document data through a content collection agent, and standardizing the initial document data to obtain input content; the input content is distributed to different intelligent agents; performing text generation on the standardized document data through a content generation agent and a multi-task parallel processing mechanism to obtain preliminary document content; performing structure planning coordination on the preliminary document content according to framework requirements through a structure planning agent to obtain a structured document draft; performing language retouching optimization on the structured document draft through a language retouching agent according to a retouching requirement to obtain retouched document content; and optimizing the different intelligent agents, integrating output results of the optimized intelligent agents to obtain final document content, and storing the final document content to obtain final document output.
Owner:SHANGHAI SOURCE CODE BANG DIGITAL TECHNOLOGY CO LTD +2

Information extraction system for unstructured documents using independent tabular and textual retrieval augmentation

A system for extracting a number of data elements from one or more data sources. The system may separate the text from the tables in a document, such that only the table data may be sent to the large language model (LLM), when the LLM only needs to review the table data. The system may include converting a PDF to text, and separating the tables form the document text using markdown language from converting the PDF. The system may form table chunks and text chunks, index the chunks using a vector embedding and store a chunk identifier, document identifier, and or a page identifier with the chunk to provide result traceability. The system may, in response to a prompt, retrieve and send the targeted table chunks or text chunks to the LLM to extract the data elements. The system populates an ontological data store based on the LLM response.
Owner:AMERICAN INTERNATIONAL GROUP INC

Multi-modal PDF document analysis method and device, equipment and medium

The invention provides a multi-mode PDF (Portable Document Format) document analysis method, device and equipment and a medium, and the method comprises the following steps: loading a PDF document and carrying out preprocessing, including page splitting and content cleaning, to generate standardized document data; dynamically extracting contents in the standardized document data, wherein the contents comprise paragraphs, pictures and table elements; processing the dynamically extracted pictures, including shielding meaningless pictures based on a preset rule, and analyzing picture contents by using a multi-modal model to generate readable picture information; processing the dynamically extracted table, including optimizing and merging the table structure into a single element format to generate structured table data; combining paragraphs, readable picture information and structured table data, converting the paragraphs, the readable picture information and the structured table data into a complete structured document format, and inserting in key positions to enhance coherence context description; and outputting the complete structured document format as a final analysis result. According to the method, the analysis precision and the knowledge extraction efficiency of the complex PDF document can be remarkably improved.
Owner:深圳市和讯华谷信息技术有限公司

Structured document generation using different embedding space regions

A method and related system for generating a document using different portions of an embedding space includes obtaining a related document based on a first text, generating first vectors in an embedding space based on the first text and second vectors in the embedding space based on the related document, and determining a first region in the embedding space based on the first vectors and a second region in the embedding space based on the second vectors. The method further includes generating a first portion of a structured document based on the first vectors and third vectors in a third region within the first region but not within the second region. The method further includes generating a second portion of the structured document based on the first and second vectors and the first portion of the structured document.
Owner:CAPITAL ONE SERVICES LLC

Question and answer method based on structured document retrieval enhancement

The invention discloses a question answering method based on structured document retrieval enhancement, and belongs to the technical field of artificial intelligence. The method comprises the following steps: performing structured information analysis and blocking processing on a structured document to construct a tree structure reflecting a hierarchical relationship of the document; based on a cross-document association relationship, combining the tree structures corresponding to the plurality of structured documents to form a multi-document tree structure network; and when a question request is received, generating corresponding retrieval enhancement information through the tree structure network, and inputting the retrieval enhancement information into the question and answer model to obtain answer information matched with the question request. According to the method, the multi-document tree structure network which takes the tree hierarchical relationship of the structured document as a core and fuses cross-document topic clustering and reference mapping is constructed, so that efficient retrieval enhanced questions and answers oriented to the structured document are realized, and the accuracy, context consistency and traceability of answers are remarkably improved.
Owner:GRG BANKING EQUIPMENT CO LTD

Character-based representation learning for information extraction using artificial intelligence techniques

Methods, apparatus, and processor-readable storage media for character-based representation learning for information extraction using artificial intelligence techniques are provided herein. An example computer-implemented method includes identifying, from unstructured documents, words and corresponding document position information using artificial intelligence-based text extraction techniques; generating an intermediate output by implementing at least one character embedding with respect to the unstructured documents using at least one artificial intelligence-based encoder; determining structure-related information for at least a portion of the unstructured documents using one or more artificial intelligence-based graph-related techniques; generating a character-based representation of at least a portion of the unstructured documents using at least one artificial intelligence-based decoder; classifying one or more portions of the character-based representation using one or more artificial intelligence-based statistical modeling techniques; and performing one or more automated actions based on the classifying.
Owner:DELL PROD LP

Bidding document compliance auditing method based on multi-modal knowledge graph

PendingCN121504484ASemantic analysisKnowledge representationEngineeringAutomated reasoning
The invention relates to the technical field of bidding document auditing, and particularly discloses a bidding document compliance auditing method based on a multi-modal knowledge graph, which comprises the following steps: extracting a standard reference file from an authoritative data source, and constructing the multi-modal knowledge graph in combination with natural language processing and image recognition technologies; through modal separation and structured processing, texts, tables and images in the bidding document file are converted into structural features, and standardized expression of an unstructured document is achieved; through semantic mapping and node matching, a one-to-one correspondence relationship is established between text modal features of a bidding document file and standard term nodes in a knowledge graph, and text difference features are extracted, so that the accuracy and pertinence of an auditing result are improved; based on an automatic reasoning mechanism of a knowledge graph logic rule, automatic compliance judgment is completed, and subjectivity and omission risks of manual auditing are reduced; the bidding document file is subjected to hierarchical labeling through cross-modal consistency verification, a compliance evaluation result is formed, and the intelligent level of auditing is improved.
Owner:江苏省设备成套股份有限公司

Bidding knowledge graph construction method and system based on OCR and NLP

The invention discloses a bidding and tendering knowledge graph construction method and system based on OCR and NLP. The method comprises the following steps: collecting a multi-source bidding document, and converting an unstructured document into text data by utilizing OCR (Optical Character Recognition); key entities, relations and attributes are extracted through the NLP technology; constructing a knowledge graph with entities as nodes and relationships as edges, and generating entity node vectors; analyzing the bidding and tendering file, extracting text demands such as qualification requirements and scoring standards, and converting the text demands into demand vectors; performing semantic matching on the demand vector and the entity node vector, and calculating the similarity; and extracting high-similarity entities and attribute relationships thereof from the knowledge graph based on a matching result, and automatically filling and generating bidding document core chapters according to a preset template. The system can realize structured storage and intelligent reuse of bidding and tendering knowledge, and improves the bidding document compiling efficiency and quality.
Owner:JINHUA BADA GRP CO LTD

Test case library construction method and system based on large model

The invention provides a test case library construction method and system based on a large model, and relates to the technical field of intelligent automobile tests.The method comprises the steps that semantic analysis of a tested function document is separated from test scene recognition, and an initial test case is generated in combination with a rule-guided large language model; systematization, intellectualization and standardization of a test case library construction process are realized; formats and content specifications output by the large model are effectively restrained through preset professional rules, and the professionality and structural consistency of the generated use cases are guaranteed; meanwhile, problems are found and fed back in time through a multi-dimensional verification and quality evaluation mechanism, self-adaptive optimization of a construction strategy is driven, and the quality and generation efficiency of the test cases are improved; finally, qualified cases are integrated into standardized records, a high-quality test case library with both manageability and technicality is formed, and efficient and automatic conversion from unstructured documents to reusable and maintainable test assets is achieved on the whole.
Owner:CHINA FAW CO LTD

RAG system intelligent evaluation method and device based on large model, and computer equipment

The invention discloses an intelligent evaluation method and device for an RAG system based on a large model, and computer equipment. The method comprises the following steps: acquiring a knowledge base source file required by a to-be-evaluated RAG system; analyzing and converting the knowledge base source file, creating a structured document block set, and executing preprocessing operation to obtain an intermediate knowledge base; automatically generating an evaluation data set including the question, the answer and the document block of the corresponding evidence according to the intermediate knowledge base based on LLM; performing performance evaluation on the to-be-evaluated RAG system by using an NLP evaluation index and an LLM-based multi-dimensional scoring mechanism based on the evaluation data set to obtain an evaluation result; and generating a report according to the evaluation result, and outputting the report. By implementing the method provided by the invention, the performance evaluation and iterative optimization of the RAG system can be comprehensively and accurately improved at low cost.
Owner:HANGZHOU FRAUDMETRIX TECH CO LTD

Electric power multi-mode corpus construction query method and system based on sliding window

The invention discloses an electric power multi-modal corpus construction query method and system based on a sliding window, which is applied to the field of electric power data query, and comprises the following steps: obtaining a structured document according to electric power multi-modal data, segmenting the structured document to obtain a plurality of segmented text blocks, and storing the segmented text blocks into a database; inputting each segmented text block into a large language model to generate a to-be-stored text vector and construct an electric power multi-mode corpus, when a user query request is received, generating a plurality of query variants according to query data, performing nearest neighbor search on each query variant in the electric power multi-mode corpus to obtain a corresponding nearest neighbor search result, and storing the nearest neighbor search result in the electric power multi-mode corpus. And fusing each nearest neighbor search result to generate an electric power related document set comprising a multi-modal association mark. According to the method, the semantic units can be accurately captured, semantic breakage is avoided, the retrieval continuity and coverage rate are improved, the comprehensiveness and context adaptability of retrieval results are improved, and the data retrieval requirement under the complex scene of the power industry is met.
Owner:STATE GRID ZHEJIANG ELECTRIC POWER CO LTD +1

Fire safety intelligent evaluation system and closed-loop management and control method

The invention discloses a fire safety intelligent evaluation system and a closed-loop management and control method, and the method comprises the steps: analyzing national standard and local standard texts through a natural language processing technology, disassembling terms into a three-stage quantifiable index system, and constructing a dynamic index library comprising at least 200 subdivisions; a dynamic rule base is generated based on a BERT-NER model, a standard parameter threshold value is automatically extracted, and version difference comparison and real-time updating are supported; deploying a fire-fighting special BERT model to analyze an unstructured document, extracting violation events and converting the violation events into structured evaluation factors, aggregating multi-source sensor data and video streams through an edge computing gateway, and adding space-time tags to realize millisecond-level real-time monitoring; generating a comprehensive score in combination with the static weight and the dynamic weight, constructing a risk prediction model based on an Attention-LSTM network, inputting historical hidden danger data, environment variables and enterprise operation data, and outputting future risk probability distribution; according to the invention, compliance management is efficient, assessment accuracy is improved, risk response is real-time, and a management and control process is digital.
Owner:SUIREN FIRE TECH CO LTD

Application system integrated management method based on artificial intelligence

The invention relates to the field of emerging software and information technology service, in particular to an artificial intelligence-based application system integrated management method, which comprises the following steps of: analyzing an API (Application Program Interface) log, structured data, an unstructured document, business work order data and business image data by pre-training a large model, and generating a unified semantic representation in combination with multi-modal contrast learning; a knowledge distillation technology is adopted to migrate the large model capability to a lightweight neural network, and a domain knowledge base is generated based on fine tuning of business work order data; a user multi-mode instruction is analyzed, a DAG service agent template library is matched, and an execution chain is dynamically combined based on a reinforcement learning strategy, so that text processing, multi-system collaborative reasoning and RPA service process automation are realized; and collecting an execution log, and updating lightweight network parameters and an intelligent agent strategy library through an online distillation algorithm to form a closed-loop optimization mechanism. According to the method, the service response efficiency can be improved, and the problems of enterprise multi-service system data splitting and process stiffness are solved.
Owner:CHANGZHOU XIAOZHI NETWORK TECHNOLOGY CO LTD

System and Method for Flowsheet Population

A method, computer program product, and computing system for flowsheet population. A structured document conforming to a schema and having a plurality of rows, each row comprising a key and a value, is processed. A plurality of key / value variations associated with an instance of information are identified. A plurality of key / value variation vectors are generated by embedding each key / value variation in a vector. The plurality of key / value variation vectors are combined into a combined vector representing the instance of information. A transcript is segmented into a segment, the segment corresponding to the instance of information. A transcript segment vector is generated by embedding the instance of information in the segment into a vector. A similarity between the transcript segment vector and the and the combined vector is determined. The instance is extracted from the transcript segment vector by processing a prompt with a generative artificial intelligence (AI) model using retrieval augmented generation (RAG). A value of a row corresponding to the key is populated with the instance of information.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Metro industry-based knowledge file reading method

The invention provides a metro industry knowledge file reading method, which belongs to the technical field of file reading, and comprises the following steps: processing input metro industry knowledge files in various formats, and outputting uniformly formatted intermediate text data; performing word segmentation and part-of-speech tagging, domain term recognition and enhancement, named entity recognition, relation extraction and key information extraction on the intermediate text data; constructing a metro field knowledge graph based on the obtained entities, relationships and key information, storing the structured key information into a structured database, and establishing a graph node and index link to obtain a knowledge base fusing structured knowledge and unstructured document indexes; and performing query understanding on the user query and performing retrieval based on the knowledge base. The query accuracy and knowledge relevance are improved, and the efficiency of obtaining knowledge by subway workers is remarkably improved.
Owner:BEIJING MASS TRANSIT RAILWAY OPERATION CORPORATION LIMITED

System and method for trade finance operations and sanctions screening process

PendingUS20250378488A1Digital data information retrievalFinanceRisk ControlDocument representation
The present invention discloses a system and method for processing trade finance documents and performing automated compliance screening. The system comprises a computing device, and a database for storing trade finance documents. The system processes documents using OCR to extract text and positional data, generating structured document representations via a layout-aware AI model. An AI classifier module categorizes documents based on content, layout, and domain-specific roles, while a semantic verification module aligns document data with master Letter of Credit templates. A rule management module validates compliance against international trade standards, and a financial crime risk control module performs real-time checks against external sanctions, vessel, and dual-use goods databases. The system further determines and reports discrepancies, anomalies, and compliance issues. The system supports heterogeneous layouts, multi-language documents, and integration with banking APIs.
Owner:CLEARTRADE AI INC

Event context reduction method and system based on vector retrieval and large language model

The invention discloses an event venation reduction method and system based on vector retrieval and a large language model, and relates to the field of natural language processing, and the method comprises the following steps: obtaining a query word of a user and an original structured document corresponding to the query word; performing semantic retrieval in the original structured document by using the query word to obtain a similar text set which is similar to the query word in terms; based on the similar text set, constructing an information extraction cue word; performing event information extraction on the similar text set by utilizing a preset large language model according to the information extraction cue word to obtain structured event information; performing time sequence reconstruction on all structured events in the structured event information to obtain an event timeline; performing causal logic verification on the event timeline to obtain a causal logic verification result; and if the causal logic verification result is passed, outputting the event timeline. According to the invention, the integrity and logic reliability of the event chain in the output event timeline are effectively improved.
Owner:DATA SPACE RES INST