Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

12 results about "Historical document" patented technology

Historical documents are original documents that contain important historical information about a person, place, or event and can thus serve as primary sources as important ingredients of the historical methodology.

Machine-learning-based system and method for automatically extracting fields from documents

A machine-learning based (ML-based) system and method for automatically extracting one or more data fields from one or more documents, are disclosed. The ML-based system includes a document obtaining subsystem to obtain documents, a document pre-processing subsystem to generate pre-processed data, a field identifying subsystem to identify data fields using a trained ML model, and a field extracting subsystem to extract financial information. The ML-based system also comprises an output subsystem to deliver the extracted data to end users via user interfaces. The ML model is trained using historical documents, labelled data fields, and features such as distance-based features, direction-based features, dimension-based features, positional features, and value-based features. The M-based system employs hyperparameter optimization, noise removal, and accuracy assessment mechanisms to enhance performance. This ML-based system provides a scalable, accurate, and automated solution for financial information extraction, ensuring efficiency, adaptability, and seamless integration with enterprise systems.
Owner:HIGHRADIUS CORP

A low-cost entity annotation method and system based on user behavior analysis

The application relates to a low-cost entity labeling method and system based on user behavior analysis, which comprises the following steps: S1, data collection: using a state machine to provide a document layout service, collecting user historical documents, entity recognition results and user revision records; S2, data labeling: generating a labeling data set according to the entity recognition results, finding out suspicious incorrect labeling by using the user revision records and reconfirming to optimize the labeling data set; S3, model updating: training an NER model by using the labeling data set, and replacing the state machine in the step S1 with the NER model when the accuracy of the NER model exceeds that of the state machine. The method and system can obtain a higher labeling accuracy under the premise of less labeling workload.
Owner:FUZHOU UNIV ZHICHENG COLLEGE

Intelligent retrieval method and system based on electricity transaction

The embodiment of the invention provides an intelligent retrieval method and system based on electricity transaction, and belongs to the technical field of electricity transaction retrieval. The intelligent retrieval method comprises the steps that historical documents of electricity transactions are acquired, and an electricity transaction knowledge base is constructed according to the historical documents; selecting a basic large model; finely adjusting the basic large model by adopting the electricity transaction knowledge base; obtaining a retrieval instruction of a current user; and obtaining a retrieval result according to the retrieval instruction and the fine-tuned basic large model. And by adopting a mode of finely adjusting the basic large model, the retrieval precision and efficiency of the power transaction can be effectively improved.
Owner:CAPITAL ELECTRIC POWER TRADING CENT CO LTD +1

Prediction and notification of agreement document expirations

A document management system can include an artificial intelligence-based document manager that can perform one or more predictive operations based on characteristics of a user, a document, a user account, or historical document activity. For instance, the document management system can apply a machine-learning model to determine how long an expiring agreement document is likely to take to renegotiate and can prompt a user to begin the renegotiation process in advance. The document management system can detect a change to language in a particular clause type and can prompt a user to update other documents that include the clause type to include the change. The document management system can determine a type of a document being worked on and can identify one or more actions that a corresponding user may want to take using a machine-learning model trained on similar documents and similar users.
Owner:DOCUSIGN INC

Text data generation method and device, equipment, storage medium and program product

PendingCN122287551AUser inputEngineering
This application provides a method, apparatus, device, storage medium, and program product for generating text data. It relates to the fintech field or other related fields. The method includes: responding to a user's input operation at a user interface, acquiring a user input command and a user identifier; acquiring historical documents corresponding to the user identifier from a pre-built historical document library; acquiring a user style feature vector based on the historical documents corresponding to the user identifier; acquiring a scene style feature vector of the scene to which the user input command belongs; acquiring a final style feature vector based on the user style feature vector and the scene style feature vector; inputting the user input command and the final style feature vector into a text generation model, causing the text generation model to generate text data; and displaying the text data at the user interface. This method improves the accuracy of generated text data and enhances the user experience.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

A document examination and approval opinion recommendation method based on semantic matching and role perception

PendingCN122087097AGuarantee normativeGuaranteed professionalismSemantic analysisDatabase modelsPersonalizationFeature extraction
This invention discloses a method for recommending official document approval opinions based on semantic matching and role awareness, comprising the following steps: S1. Constructing a historical document vector database and a structured metadata database; S2. Obtaining documents to be approved and performing feature extraction and preliminary screening; S3. Performing text block-level similarity retrieval on the documents and aggregating them; S4. Calculating job matching scores based on job information; S5. Calculating the final recommendation score and sorting the documents in descending order based on the final recommendation score, outputting the top 5 with the highest scores as the recommended results. This invention improves recommendation accuracy by using a pre-trained semantic embedding model to understand the deep semantics of document content. Furthermore, through job matching score calculation and approval node position alignment mechanisms, it deeply binds the recommended approval opinions to specific approval positions and approval process positions, achieving personalized recommendations.
Owner:QIMING INFORMATION TECH

An Automatic Mining and Classification Method for Private Data

This invention discloses an automatic mining and classification method for private data, comprising the following steps: First, receiving private unstructured documents, performing preprocessing and semantic alignment, and using sliding window technology to segment the documents into continuous text blocks; then loading a locally deployed general basic model and an industry-specific model, calculating the perplexity of each text block, and determining the scarcity of the text block by comparing the differences in the model output; next, constructing a statistical model based on the perplexity distribution of the enterprise's historical documents and dynamically setting a threshold, calculating the value score for text blocks exceeding the threshold; finally, projecting the scarcity and value score onto a preset classification decision matrix to automatically determine the secret level of the text block and trigger corresponding security handling strategies. This invention achieves automated and intelligent classification of private data, improving the efficiency and accuracy of data security management.
Owner:JIANGSU DAOYUNYIN TECH CO LTD

Machine-learning-based system and method for automatically extracting fields from documents

PendingUS20260187156A1Data fieldEngineering
A machine-learning based (ML-based) system and method for automatically extracting one or more data fields from one or more documents, are disclosed. The ML-based system includes a document obtaining subsystem to obtain documents, a document pre-processing subsystem to generate pre-processed data, a field identifying subsystem to identify data fields using a trained ML model, and a field extracting subsystem to extract financial information. The ML-based system also comprises an output subsystem to deliver the extracted data to end users via user interfaces. The ML model is trained using historical documents, labelled data fields, and features such as distance-based features, direction-based features, dimension-based features, positional features, and value-based features. The M-based system employs hyperparameter optimization, noise removal, and accuracy assessment mechanisms to enhance performance. This ML-based system provides a scalable, accurate, and automated solution for financial information extraction, ensuring efficiency, adaptability, and seamless integration with enterprise systems.
Owner:HIGHRADIUS CORP

A document semantic comparison method, device, equipment, medium and product

PendingCN122287592ALinguistic modelEngineering
This application provides a document semantic comparison method, apparatus, device, medium, and product, comprising: performing structural parsing on the documents to be compared to obtain a structured paragraph set; vectorizing and encoding each paragraph to obtain an encoding result, and performing an approximate nearest neighbor search from a historical document database based on the encoding result to obtain candidate paragraphs; using each paragraph in the structured paragraph set as a source paragraph, forming paragraph pairs with the candidate paragraphs, and performing semantic comparison on the paragraph pairs to obtain a semantic similarity score; calculating a dynamic threshold based on meta-information tags, comparing the dynamic threshold with the semantic similarity score to determine the reuse risk level, and inputting high-risk paragraph pairs into a large language model to obtain an evidence chain risk report. This improves the accuracy, processing efficiency, and interpretability of document reuse risk identification.
Owner:CHINA MOBILE ZIJIN INNOVATION INST CO LTD +2

A document abstract optimization generation method and system

PendingCN122262325ASemantic analysisBiological modelsDocument summarizationEngineering
This invention proposes a document summarization optimization generation method and system, belonging to the field of data processing technology. The method includes: based on a historical document dataset, segmenting sentence sequences and constructing a document graph structure; combining a graph neural network to obtain target node features; then combining a contrastive learning model to construct positive and negative sample pairs and calculate the contrastive loss; inputting the target node features into a preset large model to generate a candidate summary sequence; constructing a joint objective optimization function and optimizing it to obtain an optimized graph neural network, an optimized contrastive learning model, and an initial optimized large model; performing knowledge distillation on the initial optimized large model to obtain a target optimized large model; and generating an optimized document summary from the document to be processed using the optimized model. This invention achieves structured processing of documents and fully utilizes structured information through joint optimization and knowledge distillation, fully considering the logical structure of document sentences and the deep semantic relationships between sentences to generate accurate document summaries.
Owner:GUANGDONG POWER GRID CO LTD

A method and device for training a question generation model

The one or more embodiments of the specification disclose a training method of a question generation model. The method first acquires a historical complex question and a corresponding answer, and a historical document type knowledge base corresponding to the historical complex question, and the historical document type knowledge base at least includes a historical document related to the historical complex question. Then, an entity word related to the historical complex question is extracted from the historical document, a simple question is constructed based on the entity word, and a plurality of historical knowledge points are determined based on the simple question and the historical complex question. Finally, the question generation model is trained based on the historical document and the plurality of historical knowledge points through a preset loss function to obtain a trained question generation model, wherein the question generation model is used to generate a historical complex question corresponding to the historical document.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

A knowledge base retrieval method and related products

The application discloses a knowledge base retrieval method and related products, the method comprises the following steps: obtaining modification record data; sending the modification record data to the knowledge base and storing it in the database; the knowledge base comprises a plurality of documents; obtaining retrieval data; based on the retrieval data, the knowledge base is retrieved to obtain a retrieval result. The application automatically obtains the modification record data generated by the user in the collaborative document, and synchronously sends the data to the knowledge base and stores it in the structured database. The pain point that "the approval information is scattered in each independent document and is difficult to be uniformly searched" in the traditional collaborative office mode is effectively solved. Without manually opening each historical document for searching, the approval content and the responsible person accumulated for many years can be secondly traced back, the information acquisition cost is greatly reduced, the auditing efficiency and the decision support capability are improved, and the knowledge base can be quickly searched and traced back.
Owner:BEIJING COCONUT TREE INFORMATION TECH CO LTD