Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

137 results about "Document classification" patented technology

Document classification or document categorization is a problem in library science, information science and computer science. The task is to assign a document to one or more classes or categories. This may be done "manually" (or "intellectually") or algorithmically. The intellectual classification of documents has mostly been the province of library science, while the algorithmic classification of documents is mainly in information science and computer science. The problems are overlapping, however, and there is therefore interdisciplinary research on document classification.

Text classification system

To efficiently and effectively classify a large amount of document and text data by an LLM.SOLUTION: The present invention relates to a text classification system 1 which classifies documents accumulated in a document DB 14 by categories, and the text classification system has a classification processing part 12 which lets an LLM 2 proposes one or more categories based upon a classification policy specified by a user, a search processing part 13 which searches the document DB 14 for documents belonging to the respective proposed categories, and a UI processing part 11 which presents the respective categories and the numbers of documents belonging to the respective categories to the user.SELECTED DRAWING: Figure 1
Owner:NOMURA RESEARCH INSTITUTE

Document processing service using customizable schemas

Systems and methods described herein provide a document processing service using customizable schemas. First user input includes a document and second user input identifies a classification schema. The classification schema is retrieved to obtain first prompt content. Prompt data is generated. The prompt data includes the first prompt content and second prompt content that includes a document classification instruction. The prompt data and document content of the document are processed by a large language model to obtain a classification result for the document. The classification result is presented via a user interface at a user device.
Owner:SAP SE

Engineering industry document intelligent review system based on large language model

The invention relates to the technical field of artificial intelligence and engineering management, in particular to an intelligent review system for engineering industry documents based on a large language model, which comprises a multi-level review point library, a document access module, a document classification module, a large language model processing module and a review report generation module, the multi-level review point library is in four-level classification, covers review dimensions such as multi-service fields, professional classification, project stages and integrity, and provides review standards; the document access module supports multi-type document and multi-channel access; the document classification module performs automatic classification based on a BERT model of engineering corpus fine adjustment; the large language model processing module matches the examination point list, performs slice examination and outputs opinions; the review report generation module summarizes the opinions to generate a structured report, so that the whole-process intelligent review of the document is realized; according to the method, the problems of low efficiency, long time consumption of complicated documents, non-uniform standard and poor result consistency caused by difference of standard understanding of reviews in traditional manual review are solved.
Owner:POWERCHINA BEIJING ENG CORP

Literature classification method and device, electronic equipment and storage medium

The invention provides a document classification method and device, electronic equipment and a storage medium, and relates to the technical field of computers. And constructing a standard subject classification system at least comprising a standard subject identifier uniquely corresponding to the subject. Extracting a target literature meeting a preset condition; the preset condition is that at least one identifier consistent with the standard subject identifier exists in the original subject identifiers labeled for the literature in advance. And based on the standard subject identifier, processing the original subject identifier to obtain a target subject identifier of the target literature in the standard subject classification system. Equivalently, a standard subject classification system is used as a unified standard of document classification, and subject identifiers of documents from different sources are standardized, so that the target document and the target subject identifier form a high-quality training data pair. Training data is utilized to train a specified large model, a document classification model is obtained, and the classification performance of the model is improved. The subject identification of the to-be-classified literature is determined by using the literature classification model, so that the accuracy of literature classification is ensured.
Owner:ZHEJIANG LAB

Intelligent agent construction method and system based on large language model, equipment and medium

The invention provides an agent construction method and system based on a large language model, equipment and a medium, and belongs to the technical field of artificial intelligence. The method comprises the steps that a user interaction module is called to receive a first target document input by a user; calling a document analysis module to analyze the first target document to obtain a plurality of first chapter titles; calling an automatic mapping module to classify the plurality of first chapter titles based on a document classification model, and constructing a mapping relationship between the plurality of first chapter titles and a plurality of second chapter titles of a second target document based on a classification result; calling a large language model to respectively extract and summarize paragraph contents of the plurality of first chapter titles to obtain paragraph contents of a second chapter title; and calling a document output module to output the second target document according to a preset format. According to the agent construction method and system based on the large language model, the equipment and the medium provided by the invention, the accuracy of generating the review report by the agent can be improved.
Owner:ZHONGJIAO ROAD & BRIDGE (HEBEI) CO LTD

Classification and key information extraction method and system for multi-page bidding and tendering files

PendingCN121121788ACharacter and pattern recognitionCommerceData setTyping Classification
The invention discloses a multi-page bidding document-oriented classification and key information extraction method and system, and the method comprises the steps: obtaining bidding document image data containing a plurality of pages, and constructing a corresponding document representation data set; performing structured modeling on bidding document image data corresponding to each page in the document representation data set, extracting and fusing multi-modal features, and constructing unified page representation; based on the unified page representation, document module type classification is carried out on the pages, and structure category labels which all the pages belong to are obtained; on the basis of the structure category labels of the pages, in combination with the multi-modal features of the corresponding pages, key information extraction operation is executed, and target field information contained in the corresponding pages is extracted; and outputting a structured result, wherein the structured result comprises the structure category label to which each page belongs and the corresponding target field information. According to the method, accurate classification of the pages in the bidding and tendering document is realized, targeted key information extraction is performed, and the method has relatively high accuracy, maintainability and actual landing capability.
Owner:TIANJIN UNIV OF SCI & TECH

Document classification method and electronic equipment

The invention discloses a document classification method and electronic equipment, and relates to the technical field of computers.The document classification method comprises the steps that multi-dimensional information of a to-be-classified document is extracted from multiple dimensions, multi-dimensional feature vectors are extracted through different feature extraction methods according to different dimension information, and the multi-dimensional feature vectors are used for classifying the to-be-classified document; a weighted fusion method based on an attention mechanism is adopted to construct comprehensive document feature vectors, and meanwhile, importance weights of feature vectors of all dimensions are dynamically adjusted, so that features which contribute to classification greatly obtain higher weights, the effectiveness of feature fusion is improved, and the accuracy of feature fusion is improved. And finally, based on the pre-trained classification model, realizing accurate classification based on the document feature vector, thereby effectively solving the technical problem that the document classification accuracy and robustness are seriously insufficient, and achieving the technical effect of improving the document classification accuracy.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Medical document processing method and system based on double-pipeline architecture

The invention discloses a medical document processing method and system based on a double-pipeline architecture, and relates to the technical field of document processing. According to the medical document processing method based on the double-assembly-line architecture, through a closed-loop process of document classification, preprocessing, double-assembly-line directional parallel processing and hierarchical vectorization storage, precise adaptation and efficient processing of multi-format and multi-type medical documents are achieved, the information loss rate and the key information truncation rate are greatly reduced, and the medical document processing efficiency is improved. According to the method, the document processing efficiency and the data standardization degree are improved, the warehousing success rate and the data traceability of the vector library are ensured, high-quality and structured data source support is provided for subsequent medical intelligent retrieval, clinical question and answer and retrieval enhancement generation system application, the knowledge base construction and maintenance cost is remarkably reduced, and the method is suitable for popularization and application. The problems that in existing medical document processing, medical semantics are not taken into consideration, so that key clinical information is easy to cut off, and a single processing flow cannot adapt to a structured guide and an unstructured case are solved.
Owner:SONGJIANG HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIVERSITY SCHOOL OF MEDICINE +2

Method and System for Data Modeling, Document Classification and Analysis

A method is disclosed for analysing a data set to determine a first processes. First messages are provided, the first messages classified into a plurality of different classes with a plurality of different likelihoods, a single first message classified into different classes based on different criteria. From the first messages a first subset of the first messages is retrieved based on a combination of one or more classifications, a likelihood of the one or more classifications, and another classification for messages within the first subset of the first messages. The likelihood of the classifications has more than two (2) potential values.
Owner:VIGILANT AI INC

Interaction management for supervised, assisted, or autonomous modernization agents of an application modernization platform

ActiveUS20250306916A1Version controlCharacter and pattern recognitionEnterprise computingEngineering
A method of application modernization engine executed by processors. The application modernization engine provides an improved contextual document processing and specific information extraction. The method includes receiving, by document intelligence sub-system of the application modernization engine, data of an enterprise computer system generated by artificial intelligence agents. The method includes processing, by the document intelligence sub-system, the data by extracting and ingesting specific contextual information to achieve document classification and feature extraction. The method includes providing, by the document intelligence sub-system, the specific contextual information, the document classification, and the feature extraction to the artificial intelligence agents to implement processes on top of the existing enterprise computer system to achieve objectives.
Owner:FLOWX AI INC

Document Classification and Extraction

Embodiments herein extract information from documents (e.g., scanned images of documents) to fields of a database or other collection of fields. The location and contents of blocks of text within the document are detected and then applied to a trained model to map a subset of the text to a set of target fields, where the target fields include one or more sets of repeated fields (e.g., corresponding to rows of a table). This mapping is presented to a user, optionally superimposed on an image of the document, to facilitate the user providing corrective feedback to the mapping. The mapping can then be updated, and the model trained to exhibit improved accuracy, based on the corrective feedback. The corrective feedback can include indicating the extent of a table and / or rows or columns thereof, facilitating correction of large numbers of field mappings.
Owner:SERVICENOW INC

Document classification using free-form integration of machine learning models

The technology automatically classifies documents using a decision tree integrating both rule-based nodes and machine learning (ML) model-based nodes. Rule-based nodes evaluate document information against predefined rules to generate classifications, while ML model-based nodes provide classifications along with the corresponding confidence level probabilities. Upon receiving an unclassified set of documents, the technology classifies each document by traversing the decision tree. At rule-based nodes, document evaluation entails comparing outcomes of logical conditions within the node. At ML model-based nodes, the evaluation depends on confidence level probabilities meeting predefined thresholds for each node. Using the evaluations, the technology assigns a proposed classification to each document. Once all documents have been classified, the technology generates a set of classified documents.
Owner:RECORDPOINT SOFTWARE HOLDINGS PTY LTD

File management system and method based on artificial intelligence

The invention relates to the technical field of computer information, in particular to a file management system and method based on artificial intelligence, and the system comprises an ontology construction module which is used for constructing an ontology structure in the fields of finance and user archives and providing a semantic representation framework for documents and user information; the document classification and information extraction module is used for performing image classification on an input document by utilizing a convolutional neural network and an optical character recognition technology, and selecting a corresponding information extraction template based on a document category so as to extract key field information of the document; the document structure and user portrait modeling module is used for converting the key field information into document structured data and user portrait data and expressing the document structured data and the user portrait data in a JSON format; the RDF mapping module is used for mapping the data in the JSON format into resource description framework triple data according to the semantic relation of the ontology structure; the inference engine module is used for performing semantic inference and data verification on the triple data by using an SHACL rule so as to judge whether the document and the user data meet the requirements of financial laws and regulations, and inferring to generate a corresponding user portrait classification result and a document label; and the data interaction mechanism is used for realizing data asynchronous transmission based on the message queue. According to the method, automatic processing, structured modeling, compliance verification and user portrait classification of financial and administrative documents can be realized, and the automation level and the intelligent ability of document management are improved.
Owner:WENXIN SOFTWARE TECH (GUANGZHOU) CO LTD

Document classification splitting method, device, system and equipment and storage medium

The invention relates to the technical field of document processing, and provides a document classification splitting method, device, system and equipment and a storage medium, and the method comprises the following steps: obtaining a PDF format document, and judging whether the PDF format document is in a scanning PDF format; if the document in the PDF format is in a scanning PDF format, analyzing each structured content of the document in the PDF format on the basis of a multi-mode OCR (Optical Character Recognition) engine; if the document in the PDF format is not in the scanning PDF format, analyzing each structured content of the document in the PDF format based on a structured document analysis engine; sorting each structured content into a plurality of document logic units according to context semantics based on a reinforcement learning model; and for each document logic unit, fusing the text feature, the visual feature and the structural feature of the document logic unit, and carrying out classified marking on the document logic unit based on the classification model. The PDF format document can be intelligently and automatically classified and split, the splitting accuracy is high, and the splitting efficiency is high. And the splitting effect is good.
Owner:CHINA CITIC BANK CO LTD

Multimodal multitask machine learning system for document intelligence tasks

Multimodal multitask machine learning system for document intelligence tasks includes a feature extractor processing token values obtained from a document to obtain features, and a token extraction head classifying, using the features, the token values to obtain classified tokens. The classified tokens are aggregated into entities. A document classification model is executed on the features to classify the document and obtain a document label prediction. Further a confidence head model applying the document label prediction processes the entities to obtain a result.
Owner:INTUIT INC

Data Classification Models Using Feature Extraction and Clustering

A method for document classification is described. A first dataset of labeled corporate data, a second dataset of internal labeled documents for a customer, and a third dataset of unlabeled documents for the customer are obtained. A classification model is trained using the first dataset. The classification model is further trained using the second dataset. Feature extraction is performed on each of the unlabeled documents of the third dataset by vectorizing content and metadata of each unlabeled document into one or more vectors and concatenating the one or more vectors to obtain a fixed length vector. Each of the unlabeled documents of the third dataset is clustered into one or more clusters based on similarity between the fixed length vectors for each unlabeled document. The unlabeled documents in each of the clusters are automatically labeled using text summarization. The classification model is retrained using the automatically labeled documents.
Owner:PROOFPOINT INC

Document classification apparatus, method, and storage medium

According to one embodiment, a document classification apparatus includes a processing circuit. The processing circuit is configured to: acquire text content for each of logical elements for semi-structured document data including text data stored for each of the logical elements; select logical elements from the logical elements and generating logical element sets each including the logical elements; analyze text contents for the respective logical element sets and constructing respective word embedded spaces; select a first word embedded space and a second word embedded space including a common word shared with the first word embedded space from the word embedded spaces, and update the first word embedded space based on similarity to the common word in the second word embedded space; and output a classification result of the document data using the first word embedded space and embedding information of a feature quantity of a classification target.
Owner:KK TOSHIBA

Audit text classifying and filing method and system suitable for auditing large model training

The invention provides an audit text classification filing method and system suitable for audit large model training, and relates to the technical field, the method comprises the following steps: carrying out entity identification and relation extraction on an unstructured audit document to generate a knowledge graph, and extracting a sub-graph structure as a feature representation; performing node embedding learning by using a graph neural network, generating a document vector containing structure and semantic information, and determining a document attribution category through an attention mechanism; and a natural language retrieval request can be analyzed, multi-hop path reasoning is executed, and a document corresponding to an implicitly associated extension node is obtained. The audit document classification accuracy is improved, efficient association query is realized, and the audit knowledge mining capability is enhanced.
Owner:TECH TRAINING CENT OF STATE GRID HUBEI ELECTRIC POWER CO LTD

Intelligent retrieval system for unstructured documents

The application relates to the technical field of document retrieval, in particular to an intelligent retrieval system for unstructured documents. The system comprises a data acquisition module for acquiring unstructured documents; a document feature analysis module for determining representative feature values in combination with keyword semantic importance, paragraph quantity and frequency, and constructing a theme consistency feature vector based on local and global dimensional theme distribution; a document classification module for clustering by comprehensively calculating and measuring distance of themes, keywords and consistency features, and selecting representative documents to construct a knowledge graph; and a retrieval module for generating a retrieval result based on the knowledge graph in combination with a large language model. The application solves the problem of serious homogenization of unstructured document retrieval results, improves the efficiency of intelligent retrieval of unstructured documents by clustering and deduplication and combining with a knowledge graph to enhance semantic association.

Automatic folder categorization of documents based on embedded brand logo methods

A document categorization method of a document management system receives a plurality of documents imported. The method detects at least one brand logo in a document of a plurality of documents and compare the detected at least one brand logo and a brand logo lookup table to find a brand logo information of a corresponding brand logo. The method adds the brand logo information to a document metadata of the document for folder categorization. The method sends the document to destination folder based on the document class of the document metadata and creates a subfolder for placing the document in the destination folder based on the brand logo information of the document metadata. The method also removes a brand logo within a document to reduce the data size to create additional capacity in the document management system.
Owner:KYOCERA DOCUMENT SOLUTIONS INC

Machine learning techniques for context-based document classification

ActiveUS12675732B2Context basedData mining
Various embodiments of the present invention provide methods, apparatus, systems, computing devices, computing entities, and / or the like for performing context-based document classification prediction using a hierarchical attention-based keyword classifier machine learning framework. Certain embodiments of the present invention utilize systems, methods, and computer program products that perform context-based document classification prediction using at least one of techniques using contextual keyword classifications, techniques using attention-based keyword classifier machine learning framework, techniques using a greedy matching indicator, and / or the like.
Owner:UNITEDHEALTH GROUP INC

Systems and methods for providing user interfaces for configuration of a flow for extracting information from documents via a large language model

Systems and methods for providing user interfaces for configuration of a flow for extracting information from documents via a large language model are disclosed. Exemplary implementations may: present a user interface configured to obtain entry of user input from a user to select a set of exemplary documents; select one or more document classifications for the set of exemplary documents; select one or more extraction fields that correspond to individual queries; navigate between different portions of the user interface; present the set of document classifications; present a particular individual document in the user interface; present a set of extraction fields in the user interface, wherein the individual extraction fields present individual replies obtained from the large language model in reply to the individual queries; and / or perform other steps.
Owner:INSTABASE INC

Machine learning powered cloud sandbox for malware detection in portable document format (PDF) files

A cloud-based network security system (NSS) is described. The NSS uses a sandbox to safely open and extract information about a PDF file and uses machine learning algorithms to analyze the information to predict whether the PDF file contains malware. Specifically, dynamic information about the PDF file is captured while it is open in the sandbox. Static information is extracted from the PDF file as well. The dynamic and static information is input to an AI or machine learning model trained to provide an output indicating a prediction of whether the PDF file contains malware. A verdict engine uses the output from the AI or machine learning model to classify the document as malicious or clean. Security policies can then be applied based on the classification.
Owner:NETSKOPE INC

Foreign trade document classification device

The utility model discloses a foreign trade document classification device which comprises a machine frame and further comprises a fixing assembly used for loading documents; the fixing assembly comprises a bottom frame, a pressing head and plastic clamping strips, the bottom frame is placed in a groove in the surface of the rack, the plastic clamping strips are fixedly connected to the two sides of the pressing head, the plastic clamping strips are connected with the bottom frame in a clamped mode through the grooves in the two sides of the bottom frame, and a plurality of cutter rings are symmetrically and fixedly connected to the pressing head; according to the utility model, the classified receipts can be effectively and integrally fixed.
Owner:JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS

Paper classification method based on graph matching and self-supervised graph learning

The invention discloses a paper classification method based on graph matching and self-supervised graph learning, and relates to the technical field of document classification based on deep learning. According to the method, literature data is represented by adopting a literature relation graph, a graph learning model ConGM based on the literature relation graph is constructed, and reference and theme association between literatures are mined through sub-graph sampling and data enhancement, linear node matching, secondary edge alignment and double-layer negative sample selection, so that precise classification of fields to which papers belong is realized.
Owner:PEKING UNIV

Document classification method, computer equipment and storage medium

The invention discloses a document classification method, computer equipment and a storage medium. The document classification method comprises the following steps: receiving a target classification original document; specific information used for referring to a specific instance in the target classification original document is filtered out through a first generative model, common features of the category to which the target classification original document belongs are reserved, and standardized description of the target classification original document is generated; performing similarity retrieval in a rule database based on the standardized description, and recalling a plurality of candidate classification rules; and inputting the standardized description and the candidate classification rule into a second generative model to obtain a final classification result. According to the method, document expressions are aligned with classification rule expressions through semantic conversion, and the document classification accuracy is improved to a certain extent by adopting a mode of combining retrieval and reasoning.
Owner:浙江太美医疗科技股份有限公司

Document tag management methods, devices and storage media

This invention discloses a document tag management method and apparatus. The method includes: if a document event is triggered, determining the target file; matching the target file's directory and file type according to a directory policy to determine if the target file enters the tag processing stage; reading the target file's current tag information and processing the current tag information according to policy conditions. This invention, through the collaboration of directory policy matching, automatic backup, tag reading and writing, and failure retries, can significantly reduce the cost of manual tag maintenance, improve the accuracy and timeliness of document classification, and reduce the risk of data leakage. For industries such as finance, government, and manufacturing, this solution helps to improve knowledge asset management capabilities through compliance audits, and has significant economic and social benefits.
Owner:SHENZHEN LEAGSOFT TECH

Document classification method and device based on large model, intelligent agent and electronic equipment

The invention provides a document classification method and device based on a large model, an intelligent agent and electronic equipment, and relates to the technical field of artificial intelligence, in particular to the technical fields of natural language processing, large models and the like. The method can be applied to scenes of enterprise internal document management based on artificial intelligence, patient record management of medical institutions, news classification, social media topic classification, spam filtering, medical literature classification, legal document classification and the like. According to the specific implementation scheme, an input document is obtained; determining the document type of the input document, wherein the document type is divided into a streaming document and a format document; generating corresponding target prompt word information according to the document type; document classification is conducted on the input document based on the target prompt word information through the multi-modal large model, the category of the input document is obtained, and the category of the input document refers to the semantic category of the input document. The prediction performance of the document classification system can be improved, and the prediction time can be shortened.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD