Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

80 results about "Document classification" patented technology

Document classification or document categorization is a problem in library science, information science and computer science. The task is to assign a document to one or more classes or categories. This may be done "manually" (or "intellectually") or algorithmically. The intellectual classification of documents has mostly been the province of library science, while the algorithmic classification of documents is mainly in information science and computer science. The problems are overlapping, however, and there is therefore interdisciplinary research on document classification.

Literature classification method and device, electronic equipment and storage medium

The invention provides a document classification method and device, electronic equipment and a storage medium, and relates to the technical field of computers. And constructing a standard subject classification system at least comprising a standard subject identifier uniquely corresponding to the subject. Extracting a target literature meeting a preset condition; the preset condition is that at least one identifier consistent with the standard subject identifier exists in the original subject identifiers labeled for the literature in advance. And based on the standard subject identifier, processing the original subject identifier to obtain a target subject identifier of the target literature in the standard subject classification system. Equivalently, a standard subject classification system is used as a unified standard of document classification, and subject identifiers of documents from different sources are standardized, so that the target document and the target subject identifier form a high-quality training data pair. Training data is utilized to train a specified large model, a document classification model is obtained, and the classification performance of the model is improved. The subject identification of the to-be-classified literature is determined by using the literature classification model, so that the accuracy of literature classification is ensured.
Owner:ZHEJIANG LAB

Medical document processing method and system based on double-pipeline architecture

The invention discloses a medical document processing method and system based on a double-pipeline architecture, and relates to the technical field of document processing. According to the medical document processing method based on the double-assembly-line architecture, through a closed-loop process of document classification, preprocessing, double-assembly-line directional parallel processing and hierarchical vectorization storage, precise adaptation and efficient processing of multi-format and multi-type medical documents are achieved, the information loss rate and the key information truncation rate are greatly reduced, and the medical document processing efficiency is improved. According to the method, the document processing efficiency and the data standardization degree are improved, the warehousing success rate and the data traceability of the vector library are ensured, high-quality and structured data source support is provided for subsequent medical intelligent retrieval, clinical question and answer and retrieval enhancement generation system application, the knowledge base construction and maintenance cost is remarkably reduced, and the method is suitable for popularization and application. The problems that in existing medical document processing, medical semantics are not taken into consideration, so that key clinical information is easy to cut off, and a single processing flow cannot adapt to a structured guide and an unstructured case are solved.
Owner:SONGJIANG HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIVERSITY SCHOOL OF MEDICINE +2

Document classification using free-form integration of machine learning models

The technology automatically classifies documents using a decision tree integrating both rule-based nodes and machine learning (ML) model-based nodes. Rule-based nodes evaluate document information against predefined rules to generate classifications, while ML model-based nodes provide classifications along with the corresponding confidence level probabilities. Upon receiving an unclassified set of documents, the technology classifies each document by traversing the decision tree. At rule-based nodes, document evaluation entails comparing outcomes of logical conditions within the node. At ML model-based nodes, the evaluation depends on confidence level probabilities meeting predefined thresholds for each node. Using the evaluations, the technology assigns a proposed classification to each document. Once all documents have been classified, the technology generates a set of classified documents.
Owner:RECORDPOINT SOFTWARE HOLDINGS PTY LTD

Multimodal multitask machine learning system for document intelligence tasks

Multimodal multitask machine learning system for document intelligence tasks includes a feature extractor processing token values obtained from a document to obtain features, and a token extraction head classifying, using the features, the token values to obtain classified tokens. The classified tokens are aggregated into entities. A document classification model is executed on the features to classify the document and obtain a document label prediction. Further a confidence head model applying the document label prediction processes the entities to obtain a result.
Owner:INTUIT INC

Document classification apparatus, method, and storage medium

According to one embodiment, a document classification apparatus includes a processing circuit. The processing circuit is configured to: acquire text content for each of logical elements for semi-structured document data including text data stored for each of the logical elements; select logical elements from the logical elements and generating logical element sets each including the logical elements; analyze text contents for the respective logical element sets and constructing respective word embedded spaces; select a first word embedded space and a second word embedded space including a common word shared with the first word embedded space from the word embedded spaces, and update the first word embedded space based on similarity to the common word in the second word embedded space; and output a classification result of the document data using the first word embedded space and embedding information of a feature quantity of a classification target.
Owner:KK TOSHIBA

Audit text classifying and filing method and system suitable for auditing large model training

PendingCN121858741AClassification is accurate and reliableEfficient captureSemantic analysisKnowledge representationGraph neural networksRelational query
The invention provides an audit text classification filing method and system suitable for audit large model training, and relates to the technical field, the method comprises the following steps: carrying out entity identification and relation extraction on an unstructured audit document to generate a knowledge graph, and extracting a sub-graph structure as a feature representation; performing node embedding learning by using a graph neural network, generating a document vector containing structure and semantic information, and determining a document attribution category through an attention mechanism; and a natural language retrieval request can be analyzed, multi-hop path reasoning is executed, and a document corresponding to an implicitly associated extension node is obtained. The audit document classification accuracy is improved, efficient association query is realized, and the audit knowledge mining capability is enhanced.
Owner:TECH TRAINING CENT OF STATE GRID HUBEI ELECTRIC POWER CO LTD

Intelligent retrieval system for unstructured documents

The application relates to the technical field of document retrieval, in particular to an intelligent retrieval system for unstructured documents. The system comprises a data acquisition module for acquiring unstructured documents; a document feature analysis module for determining representative feature values in combination with keyword semantic importance, paragraph quantity and frequency, and constructing a theme consistency feature vector based on local and global dimensional theme distribution; a document classification module for clustering by comprehensively calculating and measuring distance of themes, keywords and consistency features, and selecting representative documents to construct a knowledge graph; and a retrieval module for generating a retrieval result based on the knowledge graph in combination with a large language model. The application solves the problem of serious homogenization of unstructured document retrieval results, improves the efficiency of intelligent retrieval of unstructured documents by clustering and deduplication and combining with a knowledge graph to enhance semantic association.

Automatic folder categorization of documents based on embedded brand logo methods

A document categorization method of a document management system receives a plurality of documents imported. The method detects at least one brand logo in a document of a plurality of documents and compare the detected at least one brand logo and a brand logo lookup table to find a brand logo information of a corresponding brand logo. The method adds the brand logo information to a document metadata of the document for folder categorization. The method sends the document to destination folder based on the document class of the document metadata and creates a subfolder for placing the document in the destination folder based on the brand logo information of the document metadata. The method also removes a brand logo within a document to reduce the data size to create additional capacity in the document management system.
Owner:KYOCERA DOCUMENT SOLUTIONS INC

Machine learning techniques for context-based document classification

ActiveUS12675732B2Context basedData mining
Various embodiments of the present invention provide methods, apparatus, systems, computing devices, computing entities, and / or the like for performing context-based document classification prediction using a hierarchical attention-based keyword classifier machine learning framework. Certain embodiments of the present invention utilize systems, methods, and computer program products that perform context-based document classification prediction using at least one of techniques using contextual keyword classifications, techniques using attention-based keyword classifier machine learning framework, techniques using a greedy matching indicator, and / or the like.
Owner:UNITEDHEALTH GROUP INC

Machine learning powered cloud sandbox for malware detection in portable document format (PDF) files

A cloud-based network security system (NSS) is described. The NSS uses a sandbox to safely open and extract information about a PDF file and uses machine learning algorithms to analyze the information to predict whether the PDF file contains malware. Specifically, dynamic information about the PDF file is captured while it is open in the sandbox. Static information is extracted from the PDF file as well. The dynamic and static information is input to an AI or machine learning model trained to provide an output indicating a prediction of whether the PDF file contains malware. A verdict engine uses the output from the AI or machine learning model to classify the document as malicious or clean. Security policies can then be applied based on the classification.
Owner:NETSKOPE INC

Paper classification method based on graph matching and self-supervised graph learning

The invention discloses a paper classification method based on graph matching and self-supervised graph learning, and relates to the technical field of document classification based on deep learning. According to the method, literature data is represented by adopting a literature relation graph, a graph learning model ConGM based on the literature relation graph is constructed, and reference and theme association between literatures are mined through sub-graph sampling and data enhancement, linear node matching, secondary edge alignment and double-layer negative sample selection, so that precise classification of fields to which papers belong is realized.
Owner:PEKING UNIV

Document classification method, computer equipment and storage medium

The invention discloses a document classification method, computer equipment and a storage medium. The document classification method comprises the following steps: receiving a target classification original document; specific information used for referring to a specific instance in the target classification original document is filtered out through a first generative model, common features of the category to which the target classification original document belongs are reserved, and standardized description of the target classification original document is generated; performing similarity retrieval in a rule database based on the standardized description, and recalling a plurality of candidate classification rules; and inputting the standardized description and the candidate classification rule into a second generative model to obtain a final classification result. According to the method, document expressions are aligned with classification rule expressions through semantic conversion, and the document classification accuracy is improved to a certain extent by adopting a mode of combining retrieval and reasoning.
Owner:浙江太美医疗科技股份有限公司

Document tag management methods, devices and storage media

This invention discloses a document tag management method and apparatus. The method includes: if a document event is triggered, determining the target file; matching the target file's directory and file type according to a directory policy to determine if the target file enters the tag processing stage; reading the target file's current tag information and processing the current tag information according to policy conditions. This invention, through the collaboration of directory policy matching, automatic backup, tag reading and writing, and failure retries, can significantly reduce the cost of manual tag maintenance, improve the accuracy and timeliness of document classification, and reduce the risk of data leakage. For industries such as finance, government, and manufacturing, this solution helps to improve knowledge asset management capabilities through compliance audits, and has significant economic and social benefits.
Owner:SHENZHEN LEAGSOFT TECH

Methods, Systems, and Computer-readable Media for Training Document Type Prediction Models, and Use Thereof for Creating Accounting Records

PendingUS20260187413A1EngineeringData mining
Methods are described that include: determining a candidate document associated with a user of an accounting system; providing the candidate document to a numerical representation generation model to generate a numerical representation of the candidate document; and providing the numerical representation to a document type attribute predictor to generate a predicted document type. The document type attribute predictor is configured to classify the document as one of a plurality of accounting document types.
Owner:XERO

A method and system for identifying and cutting a mixed bond

ActiveCN116092093Bachieve recognizabilityachieve the purpose of cuttingManufacturing computing systemsInstrumentsMinimum bounding rectangleEngineering
This invention provides a method and system for identifying and segmenting mixed-issue documents. The method includes the following steps: after detecting the documents, multiple mask data of individual documents are obtained; the mask data is filtered according to a set threshold to obtain mask data that meets certain conditions; after merging the mask data that meets the conditions, the minimum bounding rectangle is calculated; and the positions of the obtained minimum bounding rectangle are segmented using an array method. This method for identifying and segmenting mixed-issue documents detects different documents without missing any, achieving high accuracy and effectively improving the efficiency and accuracy of document classification while saving manpower, thus possessing significant application value.
Owner:金科览智科技(北京)有限公司

Information processing device, control method, program

The present invention aims to provide a mechanism that allows for the efficient verification of information related to relevant documents. [Solution] The system comprises a classification means for classifying documents into document groups, and a display control means for controlling the aggregation and display of information relating to the classified documents for each document group, wherein the display control means displays information indicating whether the document group contains multiple documents.
Owner:CANON MARKETING JAPAN INC +1

Document Classification

A computer implemented system for predicting a property or classification associated with document data. The system has a data extraction module configured to receive input document data and extract from the input document data a plurality of data sets, each data set having data of one of a plurality of data types. The system also has processing pathways each configured to process one of the plurality of data sets to generate a vector output representative of the data set processed by the processing pathway. The system has a vector concatenation layer configured to concatenate the vector outputs of each processing pathway to generate a concatenated vector, and a plurality of predictions heads. Each prediction head is configured to process the concatenated vector to generate a prediction variable indicative of a property or classification predicted to be associated with the input document data.
Owner:SAGE GLOBAL SERVICES LTD

Dynamic document classification

ActiveUS12670737B2Data ingestionData field
In an approach, a processor performs document layout analysis on a document generating a plurality of textual regions; extracts characteristics from each of the plurality of textual regions and associates the respective characteristics to the respective textual region as metadata; classifies each of the plurality of textual regions as an optical character recognition (OCR) region, non-OCR valuable region, or non-OCR non-valuable region using a classifier; performs OCR on each OCR region generating an OCR output; identifies associated constant OCR data from a constant OCR data repository for each non-OCR valuable region; merges the associated constant OCR data with the OCR output generating a complete OCR data for the received document; performs data extraction on the complete OCR data to identify data fields and key-value pairs generating extracted data; and determines whether the extracted data is valid based on a set of rules.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Grayware website detection with transfer learning

PendingUS20260189590A1Web siteData source
A model trainer obtains initial training data and refined training data to be used for training a classification model to detect grayware in Hypertext Markup Language (HTML) documents using transfer learning. The model trainer obtains the refined training data by collecting grayware HTML documents from a trusted data source(s), embedding and clustering the grayware HTML documents, and identifying and removing clusters having low confidence of corresponding to known grayware campaigns. The model trainer then trains a baseline model to classify HTML documents as grayware or benign with the initial training data, replaces the classification head of the baseline model with a new classification head to obtain a refined model, and further trains via the refined model via transfer learning with the refined training data.
Owner:PALO ALTO NETWORKS INC

An ai semantic analysis and data processing method based on a large language model

The application relates to the field of data processing, in particular to an AI semantic analysis and data processing method based on a large language model. First, professional terms and context windows in vertical field texts are extracted, hidden layer vectors are extracted by using the large language model, and spatial dispersion is calculated, and the terms are divided into three-state working modes according to semantic stability. For polysemous terms, independent semantic branches are identified through vector clustering, co-occurrence word matching and vector distance determination are used to realize accurate semantic attribution in a new context, and dynamic segmented adjustment of static term frequency (TF-IDF) weight is carried out accordingly, and finally the improved weight is injected into the downstream analysis process. The application gives the same term different weight expressions in different contexts, and the low-overhead cascade disambiguation mechanism efficiently solves the polysemy problem of terms, avoids confusion of structured data and error clustering of documents, and significantly improves the accuracy of text matching, information extraction and document classification and the like.
Owner:ZHONGNAN INFORMATION TECH (SHENZHEN) CO LTD +1

Intelligent document classification, label generation and abstract extraction method and system

The invention discloses an intelligent document classification, label generation and abstract extraction method and system, and belongs to the technical field of natural language processing and information management, and the method comprises the following steps: document preprocessing: carrying out data cleaning and standardization processing on an uploaded document; text feature extraction: extracting semantic features in a document through a deep learning model, and identifying key information in a text; document classification and label generation: automatically classifying the documents by using the classification model, and generating related labels for the documents; document abstract and key information extraction: automatically generating an abstract of the document by adopting a text abstract algorithm, and extracting core information of the document; and post-processing and optimization: verifying and optimizing the generated classifications, labels and abstracts to ensure the accuracy and correlation thereof. According to the method, manual intervention can be greatly reduced, the document management efficiency can be improved, the information retrieval accuracy is improved, and the labor cost is reduced.
Owner:INSPUR QILU SOFTWARE IND

Artificial intelligence-based document classification method and device, computer device and medium

The application is suitable for the technical field of document classification, and particularly relates to a document classification method and device based on artificial intelligence, computer equipment and medium. The application determines a primary document category by matching a to-be-classified document with a plurality of primary category document templates, and performs rough classification on the to-be-classified document. The application performs error correction processing on a preliminary text extraction result of the to-be-classified document according to a preset dictionary to obtain an optimized text extraction result, thereby improving the accuracy of character recognition. The application performs keyword matching on the optimized text extraction result to obtain document keywords, thereby improving the extraction accuracy of category semantic information. The application inputs the document keywords into a trained document classification model corresponding to the primary document category, and more accurately determines a secondary document category on the basis of the primary document category, thereby improving the accuracy of document classification. The application separates the primary document category and the secondary document category through a preset separator, thereby improving the orderliness of actual document categories.
Owner:CHINA PING AN LIFE INSURANCE CO LTD

Machine learning powered cloud sandbox for malware detection in portable document format (PDF) files

ActiveUS20260099598A1Digital data protectionPlatform integrity maintainanceEngineeringPortable document format
A cloud-based network security system (NSS) is described. The NSS uses a sandbox to safely open and extract information about a PDF file and uses machine learning algorithms to analyze the information to predict whether the PDF file contains malware. Specifically, dynamic information about the PDF file is captured while it is open in the sandbox. Static information is extracted from the PDF file as well. The dynamic and static information is input to an AI or machine learning model trained to provide an output indicating a prediction of whether the PDF file contains malware. A verdict engine uses the output from the AI or machine learning model to classify the document as malicious or clean. Security policies can then be applied based on the classification.
Owner:NETSKOPE INC

Security scheduling model response method based on sparse expert model and self-supervised learning

The invention discloses a security scheduling model response method based on a sparse expert model and self-supervised learning. The method comprises a data processing and document classification mechanism based on an intelligent algorithm, a scheduling document processing mechanism based on the sparse expert model, and sparse expert model improvement based on self-supervised learning. According to the method, only the expert module related to the current task is activated through a dynamic routing mechanism of the sparse expert model, so that the waste of computing resources is reduced, the task processing efficiency is remarkably improved, a load balancing mechanism is adopted, the load between the expert modules is ensured to be uniformly distributed under high-concurrency and complex scheduling tasks, and the task scheduling efficiency is improved. And the problem of processing delay is effectively avoided. Meanwhile, due to the introduction of the self-supervised learning technology, the semantic understanding capability of the system on the scheduling task is enhanced, and accurate analysis of task data and high-quality text generation are ensured. A sparse expert model is combined with self-supervised learning.
Owner:CHINA SOUTHERN POWER GRID COMPANY

Automated knowledge graph populator for data selection

A computing device records feature embeddings of each document as a document node of a knowledge graph and connects each document node of the knowledge graph with one or more engagement edges based on engagement telemetry data indicating a measure of engagement with the documents stored in the document datastore. The computing device trains a graph neural network using the knowledge graph populated with each document node and the one or more engagement edges. The computing device may generate a feature embedding for the document query and classify one or more documents from the document datastore as relevant to the document query using the graph neural network based on the feature embedding of the document query.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Medical record classification method, system, terminal and storage medium

The application provides a medical record classification method, system, terminal and storage medium, the method comprises the following steps: carrying out text recognition on the medical record image of a medical record to be classified to obtain a medical record text, and carrying out keyword detection on the medical record image according to the medical record text; if the keyword detection of the medical record text is qualified, carrying out keyword matching on the medical record text to obtain a keyword matching result; if the keyword matching result does not satisfy a keyword classification condition, carrying out regular matching on the medical record text to obtain a text regular matching result; if the text regular matching result does not satisfy a regular matching classification condition, carrying out document classification prediction on the medical record text to obtain a document classification prediction result; if the document classification prediction result does not satisfy a preset classification condition, carrying out target detection on the medical record image, and determining the document category of the medical record to be classified according to the target detection result. The application can effectively determine the document category of a paper medical record to be classified, and improves the accuracy of medical record classification.
Owner:BEIJING UNISOUND INFORMATION TECH CO LTD

An automobile data document classification method, device, equipment and medium

The application provides a car data document classification method and device, equipment and medium, relates to the technical field of data document classification, and obtains a text feature vector, a similarity feature vector and an initial classification feature vector corresponding to a to-be-classified car data document; generates a fusion classification feature vector corresponding to the to-be-classified car data document according to the text feature vector, the similarity feature vector and the initial classification feature vector; inputs the fusion classification feature vector and the similarity feature vector into a preset document classification model obtained by training a multi-Fisher classifier according to the fusion classification feature vectors of historical car data documents and corresponding classification labels; and obtains a classification result of the to-be-classified car data document according to an output result of the preset document classification model. The car data document classification model of the multi-classifier fusion is constructed, so that the generalization ability of the traditional classification model is improved, and the classification precision and efficiency are improved.
Owner:CHINA AUTOMOBILE INTELLECTUAL PROPERTY (GUANGZHOU) CO LTD

Systems and methods for resolving large taxonomy selection

A method for classifying a document into a hierarchical taxonomy associated with a corpus of documents, the document being associated with document information, the hierarchical taxonomy comprising a plurality of levels with each level comprising one or more nodes, each node comprising a label; the method may include inputting the taxonomy and the document information into a large language model, inputting a prompt into the large language model to cause the large language model to output one or more nodes of the taxonomy for classifying the document based on the document information, and classifying the document into each of the nodes output by the large language model.
Owner:ELSEVIER INC

Bill classification storage device

The utility model relates to the related technical field of storage devices, in particular to a bill classification storage device which comprises a classification box arranged in a box body. The handle is arranged on the front surface of the classification box; a pressing assembly is arranged in the box body, and the pressing assembly comprises a sliding groove formed in the box body; the protruding strips are arranged on the two sides of the classification box and are in sliding connection with the sliding grooves. The V-shaped groove is formed in the top of the classification box; according to the storage device, the pressing assembly can extrude the protruding block to slide while the classification box is pulled out, the pressing plate is driven to move upwards to relieve pressing on the bills, meanwhile, after the classification box is installed in the box body, the first spring rebounds to push the pressing plate to press the bills, and compared with an existing device, operation is simpler and more convenient.
Owner:CHONGQING THREE GORGES VOCATIONAL COLLEGE

Systems and methods for resolving large taxonomy selection

A method for classifying a document into a hierarchical taxonomy associated with a corpus of documents, the document being associated with document information, the hierarchical taxonomy comprising a plurality of levels with each level comprising one or more nodes, each node comprising a label; the method may include inputting the taxonomy and the document information into a large language model, inputting a prompt into the large language model to cause the large language model to output one or more nodes of the taxonomy for classifying the document based on the document information, and classifying the document into each of the nodes output by the large language model.
Owner:TABATABAEI SEYEDAMIN +5