Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

15 results about "Document clustering" patented technology

Document clustering (or text clustering) is the application of cluster analysis to textual documents. It has applications in automatic document organization, topic extraction and fast information retrieval or filtering.

Data Classification Models Using Feature Extraction and Clustering

A method for document classification is described. A first dataset of labeled corporate data, a second dataset of internal labeled documents for a customer, and a third dataset of unlabeled documents for the customer are obtained. A classification model is trained using the first dataset. The classification model is further trained using the second dataset. Feature extraction is performed on each of the unlabeled documents of the third dataset by vectorizing content and metadata of each unlabeled document into one or more vectors and concatenating the one or more vectors to obtain a fixed length vector. Each of the unlabeled documents of the third dataset is clustered into one or more clusters based on similarity between the fixed length vectors for each unlabeled document. The unlabeled documents in each of the clusters are automatically labeled using text summarization. The classification model is retrained using the automatically labeled documents.
Owner:PROOFPOINT INC

Dataset clustering and evaluation

Improved solutions for dataset clustering and evaluation are disclosed. Examples cluster a set of documents into set of clusters using a language model, in an iterative process. In second and later clustering tasks, the current cluster titles and descriptions are provided in the language model prompt, to avoid near-duplications. Upon determining that the set of clusters is sufficiently complete and representative of the set of documents, the tasking switches to classification of the set of documents into the set of clusters using a language model. Classification continues until a sufficient percentage of the set of documents is classified. Some examples use batching, to avoid overloading the language model(s). In some examples, different language models are used for clustering and classification. Some examples use intruder detection to determine the quality of the clustering. This process provides superior performance on classifying documents having little configuration control, such as website feedback from consumers.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Document processing device

PCT designated stageWO2026062738A1Natural language data processingMeta clusteringEngineering
Provided is a document processing device including a metadata assignment unit, a clustering unit, a difference extraction unit, and an output unit. The metadata assignment unit converts each of a plurality of documents into a vector on the basis of word occurrence counts in the document, and assigns, as metadata to the corresponding document, the vector and an issue time of the document. The clustering unit clusters the documents on the basis of similarity in the word occurrence counts represented by the vectors assigned by the metadata assignment unit. For each cluster formed by the clustering unit, the difference extraction unit rearranges the documents in order of issue time and extracts differences between the documents rearranged in that order. The output unit outputs the differences between the documents.
Owner:NT T INC

Expediting automated near-duplicate detection for new text documents

Techniques described herein provide for automated near-duplicate detection for new text documents given text documents that were previously processed using automated near-duplicate detection for text documents. In one example, a system can receive new documents and documents that were previously processed using a predefined processing technique for automated near-duplicate detection. The system can process the new documents and cluster the new documents into multiple predefined clusters previously identified using the predefined processing technique. For each predefined cluster including at least one new document, the system can generate document groups by determining similarity scores using the predefined processing technique as applied to the documents in the predefined clusters. The system can identify a representative document for each document group and generate an output data structure including the document groups and the representative document for each group.
Owner:SAS INSTITUTE INC

Method, device and storage medium for detecting internet harmful event

The application discloses an internet harmful event detection method, device and storage medium, constructs or updates a keyword knowledge graph, divides the knowledge graph into a plurality of subgraphs, clusters documents into harmful events by using a clustering algorithm, inserts each harmful event into a harmful story tree by updating each harmful event, or creates a new harmful story tree according to the harmful event, realizes harmful content detection and classification in mobile internet and internet, and the purpose is to find harmful events from massive webpage and document data, organizes harmful events in a reasonable harmful story tree in an online mode, proposes a two-layer webpage document clustering algorithm based on a knowledge graph, extracts harmful events from a large amount of webpage text or webpage related news, and organizes events into a story tree by using an online algorithm after new webpage and document data arrive, and is more effective than a traditional keyword-based harmful algorithm in the aspect of harmful event extraction.
Owner:ZHUHAI GAOLING INFORMATION TECH COLTD

t-sne-based high-entropy alloy document keyword implicit semantic mining method and device

The application provides a t-SNE-based high-entropy alloy document keyword implicit semantic mining method and device, and relates to the technical field of material document clustering. The method comprises the following steps: obtaining and preprocessing an entropy alloy document data set; training a high-entropy alloy document word vector; representing a high-entropy alloy document; clustering analyzing a high-entropy alloy document vector; and mining a high-entropy alloy hotspot and development trend according to a clustering analysis result. The high-entropy alloy clustering method comprehensively considers the abstract and keyword information of the high-entropy alloy document, fully mines the implicit semantic information from the high-entropy alloy abstract, fully mines the implicit semantic information of the high-entropy alloy document keyword, and through the t-SNE visualized high-entropy alloy document clustering, the distribution and clustering effect of the high-entropy alloy document can be more intuitively observed, thereby providing technical support for mining potential information such as the development trend of the high-entropy alloy.
Owner:UNIV OF SCI & TECH BEIJING

Multi-source document clustering method, system and equipment based on machine learning

The invention belongs to the technical field of natural language processing, particularly relates to a multi-source document clustering method, system and equipment based on machine learning, and aims to solve the problems of weak representation capability, poor semantic consistency, insufficient label interpretation and the like in document clustering. The method comprises the steps of obtaining a plurality of documents to be clustered; taking the document and the target document abstract word number as input data of a first language model, and generating a document abstract according with the target document abstract word number through the first language model; performing semantic analysis on the document abstract, and converting the document abstract into an embedded vector based on a semantic analysis result; dividing the document into a plurality of document clusters based on the similarity between the embedding vector of any document and the embedding vectors of other documents; generating a clustering label for representing a document clustering semantic topic; and establishing a document set under the same document cluster according to the cluster labels, and displaying the document set in a structured view. According to the scheme, the cross-domain document set can be effectively normalized and clustered.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

Method and system for extracting data from documents and automatically modifying data item of the extracted data based on guidance retrieved from feedback file

Improved techniques for extraction of data from documents, namely, from images of documents, so as to better enable software automation. The software automation can, for example, be provided by software robots of RPA systems. The improved techniques can provide automated feedback-based modification of data extracted from an image of a document through use of previously used validation guidance provided by a user in validating extracted data from a same or similar document. In one embodiment, the automated feedback-based modification can locate an appropriate feedback file through use of document fingerprints and document clusters. Then, guidance from the appropriate feedback file can be used to automatically modify at least a portion of the data extracted data from the image based on the guidance retrieved from the feedback file. Advantageously, the improved techniques can reduce the need for user participation in validation of data extracted from images of documents, and yield greater and more accurate automated data extraction.
Owner:AUTOMATION ANYWHERE INC

Tree diagram and knowledge graph retrieval enhanced large model inference method and system

The application discloses a tree diagram and knowledge graph retrieval enhanced large model reasoning method and system, which comprises the following steps: cutting a document into text blocks, identifying entities, relationships and factual descriptions of the text blocks to form a knowledge graph of the text blocks; constructing a factual summary of the document based on the knowledge graph of the text blocks, recursively clustering all documents using a Gaussian mixture model to generate a similarity tree structure diagram; obtaining a user question, selecting a reasoning path from the similarity tree structure diagram through a large language model debate iteration method, obtaining a corresponding document cluster, then pruning the cluster based on the user question and the confidence score of the document, and selecting a plurality of text blocks from the pruned document, and finally generating a reasoning result based on the knowledge graph of the selected text blocks. The application relates to the technical field of natural language processing, and can realize retrieval enhancement based on a tree diagram and a knowledge graph, accurately query relevant information when a large model answers a complex multi-hop question, and effectively guarantee reasoning accuracy.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Machine learning system for multi-domain long document clustering

A model pipeline has been created that allows an organization to gain a clear understanding of contents of a collection of documents despite varying lengths and multiple domain content. A collection of documents with content of different domains are normalized by summarizing the documents. For summarizing, the pipeline uses a language model that has been trained for text summarization across domains to a constrained summary length or length limit. The pipeline extracts embeddings of the summaries and clusters the summary embeddings. The pipeline uses the summaries to label the clusters using another language model that has been trained to generate a label or name for a cluster based on a set of summaries corresponding to a sample of cluster members. The labeled clusters can then be used to generate an organized presentation of the content of the documents in the document collection.
Owner:PALO ALTO NETWORKS INC

Encryption management method and system based on edge algorithm

The invention relates to the technical field of data processing, and discloses an encryption management method and system based on an edge algorithm, and the system comprises a collection and analysis module which extracts literature features of a target management literature and converts the literature features into domain feature vectors; the key generation module screens edge nodes from three dimensions of identity legality, security creditworthiness and resource capability to determine a target edge node, performs literature clustering processing on a domain feature vector to determine a domain subset, and constructs an edge node trust chain based on the target edge node and the domain subset; encrypting literatures in the domain subset based on a symmetric encryption algorithm of an edge node trust chain, and deriving a basic key based on a unique identifier of the literatures to determine an encryption key; the encryption adjustment module adjusts the storage position of the literature in the domain subset based on the encryption management strategy; and the encryption management module adjusts the encryption backup mode based on the storage data volume. According to the invention, the reliability of encryption management is ensured.
Owner:SHENZHEN KELVIN TECH CO LTD

Topic generation device and topic generation method

Improve the quality of generated topics by utilizing metadata attached to document data. [Solution] The topic generation device (1) includes an acquisition unit (11) that acquires document data relating to multiple documents, a first feature generation unit (12) that generates first feature quantities for multiple documents, a knowledge graph generation unit (13) that generates a knowledge graph, a second feature generation unit (14) that generates second feature quantities for multiple documents based on the knowledge graph, a feature merging unit (15) that combines the first feature quantities and the second feature quantities to generate a first combined feature quantity, a clustering unit (16) that generates multiple document clusters, and a topic generation unit (17) that generates topics for multiple documents.
Owner:UNIV OF TSUKUBA +1

Document clustering and ranking method, system, device and medium based on language large model

The application discloses a document clustering and sorting method, system, device and medium based on a language large model, wherein the method comprises the following steps: collecting document data for structured processing and preprocessing; inputting the document content into the language large model to obtain vectorized representation; using a clustering algorithm on the vectorized document content to obtain a document cluster and a similarity matrix in the document cluster, sorting the documents in each document cluster according to the weighted sum of the similarity matrix, and taking the top ten document titles as seed document titles; counting the number of each level document in the document cluster, the total number of documents and the weighted sum of the document cluster correlation coefficient, and calculating the weighted sum of the three indexes to obtain the final score of each document cluster, and sorting according to the score; inputting the seed document title and the set prompt into the language large model to generate a short sentence as the class label of the document cluster. The application can make the document vectorization more accurate, the class sorting more scientific, and the generation of the class label more specific and automatic.
Owner:FUJIAN YIRONG INFORMATION TECH +1

AI semantic analysis and data processing method based on large language model

The invention relates to the field of data processing, in particular to an AI semantic analysis and data processing method based on a large language model, which comprises the following steps: firstly, extracting professional terms and context windows in a vertical field text, extracting hidden layer vectors by using the large language model, calculating spatial dispersion, and dividing the terms into three-state working modes according to semantic stability; aiming at ambiguous terms, independent semantic branches are identified through vector clustering, accurate semantic affiliation under a new context is realized by adopting co-occurrence word matching and vector distance judgment, dynamic segmentation adjustment is performed on static word frequency (TF-IDF) weights according to the accurate semantic affiliation, and finally, the improved weights are injected into a downstream analysis process. According to the method, the same term is endowed with differentiated weight expressions in different contexts, the problem of term ambiguity is efficiently solved through a low-overhead cascade disambiguation mechanism, structured data confusion and document error clustering are avoided, and the accuracy of tasks such as text matching, information extraction and document classification is remarkably improved.
Owner:ZHONGNAN INFORMATION TECH (SHENZHEN) CO LTD +1