Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

8 results about "Document clustering" patented technology

Document clustering (or text clustering) is the application of cluster analysis to textual documents. It has applications in automatic document organization, topic extraction and fast information retrieval or filtering.

Document processing device

PCT designated stageWO2026062738A1Natural language data processingMeta clusteringEngineering
Provided is a document processing device including a metadata assignment unit, a clustering unit, a difference extraction unit, and an output unit. The metadata assignment unit converts each of a plurality of documents into a vector on the basis of word occurrence counts in the document, and assigns, as metadata to the corresponding document, the vector and an issue time of the document. The clustering unit clusters the documents on the basis of similarity in the word occurrence counts represented by the vectors assigned by the metadata assignment unit. For each cluster formed by the clustering unit, the difference extraction unit rearranges the documents in order of issue time and extracts differences between the documents rearranged in that order. The output unit outputs the differences between the documents.
Owner:NT T INC

Method, device and storage medium for detecting internet harmful event

The application discloses an internet harmful event detection method, device and storage medium, constructs or updates a keyword knowledge graph, divides the knowledge graph into a plurality of subgraphs, clusters documents into harmful events by using a clustering algorithm, inserts each harmful event into a harmful story tree by updating each harmful event, or creates a new harmful story tree according to the harmful event, realizes harmful content detection and classification in mobile internet and internet, and the purpose is to find harmful events from massive webpage and document data, organizes harmful events in a reasonable harmful story tree in an online mode, proposes a two-layer webpage document clustering algorithm based on a knowledge graph, extracts harmful events from a large amount of webpage text or webpage related news, and organizes events into a story tree by using an online algorithm after new webpage and document data arrive, and is more effective than a traditional keyword-based harmful algorithm in the aspect of harmful event extraction.
Owner:ZHUHAI GAOLING INFORMATION TECH COLTD

Method and system for extracting data from documents and automatically modifying data item of the extracted data based on guidance retrieved from feedback file

Improved techniques for extraction of data from documents, namely, from images of documents, so as to better enable software automation. The software automation can, for example, be provided by software robots of RPA systems. The improved techniques can provide automated feedback-based modification of data extracted from an image of a document through use of previously used validation guidance provided by a user in validating extracted data from a same or similar document. In one embodiment, the automated feedback-based modification can locate an appropriate feedback file through use of document fingerprints and document clusters. Then, guidance from the appropriate feedback file can be used to automatically modify at least a portion of the data extracted data from the image based on the guidance retrieved from the feedback file. Advantageously, the improved techniques can reduce the need for user participation in validation of data extracted from images of documents, and yield greater and more accurate automated data extraction.
Owner:AUTOMATION ANYWHERE INC

Machine learning system for multi-domain long document clustering

A model pipeline has been created that allows an organization to gain a clear understanding of contents of a collection of documents despite varying lengths and multiple domain content. A collection of documents with content of different domains are normalized by summarizing the documents. For summarizing, the pipeline uses a language model that has been trained for text summarization across domains to a constrained summary length or length limit. The pipeline extracts embeddings of the summaries and clusters the summary embeddings. The pipeline uses the summaries to label the clusters using another language model that has been trained to generate a label or name for a cluster based on a set of summaries corresponding to a sample of cluster members. The labeled clusters can then be used to generate an organized presentation of the content of the documents in the document collection.
Owner:PALO ALTO NETWORKS INC

Encryption management method and system based on edge algorithm

The invention relates to the technical field of data processing, and discloses an encryption management method and system based on an edge algorithm, and the system comprises a collection and analysis module which extracts literature features of a target management literature and converts the literature features into domain feature vectors; the key generation module screens edge nodes from three dimensions of identity legality, security creditworthiness and resource capability to determine a target edge node, performs literature clustering processing on a domain feature vector to determine a domain subset, and constructs an edge node trust chain based on the target edge node and the domain subset; encrypting literatures in the domain subset based on a symmetric encryption algorithm of an edge node trust chain, and deriving a basic key based on a unique identifier of the literatures to determine an encryption key; the encryption adjustment module adjusts the storage position of the literature in the domain subset based on the encryption management strategy; and the encryption management module adjusts the encryption backup mode based on the storage data volume. According to the invention, the reliability of encryption management is ensured.
Owner:SHENZHEN KELVIN TECH CO LTD

Topic generation device and topic generation method

Improve the quality of generated topics by utilizing metadata attached to document data. [Solution] The topic generation device (1) includes an acquisition unit (11) that acquires document data relating to multiple documents, a first feature generation unit (12) that generates first feature quantities for multiple documents, a knowledge graph generation unit (13) that generates a knowledge graph, a second feature generation unit (14) that generates second feature quantities for multiple documents based on the knowledge graph, a feature merging unit (15) that combines the first feature quantities and the second feature quantities to generate a first combined feature quantity, a clustering unit (16) that generates multiple document clusters, and a topic generation unit (17) that generates topics for multiple documents.
Owner:UNIV OF TSUKUBA +1

Document clustering and ranking method, system, device and medium based on language large model

The application discloses a document clustering and sorting method, system, device and medium based on a language large model, wherein the method comprises the following steps: collecting document data for structured processing and preprocessing; inputting the document content into the language large model to obtain vectorized representation; using a clustering algorithm on the vectorized document content to obtain a document cluster and a similarity matrix in the document cluster, sorting the documents in each document cluster according to the weighted sum of the similarity matrix, and taking the top ten document titles as seed document titles; counting the number of each level document in the document cluster, the total number of documents and the weighted sum of the document cluster correlation coefficient, and calculating the weighted sum of the three indexes to obtain the final score of each document cluster, and sorting according to the score; inputting the seed document title and the set prompt into the language large model to generate a short sentence as the class label of the document cluster. The application can make the document vectorization more accurate, the class sorting more scientific, and the generation of the class label more specific and automatic.
Owner:FUJIAN YIRONG INFORMATION TECH +1

AI semantic analysis and data processing method based on large language model

The invention relates to the field of data processing, in particular to an AI semantic analysis and data processing method based on a large language model, which comprises the following steps: firstly, extracting professional terms and context windows in a vertical field text, extracting hidden layer vectors by using the large language model, calculating spatial dispersion, and dividing the terms into three-state working modes according to semantic stability; aiming at ambiguous terms, independent semantic branches are identified through vector clustering, accurate semantic affiliation under a new context is realized by adopting co-occurrence word matching and vector distance judgment, dynamic segmentation adjustment is performed on static word frequency (TF-IDF) weights according to the accurate semantic affiliation, and finally, the improved weights are injected into a downstream analysis process. According to the method, the same term is endowed with differentiated weight expressions in different contexts, the problem of term ambiguity is efficiently solved through a low-overhead cascade disambiguation mechanism, structured data confusion and document error clustering are avoided, and the accuracy of tasks such as text matching, information extraction and document classification is remarkably improved.
Owner:ZHONGNAN INFORMATION TECH (SHENZHEN) CO LTD +1