Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

15 results about "Document similarity" patented technology

Document similarity (or distance between documents) is a one of the central themes in Information Retrieval. How humans usually define how similar are documents? Usually documents treated as similar if they are semantically close and describe similar concepts. On other hand “similarity” can be used in context of duplicate detection.

A text classification method based on document similarity

ActiveCN122285904BFeature vectorDocument similarity
The application discloses a text classification method based on document similarity, and belongs to the technical field of document processing. The method updates the word set of a document sentence through iterative optimization of a word segmentation mechanism and generates a first sentence vector, and generates a first similarity vector in combination with a sentence weight and the first sentence vector; a plurality of theme words are extracted from the word set according to an entity weight and a theme feature vector is generated, a knowledge supplement vector is extracted from a knowledge base based on the context vector of each theme word and the entity weight, a theme enhancement vector is generated in combination with the theme feature vector and the knowledge supplement vector, and a second similarity vector is generated based on the theme enhancement vector; after the first similarity vector and the second similarity vector are compared and optimized in terms of difference, a document similarity vector is obtained and a text category is output. Through iterative word segmentation optimization, theme-guided knowledge enhancement and a similarity comparison mechanism, the application can effectively improve the reliability and accuracy of text classification.
Owner:JIANGXI MECHANICAL & ELECTRICAL VOCATIONAL & TECH COLLEGE

Knowledge management method and system based on document similarity detection and multi-modal knowledge base

PendingCN121998054ASemantic analysisInference methodsDocument similarityEngineering
The invention discloses a knowledge management method and system based on document similarity detection and a multi-modal knowledge base. The method comprises the steps of obtaining a target document; when a target document is put in storage, a similarity detection method based on a resource self-adaption mechanism is adopted, and existing documents are screened out from the multi-modal knowledge base to serve as a candidate document set; a similarity detection method based on a precision and efficiency optimal mechanism is adopted, and the candidate document set is used for conducting duplicate removal on the target document or storing the target document which is judged to be a new document; obtaining a question asked by the user; based on the multi-mode knowledge base subjected to duplicate removal or updating, multi-path mixed retrieval is conducted on the problem, a multi-path mixed retrieval result is subjected to two-stage sorting to obtain a rearrangement result, reference data are selected from the rearrangement result, the problem, cue words, context and the reference data are assembled into a large model input instruction according to a preset cue word template, and the large model input instruction is input into the multi-mode knowledge base. An answer generated by the large model is obtained and returned to the user; and comprehensive improvement of the knowledge utilization rate, the retrieval precision and the efficiency is realized.
Owner:BEIJING SIFANG JIBAO ENG TECH +1

A generative multi-document summarization method using entity explicit graph

The present application provides a generative multi-document summarization method using Entity explicit graph to solve the problems of the prior art. The Entity graph based on SimBert is used instead of the traditional document similarity graph to express the connection between documents. Entity extraction is performed on each paragraph, and the Entity of each paragraph is spliced into a sentence. Then, the cosine similarity between all paragraphs is calculated using SimBert. After threshold filtering, an explicit graph representing the connection between paragraphs is obtained. The graphing method based on neural networks and aimed at the entities within the document is more effective than the inflexible rules defined by humans. In addition, the present application improves the attention mechanism and hierarchical graph attention mechanism of the graph perception, introduces a gating mechanism and a residual connection, so that the explicit graph and implicit graph information can be better integrated, and the implicit relationship learned by the attention mechanism is reserved with a residual path to ensure that the information learned by the network occupies a more important position, thereby guiding the generative multi-document summarization.
Owner:WUHAN UNIV

Document similarity determination method, computer-readable storage medium, electronic device, and computer program product

PendingCN122311165ASolve the problem of inaccurate similarityimprove accuracyDocument similarityDegree of similarity
This application provides a document similarity determination method, a computer-readable storage medium, an electronic device, and a computer program product. The method includes: determining the semantic similarity between a first document and a second document; determining the structural similarity between the first document and the second document; determining the weights of the semantic similarity and the structural similarity; and determining the similarity between the first document and the second document based on the semantic similarity, the weights of the semantic similarity, the structural similarity, and the weights of the structural similarity. Therefore, it can at least solve the problem in related technologies where document similarity determined based on extracted semantic information is not accurate enough, enabling a more comprehensive assessment of the similarity between documents and improving the accuracy of similarity determination.
Owner:ZTE CORP

A text classification method based on document similarity

PendingCN122285904AFeature vectorDocument similarity
This invention discloses a text classification method based on document similarity, belonging to the field of document processing technology. The method iteratively optimizes the word segmentation mechanism to update the word set of document sentences and generates a first sentence vector. A first similarity vector is generated by combining sentence weights and the first sentence vector. Multiple topic words are extracted from the word set based on entity weights to generate topic feature vectors. Knowledge supplement vectors are extracted from a knowledge base based on the context vectors of each topic word and entity weights. A topic enhancement vector is generated by combining the topic feature vector and the knowledge supplement vector. A second similarity vector is generated based on the topic enhancement vector. The first and second similarity vectors are then optimized through difference comparison to obtain the document similarity vector, and the text category is output. This invention, through iterative word segmentation optimization, topic-guided knowledge enhancement, and a similarity comparison mechanism, can effectively improve the reliability and accuracy of text classification.
Owner:JIANGXI MECHANICAL & ELECTRICAL VOCATIONAL & TECH COLLEGE

Document similarity detection method based on multi-modal content fusion

ActiveCN121882018ASemantic analysisInference methodsDocument similaritySemantic feature
The invention relates to the technical field of data processing and information security, in particular to a document similarity detection method based on multi-modal content fusion, which comprises the following steps: a multi-modal feature extraction step: analyzing a target document into text, image and video data, extracting deterministic hash fingerprints and constructing semantic feature vectors; a difference entropy value calculation step: calculating the Hash similarity, calculating the semantic similarity when the Hash similarity is lower than a threshold value, and mapping the difference between the Hash similarity and the semantic similarity into a modal decision confidence entropy; a game strategy calculation step: collecting system calculation resource load data, and calculating a cost perception factor by using a dynamic game strategy model in combination with confidence entropy; a dynamic fusion and survival step: distributing an asymmetric fusion weight based on the perception factor, and generating a detection result; when the load exceeds the limit and the confidence entropy indicates a high risk, triggering a degradation survival mechanism; according to the invention, the recall ratio and throughput can be automatically balanced in a resource-limited scene, and system avalanche caused by single-point attack is prevented.
Owner:CHENGDU YOUA NETWORK TECH CO LTD

Automatic duplicate checking and rewriting method and device for document and program product

PendingCN121328511ANatural language data processingLinguistic modelDocument similarity
The invention relates to the technical field of artificial intelligence, and discloses an automatic document duplicate checking and rewriting method and device and a program product, and the method comprises the steps: obtaining a target document, carrying out duplicate checking comparison on the target document according to a reference document, generating a corresponding document similarity, and carrying out the matching of terms in the target document according to a user-defined term library. Proprietary terms in the target document are determined, other contents except the proprietary terms in the target document are rewritten according to the language model, and a rewritten document is obtained. According to the method, the corresponding document similarity is generated through automatic duplicate checking comparison, the special terms in the document are protected, then other contents except the special terms in the target document are rewritten according to the language model, and the two steps of duplicate checking and rewriting are connected, so that the number of times of repeatedly switching tools by a user is reduced, the labor cost and the knowledge cost are reduced, and the user experience is improved. The process of duplicate checking and rewriting is shortened, and the efficiency of duplicate checking and rewriting is improved.
Owner:GLODON CO LTD

Dual-mode intelligent duplicate checking and similarity quantitative evaluation system for science and technology project management

The invention belongs to the technical field of document similarity detection, knowledge management and project management informatization, and particularly relates to a science and technology project management-oriented bimodal intelligent duplicate checking and similarity quantitative evaluation system, which comprises the following steps of: firstly, carrying out preprocessing and feature extraction on data of three modalities of a text, a graphic representation and a table, and associating traceability meta-information; then mapping entities in texts, diagrams and tables to a unified project semantic map to form a node set; screening key nodes as anchor points through node centrality and embedded clustering density, and establishing a reverse index; meanwhile, deep local matching is carried out in the alignment stage; on the basis of the method, complementary fusion of text, graphical representation, structure and table information is realized through construction of a multi-modal semantic map and an anchor point driven local matching mechanism, and the accuracy and robustness of similarity judgment are remarkably improved.
Owner:FOSHAN POWER SUPPLY BUREAU GUANGDONG POWER GRID

Document retrieval method based on multi-field information and outlier detection

The application relates to a document retrieval method based on multi-field information and outlier detection. The method comprises the following steps: acquiring document query information used for retrieving a document, and acquiring field similarity between the document query information and different field information of each candidate document; if at least one high outlier exists in current field similarity between the document query information and any field information of a current candidate document, obtaining document similarity between the document query information and the current candidate document according to each high outlier; if no high outlier exists in the current field similarity, obtaining the document similarity between the document query information and the current candidate document according to at least one current field similarity; and obtaining a document retrieval result corresponding to the document query information according to the document similarity between the document query information and each candidate document. The method can improve the relevance of the document retrieval result.
Owner:SHENZHEN LANLING SOFTWARE CO LTD

Policy-aware knowledge base deduplication

PendingUS20260252540A1Document similarityData mining
Auditable removal of duplicate records from a knowledge base is disclosed. A similarity graph is constructed with documents of the knowledge base as nodes, with similarity edges between the nodes representing inter-document similarity exceeding a similarity threshold. The documents of the similarity graph are clustered by similarity to provide a plurality of clusters within the similarity graph. A set of safeguarded documents of the knowledge base is provided and, for at least one cluster of the plurality of clusters, a representative document is selected, and / or a non-representative, non-safeguarded document is removed from the cluster. A retain set of retained documents and / or a prune set of removed documents may be constructed to improve knowledge base health and output quality of downstream retrieval-augmented generation.
Owner:CIBC

Retrieval enhancement generation system knowledge poisoning attack defense method

The invention discloses a knowledge poisoning attack defense method for a retrieval enhancement generation system, and relates to the technical field of artificial intelligence security and graph neural networks. The method solves the problems that an existing defense method depends on isolated detection of single document content and ignores the structural relation between documents and the interactive context of the documents and user query, so that high retrieval relevance but low semantic consistency of malicious documents is difficult to capture and the like. According to the method, query and retrieval documents are modeled into a heterogeneous graph structure, semantic embedding and graph structure features are fused, and a global graph containing query-document edges and document-document similar edges is constructed; document nodes are classified through an edge type perceived graph attention network, and malicious injection documents are recognized. The defense capability of the RAG system on knowledge poisoning attacks is improved, and safe and reliable external knowledge input is provided for a large language model.
Owner:CHANGCHUN UNIV OF SCI & TECH

Method and system for using robotic process automation to provide real-time case assistance to client support professionals

PendingUS20260044798A1Office automationKnowledge representationData packDocument similarity
A case assistant is provided to client support professionals, which utilizes robotic process automation (RPA) technologies to analyze large amounts of data related to historical client cases that are similar to current open cases, data related to skilled experts associated with similar client cases, and data related to business exceptions. Several processes are utilized to provide this data to client support professionals, including a document similarity finder that utilizes a vector data collector, a tokenizer, a stop word remover, a relevance finder, and a similarity finder, several of which utilize a variety of machine learning technologies. Additional processes include a skilled experts finder and a business exceptions finder.
Owner:RIMINI STREET INC

A retrieval augmented generation system knowledge poisoning attack defense method

The application discloses a retrieval enhanced generation system knowledge poisoning attack defense method, relates to the technical field of artificial intelligence security and graph neural networks, and solves the problems that existing defense methods depend on isolated detection of single document content, ignore the structural relationship between documents and the interactive context of documents and user queries, and thus it is difficult to capture high retrieval relevance but low semantic consistency of malicious documents. The application models queries and retrieval documents as a heterogeneous graph structure, fuses semantic embedding and graph structure features, constructs a global graph containing query-document edges and document-document similarity edges, classifies document nodes through an edge type-aware graph attention network, and identifies malicious injected documents. The application improves the defense capability of RAG systems against knowledge poisoning attacks and provides safe and reliable external knowledge input for large language models.
Owner:CHANGCHUN UNIV OF SCI & TECH

Bidding document similarity analysis method and system based on OCR (Optical Character Recognition)

PendingCN121767833AOvercome the limitations of a single dimensionExplicit explainabilityCharacter and pattern recognitionDocument similarityThresholding
The invention discloses a bidding document similarity analysis method and system based on OCR (optical character recognition), belongs to the technical field of document intelligent recognition and electronic bidding and tendering anti-cheating, and aims to solve the problems of low precision and poor efficiency of bid and bid behavior recognition due to neglect of non-text features and deep semantics in an existing detection method. The method comprises the following steps of: synchronously extracting text contents, format structures and image elements by carrying out multi-mode OCR (Optical Character Recognition) on a bidding file; respectively extracting text semantic features, format structure features and image features based on the text semantic features, the format structure features and the image features; and the comprehensive similarity between the bidding files is calculated by fusing the multi-dimensional features, and similarity grading judgment is carried out according to a dynamic threshold value. According to the method, accurate and efficient detection of multi-dimensional similarities of texts, formats and images is realized, and the accuracy and automation level of bidding document similarities analysis are remarkably improved.
Owner:许敏星