Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

29 results about "Document similarity" patented technology

Document similarity (or distance between documents) is a one of the central themes in Information Retrieval. How humans usually define how similar are documents? Usually documents treated as similar if they are semantically close and describe similar concepts. On other hand “similarity” can be used in context of duplicate detection.

Information management method and system based on electronic bidding transaction platform

ActiveCN120672472AFinanceSemantic analysisRisk preventionDocument similarity
The invention discloses an information management method and system based on an electronic bidding transaction platform, and relates to the technical field of information management. By constructing a multi-dimensional and multi-level information management system, deep mining and intelligent identification of potential association relationships and behavior patterns among bidding units are realized, so that the capability of the electronic bidding platform in the aspects of risk prevention and control and abnormal behavior identification is effectively improved. By fusing multi-source heterogeneous data such as industry and commerce registration information, bidding history, project contracts and the like and introducing atlas modeling and semantic analysis technologies, key risk factors such as legal person penetration, historical cooperation, high bidding document similarity and the like can be accurately identified, then the risk association strength between bidding units is quantitatively evaluated, and the risk assessment efficiency is improved. And a visual map and an auxiliary auditing mechanism are formed, so that the judgment efficiency of supervisors is improved.
Owner:GRASP NETWORK TECH CO LTD +1

Retrieval enhancement method for threat intelligence mapping knowledge domain

The retrieval enhancement method for the threat intelligence knowledge graph provided by the invention comprises the following steps: preprocessing threat intelligence information to obtain block intelligence documents; performing entity extraction on the partitioned intelligence document to obtain an intelligence entity group; performing expansion classification and entity attribute information completion on the entities in the intelligence entity group to obtain completed intelligence entities, determining a relationship between the intelligence entities, and establishing a threat intelligence knowledge graph according to the relationship between the intelligence entities; performing entity matching and block intelligence document similarity matching according to question and answer statement input to obtain a related intelligence document block set; weighting, mixing and rearranging the partitioned intelligence documents in the related intelligence document block set to obtain a final related intelligence document block set, and generating an input answer corresponding to the question and answer statement according to the final related intelligence document block set. By applying the method, high-dimensional semantic hierarchy representation can be better carried out on heterogeneous data source intelligence, semantic information can be well understood and inquired, and therefore the quality of large model generation content is enhanced.
Owner:GUANGZHOU UNIVERSITY

Document retrieval method based on multi-field information and outlier detection

The invention relates to a document retrieval method based on multi-field information and outlier detection. The method comprises the following steps: acquiring document query information for retrieving documents, and acquiring field similarity between the document query information and different field information of each candidate document; if at least one high outlier exists in the current field similarity between the document query information and any one piece of field information of the current candidate document, obtaining the document similarity between the document query information and the current candidate document according to each high outlier; if the high outlier does not exist in the current field similarity, obtaining the document similarity between the document query information and the current candidate document according to at least one current field similarity; and obtaining a document retrieval result corresponding to the document query information according to the document similarity between the document query information and each candidate document. By adopting the method, the correlation of document retrieval results can be improved.
Owner:SHENZHEN LANLING SOFTWARE CO LTD

Secure distributed document discovery via vector similarity

Secure distributed document discovery via vector similarity is performed by a data discovery service. The service receives a request to perform a search using input data. The service sends the input data and an embedding protocol to different data repositories that comprise vector databases. The service receives a results from the different data repositories; each result identifies documents and corresponding document similarity scores. The service generates an overall ranking of the documents according to the document similarity scores. The service returns a query result that identifies the at least a portion of the documents and indicates the ranking of the identified documents.
Owner:AMAZON TECH INC

Multi-modal document retrieval method and device and electronic equipment

The invention discloses a multi-mode document retrieval method and device and electronic equipment. The method comprises the following steps: acquiring document retrieval demand parameters; determining a document retrieval vector corresponding to the document retrieval demand parameter; a target document set is called, the target document set comprises to-be-selected documents corresponding to multiple to-be-selected document vectors respectively, the multiple to-be-selected document vectors are determined according to corresponding document fusion features, and the corresponding document fusion features are obtained according to the corresponding multi-modal data; respectively determining document similarities between the document retrieval vector and the plurality of document vectors to be selected; and according to the target document set and the plurality of document similarities, determining a target document corresponding to the document retrieval requirement parameter from the plurality of to-be-selected documents. According to the method and the device, the technical problem that a retrieval result is inaccurate when a multi-modal document is retrieved in related technologies due to complexity of multi-modal data is solved.
Owner:STATE GRID BEIJING ELECTRIC POWER CO +2

A text classification method based on document similarity

ActiveCN122285904BFeature vectorDocument similarity
The application discloses a text classification method based on document similarity, and belongs to the technical field of document processing. The method updates the word set of a document sentence through iterative optimization of a word segmentation mechanism and generates a first sentence vector, and generates a first similarity vector in combination with a sentence weight and the first sentence vector; a plurality of theme words are extracted from the word set according to an entity weight and a theme feature vector is generated, a knowledge supplement vector is extracted from a knowledge base based on the context vector of each theme word and the entity weight, a theme enhancement vector is generated in combination with the theme feature vector and the knowledge supplement vector, and a second similarity vector is generated based on the theme enhancement vector; after the first similarity vector and the second similarity vector are compared and optimized in terms of difference, a document similarity vector is obtained and a text category is output. Through iterative word segmentation optimization, theme-guided knowledge enhancement and a similarity comparison mechanism, the application can effectively improve the reliability and accuracy of text classification.
Owner:JIANGXI MECHANICAL & ELECTRICAL VOCATIONAL & TECH COLLEGE

Document recognition method and device, computer device and storage medium

The application relates to a method and device for identifying official documents, computer equipment and a storage medium. The method comprises the following steps: obtaining a to-be-identified file; extracting text information in the to-be-identified file; matching the text information with reference official document elements in a reference official document element set to obtain a first official document element set, wherein the first official document element set is composed of first official document elements matched successfully with the reference official document elements in the text information; calculating a target confidence degree corresponding to the to-be-identified file according to a first confidence degree corresponding to the first official document elements in the first official document element set; and determining that the to-be-identified file is an official document if the target confidence degree meets an official document similarity condition. The method can realize automatic identification of official documents and improve the identification efficiency of official documents.
Owner:QI AN XIN TECHNOLOGY GROUP INC +1

Knowledge management method and system based on document similarity detection and multi-modal knowledge base

PendingCN121998054ASemantic analysisInference methodsDocument similarityEngineering
The invention discloses a knowledge management method and system based on document similarity detection and a multi-modal knowledge base. The method comprises the steps of obtaining a target document; when a target document is put in storage, a similarity detection method based on a resource self-adaption mechanism is adopted, and existing documents are screened out from the multi-modal knowledge base to serve as a candidate document set; a similarity detection method based on a precision and efficiency optimal mechanism is adopted, and the candidate document set is used for conducting duplicate removal on the target document or storing the target document which is judged to be a new document; obtaining a question asked by the user; based on the multi-mode knowledge base subjected to duplicate removal or updating, multi-path mixed retrieval is conducted on the problem, a multi-path mixed retrieval result is subjected to two-stage sorting to obtain a rearrangement result, reference data are selected from the rearrangement result, the problem, cue words, context and the reference data are assembled into a large model input instruction according to a preset cue word template, and the large model input instruction is input into the multi-mode knowledge base. An answer generated by the large model is obtained and returned to the user; and comprehensive improvement of the knowledge utilization rate, the retrieval precision and the efficiency is realized.
Owner:BEIJING SIFANG JIBAO ENG TECH +1

Similarity searching across digital standards

One embodiment provides a method for identifying similar objects by performing document attribute comparisons, the method including: providing a digital standard system that includes a user interface and a data store of digital standards; receiving a request for a similarity comparison; performing the similarity comparison; generating a document similarity score for each of the digital standards within a group of digital standards; and displaying at least one of the digital standards from the group based upon the document similarity score. Other aspects are described and claimed.
Owner:SAE INT

A method for judging document similarity based on image, video and text content

ActiveCN115221856BNatural language data processingDocument similarityDegree of similarity
The present invention discloses a method for determining document similarity based on image, video, and text content simultaneously, comprising the following steps: S1: selecting two documents D1 and D2 of the same type, performing a hash calculation on them, and determining whether D1 and D2 are duplicate documents; if D1 and D2 are not duplicate documents, performing text, image, and video content similarity calculations on documents D1 and D2, respectively, setting weights for the text, image, and video similarities of the documents, obtaining document similarity through weighted calculation, and comparing the document similarity with a preset threshold to draw a conclusion on document similarity. The present invention determines both text similarity and image and video similarity, calculates document similarity after comprehensive analysis, and prompts a human to perform a recheck for subsequent processing.
Owner:CHINA TELECOM CORP LTD

A smart legal query method based on a multi-round pruning Skyline algorithm

The application discloses a kind of wisdom legal inquiry method based on multi-wheel pruning Skyline algorithm, comprising: local equipment of legal department obtains local legal case according to the query command issued by central server;The local legal case obtained is uploaded to the central server in encrypted form, the central server carries out integration analysis, and the legal case with higher comprehensive document similarity is sent to local equipment.The scheduling strategy module of the central server obtains the physical node of local legal department or the corresponding legal department priority information according to the attributes such as name, gender, native place, address, case name, case-cracking event, capture event and report event in legal case.The application can make legal information retrieval personnel free from reading a large number of cases, save a lot of time and manpower and material resources, especially provides technical guarantee for finding similar cases, case-cracking clues and inducing crime trend.
Owner:DALIAN UNIV

A generative multi-document summarization method using entity explicit graph

The present application provides a generative multi-document summarization method using Entity explicit graph to solve the problems of the prior art. The Entity graph based on SimBert is used instead of the traditional document similarity graph to express the connection between documents. Entity extraction is performed on each paragraph, and the Entity of each paragraph is spliced into a sentence. Then, the cosine similarity between all paragraphs is calculated using SimBert. After threshold filtering, an explicit graph representing the connection between paragraphs is obtained. The graphing method based on neural networks and aimed at the entities within the document is more effective than the inflexible rules defined by humans. In addition, the present application improves the attention mechanism and hierarchical graph attention mechanism of the graph perception, introduces a gating mechanism and a residual connection, so that the explicit graph and implicit graph information can be better integrated, and the implicit relationship learned by the attention mechanism is reserved with a residual path to ensure that the information learned by the network occupies a more important position, thereby guiding the generative multi-document summarization.
Owner:WUHAN UNIV

Document similarity determination method, computer-readable storage medium, electronic device, and computer program product

PendingCN122311165ASolve the problem of inaccurate similarityimprove accuracyDocument similarityDegree of similarity
This application provides a document similarity determination method, a computer-readable storage medium, an electronic device, and a computer program product. The method includes: determining the semantic similarity between a first document and a second document; determining the structural similarity between the first document and the second document; determining the weights of the semantic similarity and the structural similarity; and determining the similarity between the first document and the second document based on the semantic similarity, the weights of the semantic similarity, the structural similarity, and the weights of the structural similarity. Therefore, it can at least solve the problem in related technologies where document similarity determined based on extracted semantic information is not accurate enough, enabling a more comprehensive assessment of the similarity between documents and improving the accuracy of similarity determination.
Owner:ZTE CORP

A text classification method based on document similarity

PendingCN122285904AFeature vectorDocument similarity
This invention discloses a text classification method based on document similarity, belonging to the field of document processing technology. The method iteratively optimizes the word segmentation mechanism to update the word set of document sentences and generates a first sentence vector. A first similarity vector is generated by combining sentence weights and the first sentence vector. Multiple topic words are extracted from the word set based on entity weights to generate topic feature vectors. Knowledge supplement vectors are extracted from a knowledge base based on the context vectors of each topic word and entity weights. A topic enhancement vector is generated by combining the topic feature vector and the knowledge supplement vector. A second similarity vector is generated based on the topic enhancement vector. The first and second similarity vectors are then optimized through difference comparison to obtain the document similarity vector, and the text category is output. This invention, through iterative word segmentation optimization, topic-guided knowledge enhancement, and a similarity comparison mechanism, can effectively improve the reliability and accuracy of text classification.
Owner:JIANGXI MECHANICAL & ELECTRICAL VOCATIONAL & TECH COLLEGE

Document similarity detection method based on multi-modal content fusion

ActiveCN121882018ASemantic analysisInference methodsDocument similaritySemantic feature
The invention relates to the technical field of data processing and information security, in particular to a document similarity detection method based on multi-modal content fusion, which comprises the following steps: a multi-modal feature extraction step: analyzing a target document into text, image and video data, extracting deterministic hash fingerprints and constructing semantic feature vectors; a difference entropy value calculation step: calculating the Hash similarity, calculating the semantic similarity when the Hash similarity is lower than a threshold value, and mapping the difference between the Hash similarity and the semantic similarity into a modal decision confidence entropy; a game strategy calculation step: collecting system calculation resource load data, and calculating a cost perception factor by using a dynamic game strategy model in combination with confidence entropy; a dynamic fusion and survival step: distributing an asymmetric fusion weight based on the perception factor, and generating a detection result; when the load exceeds the limit and the confidence entropy indicates a high risk, triggering a degradation survival mechanism; according to the invention, the recall ratio and throughput can be automatically balanced in a resource-limited scene, and system avalanche caused by single-point attack is prevented.
Owner:CHENGDU YOUA NETWORK TECH CO LTD

A legal and regulatory retrieval system and method based on semantic information weighting

ActiveCN116340485BSemantic analysisMachine learningAnalytic modelDocument similarity
The present invention relates to a legal and regulatory retrieval system and method based on semantic information weighting, belonging to the field of information. The system comprises an index generation module, a domain analysis module, a semantic weight assignment module, and a score normalization module. The index generation module is used to segment and index the data in the database by introducing an IK Analyzer, generate an index information table, and merge the data to form an index database; the domain analysis module is used to input the query statement input by the user into a trained domain analysis model for domain analysis, and match the domain analysis results with the index database to obtain domain-related temporary index information; the semantic weight assignment module is used to assign weights to each term in the query statement, and incorporate the weights assigned to each term into the document similarity calculation to comprehensively obtain the document score; and the score normalization module is used to normalize the document score. The present invention effectively improves the retrieval accuracy and optimizes the retrieval results.
Owner:SHANDONG UNIV

Similar document search device and program

PendingJP2025132027AMetadata based other databases retrievalDocument similarityDocumentation
To provide a similar document search device and a program which allow search of a document which fits a context.SOLUTION: A similar document search server 1 includes: an item-set storage unit 23 for storing, in association with stored documents, a plurality of item sets each being a combination of one item name and one item value corresponding thereto; an item-set extraction unit 14 for extracting a plurality of item sets from a target document image; and a similarity calculation unit 15 for calculating document similarity by using the plurality of item sets extracted by the item-set extraction unit 14 and the plurality of item sets corresponding to the plurality of stored documents stored in the item-set storage unit 23.SELECTED DRAWING: Figure 1
Owner:DAI NIPPON PRINTING CO LTD

Automatic duplicate checking and rewriting method and device for document and program product

PendingCN121328511ANatural language data processingLinguistic modelDocument similarity
The invention relates to the technical field of artificial intelligence, and discloses an automatic document duplicate checking and rewriting method and device and a program product, and the method comprises the steps: obtaining a target document, carrying out duplicate checking comparison on the target document according to a reference document, generating a corresponding document similarity, and carrying out the matching of terms in the target document according to a user-defined term library. Proprietary terms in the target document are determined, other contents except the proprietary terms in the target document are rewritten according to the language model, and a rewritten document is obtained. According to the method, the corresponding document similarity is generated through automatic duplicate checking comparison, the special terms in the document are protected, then other contents except the special terms in the target document are rewritten according to the language model, and the two steps of duplicate checking and rewriting are connected, so that the number of times of repeatedly switching tools by a user is reduced, the labor cost and the knowledge cost are reduced, and the user experience is improved. The process of duplicate checking and rewriting is shortened, and the efficiency of duplicate checking and rewriting is improved.
Owner:GLODON CO LTD

Methods and systems for transfer learning of deep learning model based on document similarity learning

ActiveUS12469322B2Mathematical modelsSemantic analysisDocument similarityDocumentation
Disclosed is a method and system for transfer learning of a deep learning model based on a document similarity learning. A transfer learning method may include pre-training, by the at least one processor, a similarity model to output a similarity between documents, generating, by the at least one processor, a fine tuning model by replacing a first output function of the pre-trained similarity model with a second output function, and training, by the at least one processor, the fine tuning model to output a score for a document input to the fine tuning model.
Owner:NAVER CORP

Dual-mode intelligent duplicate checking and similarity quantitative evaluation system for science and technology project management

The invention belongs to the technical field of document similarity detection, knowledge management and project management informatization, and particularly relates to a science and technology project management-oriented bimodal intelligent duplicate checking and similarity quantitative evaluation system, which comprises the following steps of: firstly, carrying out preprocessing and feature extraction on data of three modalities of a text, a graphic representation and a table, and associating traceability meta-information; then mapping entities in texts, diagrams and tables to a unified project semantic map to form a node set; screening key nodes as anchor points through node centrality and embedded clustering density, and establishing a reverse index; meanwhile, deep local matching is carried out in the alignment stage; on the basis of the method, complementary fusion of text, graphical representation, structure and table information is realized through construction of a multi-modal semantic map and an anchor point driven local matching mechanism, and the accuracy and robustness of similarity judgment are remarkably improved.
Owner:FOSHAN POWER SUPPLY BUREAU GUANGDONG POWER GRID

Document retrieval method based on multi-field information and outlier detection

The application relates to a document retrieval method based on multi-field information and outlier detection. The method comprises the following steps: acquiring document query information used for retrieving a document, and acquiring field similarity between the document query information and different field information of each candidate document; if at least one high outlier exists in current field similarity between the document query information and any field information of a current candidate document, obtaining document similarity between the document query information and the current candidate document according to each high outlier; if no high outlier exists in the current field similarity, obtaining the document similarity between the document query information and the current candidate document according to at least one current field similarity; and obtaining a document retrieval result corresponding to the document query information according to the document similarity between the document query information and each candidate document. The method can improve the relevance of the document retrieval result.
Owner:SHENZHEN LANLING SOFTWARE CO LTD

Policy-aware knowledge base deduplication

PendingUS20260252540A1Document similarityData mining
Auditable removal of duplicate records from a knowledge base is disclosed. A similarity graph is constructed with documents of the knowledge base as nodes, with similarity edges between the nodes representing inter-document similarity exceeding a similarity threshold. The documents of the similarity graph are clustered by similarity to provide a plurality of clusters within the similarity graph. A set of safeguarded documents of the knowledge base is provided and, for at least one cluster of the plurality of clusters, a representative document is selected, and / or a non-representative, non-safeguarded document is removed from the cluster. A retain set of retained documents and / or a prune set of removed documents may be constructed to improve knowledge base health and output quality of downstream retrieval-augmented generation.
Owner:CIBC

Retrieval enhancement generation system knowledge poisoning attack defense method

The invention discloses a knowledge poisoning attack defense method for a retrieval enhancement generation system, and relates to the technical field of artificial intelligence security and graph neural networks. The method solves the problems that an existing defense method depends on isolated detection of single document content and ignores the structural relation between documents and the interactive context of the documents and user query, so that high retrieval relevance but low semantic consistency of malicious documents is difficult to capture and the like. According to the method, query and retrieval documents are modeled into a heterogeneous graph structure, semantic embedding and graph structure features are fused, and a global graph containing query-document edges and document-document similar edges is constructed; document nodes are classified through an edge type perceived graph attention network, and malicious injection documents are recognized. The defense capability of the RAG system on knowledge poisoning attacks is improved, and safe and reliable external knowledge input is provided for a large language model.
Owner:CHANGCHUN UNIV OF SCI & TECH

Normalizing disparate inputs between electronic documents

InactiveUS20250335787A1Mathematical modelsKnowledge based modelsElectronic documentDocument similarity
A computer-implemented method of normalizing disparate inputs between electronic documents, including determining, for each feature of each computing component, an occurrence probability of the feature across the electronic documents; identifying a predefined weight of each feature of each computing component; calculating a heuristic weight of the feature based on i) the predefined weight of the feature and ii) the occurrence probability of the feature; determining a minimum and a maximum heuristic weight of each of the features of the computing component; determining a computing component similarity ratio of the computing component between any subset of the electronic documents based on the minimum and the maximum heuristic weight of the computing component of each electronic document of the subset; determining a document similarity ratio between a particular electronic document and another electronic document based on the computing component similarity ratio of each computing component shared by the electronic documents.
Owner:DELL PROD LP

Method and system for using robotic process automation to provide real-time case assistance to client support professionals

PendingUS20260044798A1Office automationKnowledge representationData packDocument similarity
A case assistant is provided to client support professionals, which utilizes robotic process automation (RPA) technologies to analyze large amounts of data related to historical client cases that are similar to current open cases, data related to skilled experts associated with similar client cases, and data related to business exceptions. Several processes are utilized to provide this data to client support professionals, including a document similarity finder that utilizes a vector data collector, a tokenizer, a stop word remover, a relevance finder, and a similarity finder, several of which utilize a variety of machine learning technologies. Additional processes include a skilled experts finder and a business exceptions finder.
Owner:RIMINI STREET INC

A retrieval augmented generation system knowledge poisoning attack defense method

The application discloses a retrieval enhanced generation system knowledge poisoning attack defense method, relates to the technical field of artificial intelligence security and graph neural networks, and solves the problems that existing defense methods depend on isolated detection of single document content, ignore the structural relationship between documents and the interactive context of documents and user queries, and thus it is difficult to capture high retrieval relevance but low semantic consistency of malicious documents. The application models queries and retrieval documents as a heterogeneous graph structure, fuses semantic embedding and graph structure features, constructs a global graph containing query-document edges and document-document similarity edges, classifies document nodes through an edge type-aware graph attention network, and identifies malicious injected documents. The application improves the defense capability of RAG systems against knowledge poisoning attacks and provides safe and reliable external knowledge input for large language models.
Owner:CHANGCHUN UNIV OF SCI & TECH

Bidding document similarity analysis method and system based on OCR (Optical Character Recognition)

PendingCN121767833AOvercome the limitations of a single dimensionExplicit explainabilityCharacter and pattern recognitionDocument similarityThresholding
The invention discloses a bidding document similarity analysis method and system based on OCR (optical character recognition), belongs to the technical field of document intelligent recognition and electronic bidding and tendering anti-cheating, and aims to solve the problems of low precision and poor efficiency of bid and bid behavior recognition due to neglect of non-text features and deep semantics in an existing detection method. The method comprises the following steps of: synchronously extracting text contents, format structures and image elements by carrying out multi-mode OCR (Optical Character Recognition) on a bidding file; respectively extracting text semantic features, format structure features and image features based on the text semantic features, the format structure features and the image features; and the comprehensive similarity between the bidding files is calculated by fusing the multi-dimensional features, and similarity grading judgment is carried out according to a dynamic threshold value. According to the method, accurate and efficient detection of multi-dimensional similarities of texts, formats and images is realized, and the accuracy and automation level of bidding document similarities analysis are remarkably improved.
Owner:许敏星

Document vector similarity calculation method combined with semantic feature enhancement

PendingCN121052237ASemantic analysisEnergy efficient computingDocument similarityLexical item
The invention is suitable for the technical field of natural language processing and information retrieval, and provides a document vector similarity calculation method combined with semantic feature enhancement, which comprises the following steps of: firstly, carrying out standardization processing and vocabulary screening on an original document set, and constructing an initial TF-IDF matrix; then mapping the lexical items into a semantic structure formed by multiple classes of primitives, calculating semantic similarity among the lexical items and constructing a semantic similarity matrix; and after the matrix is subjected to sparse processing, generating a lexical item-document matrix with enhanced semantics, fusing the lexical item-document matrix with the original matrix, and calculating the final document similarity. In this way, the expression ability of document similarity calculation on the semantic level is improved, the calculation result is closer to the real semantic association, and higher accuracy and robustness are achieved.
Owner:CHINA UNIV OF MINING & TECH