Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

79 results about "Duplicate detection" patented technology

Duplicate Detector is smart enough to find duplicate or similar records in many situations. Duplicate Detector for SugarCRM works on any existing or custom field of type varchar, name or phone. It prompts the user if the value has already been used while they are creating records in the edit view or quick create mode.

Duplication check system and method for paper generated by artificial intelligence

A duplication check system and method for paper generated by artificial intelligence includes steps: S1: the user uploading the academic paper to be detected to a system, and the system automatically extracting the title, the abstract, and the headline of each paragraph of the paper; S2: fusing the title, the abstract, and the headline of each paragraph of the paper with the contextual information of the paper and extracting theme features; S3: after the different themes of the paper are extracted, repeatedly using similar tones for each theme in all different AI tools to propose integrate text requirements, searching each theme for times of the number of repetitions of integration in each AI tool until no new content is obtained, matching all the obtained texts with the paper to be duplication checked, based on natural language understanding, for duplication check, and marking the matching repeated parts and indicating the sources.
Owner:WU JIANG

Multi-terminal code repetition detection and component reconstruction method and device, equipment and medium

The invention relates to the technical field of code analysis, and discloses a multi-terminal code repetition detection and component reconstruction method, device, equipment and medium, and the method comprises the steps: carrying out code analysis on a pre-acquired project code file, and identifying a code mapping relation to obtain an abstract syntax tree and a cross-platform code mapping relation, and according to the abstract syntax tree and the cross-platform code mapping relationship, carrying out repetition logic detection on the project code file to obtain a repetition detection result, according to the repetition detection result, confirming a repetition code in the project code file, carrying out cross-platform code optimization on the repetition code to obtain an optimized code, and carrying out cross-platform code optimization on the optimized code. The method comprises the steps of obtaining duplicated codes, extracting reusable components on the basis of the duplicated codes, generating a multi-end adaptive component template according to a preset condition compiling logic, and reconstructing the optimized codes according to the reusable components and the multi-end adaptive component template to obtain reconstructed codes, so that the code duplicating detection and reconstruction efficiency is improved.
Owner:SHENZHEN LEXIN SOFTWARE TECH CO LTD

Text duplicate checking method, vehicle, computer readable storage medium and computer program product

The invention discloses a text duplicate checking method, a vehicle, a computer readable storage medium and a computer program product, and relates to the technical field of information processing. The method comprises the following steps: carrying out standardized preprocessing on an input text to be subjected to duplicate checking to obtain a current text fragment; performing semantic coding processing on the current text fragment by utilizing the target coding model to obtain a query vector; a similarity retrieval interface corresponding to the target search engine is called, a plurality of candidate vectors corresponding to the query vector are recalled from a semantic vector index, and the semantic vector index is constructed based on a historical text set and a graph index algorithm in the target search engine; sorting analysis is conducted on the multiple candidate vectors based on the multi-level duplicate judgment threshold values, a duplicate checking result is obtained, and the duplicate checking result is used for representing the duplicate situation between the historical texts corresponding to the multiple candidate vectors and the input text. The technical problems of low accuracy and low efficiency of a text duplicate checking method in related technologies are solved.
Owner:GUANGZHOU XIAOPENG MOTORS TECH CO LTD

Fuzzy testing and compliance evaluation platform and method for cloud password service

The invention discloses a cloud password service-oriented fuzzy test and compliance evaluation platform and method, and relates to the technical field of cloud computing and password security evaluation. The evaluation platform comprises a demand analysis module, a test case generation module, a fuzzy test execution module, a compliance detection module, a result fusion analysis module and a report generation module, and the result fusion analysis module is connected with the fuzzy test execution module and the compliance detection module. And the cloud password service detection module is used for carrying out association analysis on the response data, the abnormal log and the compliance detection result acquired by the fuzzy test, identifying the security vulnerability type and the compliance defect level of the cloud password service, and removing a repeated detection result. Fuzzy testing and compliance detection are organically fused, correlation analysis of the two results is achieved through the result fusion analysis module, the problems that in the prior art, the two tests are separated, and the results are difficult to integrate are solved, and the safety compliance condition of the cloud password service can be comprehensively reflected.
Owner:FUJIAN ZHIAN INFORMATION TECHNOLOGY CO LTD

Method and system for wireless monitoring of communication data for charging piles

This invention discloses a wireless monitoring method and system for communication data of charging piles. The monitoring method includes: receiving electromagnetic signals radiated from a CAN bus; performing analog-to-digital conversion on the electromagnetic signals to obtain a digitized signal to be processed; detecting the signal amplitude of sampled data in the signal to be processed in a time sequence; when the signal amplitude of any sampled data exceeds a threshold, inverting the level state of the sampled data to obtain feature data; filling a corresponding number of follower data between two adjacent feature data to obtain a monitoring signal including feature data and follower data; and performing frame parsing on the monitoring signal based on the CAN bus data frame transmission structure to obtain communication data. The above monitoring method for detecting communication data of charging piles has low detection cost, avoids repeated data detection, and offers fast detection speed and high accuracy.
Owner:KAIYUAN WANGAN INTERNET OF THINGS TECH (WUHAN) CO LTD

Deduplication methods and equipment for detection targets

This application provides a method and apparatus for deduplicating detected targets. The method includes: acquiring suspicious items from at least two human body images of the same scanned object from different directions; defining suspicious items located at the edge of the human body contour in the first human body image from the at least two different directions as targets to be deduplicated; deduplicating suspicious items in a second human body image from the at least two different directions, based on the targets to be deduplicated, to obtain the number of duplicate suspicious items in the second human body image; and determining the total number of suspicious items carried by the scanned object based on the number of duplicates and the number of suspicious items in the first human body image. This method can effectively identify and remove the same suspicious item repeatedly detected in images from different directions, avoiding duplicate counting and improving security inspection efficiency.
Owner:HANGZHOU RAYIN TECH CO LTD

Online man-machine conversation optimization method, system and device, electronic equipment, storage medium and program product

The embodiment of the invention provides an online man-machine conversation optimization method, system and device, electronic equipment, a storage medium and a program product. In the scheme provided by the embodiment, when a first dialogue generation model executes a streaming response based on user input data, a currently generated first dialogue text fragment is subjected to repeated detection, and when it is detected that repeated content exists in the first dialogue text fragment, the first dialogue generation model is controlled to interrupt execution of the streaming response; in addition, reply generation constraint information is determined according to the repeated content in the first reply text segment, and then reply generation cue words are generated based on the reply generation constraint information and the non-repeated content in the first reply text segment. And after the dialogue generation cue word is input into the first dialogue generation model, the first dialogue generation model is triggered to determine a dialogue generation starting point according to the non-repeated content in the first dialogue text fragment, and streaming response is continuously executed from the dialogue generation starting point.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Approximate repetition detection method and system fusing local retrieval and multi-dimensional decision

The invention discloses an approximate repetition detection method and system fusing local retrieval and multi-dimensional decision. The method comprises the steps that a to-be-detected document is preprocessed, and text content and service meta-information are extracted; utilizing MinHash to generate a compact signature; quickly recalling a candidate set from a massive historical library through an LSH index; adopting a paragraph-level weighted Jaccard algorithm to accurately calculate the structural similarity between the candidate document and the document to be detected; when the similarity exceeds a threshold value, intelligent decision making is carried out according to a preset priority chain considering multi-dimensional business rules such as release mechanism levels and time, and finally reserved documents are determined; and recording the structured audit log in the whole process for tracing. The method realizes approximate repetition detection with high recall, high precision, low time consumption and interpretable and auditing decision process, and is especially suitable for massive de-duplication scenes of strong normative texts such as government affair official documents, news announcements and the like.
Owner:BEIJING FANGCUN WUYOU TECH DEV CO LTD

Structured output of duplicate or near-duplicate text documents identified using automated near-duplicate detection for text documents

Techniques described herein provide for generation of structured output for documents identified using automated near-duplicate detection. In one example, a system can receive a set of documents including at least one pair of similar documents determined to be similar to one another based on similarity scores generated using a predefined similarity scoring technique. The system can generate document groups by merging together pairs of documents that share at least one document. The system can, for each of the document groups, identify a representative document for the document group. The system can generate an output for display including a section for each document group, in which each section includes the representative document for the document group and, for each document in the document group, the similarity score relative to the representative document for the document group.
Owner:SAS INSTITUTE INC

A multi-source data fusion method and device, electronic equipment and storage medium

The application discloses a multi-source data fusion method and device, electronic equipment and storage medium. The method comprises the following steps: acquiring source data in at least one source database and mother database data in a data mother database; performing quality evaluation on the source data and the mother database data to obtain a quality evaluation result; for any source data, performing a duplicate detection process on the source data and the mother database data to obtain a duplicate detection result; and performing data fusion on the source data and the mother database data based on the duplicate detection result and the quality evaluation result. The application increases the consideration of the quality of source data and mother database data, avoids the influence of low-quality data on fused data, and improves the data quality of the mother database data after fusion.
Owner:CHINA POST INFORMATION TECH (BEIJING CO LTD

UAV path planning method, device, equipment and medium for underground leak detection

The present disclosure provides a method, apparatus, device, and medium for drone path planning for underground leak detection. The method comprises: obtaining a grid map of an area to be detected, assigning each micro-area in the grid map to multiple drones, obtaining an initial path for each drone, utilizing a reinforcement learning algorithm to dynamically adjust the initial path for each drone based on at least one of the following: pheromone concentration in each micro-area, window time in each micro-area, location information of each drone, power information of each drone, and obstacle information, thereby obtaining a dynamically adjusted path corresponding to each drone. Based on the dynamically adjusted path, each drone is controlled to perform underground leak detection in the area to be detected. As can be seen, the embodiments of the present disclosure can effectively avoid duplicate detection and missed areas between drones, thereby improving the comprehensiveness and efficiency of underground leak detection.
Owner:CHINA RAILWAY 19 BUREAU GRP CO LTD +2

Method for duplicate detection and storage based on homomorphic encryption and simhash

The application discloses a homomorphic encryption and Simhash-based ciphertext duplicate checking and storage method, which comprises RSA data encryption and decryption, chameleon hash calculation, secret sharing calculation, simhash calculation and a homomorphic encryption method. The ciphertext state verification can be realized, and the security of the file is enhanced. In the encryption and decryption and chameleon hash calculation process, the random number of the algorithm is deleted, so that the ciphertext of the same file after encryption is the same, and the same chameleon hash is generated. The file is transmitted to the IPFS to realize distributed storage, and the IPFS can ensure that the file corresponds to a unique storage address according to content hash addressing. The data ciphertext is calculated by using the homomorphic encryption technology, so that the security of the data is ensured.
Owner:XIAN UNIV OF TECH

Vehicle-mounted bus communication database file detection method and device and storage medium

The application relates to a vehicle-mounted bus communication database file detection method and device and a storage medium, and comprises the following steps: acquiring a vehicle-mounted bus communication database file to be detected; wherein the vehicle-mounted bus communication database file comprises a DBC file and / or an LDF file; analyzing the vehicle-mounted bus communication database file, storing a message corresponding to the vehicle-mounted bus communication database file read by the vehicle-mounted bus communication database file into a message list; polling the message list, detecting the message according to multiple preset detection items; and outputting a test report according to each detection result. The method can detect the message corresponding to the vehicle-mounted bus communication database file read according to the preset detection items, such as message length matching detection, detection of whether a signal initial value is in an expected interval and signal position repetition detection in the same message, so that the detection of the vehicle-mounted bus communication database file is efficiently and quickly completed, and support is provided for subsequent integrated development.
Owner:CHONGQING CHANGAN TECH CO LTD

Deepfake detection

Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.
Owner:PINDROP SECURITY INC

Weak event calibration and sensing method and system for emergency scene

The invention discloses a weak event space-time calibration and metastable state heat perception method and system for an emergency scene. The method comprises the following steps that: a client acquires an event type, an event timestamp, a position coordinate and media information and uploads the information to a server; the server performs time and space dimension repeated detection on the events, and writes the de-duplicated events into a buffer area; calculating a time offset based on the client event time and the server receiving time, and carrying out unified calibration on the event time; re-estimating the confidence coefficient of the event position in combination with the regional positioning quality; counting an event position offset trend in a preset time window, identifying the overall positioning drift of the region, and performing position correction when the overall positioning drift exceeds a threshold value; self-adaptively constructing a metastable state time window for weak event aggregation according to the change of the number of events in the window; and based on the corrected position, the revaluation confidence and the distance attenuation weight, performing weighted stacking on the events in the window to generate a dynamic popularity value.
Owner:BEIJING BIAOYANG CROSSING TECH CO LTD

Inspection drawing behavior detection method and device, electronic equipment and storage medium

The invention discloses an inspection drawing behavior detection method and device, electronic equipment and a storage medium, and relates to the technical field of image processing, and the method comprises the steps: carrying out the geographic position verification of a shooting position of a to-be-detected inspection picture based on the merchant information associated with the to-be-detected inspection picture; if it is judged that the geographic position verification is passed, basic features of the to-be-detected inspection picture are extracted, and preliminary repeated detection is carried out based on the basic features; if it is judged that the preliminary repeated detection is passed, extracting an image feature vector of the to-be-detected inspection picture based on a deep learning model, and extracting a text feature vector in the to-be-detected inspection picture in combination with a text recognition technology; and performing similarity comparison on the image feature vector and the text feature vector with a feature vector of a historical inspection picture stored in a vector database, and judging whether the inspection picture to be detected has a picture sleeving behavior or not based on a similarity comparison result. According to the invention, through comprehensive analysis of the geographic position and the multi-modal features, the potential graph sleeving behavior can be accurately identified.
Owner:CHINA UNIONPAY MERCHANT SERVICES CO LTD

Semantic focusing test question duplicate checking method based on large language model

The invention discloses a semantic focusing test question duplicate checking method based on a large language model, and the method comprises the steps: completing the construction of a question corpus and metadata labeling based on corpus construction and text standardization; mapping the topics in the corpus into dense vectors by adopting a semantic vectorization representation strategy, and constructing an offline semantic vector library; a semantic vector recall-SimHash denoising-Reranker model rearrangement screening mechanism is provided, and the problem that synonym rewriting cannot be recognized in traditional literal comparison is solved; a multi-level screening-large model deep judgment-online threshold value self-adaptive cooperation framework is provided, and semantic-level accurate duplicate checking is realized through real-time feedback continuous iteration; by starting a review mechanism, a vector recall threshold value and a large model deep judgment threshold value are dynamically adjusted according to data in a manual review library, so that the accuracy and efficiency of duplicate checking are maximized. According to the method, the problems of synonymous rewriting missing net, short text representation failure and static threshold false alarm / missing alarm are solved.
Owner:HARBIN INST OF TECH

Bidding document duplicate checking method, system and equipment and medium

The invention provides a bidding document duplicate checking method, system and device and a medium. The method comprises the steps of obtaining file content of a bidding document; the bidding document comprises a document to be subjected to duplicate checking and a comparison document; extracting sub-strings of the file content; performing bucket dividing processing on the sub-character strings, and mapping the sub-character strings meeting a preset condition into the same hash bucket; calculating the similarity between the sub-character strings of the document to be subjected to duplicate checking and the sub-character strings of the comparison document in the same hash bucket, and determining the pairing relationship between the sub-character strings; performing pairing and merging according to the pairing relationship to obtain a maximum repeated fragment; and calculating a content repetition rate according to the maximum repetition fragment so as to determine a duplicate checking result of the bidding document. Similar sub-strings are concentrated in the same hash bucket through a bucket dividing strategy, so that rapid repeated detection of large-scale text data is realized, and the duplicate checking response speed is increased; and determining the maximum duplicate fragment according to the similarity and the pairing relationship, thereby solving the problem that duplicate fragments with a small number of altered characters cannot be identified.
Owner:CHINA THREE GORGES CORPORATION

Expediting automated near-duplicate detection for new text documents

Techniques described herein provide for automated near-duplicate detection for new text documents given text documents that were previously processed using automated near-duplicate detection for text documents. In one example, a system can receive new documents and documents that were previously processed using a predefined processing technique for automated near-duplicate detection. The system can process the new documents and cluster the new documents into multiple predefined clusters previously identified using the predefined processing technique. For each predefined cluster including at least one new document, the system can generate document groups by determining similarity scores using the predefined processing technique as applied to the documents in the predefined clusters. The system can identify a representative document for each document group and generate an output data structure including the document groups and the representative document for each group.
Owner:SAS INSTITUTE INC

Information processing system, information processing method, and information processing program

This system more accurately detects duplicate evidence and prevents the inclusion of duplicate documents. [Solution] The duplicate detection unit P5 detects a duplicate between the source document information Ss and the referenced document information Sd when confirmed information F included in the source document information Ss is input, and the confirmed information F included in the source document information Ss is common with at least one of the confirmed information F or estimated information U included in the referenced document information Sd.
Owner:FREEE

A railway line video device extraction method and device based on unified identification of multiple types of devices and time sequence fusion

The present application relates to the technical field of railway detection, and discloses a railway line video equipment extraction method and device based on unified identification and time sequence fusion of multiple types of equipment. It aims to solve the technical problems of lack of unified identification ability of multiple types of key equipment, high false detection, serious repeated detection, low detection efficiency and insufficient robustness caused by single frame detection in the prior art. The present application includes the following steps: video acquisition and image preprocessing; generation of equipment candidate area based on prior knowledge of the railway industry; unified identification of multiple types of key equipment; construction and association of time sequence trajectories of continuous frame detection results; equipment existence confirmation based on time sequence fusion; equipment level result output and spatial deduplication. From the complete processing link of vehicle-mounted video data to equipment level structured results. The present application realizes unified identification and extraction of multiple types of railway 2C key equipment, significantly reduces false detection and repeated detection by using time sequence fusion, and improves detection efficiency and robustness by introducing prior knowledge of the railway industry.
Owner:ZHENGZHOU PANHUI ELECTRONIC TECHNOLOGY CO LTD

Method and apparatus for determining duplicate data, and device

PCT designated stageWO2026036630A1Bitwise operationEngineering
The embodiments of the present application relate to the field of big data. Provided are a method and apparatus for determining duplicate data, and a device. The method comprises: on the basis of a merged computing node, acquiring bitmap arrays and single-machine duplicate sets of M computing nodes; performing pairwise bitwise operation processing on the bitmap arrays of the M computing nodes, so as to obtain M duplicate values; acquiring local reverse dictionaries corresponding to the computing nodes, and on the basis of the local reverse dictionaries, performing secondary duplicate check processing on the duplicate values, so as to obtain multi-machine duplicate sets corresponding to the duplicate values; and merging the multi-machine duplicate sets and the single-machine duplicate sets, so as to obtain a global duplicate set. The method in the present application improves data duplicate detection efficiency.
Owner:CHINA UNIONPAY

Systems and method of managing documents

Systems and methods of managing documents associated with resources. The system includes a communication module, a processor, and a memory. The memory stores duplicate detection data and instructions that, when executed, configure the processor to: receive, via the communication module and from a client device, an image of the subject document; extract a document identifier from the image of the subject document; obtain a date associated with the subject document; determine that the subject document is unique by comparing a set of validation data values with the duplicate detection data; and in response to determining that the subject document is unique, transmit, to the client device, a provisional acceptance notification and provisionally allocate the resource associated with the subject document to a data record corresponding to a second identifier. Extracting the date from the image of the subject document may include using image recognition.
Owner:THE TORONTO DOMINION BANK

Graph node hybrid traversal ordering method and system for workflow visualization

ActiveCN122114587BSorting algorithmAlgorithm
The application discloses a kind of graph node mixed traversal sorting method and system of workflow visualization, it is related to data visualization technical field.The method includes: obtaining workflow graph and initializing data structure;From start node, perform depth-first mixed traversal, repeat detection and block to node in the traversal process;After traversal is completed, the coordinates of each node are calculated, and the secondary sorting is executed to the node list, finally, continuous serial number is distributed and the sorting result is output.The application realizes zero repeat processing by throttling blocking mechanism, realizes zero data backflow by branch anti-backflow control mechanism, improves the sorting accuracy by secondary sorting algorithm, has excellent real-time response performance, and can be widely applied in workflow visualization, ETL task arrangement, low-code development platform and other fields.
Owner:SHENZHEN FENXIANG INTERNET TECH CO LTD

Near-duplicate detection of images for training or validation of machine learning models

A system filters near-duplicate images to generate data for training or validation of a machine learning model. The system receives a set of images and generates feature vectors from the images. The system clusters the feature vectors. For each cluster of feature vectors, the system determines near-duplicate pairs of images. The system may generate a cost matrix representing a linear assignment problem and find near-duplicate pairs of images by solving the linear assignment problem. The system filters images from the set of images based on the near-duplicate pairs of images. The system uses the filtered set of images for training or validation of the machine learning model.
Owner:LANDINGAI INC

An electronic system for automated near duplicate detection using locality sensitive hashing (LSH) and corresponding method

Proposed is a digital system using Locality Sensitive Hashing (LSH) for near duplicate detection on weighted datasets providing a robust technical solution to the problem by improving (i) efficiency by allowing for fast and accurate identification of near duplicates in large datasets; (ii) accuracy by extending the hashes calculated via 5 the LSH approach with an insurance-specific weight, thus making sure that the concept of similarity between documents is calculated on a technical and industry-relevant basis; (iii) cost reduction by reducing redundant data lowers storage and processing costs; (iv) data integrity by improved accuracy and consistency in records enhance overall data integrity; and (v) regulatory compliance matching by automated data 10 management practices supporting compliance with industry regulations.
Owner:SWISS REINSURANCE CO LTD

A method and system for detecting dense semi-transparent shrimp fry and a storage medium

The present application provides a kind of dense translucent fry detection method, system and storage medium, it is related to computer vision technical field, the present application is by the transparency of target quantification, and as the basis, in feature fusion stage to weak signal is compensated and enhanced, while in model training stage to loss function is self-adaptive adjustment, to systematically solve the bottleneck problem such as missing detection, recheck and inaccurate positioning that existing technology faces when processing dense, small, translucent target.The detection recall rate of low contrast, translucent target is improved, the missing detection is effectively reduced;The detection accuracy and robustness in high-density, overlapping scene are improved, and the repeated detection is reduced;The positioning accuracy of model to edge fuzzy target is enhanced;By introducing physical priori knowledge, the overall generalization ability of model and applicability in complex underwater environment are improved.
Owner:YANGTZE DELTA REGION INST OF UNIV OF ELECTRONICS SCI & TECH OF CHINE (HUZHOU)

Machine-learning based data entry duplication detection and mitigation and methods thereof

Systems and methods of the present disclosure enable a processor to automatically detect duplicate data entries by receiving data entries associated with a user, where each data entry includes a value, a time, an entity identifier, and a location. Pairs of similar data entries are determined by matching the entity identifier and the location pairs data entries. Candidate duplicate data entries are determined based on a proximity in time between data entries of the similar data entries. For each candidate duplicate data entry, a feature vector is generated including the entity identifier, location, value and time, and each feature vector is submitted to a duplicate classification model to automatically determine duplicate data entries from the candidate duplicate data entries, the duplicate classification model being trained according to a historical dispute entries.
Owner:CAPITAL ONE SERVICES LLC