Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

112 results about "Duplicate detection" patented technology

Duplicate Detector is smart enough to find duplicate or similar records in many situations. Duplicate Detector for SugarCRM works on any existing or custom field of type varchar, name or phone. It prompts the user if the value has already been used while they are creating records in the edit view or quick create mode.

Structured output of duplicate or near-duplicate text documents identified using automated near-duplicate detection for text documents

Techniques described herein provide for generation of structured output for documents identified using automated near-duplicate detection. In one example, a system can receive a set of documents including at least one pair of similar documents determined to be similar to one another based on similarity scores generated using a predefined similarity scoring technique. The system can generate document groups by merging together pairs of documents that share at least one document. The system can, for each of the document groups, identify a representative document for the document group. The system can generate an output for display including a section for each document group, in which each section includes the representative document for the document group and, for each document in the document group, the similarity score relative to the representative document for the document group.
Owner:SAS INSTITUTE INC

Automated near-duplicate detection for text documents

Techniques described herein provide for automated detection of near-duplicate documents. In one example, a system can cluster documents into a set of clusters based on character frequencies associated with the documents. For a given cluster, the system can generate first similarity scores associated with every pair of documents in the cluster. The system can then select a filtered group of documents associated with first similarity scores that meet or exceed a first predefined similarity threshold. Next, the system can convert the filtered group of documents into matrix representations. The system can generate second similarity scores for every pair of matrix representations. The system can then identify documents, from among the filtered group of documents, associated with second similarity scores that meet or exceed a second predefined similarity threshold. The identified documents can be duplicate or near-duplicate text documents.
Owner:SAS INSTITUTE INC

Target intelligent detection method and device based on computer vision

The invention discloses an intelligent target detection method and device based on computer vision, and relates to the technical field of visual recognition. The method comprises the following specific implementation steps: step 1, detecting and positioning an area possibly containing a small target in an image by using saliency; step 2, carrying out saliency map fusion; step 3, carrying out super-separation treatment; step 4, adaptively cutting the image, and segmenting the large image by adopting a sliding window; step 5, combining repeated detection results; step 6, carrying out weighted fusion; according to the method, a set of complete processing flow from global saliency detection, candidate region screening, local super-resolution reconstruction, adaptive cutting to result fusion is formed through Step1 to Step6, and the accuracy and precision of small target detection are improved.
Owner:INNER MONGOLIA UNIVERSITY

Duplication check system and method for paper generated by artificial intelligence

A duplication check system and method for paper generated by artificial intelligence includes steps: S1: the user uploading the academic paper to be detected to a system, and the system automatically extracting the title, the abstract, and the headline of each paragraph of the paper; S2: fusing the title, the abstract, and the headline of each paragraph of the paper with the contextual information of the paper and extracting theme features; S3: after the different themes of the paper are extracted, repeatedly using similar tones for each theme in all different AI tools to propose integrate text requirements, searching each theme for times of the number of repetitions of integration in each AI tool until no new content is obtained, matching all the obtained texts with the paper to be duplication checked, based on natural language understanding, for duplication check, and marking the matching repeated parts and indicating the sources.
Owner:WU JIANG

Text string comparison for duplicate or near-duplicate text documents identified using automated near-duplicate detection for text documents

ActiveUS20250231993A1Pattern recognitionBoilerplate text
Techniques described herein provide for text string comparison for documents identified using automated near-duplicate detection. In one example, a system can receive a pair of documents. The system can extract text strings from the documents. The system can normalize the extracted text strings using a predefined normalization scheme. The system can identify boilerplate text segments in the normalized text strings. The system can remove the boilerplate text segments from the normalized text strings to generate filtered text strings. The system can divide the filtered text strings by identifying section indicators. The system can, for each section, generate groupings of text strings and determine a similarity score between each pair of corresponding groupings to identify matching groupings of text strings. The system can generate an output for display showing the visual indications of the matched groupings of text strings.
Owner:SAS INSTITUTE INC

Expediting automated near-duplicate detection for new text documents

Techniques described herein provide for automated near-duplicate detection for new text documents given text documents that were previously processed using automated near-duplicate detection for text documents. In one example, a system can receive new documents and documents that were previously processed using a predefined processing technique for automated near-duplicate detection. The system can process the new documents and cluster the new documents into multiple predefined clusters previously identified using the predefined processing technique. For each predefined cluster including at least one new document, the system can generate document groups by determining similarity scores using the predefined processing technique as applied to the documents in the predefined clusters. The system can identify a representative document for each document group and generate an output data structure including the document groups and the representative document for each group.
Owner:SAS INSTITUTE INC

Deadlock detection method, device and storage medium

Provided in the embodiments of the present disclosure are a deadlock detection method, a device and a storage medium. In the embodiments of the present disclosure, a detection mechanism in which a triggering entity is also responsible for execution is provided, and each node in a distributed system can serve as an executor for deadlock detection, so that deadlock detection tasks in the distributed system can be executed in a distributed manner, rather than being concentrated on a single node, thereby effectively accommodating the scalability of the distributed system and preventing the detection efficiency from being affected as the system scales up. Moreover, a detection mechanism similar to optimistic locking is further provided, so as to prevent problems such as redundant detection that may occur between the deadlock detection tasks distributed on the nodes, thereby ensuring the accuracy of deadlock detection. In addition, a single deadlock detection task can detect all deadlocks in the system in one step, eliminating the need for serial detection and thus further improving the efficiency of deadlock detection in the distributed system.
Owner:CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD

Multi-terminal code repetition detection and component reconstruction method and device, equipment and medium

The invention relates to the technical field of code analysis, and discloses a multi-terminal code repetition detection and component reconstruction method, device, equipment and medium, and the method comprises the steps: carrying out code analysis on a pre-acquired project code file, and identifying a code mapping relation to obtain an abstract syntax tree and a cross-platform code mapping relation, and according to the abstract syntax tree and the cross-platform code mapping relationship, carrying out repetition logic detection on the project code file to obtain a repetition detection result, according to the repetition detection result, confirming a repetition code in the project code file, carrying out cross-platform code optimization on the repetition code to obtain an optimized code, and carrying out cross-platform code optimization on the optimized code. The method comprises the steps of obtaining duplicated codes, extracting reusable components on the basis of the duplicated codes, generating a multi-end adaptive component template according to a preset condition compiling logic, and reconstructing the optimized codes according to the reusable components and the multi-end adaptive component template to obtain reconstructed codes, so that the code duplicating detection and reconstruction efficiency is improved.
Owner:SHENZHEN LEXIN SOFTWARE TECH CO LTD

Text duplicate checking method, vehicle, computer readable storage medium and computer program product

The invention discloses a text duplicate checking method, a vehicle, a computer readable storage medium and a computer program product, and relates to the technical field of information processing. The method comprises the following steps: carrying out standardized preprocessing on an input text to be subjected to duplicate checking to obtain a current text fragment; performing semantic coding processing on the current text fragment by utilizing the target coding model to obtain a query vector; a similarity retrieval interface corresponding to the target search engine is called, a plurality of candidate vectors corresponding to the query vector are recalled from a semantic vector index, and the semantic vector index is constructed based on a historical text set and a graph index algorithm in the target search engine; sorting analysis is conducted on the multiple candidate vectors based on the multi-level duplicate judgment threshold values, a duplicate checking result is obtained, and the duplicate checking result is used for representing the duplicate situation between the historical texts corresponding to the multiple candidate vectors and the input text. The technical problems of low accuracy and low efficiency of a text duplicate checking method in related technologies are solved.
Owner:GUANGZHOU XIAOPENG MOTORS TECH CO LTD

Duplicate checking small language model training method combined with multi-level knowledge distillation

PendingCN120562402ASemantic analysisBiological modelsAlgorithmPlagiarism detection
The embodiment of the invention provides a duplicate checking small language model training method combined with multilevel knowledge distillation, and the method comprises the steps: obtaining duplicate checking sample pairs, and determining the complexity of the duplicate checking sample pairs according to the text features of the duplicate checking sample pairs; according to the complexity of the duplicate checking sample pair, determining distillation levels of the teacher model, and determining a weighting coefficient of each distillation level of the teacher model; determining the distillation loss between the teacher model and the student model according to the weighting coefficient of each network distillation of the teacher model, the first output result of each distillation level of the teacher model and the second output result of each distillation level of the student model; according to the distillation loss between the teacher model and the student model, updating parameters of the student model; and repeating the above steps until the updated student model meets the preset condition, and taking the updated student model as a duplicate checking small language model, thereby realizing a low-power-consumption and high-precision duplicate checking effect.
Owner:CHINA THREE GORGES CORPORATION

Fuzzy testing and compliance evaluation platform and method for cloud password service

The invention discloses a cloud password service-oriented fuzzy test and compliance evaluation platform and method, and relates to the technical field of cloud computing and password security evaluation. The evaluation platform comprises a demand analysis module, a test case generation module, a fuzzy test execution module, a compliance detection module, a result fusion analysis module and a report generation module, and the result fusion analysis module is connected with the fuzzy test execution module and the compliance detection module. And the cloud password service detection module is used for carrying out association analysis on the response data, the abnormal log and the compliance detection result acquired by the fuzzy test, identifying the security vulnerability type and the compliance defect level of the cloud password service, and removing a repeated detection result. Fuzzy testing and compliance detection are organically fused, correlation analysis of the two results is achieved through the result fusion analysis module, the problems that in the prior art, the two tests are separated, and the results are difficult to integrate are solved, and the safety compliance condition of the cloud password service can be comprehensively reflected.
Owner:FUJIAN ZHIAN INFORMATION TECHNOLOGY CO LTD

Repeated data detection method and device, storage medium and computer program product

The invention discloses a duplicated data detection method and device, a storage medium and a computer program product, and relates to the technical field of data processing, a preset bit array is constructed through a plurality of preset functions and stored data blocks to serve as the basis of data duplicated detection of to-be-detected data blocks, and the data duplicated detection efficiency is improved. The to-be-detected data blocks which may be repeated are preliminarily screened out through the preset bit array, and then the to-be-detected data blocks which may be repeated are further subjected to refined detection through the preset hash table, so that layered detection of data is realized, and compared with hash table look-up traversal comparison of all the to-be-detected data blocks, the detection efficiency is effectively improved, and the detection time is shortened. And meanwhile, the detection accuracy is ensured.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Deadlock detection method and device and storage medium

The embodiment of the invention provides a deadlock detection method and device and a storage medium. In the embodiment of the invention, a detection mechanism that who triggers execution is provided, and each node in the distributed system can serve as an execution main body of deadlock detection, so that the deadlock detection tasks in the distributed system can be executed in a distributed manner and are not concentrated at a single node any more, the expandability of the distributed system can be effectively adapted, and the expandability of the distributed system is improved. And the detection efficiency is not influenced by the scale expansion of the system. Moreover, the invention also provides a detection mechanism similar to an optimistic lock, so as to avoid the problems of repeated detection and the like possibly occurring between deadlock detection tasks distributed on each node, and further ensure the accuracy of deadlock detection. Besides, a single deadlock detection task can detect all deadlocks in the system at one time, and serial detection is not needed any more, so that the deadlock detection efficiency in the distributed system can be further improved.
Owner:HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD

Duplicate checking method and system for products in various stages of software research and development

The invention discloses a duplicate checking method and system for products in all stages of software research and development, and belongs to the technical field of software engineering. In order to solve the problems of high efficiency and accuracy of code, document and function duplicate checking, the technical means of abstract syntax tree analysis, hash index construction, deep learning semantic analysis, image feature extraction, table structure comparison, semantic matching calculation and the like are mainly adopted. According to the method, code-level structured duplicate checking, document-level multi-modal comparison and functional-level semantic analysis can be realized, and the accuracy and efficiency of product duplicate checking in the software development process are improved.
Owner:INST OF SOFTWARE - CHINESE ACAD OF SCI

Automatic safety facility examination system

The invention relates to the technical field of automatic safety management, in particular to an automatic safety facility review system which comprises an equipment dependence evaluation module, a risk weight calculation module, a fault trend analysis module, a review priority adjustment module and a fault diagnosis decision module. According to the method, key data such as steel wire rope tension, lifting hook bearing and winch frequency are synchronously collected, linkage frequency and response delay are accurately reflected, the dynamic sensing capacity of a system is improved, and after linkage strength between facilities is recognized, key nodes are clearly examined, resource allocation is optimized, and potential hazard recognition efficiency is improved based on abnormal fluctuation and dependence degree; by extracting the periodic trend of an alarm event, defining a hidden danger development path, adjusting the inspection period and improving the pertinence of inspection, through fault fluctuation homology analysis, optimizing a fault traceability path, improving the hidden danger treatment precision, constructing closed-loop logic, improving the inspection precision and reducing resource consumption and repeated detection burden.
Owner:SHENZHEN PARAMOUNT TECH CO LTD

Method and system for wireless monitoring of communication data for charging piles

This invention discloses a wireless monitoring method and system for communication data of charging piles. The monitoring method includes: receiving electromagnetic signals radiated from a CAN bus; performing analog-to-digital conversion on the electromagnetic signals to obtain a digitized signal to be processed; detecting the signal amplitude of sampled data in the signal to be processed in a time sequence; when the signal amplitude of any sampled data exceeds a threshold, inverting the level state of the sampled data to obtain feature data; filling a corresponding number of follower data between two adjacent feature data to obtain a monitoring signal including feature data and follower data; and performing frame parsing on the monitoring signal based on the CAN bus data frame transmission structure to obtain communication data. The above monitoring method for detecting communication data of charging piles has low detection cost, avoids repeated data detection, and offers fast detection speed and high accuracy.
Owner:KAIYUAN WANGAN INTERNET OF THINGS TECH (WUHAN) CO LTD

Deduplication methods and equipment for detection targets

This application provides a method and apparatus for deduplicating detected targets. The method includes: acquiring suspicious items from at least two human body images of the same scanned object from different directions; defining suspicious items located at the edge of the human body contour in the first human body image from the at least two different directions as targets to be deduplicated; deduplicating suspicious items in a second human body image from the at least two different directions, based on the targets to be deduplicated, to obtain the number of duplicate suspicious items in the second human body image; and determining the total number of suspicious items carried by the scanned object based on the number of duplicates and the number of suspicious items in the first human body image. This method can effectively identify and remove the same suspicious item repeatedly detected in images from different directions, avoiding duplicate counting and improving security inspection efficiency.
Owner:HANGZHOU RAYIN TECH CO LTD

Online man-machine conversation optimization method, system and device, electronic equipment, storage medium and program product

The embodiment of the invention provides an online man-machine conversation optimization method, system and device, electronic equipment, a storage medium and a program product. In the scheme provided by the embodiment, when a first dialogue generation model executes a streaming response based on user input data, a currently generated first dialogue text fragment is subjected to repeated detection, and when it is detected that repeated content exists in the first dialogue text fragment, the first dialogue generation model is controlled to interrupt execution of the streaming response; in addition, reply generation constraint information is determined according to the repeated content in the first reply text segment, and then reply generation cue words are generated based on the reply generation constraint information and the non-repeated content in the first reply text segment. And after the dialogue generation cue word is input into the first dialogue generation model, the first dialogue generation model is triggered to determine a dialogue generation starting point according to the non-repeated content in the first dialogue text fragment, and streaming response is continuously executed from the dialogue generation starting point.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Program Segmentation of Linear Transmission

Content streams may be segmented to provide automatic extraction and storage of content items without intervening commercials or other unrelated content. These content items may then be stored in a database and made accessible to subscribers through, for example, an on-demand service. Automatic segmentation may include the identification of program boundaries, segmentation of a content stream based on the boundaries and the subsequent classification of the segments into content types. For example, audio and video duplication detection may be used to identify commercials since commercials tend to repeat frequently over a relatively short amount of time. A system may further identify an end of program indicator in a video stream to determine when a program ends. Accordingly, if a program ends after a scheduled end time, a recording device (e.g., the program is being recorded) may automatically extend the recording time to capture the entire program.
Owner:COMCAST CABLE COMM LLC

Approximate repetition detection method and system fusing local retrieval and multi-dimensional decision

The invention discloses an approximate repetition detection method and system fusing local retrieval and multi-dimensional decision. The method comprises the steps that a to-be-detected document is preprocessed, and text content and service meta-information are extracted; utilizing MinHash to generate a compact signature; quickly recalling a candidate set from a massive historical library through an LSH index; adopting a paragraph-level weighted Jaccard algorithm to accurately calculate the structural similarity between the candidate document and the document to be detected; when the similarity exceeds a threshold value, intelligent decision making is carried out according to a preset priority chain considering multi-dimensional business rules such as release mechanism levels and time, and finally reserved documents are determined; and recording the structured audit log in the whole process for tracing. The method realizes approximate repetition detection with high recall, high precision, low time consumption and interpretable and auditing decision process, and is especially suitable for massive de-duplication scenes of strong normative texts such as government affair official documents, news announcements and the like.
Owner:BEIJING FANGCUN WUYOU TECH DEV CO LTD

Structured output of duplicate or near-duplicate text documents identified using automated near-duplicate detection for text documents

Techniques described herein provide for generation of structured output for documents identified using automated near-duplicate detection. In one example, a system can receive a set of documents including at least one pair of similar documents determined to be similar to one another based on similarity scores generated using a predefined similarity scoring technique. The system can generate document groups by merging together pairs of documents that share at least one document. The system can, for each of the document groups, identify a representative document for the document group. The system can generate an output for display including a section for each document group, in which each section includes the representative document for the document group and, for each document in the document group, the similarity score relative to the representative document for the document group.
Owner:SAS INSTITUTE INC

A multi-source data fusion method and device, electronic equipment and storage medium

The application discloses a multi-source data fusion method and device, electronic equipment and storage medium. The method comprises the following steps: acquiring source data in at least one source database and mother database data in a data mother database; performing quality evaluation on the source data and the mother database data to obtain a quality evaluation result; for any source data, performing a duplicate detection process on the source data and the mother database data to obtain a duplicate detection result; and performing data fusion on the source data and the mother database data based on the duplicate detection result and the quality evaluation result. The application increases the consideration of the quality of source data and mother database data, avoids the influence of low-quality data on fused data, and improves the data quality of the mother database data after fusion.
Owner:CHINA POST INFORMATION TECH (BEIJING CO LTD

UAV path planning method, device, equipment and medium for underground leak detection

The present disclosure provides a method, apparatus, device, and medium for drone path planning for underground leak detection. The method comprises: obtaining a grid map of an area to be detected, assigning each micro-area in the grid map to multiple drones, obtaining an initial path for each drone, utilizing a reinforcement learning algorithm to dynamically adjust the initial path for each drone based on at least one of the following: pheromone concentration in each micro-area, window time in each micro-area, location information of each drone, power information of each drone, and obstacle information, thereby obtaining a dynamically adjusted path corresponding to each drone. Based on the dynamically adjusted path, each drone is controlled to perform underground leak detection in the area to be detected. As can be seen, the embodiments of the present disclosure can effectively avoid duplicate detection and missed areas between drones, thereby improving the comprehensiveness and efficiency of underground leak detection.
Owner:CHINA RAILWAY 19 BUREAU GRP CO LTD +2

Method for duplicate detection and storage based on homomorphic encryption and simhash

The application discloses a homomorphic encryption and Simhash-based ciphertext duplicate checking and storage method, which comprises RSA data encryption and decryption, chameleon hash calculation, secret sharing calculation, simhash calculation and a homomorphic encryption method. The ciphertext state verification can be realized, and the security of the file is enhanced. In the encryption and decryption and chameleon hash calculation process, the random number of the algorithm is deleted, so that the ciphertext of the same file after encryption is the same, and the same chameleon hash is generated. The file is transmitted to the IPFS to realize distributed storage, and the IPFS can ensure that the file corresponds to a unique storage address according to content hash addressing. The data ciphertext is calculated by using the homomorphic encryption technology, so that the security of the data is ensured.
Owner:XIAN UNIV OF TECH

Call Detail Record Duplicate Detection Method, Device, and Computer-Readable Storage Medium

The present invention discloses a method for duplicate check of call records, a device for duplicate check of call records, and a computer-readable storage medium. The method includes: obtaining keyword fields associated with call records to be checked for duplicates corresponding to the current period and a preset field combination; generating an initial keyword field combination based on the preset field combination and the keyword fields; determining a risk coefficient corresponding to each initial keyword field combination according to a preset keyword field duplicate check risk coefficient function model; selecting a target keyword field combination from the initial keyword field combinations according to the risk coefficient; and performing duplicate check of the call records to be checked for duplicates corresponding to the current period based on the target keyword field combination. The present invention aims to achieve the effect of improving the efficiency of duplicate check of call records.
Owner:CHINA MOBILE GROUP ZHEJIANG +1

Vehicle-mounted bus communication database file detection method and device and storage medium

The application relates to a vehicle-mounted bus communication database file detection method and device and a storage medium, and comprises the following steps: acquiring a vehicle-mounted bus communication database file to be detected; wherein the vehicle-mounted bus communication database file comprises a DBC file and / or an LDF file; analyzing the vehicle-mounted bus communication database file, storing a message corresponding to the vehicle-mounted bus communication database file read by the vehicle-mounted bus communication database file into a message list; polling the message list, detecting the message according to multiple preset detection items; and outputting a test report according to each detection result. The method can detect the message corresponding to the vehicle-mounted bus communication database file read according to the preset detection items, such as message length matching detection, detection of whether a signal initial value is in an expected interval and signal position repetition detection in the same message, so that the detection of the vehicle-mounted bus communication database file is efficiently and quickly completed, and support is provided for subsequent integrated development.
Owner:CHONGQING CHANGAN TECH CO LTD

Deepfake detection

Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.
Owner:PINDROP SECURITY INC

Weak event calibration and sensing method and system for emergency scene

The invention discloses a weak event space-time calibration and metastable state heat perception method and system for an emergency scene. The method comprises the following steps that: a client acquires an event type, an event timestamp, a position coordinate and media information and uploads the information to a server; the server performs time and space dimension repeated detection on the events, and writes the de-duplicated events into a buffer area; calculating a time offset based on the client event time and the server receiving time, and carrying out unified calibration on the event time; re-estimating the confidence coefficient of the event position in combination with the regional positioning quality; counting an event position offset trend in a preset time window, identifying the overall positioning drift of the region, and performing position correction when the overall positioning drift exceeds a threshold value; self-adaptively constructing a metastable state time window for weak event aggregation according to the change of the number of events in the window; and based on the corrected position, the revaluation confidence and the distance attenuation weight, performing weighted stacking on the events in the window to generate a dynamic popularity value.
Owner:BEIJING BIAOYANG CROSSING TECH CO LTD

Inspection drawing behavior detection method and device, electronic equipment and storage medium

The invention discloses an inspection drawing behavior detection method and device, electronic equipment and a storage medium, and relates to the technical field of image processing, and the method comprises the steps: carrying out the geographic position verification of a shooting position of a to-be-detected inspection picture based on the merchant information associated with the to-be-detected inspection picture; if it is judged that the geographic position verification is passed, basic features of the to-be-detected inspection picture are extracted, and preliminary repeated detection is carried out based on the basic features; if it is judged that the preliminary repeated detection is passed, extracting an image feature vector of the to-be-detected inspection picture based on a deep learning model, and extracting a text feature vector in the to-be-detected inspection picture in combination with a text recognition technology; and performing similarity comparison on the image feature vector and the text feature vector with a feature vector of a historical inspection picture stored in a vector database, and judging whether the inspection picture to be detected has a picture sleeving behavior or not based on a similarity comparison result. According to the invention, through comprehensive analysis of the geographic position and the multi-modal features, the potential graph sleeving behavior can be accurately identified.
Owner:CHINA UNIONPAY MERCHANT SERVICES CO LTD