Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

22 results about "Duplicate content" patented technology

Duplicate content is a term used in the field of search engine optimization to describe content that appears on more than one web page. The duplicate content can be substantial parts of the content within or across domains and can be either exactly duplicate or closely similar. When multiple pages contain essentially the same content, search engines such as Google and Bing can penalize or cease displaying the copying site in any relevant search results.

Aviation text content cleaning and labeling method, system and equipment and medium

PendingCN121543549ANatural language analysisBiological modelsDuplicate contentAviation
The invention relates to the technical field of aeronautical text data processing, and discloses an aeronautical text content cleaning and labeling method, system, device and medium wherein the method comprises: noise filtering: identifying and removing noise in an aeronautical text in combination with static cleaning and a general large model; format standardization: converting the aviation text after noise removal into a standardized format text; duplicate removal and error correction: detecting duplicate contents based on a hash algorithm, and correcting spelling errors and grammar errors based on a general large model to obtain an aviation text subjected to duplicate removal and error correction; entity identification: key entities are extracted based on the general large model, and the extracted key entities are labeled; active learning: screening high-value samples in the marked key entities based on an uncertainty query strategy; and dynamic optimization: performing verification and iterative optimization on the marking result in combination with the aviation knowledge base. According to the method, the automation level, the labeling accuracy and the system self-adaptive capability of aviation text processing can be remarkably improved.
Owner:四川腾盾科技有限公司 +1

Online man-machine conversation optimization method, system and device, electronic equipment, storage medium and program product

The embodiment of the invention provides an online man-machine conversation optimization method, system and device, electronic equipment, a storage medium and a program product. In the scheme provided by the embodiment, when a first dialogue generation model executes a streaming response based on user input data, a currently generated first dialogue text fragment is subjected to repeated detection, and when it is detected that repeated content exists in the first dialogue text fragment, the first dialogue generation model is controlled to interrupt execution of the streaming response; in addition, reply generation constraint information is determined according to the repeated content in the first reply text segment, and then reply generation cue words are generated based on the reply generation constraint information and the non-repeated content in the first reply text segment. And after the dialogue generation cue word is input into the first dialogue generation model, the first dialogue generation model is triggered to determine a dialogue generation starting point according to the non-repeated content in the first dialogue text fragment, and streaming response is continuously executed from the dialogue generation starting point.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Multi-document comparison method combining similarity calculation and accurate matching

The invention discloses a multi-document comparison method combining similarity calculation and accurate matching, and belongs to the technical field of information processing, and the method comprises the following steps: S1, text preprocessing, S2, blocking and positioning, S3, vector embedding, S4, similarity matrix construction, S5, Top-K pairing extraction, S6, similarity classification, S7, accurate matching, S8, visual highlighting, and S9, similarity evaluation. Through a sentence-level semantic vector and Top-K matching mechanism, the influence of document structure change is effectively overcome, the detection accuracy and robustness of content reconstruction type plagiarism are remarkably improved, matching dislocation caused by repeated content in a long document is avoided, the comparison stability is ensured, the problem of full-text vector semantic dilution is solved by embedding fine grit according to sentences, and the accuracy and robustness of content reconstruction type plagiarism detection are improved. According to the method, the recognition sensitivity of sentence pattern rewriting and local plagiarism is enhanced, and meanwhile, the abstract similarity is converted into traceable evidence by combining double-layer quantitative indexes with visual highlight positioning, so that the result interpretability and the practical value are greatly improved.
Owner:商飞软件有限公司

A test case processing method, device, equipment and medium

Embodiments of the present disclosure relate to a test case processing method, device, equipment and medium, wherein the method comprises: obtaining a first test case file in table format; performing format conversion on first data corresponding to the first test case file to obtain second data in mind map data structure; performing deduplication and integration processing on element node data of multiple test cases in the second data to obtain third data; and converting the third data into a second test case file in mind map format through a mind map tool. With the above technical solution, in the conversion process of the test case file from the table format to the mind map format, through the deduplication and integration processing, the data redundancy problem existing in the converted mind map format is avoided, the user can quickly find the required data, the readability of the data is improved, and in subsequent maintenance, the repeated work of repeated content is avoided, and the maintenance efficiency is improved.
Owner:DOUYIN VISION CO LTD

Power planning report three-level title structured analysis generation method and device and medium

The invention relates to an electric power planning report three-level title structured analysis generation method and device and a medium, and the method calls a large language model to execute the following steps: obtaining a to-be-processed electric power planning report and carrying out report paragraph text extraction to obtain an original text; segmenting the original text into overlapped text blocks, wherein the overlapped part comprises a three-level title context; constructing a power planning professional cue word, for each text block, identifying a three-level title in a specific format, classifying sub-section contents into a corresponding three-level title text, and outputting an analysis result; and detecting repeated contents in analysis results of adjacent text blocks, eliminating repetition based on text tail fix matching, integrating the analysis results of all the text blocks, generating a structured result containing three-level titles, corresponding texts, analysis time and text statistical information, and performing structural optimization. Compared with the prior art, the method has the advantages of high three-level title recognition accuracy, high structured data format compliance rate and the like.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO +1

system

Provide a system. 【Solution means】 Means for collecting existing documents and converting them into text, Means for analyzing text data, classifying and tagging it, Means for detecting and integrating duplicate content, Means for generating a unified document, Means for collecting business result data and matching it with the document content, Means for automatically updating the document, Means for verifying user authentication and permissions, Means for allowing users to provide feedback, A system including means for revising the document based on the feedback.
Owner:SOFTBANK GROUP CORP

Remediation of unstructured data using artificial intelligence

The systems and methods disclosed herein obtain (e.g., via a user interface) a collection of unstructured data, where each document includes a content set. Using a first AI model set, multiple summaries are generated by categorizing each document into clusters based on vector comparisons of content sets and summarizing the content for each cluster. A second AI model set (same as or different from the first AI model set) identifies duplicate content within the unstructured data by generating similarity values between pairs of summaries and determining if the similarity values meet a predefined threshold. A report is generated (e.g., on the user interface) indicating the duplicate content sets and / or the collection of unstructured data.
Owner:CITIBANK N A

System and method for podcast repetitive content detection

In one aspect, a method includes detecting a fingerprint match between query fingerprint data representing at least one audio segment within podcast content and reference fingerprint data representing known repetitive content within other podcast content, detecting a feature match between a set of audio features across multiple time-windows of the podcast content, and detecting a text match between at least one query text sentences from a transcript of the podcast content and reference text sentences, the reference text sentences comprising text sentences from the known repetitive content within the other podcast content. The method also includes responsive to the detections, generating sets of labels identifying potential repetitive content within the podcast content. The method also includes selecting, from the sets of labels, a consolidated set of labels identifying segments of repetitive content within the podcast content, and responsive to selecting the consolidated set of labels, performing an action.
Owner:GRACENOTE INC

Method and apparatus for optimizing machine learning models for text expansion

ActiveCN119720964Bquality improvementReduce the problem of low qualityNatural language data processingDuplicate contentData set
The application discloses an optimization method and device for a machine learning model for text expansion, the method comprising: inputting a first text into a preset first machine learning model for expansion to obtain a second text; obtaining a sample set of repeated texts in the second text, and comparing the sample set with a first preset text to obtain a paired data set; training a reward model according to the paired data set, and optimizing the first machine learning model according to the reward model to obtain a second machine learning model; evaluating whether the repeated problem of the output text of the second machine learning model is alleviated, and determining the second machine learning model as a target machine learning model in the case that the repeated problem of the output text of the second machine learning model is alleviated. The method can significantly improve the quality of the expanded text, reduce the problem of low quality of the expanded text caused by repeated content, improve the user experience, and improve the efficiency.
Owner:BEIJING VISCOSE ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

User interfaces for transferring content on electronic devices

PendingUS20260093564A1Interprogram communicationTransmissionDuplicate contentEngineering
In some embodiments, an electronic device detects that the user of the electronic device has a first subscription to a first content application and a second subscription to a second content application. In some embodiments, an electronic device displays a content user interface of the first content application. In some embodiments, after receiving a request (e.g., an input) to initiate a process to duplicate content items associated with a second user profile of the second application to a first user profile of the first application, the electronic device saves content items to the first user profile in the first application that meet one or more criteria. In some embodiments, the electronic device displays one or more visual indications of content items of the first content application that do not meet the criteria in a review user interface.
Owner:APPLE INC

Automatic migration control method for alignment of containerized mirror image version and database version

The invention provides an automatic migration control method for alignment of a containerized mirror image version and a database version. The method comprises the following sub-steps: S1, obtaining application mirror image version information; s2, obtaining a current database version number; s3, performing version consistency judgment; s4, loading a database migration script based on a migration mode, and generating a database migration script execution queue; s5, executing database migration and recording a result; s6, performing exception handling and start control; according to the method, migration scripts are sequenced according to a version identifier incremental rule, an execution queue is generated, and the script execution sequence is guaranteed to be compliant; repeated content is removed through difference comparison, and conflicts caused by independent maintenance of scripts of multiple projects are avoided; according to the method, the database migration process and the containerized application starting process are combined, triggering and execution of database migration are achieved, the manual intervention requirement is reduced, the risk of manual misoperation is reduced, and the automation efficiency of the system deployment and upgrading process is improved.
Owner:NANJING WIT SCI & TECH CO LTD

Remediation of unstructured data using artificial intelligence

The systems and methods disclosed herein obtain (e.g., via a user interface) a collection of unstructured data, where each document includes a content set. Using a first AI model set, multiple summaries are generated by categorizing each document into clusters based on vector comparisons of content sets and summarizing the content for each cluster. A second AI model set (same as or different from the first AI model set) identifies duplicate content within the unstructured data by generating similarity values between pairs of summaries and determining if the similarity values meet a predefined threshold. A report is generated (e.g., on the user interface) indicating the duplicate content sets and / or the collection of unstructured data.
Owner:CITIBANK N A

Automatic integration method and device of vehicle domain controller PDC and medium

PendingCN121957603ADecompilation/disassemblyVersion controlData packDuplicate content
The invention provides an automatic integration method and device of a vehicle domain controller PDC and a medium, and belongs to the technical field of vehicles. According to the method, multiple software data packets are decompressed step by step, file structure flattening operation is executed, tedious manual preparation and configuration are replaced, it is ensured that assemblies from different sources can be subjected to initial integration rapidly and correctly, and integration efficiency and compatibility are greatly improved; multi-level screening and duplicate removal are performed on duplicate files, duplicate contents and duplicate symbols, link conflicts are avoided from the source, compiling failures or undefined behaviors during operation caused by file redundancy and function duplicate definition are effectively prevented, and code quality and system reliability are improved; a function call relation graph and a stack allocation model are constructed, the maximum stack use depth of each function is calculated, a compiling environment is adapted, subpackage packaging operation is executed, software package integration efficiency and integration reliability are improved, and resource waste is reduced.
Owner:CHINA FAW CO LTD

Data processing method, data processing device, electronic equipment, medium and product

PendingCN122287899ADuplicate contentTheoretical computer science
This application provides a data processing method, data processing device, electronic device, medium, and product. The method is applied to the field of large-scale intelligent agent technology. The data processing method includes: logically parsing a target JSON text to obtain a Boolean logical expression of the target JSON text, where the Boolean logical expression includes the logical conditions of various business rules; simplifying the Boolean logical expression based on the logical conditions of M business rules to obtain a simplified logical expression; and removing duplicates from the simplified logical expression to obtain the target logical expression of the target JSON text. This method can effectively reduce the number of tokens processed by large models by preprocessing the JSON text, merging duplicate content, and simplifying the JSON text, thereby improving the inference speed of large models.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Method for perceiving self appearance, clothes and scene by digital separate body and related product

ActiveCN121745149ADigital data protectionBiological modelsDuplicate contentTimestamp
The invention discloses a method for perceiving self appearance, clothes and scenes through digital separate bodies and a related product. The method comprises the following steps: acquiring digital copy content associated with a digital copy unique identifier; performing legality verification and normalization processing on the digital duplicate content to obtain a standardized material; based on a preset appearance-clothing-scene field set, constructing a structured cue word and a structured constraint, and calling a multi-modal analysis model to reasone the standardized material to obtain a candidate structured result; performing grammar verification, field verification and value domain consistency verification on the candidate structured result to obtain a target structured result; and storing the target structured result and the unique identifier of the digital duplicate in an appearance-clothing-scene setting library in an associated manner, and writing a version number and an update timestamp. By means of the technical scheme, unified processing and field-level structured output of multi-format input are achieved, and uncertainty caused by manual participation and secondary analysis is reduced.
Owner:LIANGSHENG DIGITAL CREATIVE DESIGN (HANGZHOU) CO LTD

Code repeatability detection method and related device

PendingCN121387699AError detection/correctionProgramming languageDuplicate content
The invention discloses a code repetition detection method and a related device, relates to the field of software development, and can obtain an online or local modified file list based on a state monitoring instruction. And identifying the type of each code file in the changed file list to obtain the file type of each code file. And according to a repetition degree detection strategy corresponding to the file type, detecting repeated contents of the code file and the target code file to obtain repetition detection result data representing the repeated contents of the code file. And finally, summarizing the repeated detection result data according to a preset generation template to obtain a repetition degree detection report. According to the method, checking can be carried out before local codes are submitted, and the code repetition degree of code changing of developers can be reduced so as to extract or quote public files in advance. And the codes are compared with the target branch codes before being merged, so that developers can be guided to process repeated parts, and the code repetition degree of front-end projects is reduced to a greater extent.
Owner:AGRICULTURAL BANK OF CHINA

De-duplication method for video material

The invention relates to a video material deduplication method, which establishes an accurate range for subsequent processing by obtaining a script fragment and a semantic matched candidate video material fragment. The label matching degree is calculated based on the matching condition of the candidate video material segments and the scripts in the multi-level semantic label system, and it is ensured that the materials meet the script requirements in the subject and content categories. De-duplication related scores are calculated by analyzing the use history of candidate video material segments, the probability of repeatedly using the same or similar materials is effectively reduced through diversity scores, new materials are actively encouraged to be used through cold start scores, and the repetition risk and novelty of the materials are quantitatively evaluated from different dimensions. Finally, the target material is selected according to the tag matching degree and the weighting result of the deduplication correlation score, the content matching requirement and the deduplication target are organically combined, the repeated content is automatically avoided from the source in the material matching stage, a post-event manual auditing mode which is low in efficiency and high in subjectivity is replaced, and the intelligent level of video production is improved.
Owner:LAINENG (HANGZHOU) E-COMMERCE CO LTD

Attribute weighted fusion-based multi-source heterogeneous document deduplication method and system, equipment and medium

The invention relates to the field of data cleaning, and discloses a multi-source heterogeneous document deduplication method and system based on attribute weighted fusion, equipment and a medium, the method is characterized in that repeatability judgment is carried out on two documents based on a key attribute set and a mapping table, so that the repeatability judgment process focuses on and weights the core attributes of the documents, and the data cleaning efficiency is improved. According to the method, the duplicate removal problem is converted from text similarity comparison to discriminant force weighted fusion decision of key attributes, non-key differences among documents are ignored, recognition of repeated contents of the documents is more accurate, accuracy and robustness of repeatability judgment are improved, and a weight mechanism allows that the repeated contents of the documents can be more accurately recognized when part of key attributes are missing or errors exist. Reliable judgment can still be made based on existing key attributes, harsh dependence on information integrity is reduced, and high-precision repeated content recognition can still be achieved in an information missing scene. Besides, by configuring different key attribute sets and mapping tables, the method can be quickly and flexibly adapted to various vertical fields, and the universality and generalization ability are improved.
Owner:HUNAN XINGHAN DIGITAL INTELLIGENCE TECH CO LTD

A method for intelligent content filtering based on a specified theme scenario

The application relates to a method for intelligent content filtering based on a specified theme scene, which realizes accurate filtering of irrelevant content of ASR output text in the specified theme scene. The method comprises the following steps: optimizing a zero-shot classification model to obtain an optimized zero-shot classification model; performing theme classification on first ASR output text by using the optimized zero-shot classification model to obtain first irrelevant content text; constructing a dynamic word library, calculating the relevance of the first ASR output text and the dynamic word library, and screening second irrelevant content text according to the relevance; comparing and repeating the first irrelevant content text and the second irrelevant content text to obtain third irrelevant content text; and removing the third irrelevant content text from the first ASR output text to obtain filtered first ASR output text.
Owner:GUANGZHOU YUECHUANG ZHISHU INFORMATION TECH CO LTD

Construction content similarity retrieval and verification duplicate checking method and device and medium

PendingCN121682325ADuplicate contentSemantic representation
The invention relates to a similarity retrieval and verification duplicate checking method and device for construction content and a medium, and the method comprises the following steps: obtaining an electric power construction project text to be subjected to duplicate checking and a plurality of historical electric power construction project texts, respectively carrying out document preprocessing, extracting the construction content from the texts, and carrying out standardized semantic representation; obtaining a to-be-queried vector and a plurality of historical vectors; performing semantic similarity retrieval based on the to-be-queried vector and the historical vectors to obtain semantic similarity corresponding to each historical vector, and screening candidate repeated vectors; based on texts corresponding to the candidate repetitive vectors and the to-be-queried vector, searching a corresponding sub-graph in the knowledge graph, performing entity matching degree, relation matching degree, attribute consistency and structural similarity judgment, and determining logic repetitive degree; and judging whether the construction content is repeated or not based on the semantic similarity and the logic repetition degree. Compared with the prior art, the method has the advantages of being capable of solving the problem of repeated content missing detection caused by deliberately changing the expression mode and the like.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO

Corpus data screening method, electronic device, medium and product

PendingCN122285847ADuplicate contentDegree of similarity
This application discloses a corpus data screening method, electronic device, medium, and product. This application relates to the field of artificial intelligence technology. The method includes: acquiring a set of texts to be screened; and screening the set of texts to be screened based on the semantic similarity between texts in the set and a dynamic repeating semantic judgment threshold to obtain a target retention set. The dynamic repeating semantic judgment threshold is determined based on the semantic distribution of the set of texts to be screened. This application uses the semantic distribution as the basis for determining the threshold, enabling the threshold for determining whether semantic repetition exists to adaptively adjust according to the semantic distribution of different regions in the set of texts to be screened. This dynamic adjustment of the threshold maintains the integrity of key semantic differences in the set of texts to be screened while ensuring the ability to remove substantially repeated content, thus achieving a good deduplication effect.
Owner:ZTE CORP