Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

38 results about "Duplicate content" patented technology

Duplicate content is a term used in the field of search engine optimization to describe content that appears on more than one web page. The duplicate content can be substantial parts of the content within or across domains and can be either exactly duplicate or closely similar. When multiple pages contain essentially the same content, search engines such as Google and Bing can penalize or cease displaying the copying site in any relevant search results.

Task processing method, text processing method, automatic question-answering method, task processing model training method, information processing method based on task processing model, and cloud training platform

PCT designated stageWO2025196533A1Digital data information retrievalSemantic analysisInformation processingDuplicate content
Provided in the embodiments of the present disclosure are a task processing method, a text processing method, an automatic question-answering method, a task processing model training method, an information processing method based on a task processing model, and a cloud training platform, which are applied to the technical field of computers. The task processing method comprises: acquiring task data of a target task; and inputting the task data into a task processing model, so as to obtain a task processing result of the target task, wherein the task processing model is obtained by means of performing training on the basis of a plurality of pieces of target sample data and target sample results of the plurality of pieces of target sample data, the target sample data is obtained by means of performing screening on the basis of detection results of repeated content of a plurality of pieces of sample data, and the content of the target sample results is not repeated. Target sample data is obtained by means of performing screening on the basis of detection results of repeated content, and the content of target sample results is not repeated, such that the repetition illusion of a model is reduced without damaging the processing capability of the model itself and relying on external knowledge, thereby improving the accuracy and integrity of a task processing result.
Owner:CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD

Aviation text content cleaning and labeling method, system and equipment and medium

PendingCN121543549ANatural language analysisBiological modelsDuplicate contentAviation
The invention relates to the technical field of aeronautical text data processing, and discloses an aeronautical text content cleaning and labeling method, system, device and medium wherein the method comprises: noise filtering: identifying and removing noise in an aeronautical text in combination with static cleaning and a general large model; format standardization: converting the aviation text after noise removal into a standardized format text; duplicate removal and error correction: detecting duplicate contents based on a hash algorithm, and correcting spelling errors and grammar errors based on a general large model to obtain an aviation text subjected to duplicate removal and error correction; entity identification: key entities are extracted based on the general large model, and the extracted key entities are labeled; active learning: screening high-value samples in the marked key entities based on an uncertainty query strategy; and dynamic optimization: performing verification and iterative optimization on the marking result in combination with the aviation knowledge base. According to the method, the automation level, the labeling accuracy and the system self-adaptive capability of aviation text processing can be remarkably improved.
Owner:四川腾盾科技有限公司 +1

Online man-machine conversation optimization method, system and device, electronic equipment, storage medium and program product

The embodiment of the invention provides an online man-machine conversation optimization method, system and device, electronic equipment, a storage medium and a program product. In the scheme provided by the embodiment, when a first dialogue generation model executes a streaming response based on user input data, a currently generated first dialogue text fragment is subjected to repeated detection, and when it is detected that repeated content exists in the first dialogue text fragment, the first dialogue generation model is controlled to interrupt execution of the streaming response; in addition, reply generation constraint information is determined according to the repeated content in the first reply text segment, and then reply generation cue words are generated based on the reply generation constraint information and the non-repeated content in the first reply text segment. And after the dialogue generation cue word is input into the first dialogue generation model, the first dialogue generation model is triggered to determine a dialogue generation starting point according to the non-repeated content in the first dialogue text fragment, and streaming response is continuously executed from the dialogue generation starting point.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Multi-document comparison method combining similarity calculation and accurate matching

The invention discloses a multi-document comparison method combining similarity calculation and accurate matching, and belongs to the technical field of information processing, and the method comprises the following steps: S1, text preprocessing, S2, blocking and positioning, S3, vector embedding, S4, similarity matrix construction, S5, Top-K pairing extraction, S6, similarity classification, S7, accurate matching, S8, visual highlighting, and S9, similarity evaluation. Through a sentence-level semantic vector and Top-K matching mechanism, the influence of document structure change is effectively overcome, the detection accuracy and robustness of content reconstruction type plagiarism are remarkably improved, matching dislocation caused by repeated content in a long document is avoided, the comparison stability is ensured, the problem of full-text vector semantic dilution is solved by embedding fine grit according to sentences, and the accuracy and robustness of content reconstruction type plagiarism detection are improved. According to the method, the recognition sensitivity of sentence pattern rewriting and local plagiarism is enhanced, and meanwhile, the abstract similarity is converted into traceable evidence by combining double-layer quantitative indexes with visual highlight positioning, so that the result interpretability and the practical value are greatly improved.
Owner:商飞软件有限公司

Project budget generation method and apparatus, terminal device, and storage medium

ActiveCN116862444BFinanceOffice automationDuplicate contentTerminal equipment
The application discloses a project budget generation method and device, terminal equipment and a storage medium, and relates to the technical field of data processing. The method comprises the following steps: acquiring project construction content; performing repeated content processing on the project construction content to obtain de-duplicated project construction content; performing budget estimation on the de-duplicated project construction content to obtain a budget amount of the project; and performing budget examination on the budget amount of the project to obtain a final budget result of the project. The application can perform repeated content processing on large-scale and complex information system projects through an automatic means, and thus the processing efficiency can be improved.
Owner:CHINA MOBILE INFORMATION TECHNOLOGY CO LTD +1

Task processing method, text processing method, automatic question answering method, task processing model training method, information processing method based on task processing model and cloud training platform

The embodiment of the invention provides a task processing method, a text processing method, an automatic question answering method, a task processing model training method, an information processing method based on a task processing model and a cloud training platform, which are applied to the technical field of computers, and the task processing method comprises the following steps: obtaining task data of a target task; the task data are input into a task processing model, a task processing result of the target task is obtained, the task processing model is obtained through training on the basis of multiple pieces of target sample data and target sample results of the multiple pieces of target sample data, and the target sample data are obtained through screening on the basis of repeated content detection results of the multiple pieces of sample data; the target sample result content is not repeated. The target sample data is obtained by screening based on the repeated content detection result, and the content of the target sample result is not repeated, so that the repeated illusion of the model is relieved on the premise of not damaging the processing capability of the model and not depending on external knowledge, and the accuracy and integrity of the task processing result are improved.
Owner:HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD

Method, computer device, and computer program for eliminating duplicate content in large-scale recommendation system

To provide a method, computer device, and computer program for eliminating duplicate content in a large-scale recommendation system.SOLUTION: A method for eliminating duplicate content in a large-scale recommendation system includes steps of: generating, for each piece of video content, a video hash using multiple segmented images generated from the video content; and identifying, based on the video hash, duplicate video content and eliminating duplication.SELECTED DRAWING: Figure 3
Owner:LINE PLUS

A test case processing method, device, equipment and medium

Embodiments of the present disclosure relate to a test case processing method, device, equipment and medium, wherein the method comprises: obtaining a first test case file in table format; performing format conversion on first data corresponding to the first test case file to obtain second data in mind map data structure; performing deduplication and integration processing on element node data of multiple test cases in the second data to obtain third data; and converting the third data into a second test case file in mind map format through a mind map tool. With the above technical solution, in the conversion process of the test case file from the table format to the mind map format, through the deduplication and integration processing, the data redundancy problem existing in the converted mind map format is avoided, the user can quickly find the required data, the readability of the data is improved, and in subsequent maintenance, the repeated work of repeated content is avoided, and the maintenance efficiency is improved.
Owner:DOUYIN VISION CO LTD

Power planning report three-level title structured analysis generation method and device and medium

The invention relates to an electric power planning report three-level title structured analysis generation method and device and a medium, and the method calls a large language model to execute the following steps: obtaining a to-be-processed electric power planning report and carrying out report paragraph text extraction to obtain an original text; segmenting the original text into overlapped text blocks, wherein the overlapped part comprises a three-level title context; constructing a power planning professional cue word, for each text block, identifying a three-level title in a specific format, classifying sub-section contents into a corresponding three-level title text, and outputting an analysis result; and detecting repeated contents in analysis results of adjacent text blocks, eliminating repetition based on text tail fix matching, integrating the analysis results of all the text blocks, generating a structured result containing three-level titles, corresponding texts, analysis time and text statistical information, and performing structural optimization. Compared with the prior art, the method has the advantages of high three-level title recognition accuracy, high structured data format compliance rate and the like.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO +1

A multi-document summary generation method for crisis help information

The present invention discloses a multi-document summary generation method for crisis help information; the method is divided into: an extractive summary stage and a generative summary stage. In the extractive summary stage, the sentence graph structure is used to fuse contextual information in a deep neural network, and the subgraph structure is used to optimize the information extraction to effectively extract key information from numerous documents; in the generative summary stage, the powerful sequence generation ability of the BART model and the characteristics of the pointer generation network in accurately copying key information and avoiding duplicate content generation are used to further streamline and summarize the extracted information, and generate an accurate and compact summary of the crisis help information. The method of the present invention is suitable for processing and summarizing massive help information in crisis scenarios. The two-stage method of the present invention can quickly and accurately provide clear information summaries for rescue and emergency management, greatly improving the efficiency and effectiveness of crisis response.
Owner:FUDAN UNIVERSITY

Electronic medical record data governance method based on agent mechanism and related device

This application discloses an agent-based electronic medical record (EMR) data governance method and related apparatus, relating to the field of data processing technology. After acquiring the EMR to be governed, the method first processes each chapter of the EMR by removing garbled characters and duplicate content errors, obtaining the chapter governance result, thus obtaining the chapter-governed EMR. Then, for each chapter in the chapter-governed EMR, consecutive error-free EMR text in each field is copied, and EMR text with inconsistencies between field names and content is generated, obtaining the chapter's field governance result. The field governance results of all chapters are combined to obtain the field-governed EMR, which is the data-governed EMR. This solution can both handle errors in EMR data and improve the efficiency and quality of EMR data governance.
Owner:TSINGHUA UNIVERSITY +1

system

Provide a system. 【Solution means】 Means for collecting existing documents and converting them into text, Means for analyzing text data, classifying and tagging it, Means for detecting and integrating duplicate content, Means for generating a unified document, Means for collecting business result data and matching it with the document content, Means for automatically updating the document, Means for verifying user authentication and permissions, Means for allowing users to provide feedback, A system including means for revising the document based on the feedback.
Owner:SOFTBANK GROUP CORP

Scientific and technological document duplicate checking method and device, equipment and medium

The invention provides a science and technology document duplicate checking method and device, equipment and a medium, and the method comprises the steps: obtaining a fine tuning sample set according to an obtained sample science and technology document set, carrying out the fine tuning of a candidate large model, and obtaining a fine-tuned target large model; constructing a corresponding target science and technology knowledge graph, and obtaining a similar document set of the document object to be subjected to duplicate checking based on the target science and technology knowledge graph; and calling the model capability of the target large model, and obtaining a target duplicate checking result of the document object based on the similar document set through the model capability. According to the technical scheme, the document range for duplicate checking of the document object is narrowed through the target science and technology knowledge graph, the adaptation degree of the target large model and the industry to which the science and technology document belongs is improved through fine adjustment of the large model, the target duplicate checking result of the science and technology document is obtained based on the target large model, the duplicate checking efficiency and accuracy of the science and technology document are improved, and the practicability is high. And the duplicate checking method of the science and technology document is optimized.
Owner:HUANENG CLEAN ENERGY RES INST +2

Remediation of unstructured data using artificial intelligence

The systems and methods disclosed herein obtain (e.g., via a user interface) a collection of unstructured data, where each document includes a content set. Using a first AI model set, multiple summaries are generated by categorizing each document into clusters based on vector comparisons of content sets and summarizing the content for each cluster. A second AI model set (same as or different from the first AI model set) identifies duplicate content within the unstructured data by generating similarity values between pairs of summaries and determining if the similarity values meet a predefined threshold. A report is generated (e.g., on the user interface) indicating the duplicate content sets and / or the collection of unstructured data.
Owner:CITIBANK N A

Intelligent management system for water conservancy project construction plans and construction organization design method

The present invention discloses an intelligent management system for water conservancy project construction plans, including a client installed in a local computer Windows system and a server installed in a server, wherein the client includes a desktop and a WEB client; a VSTO plug-in is used to embed the desktop client into Word or WPS; the desktop client has the functions of creating, opening, editing and saving new construction group documents; the server has a database and a template library, and the template library is used to store construction group template documents pre-made by users. The present invention also discloses a corresponding construction organization design method. The present invention can effectively reuse the content of historical construction group documents, and the results of one compilation work can be conveniently reused in subsequent projects, thereby improving the efficiency and consistency of compiling construction organization design documents, and solving problems such as inconsistent content formats, large compilation workload, confusing node levels, inconsistent expressions of repeated content, inconsistent node dates and difficulty in modification, and time-consuming and labor-intensive formula input.
Owner:HENAN PROVINCIAL WATER CONSERVANCY FIRST ENG BUREAU +1

System and method for podcast repetitive content detection

In one aspect, a method includes detecting a fingerprint match between query fingerprint data representing at least one audio segment within podcast content and reference fingerprint data representing known repetitive content within other podcast content, detecting a feature match between a set of audio features across multiple time-windows of the podcast content, and detecting a text match between at least one query text sentences from a transcript of the podcast content and reference text sentences, the reference text sentences comprising text sentences from the known repetitive content within the other podcast content. The method also includes responsive to the detections, generating sets of labels identifying potential repetitive content within the podcast content. The method also includes selecting, from the sets of labels, a consolidated set of labels identifying segments of repetitive content within the podcast content, and responsive to selecting the consolidated set of labels, performing an action.
Owner:GRACENOTE INC

Statistical image duplicate checking method and system based on large language model

The invention discloses a statistical image duplicate checking method and system based on a large language model, and the method comprises the steps: constructing a large language model architecture for converting a statistical image into a text, carrying out the analysis of the statistical image based on the large language model, converting the statistical image into a structured text, and carrying out the duplicate checking of the statistical image. And recognizing the text part and the numerical part of the statistical image structured text, and further calculating the similarity of the text and the numerical part to realize duplicate checking of the statistical image. According to the method, the image content can be efficiently retrieved and compared, the accuracy and efficiency of statistical image duplicate checking are improved, and the method has important application value in the aspects of image copyright protection, content auditing and the like.
Owner:PEKING UNIV +1

Method and apparatus for optimizing machine learning models for text expansion

ActiveCN119720964Bquality improvementReduce the problem of low qualityNatural language data processingDuplicate contentData set
The application discloses an optimization method and device for a machine learning model for text expansion, the method comprising: inputting a first text into a preset first machine learning model for expansion to obtain a second text; obtaining a sample set of repeated texts in the second text, and comparing the sample set with a first preset text to obtain a paired data set; training a reward model according to the paired data set, and optimizing the first machine learning model according to the reward model to obtain a second machine learning model; evaluating whether the repeated problem of the output text of the second machine learning model is alleviated, and determining the second machine learning model as a target machine learning model in the case that the repeated problem of the output text of the second machine learning model is alleviated. The method can significantly improve the quality of the expanded text, reduce the problem of low quality of the expanded text caused by repeated content, improve the user experience, and improve the efficiency.
Owner:BEIJING VISCOSE ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

User interfaces for transferring content on electronic devices

PendingUS20260093564A1Interprogram communicationTransmissionDuplicate contentEngineering
In some embodiments, an electronic device detects that the user of the electronic device has a first subscription to a first content application and a second subscription to a second content application. In some embodiments, an electronic device displays a content user interface of the first content application. In some embodiments, after receiving a request (e.g., an input) to initiate a process to duplicate content items associated with a second user profile of the second application to a first user profile of the first application, the electronic device saves content items to the first user profile in the first application that meet one or more criteria. In some embodiments, the electronic device displays one or more visual indications of content items of the first content application that do not meet the criteria in a review user interface.
Owner:APPLE INC

Automatic migration control method for alignment of containerized mirror image version and database version

The invention provides an automatic migration control method for alignment of a containerized mirror image version and a database version. The method comprises the following sub-steps: S1, obtaining application mirror image version information; s2, obtaining a current database version number; s3, performing version consistency judgment; s4, loading a database migration script based on a migration mode, and generating a database migration script execution queue; s5, executing database migration and recording a result; s6, performing exception handling and start control; according to the method, migration scripts are sequenced according to a version identifier incremental rule, an execution queue is generated, and the script execution sequence is guaranteed to be compliant; repeated content is removed through difference comparison, and conflicts caused by independent maintenance of scripts of multiple projects are avoided; according to the method, the database migration process and the containerized application starting process are combined, triggering and execution of database migration are achieved, the manual intervention requirement is reduced, the risk of manual misoperation is reduced, and the automation efficiency of the system deployment and upgrading process is improved.
Owner:NANJING WIT SCI & TECH CO LTD

Remediation of unstructured data using artificial intelligence

The systems and methods disclosed herein obtain (e.g., via a user interface) a collection of unstructured data, where each document includes a content set. Using a first AI model set, multiple summaries are generated by categorizing each document into clusters based on vector comparisons of content sets and summarizing the content for each cluster. A second AI model set (same as or different from the first AI model set) identifies duplicate content within the unstructured data by generating similarity values between pairs of summaries and determining if the similarity values meet a predefined threshold. A report is generated (e.g., on the user interface) indicating the duplicate content sets and / or the collection of unstructured data.
Owner:CITIBANK N A

Automatic integration method and device of vehicle domain controller PDC and medium

PendingCN121957603ADecompilation/disassemblyVersion controlData packDuplicate content
The invention provides an automatic integration method and device of a vehicle domain controller PDC and a medium, and belongs to the technical field of vehicles. According to the method, multiple software data packets are decompressed step by step, file structure flattening operation is executed, tedious manual preparation and configuration are replaced, it is ensured that assemblies from different sources can be subjected to initial integration rapidly and correctly, and integration efficiency and compatibility are greatly improved; multi-level screening and duplicate removal are performed on duplicate files, duplicate contents and duplicate symbols, link conflicts are avoided from the source, compiling failures or undefined behaviors during operation caused by file redundancy and function duplicate definition are effectively prevented, and code quality and system reliability are improved; a function call relation graph and a stack allocation model are constructed, the maximum stack use depth of each function is calculated, a compiling environment is adapted, subpackage packaging operation is executed, software package integration efficiency and integration reliability are improved, and resource waste is reduced.
Owner:CHINA FAW CO LTD

Data processing method, data processing device, electronic equipment, medium and product

PendingCN122287899ADuplicate contentTheoretical computer science
This application provides a data processing method, data processing device, electronic device, medium, and product. The method is applied to the field of large-scale intelligent agent technology. The data processing method includes: logically parsing a target JSON text to obtain a Boolean logical expression of the target JSON text, where the Boolean logical expression includes the logical conditions of various business rules; simplifying the Boolean logical expression based on the logical conditions of M business rules to obtain a simplified logical expression; and removing duplicates from the simplified logical expression to obtain the target logical expression of the target JSON text. This method can effectively reduce the number of tokens processed by large models by preprocessing the JSON text, merging duplicate content, and simplifying the JSON text, thereby improving the inference speed of large models.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Method for perceiving self appearance, clothes and scene by digital separate body and related product

ActiveCN121745149ADigital data protectionBiological modelsDuplicate contentTimestamp
The invention discloses a method for perceiving self appearance, clothes and scenes through digital separate bodies and a related product. The method comprises the following steps: acquiring digital copy content associated with a digital copy unique identifier; performing legality verification and normalization processing on the digital duplicate content to obtain a standardized material; based on a preset appearance-clothing-scene field set, constructing a structured cue word and a structured constraint, and calling a multi-modal analysis model to reasone the standardized material to obtain a candidate structured result; performing grammar verification, field verification and value domain consistency verification on the candidate structured result to obtain a target structured result; and storing the target structured result and the unique identifier of the digital duplicate in an appearance-clothing-scene setting library in an associated manner, and writing a version number and an update timestamp. By means of the technical scheme, unified processing and field-level structured output of multi-format input are achieved, and uncertainty caused by manual participation and secondary analysis is reduced.
Owner:LIANGSHENG DIGITAL CREATIVE DESIGN (HANGZHOU) CO LTD

A knowledge base management system based on large model

The present invention relates to the field of knowledge base management technology, and in particular to a knowledge base management system based on a large model, the system comprising a data acquisition module, a data processing module, a knowledge base construction module, and a verification module. The present invention continuously scans text data to remove duplicate content and avoid word segmentation errors, classifies source data in the social security field to collect unstructured data therein for data conversion, that is, by processing the text in the unstructured data, identifying entities and concepts in the text, and identifying and recording the relationships between entities, thereby improving data availability and query efficiency, verifying the accuracy of word segmentation processing by combining user feedback data, and automatically identifying whether erroneous word segmentation results occur by recording user query data, thereby adjusting the word segmentation processing process and model training parameters to achieve effective management of the social security service system.
Owner:ZHONGHE YUNKE INFORMATION TECH GRP CO LTD

Method for generating code coverage rate, related device and computer program product

PendingCN120705052AError detection/correctionDuplicate contentTest script
The invention provides a method for generating a code coverage rate, a related device and a computer program product, and the method comprises the steps: integrating a test target file through different test scripts, and obtaining a code coverage rate product corresponding to each test script; then, combining the code coverage rate products to obtain an integrated code coverage rate product; repeated content in the integrated code coverage rate product is removed, and a target code coverage rate product is obtained; and based on the target file and the target code coverage rate product, generating a code coverage rate result for the target file. Therefore, not only can the integrated code coverage rate product be utilized to supervise and manage the overall condition of the integration test, but also redundant data in the code coverage rate product can be reduced in a combination and duplicate removal mode, and a code coverage rate result with higher quality and lighter weight is provided; therefore, the user can supervise the condition of the integration test more simply, conveniently and efficiently by using the code coverage rate result.
Owner:SHANGHAI HODE INFORMATION TECH CO LTD

Intelligent content filtering method based on specified theme scene

The invention relates to an intelligent content filtering method based on a specified theme scene. Accurate filtering of irrelevant content on an ASR output text in the specified theme scene is achieved. The method comprises the following steps: optimizing a zero sample classification model to obtain an optimized zero sample classification model; performing subject classification on the first ASR output text by adopting an optimized zero sample classification model to obtain a first irrelevant content text; constructing a dynamic word bank, calculating the correlation between the first ASR output text and the dynamic word bank, and screening according to the correlation to obtain a second irrelevant content text; performing repeated content comparison on the first irrelevant content text and the second irrelevant content text to obtain a third irrelevant content text; and removing the third irrelevant content text from the first ASR output text to obtain the filtered first ASR output text.
Owner:GUANGZHOU YUECHUANG ZHISHU INFORMATION TECH CO LTD

Code repeatability detection method and related device

PendingCN121387699AError detection/correctionProgramming languageDuplicate content
The invention discloses a code repetition detection method and a related device, relates to the field of software development, and can obtain an online or local modified file list based on a state monitoring instruction. And identifying the type of each code file in the changed file list to obtain the file type of each code file. And according to a repetition degree detection strategy corresponding to the file type, detecting repeated contents of the code file and the target code file to obtain repetition detection result data representing the repeated contents of the code file. And finally, summarizing the repeated detection result data according to a preset generation template to obtain a repetition degree detection report. According to the method, checking can be carried out before local codes are submitted, and the code repetition degree of code changing of developers can be reduced so as to extract or quote public files in advance. And the codes are compared with the target branch codes before being merged, so that developers can be guided to process repeated parts, and the code repetition degree of front-end projects is reduced to a greater extent.
Owner:AGRICULTURAL BANK OF CHINA