Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

26 results about "Plagiarism detection" patented technology

Plagiarism detection is the process of locating instances of plagiarism within a work or document. The widespread use of computers and the advent of the Internet have made it easier to plagiarize the work of others.

User generated content identification method and device, electronic equipment and storage medium

The embodiment of the invention provides a user generated content identification method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring a multi-modal feature of a target work to be identified; based on the multi-modal features of the target works, a candidate work set with the similarity meeting a preset condition is retrieved from a pre-constructed multi-modal feature library, the multi-modal feature library comprises multi-modal features of historical works, the works comprise images containing user generated content, and the multi-modal features comprise fusion of visual features and semantic features of the works; screening out a reference work with a preset interaction behavior from the candidate work set; obtaining a difference degree influence factor between the target work and the reference work; and determining the originality of the target work based on the similarity and difference influence factors of the target work and the reference work. According to the UGC plagiarism detection method, the multi-modal vector database of image and decal semantic fusion is constructed, and multiple difference influence factors are introduced for comprehensive judgment, so that the accuracy and the processing efficiency of UGC plagiarism detection are remarkably improved, and the problem of missing detection of a traditional method in a light false change and local shielding scene is solved.
Owner:NETEASE (HANGZHOU) NETWORK CO LTD

Intelligent comparison and analysis method and system for similarity of examination answer codes

The invention provides an examination answer code similarity intelligent comparison and analysis method and system, and relates to the technical field of educational informationization, and the method comprises the steps: adaptively determining an optimal instrumentation position, collecting an execution path sequence and a memory access mode, and constructing a dynamic behavior feature vector; extracting static structure features and fusing the static structure features with the dynamic features; constructing a similarity measurement model by adopting a feature coding network and an adversarial discrimination network comprising three sub-networks; constructing a multi-level adversarial sample library based on a heuristic rule for adversarial training; and finally identifying and positioning the suspected plagiarism code snippets. According to the method, various plagiarism deformation strategies can be effectively identified, and the accuracy of code plagiarism detection is improved.
Owner:ATA ONLINE (BEIJING) EDUCATION TECH LTD

Automatic subjective question correcting method based on multi-technology fusion

The invention relates to the technical field of intelligent education systems, in particular to a subjective question automatic correction method based on multi-technology fusion, which comprises the following steps of: firstly, identifying whether a question is a strong proposition or not through a regular expression, and then performing knowledge point extraction, off-question judgment and double similarity evaluation on the strong proposition question; plagiarism detection, cosine similarity calculation and user-defined rounding are carried out on all the questions, and finally personalized comments are generated and correction results are output. According to the method, through double similarity evaluation, semantic relevance between answers and standard answers as well as questions is comprehensively considered, the misjudgment rate is remarkably reduced, meanwhile, through strong proposition recognition, knowledge point extraction and double similarity evaluation, the pertinence and accuracy of correction are remarkably improved, whether the answers of students meet the requirements of the questions or not can be deeply analyzed, and the correctness of the questions is improved. Misjudgment caused by insufficient semantic understanding in the prior art is avoided, and the method is particularly suitable for accurate scoring of strong proposition questions in subjective questions.
Owner:BEIJING XUECHENG GUILAI EDUCATION TECH CO LTD

Duplicate checking small language model training method combined with multi-level knowledge distillation

PendingCN120562402ASemantic analysisBiological modelsAlgorithmPlagiarism detection
The embodiment of the invention provides a duplicate checking small language model training method combined with multilevel knowledge distillation, and the method comprises the steps: obtaining duplicate checking sample pairs, and determining the complexity of the duplicate checking sample pairs according to the text features of the duplicate checking sample pairs; according to the complexity of the duplicate checking sample pair, determining distillation levels of the teacher model, and determining a weighting coefficient of each distillation level of the teacher model; determining the distillation loss between the teacher model and the student model according to the weighting coefficient of each network distillation of the teacher model, the first output result of each distillation level of the teacher model and the second output result of each distillation level of the student model; according to the distillation loss between the teacher model and the student model, updating parameters of the student model; and repeating the above steps until the updated student model meets the preset condition, and taking the updated student model as a duplicate checking small language model, thereby realizing a low-power-consumption and high-precision duplicate checking effect.
Owner:CHINA THREE GORGES CORPORATION

Science and technology project declaration duplicate checking method and system based on large language model and Jaccard similarity coefficient

The invention discloses a science and technology project declaration duplicate checking method and system based on a large language model and a Jaccard similarity coefficient, belongs to the technical field of natural language processing, and aims to solve the technical problem of how to improve the accuracy and efficiency of science and technology project declaration duplicate checking. According to the technical scheme, the method comprises the steps that document core content is extracted, wherein the core content of a science and technology project document to be subjected to duplicate checking is extracted through a large language model subjected to data training of a plurality of science and technology projects; splitting document segments: splitting the document core content into a plurality of document segments based on a natural language processing technology; document vectorization storage: converting document fragments into vectors through a text embedding model, and storing the vectors in a vector database; calculating a vector distance and retrieving historical items: calculating the Euclidean distance or cosine similarity between the vector of the item document fragment to be subjected to duplicate checking and the vector of the historical item document fragment, and extracting topK historical items with the closest distance; calculating a Jaccard similarity coefficient; aggregating the document similarity; and generating a duplicate checking result.
Owner:INSPUR SOFTWARE TECH CO LTD

Game model plagiarism detection method based on geometric features of three-dimensional model

The invention discloses a game model plagiarism detection method based on geometric features of a three-dimensional model, and the method comprises the following steps: S1, input processing: obtaining an original three-dimensional game model and a suspected infringement three-dimensional model, and taking the original three-dimensional game model and the suspected infringement three-dimensional model as input data; s2, point cloud preprocessing: respectively carrying out approximate surface uniform sampling on the two three-dimensional models, and sampling model data into a point cloud form; and S3, feature extraction: inputting the two preprocessed point clouds into a twinning neural network sharing the weight, and using a feature extractor to extract multi-level feature representations of the two point clouds. The method relates to the technical field of copyright protection of three-dimensional models, and has the beneficial effects that efficient three-dimensional model similarity detection is realized: a twin network architecture is introduced, the structure of the twin network architecture is improved, point cloud features are extracted in combination with a deep learning method, and the detection efficiency and accuracy of the similarity of the three-dimensional game models are improved.
Owner:YANBIAN UNIV

File code similarity analysis method based on directed graph isomorphism

The invention discloses a file code similarity analysis method, which is based on a directed graph isomorphism theory and comprises the following steps of: A, extracting characteristics of an open source code file, namely a Call Graph (Call Graph for short), and establishing a file sample library; b, extracting a function call relation graph of the file to be analyzed; c, executing the standardization work of the graph, and preprocessing the graph according to the related definition of the tree structure; d, extracting the maximum common subgraph of the function call relation graph of the file to be analyzed and the sample library file; and E, calculating a relationship between the maximum common subgraph and the function call relationship graph to obtain a similar result, and completing file code similarity analysis. According to the method, the similarity of the file codes is analyzed by depending on the function call relation graph, the internal characteristics and higher-level logic characteristics of the file codes are considered, and higher accuracy and efficiency are achieved. The method is suitable for occasions such as source code plagiarism detection.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Multi-document comparison method combining similarity calculation and accurate matching

The invention discloses a multi-document comparison method combining similarity calculation and accurate matching, and belongs to the technical field of information processing, and the method comprises the following steps: S1, text preprocessing, S2, blocking and positioning, S3, vector embedding, S4, similarity matrix construction, S5, Top-K pairing extraction, S6, similarity classification, S7, accurate matching, S8, visual highlighting, and S9, similarity evaluation. Through a sentence-level semantic vector and Top-K matching mechanism, the influence of document structure change is effectively overcome, the detection accuracy and robustness of content reconstruction type plagiarism are remarkably improved, matching dislocation caused by repeated content in a long document is avoided, the comparison stability is ensured, the problem of full-text vector semantic dilution is solved by embedding fine grit according to sentences, and the accuracy and robustness of content reconstruction type plagiarism detection are improved. According to the method, the recognition sensitivity of sentence pattern rewriting and local plagiarism is enhanced, and meanwhile, the abstract similarity is converted into traceable evidence by combining double-layer quantitative indexes with visual highlight positioning, so that the result interpretability and the practical value are greatly improved.
Owner:商飞软件有限公司

System, verification system and method for generating software code certification authentication credentials

The application discloses a system, a verification system and a method for generating software code copyright authentication credentials. The system comprises an AI semantic analysis module for analyzing target code using an AI code large model to generate a high-dimensional semantic embedding vector; a semantic fingerprint generation module for hashing the vector to obtain a binary semantic fingerprint; an extraction module for extracting a plurality of structural features from an abstract syntax tree; a structural fingerprint generation module for hashing the plurality of structural features to obtain a structural fingerprint; a core package construction module for constructing a code DNA core package; a signature module for encrypting the core package hash value using a private key; a trusted timestamp acquisition module for obtaining a timestamp authentication hash value by overall hashing, and obtaining a trusted timestamp certificate therefrom; and the code DNA core package, the developer signature and the trusted timestamp certificate jointly serving as the copyright authentication credentials. The application can improve the accuracy of code plagiarism detection and provide reliable copyright protection.
Owner:BEIJING UNITED TRUST TECH SERVICE CO LTD

Scientific and technological achievement duplicate checking method, device and equipment and readable storage medium

The invention provides a scientific and technological achievement duplicate checking method, device and equipment and a readable storage medium, and relates to the technical field of text duplicate checking, and the method comprises the steps of obtaining an existing research achievement and a research report to be subjected to duplicate checking; establishing a first mapping relation and a second mapping relation; based on the first mapping relationship and the second mapping relationship, constructing an original heterogeneous graph taking the research report, the research result, the paragraph and the word as nodes; constructing an initial adjacency matrix and an initial feature matrix of all nodes according to the original heterogeneous graph; building a graph convolution model, and inputting the initial adjacent matrix and the initial feature matrix into the graph convolution model for convolution training to obtain a learned feature matrix; the first vector code of the research result and the second vector code of the research report are extracted from the feature matrix, the similarity of the first vector code and the second vector code is calculated to form a duplicate checking report, dependence on manual annotation data is reduced, and duplicate checking accuracy and efficiency are remarkably improved.
Owner:CHINA STATE RAILWAY GRP CO LTD +2

Document duplicate checking method, device, equipment, medium and program product

PendingCN122309698APlagiarism detectionEngineering
This application provides a document plagiarism detection method, apparatus, device, medium, and program product, relating to the field of artificial intelligence technology, for improving the efficiency of document plagiarism detection. The specific technical solution is as follows: The document to be checked is input into a plagiarism detection model; the model extracts content from the document to obtain M primary key content texts corresponding to M plagiarism detection dimensions. The M plagiarism detection dimensions include at least: research objectives, research plan, research results, and research content; the model calculates the similarity between the M content text pairs corresponding to the M plagiarism detection dimensions, obtaining M primary similarity values; based on the M primary similarity values, the model outputs the plagiarism detection result of the document to be checked. This application is applied to document plagiarism detection scenarios.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD +1

Method, device and equipment for evaluating before patent application and storage medium

The invention discloses an assessment method and device before patent application, equipment and a storage medium, and the method comprises the steps: obtaining a to-be-assessed text in response to an assessment request, and carrying out the format detection of the to-be-assessed text based on a preset text format template, and obtaining a format detection result; according to the automatic generation detection data, the text quality detection data, the text duplicate checking data, the technical field detection data and / or the public opinion detection data corresponding to the to-be-evaluated text, obtaining an application detection result of the to-be-evaluated text; comparing the similarity of the text content between the to-be-assessed text and the similar text to obtain a new creative assessment result of the to-be-assessed text; and integrating the format detection result, the application detection result and the creativity evaluation result, and determining a pre-patent application evaluation result corresponding to the to-be-evaluated text. The technical problems that an existing evaluation method before patent application has high skill requirements on personnel during manual execution, the working efficiency is low, and only single text format evaluation can be performed during machine execution are solved.
Owner:GUANGZHOU OURCHEM INFORMATION CONSULTING

Method and device for calculating similarity, equipment, medium and program product

The invention discloses a similarity calculation method and device, equipment, a medium and a program product, and is applied to the technical field of network security. The method comprises the steps of obtaining a first binary program and obtaining a second binary program; mapping the first binary program into a plurality of first function vectors, and mapping the second binary program into a plurality of second function vectors; determining a first vector representation of the first binary program based on the number of first function vectors, and determining a second vector representation of the second binary program based on the number of second function vectors; a similarity between the first binary program and the second binary program is calculated based on the first vector representation and the second vector representation. According to the method, the similarity between the first binary program and the second binary program can be accurately determined, and the method can be applied to application scenes such as malicious software detection and software plagiarism detection.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

A Text Plagiarism Detection Method Combining Gene Mapping and Zero-Knowledge Proofs

This application relates to the technical field of text plagiarism detection, and discloses a text plagiarism detection method combining gene mapping and zero-knowledge proof. The method includes: segmenting the text into plot exons and transition introns, extracting structural features through dependency parsing, decomposing it into three-stage plot units, constructing a directed acyclic graph with plot units as nodes and logical dependencies as edges, and assigning dynamic weights; calculating structural and semantic similarity using graph neural networks and BERT models respectively, analyzing narrative logic similarity, and weighted fusion to obtain a comprehensive similarity; hashing the text gene map segments and uploading it to the blockchain to form a four-layer association chain, recording structural changes through smart contracts, and designing a zero-knowledge proof verification mechanism; setting multi-level judgment thresholds, combining multi-dimensional similarity to perform suspected content and plot plagiarism judgments, and updating the model and threshold strategy based on user feedback. This application can meet the requirements of accurate, secure, and flexible text content structure plagiarism judgment.
Owner:BEIJING BANGCLE TECH CO LTD

Article duplicate checking method for preventing generation of duplicate

The invention belongs to the technical field related to artificial intelligence, and particularly relates to an article duplicate checking method for preventing duplicate generation, which comprises the following steps that: duplicate checking is carried out by using text meaning segments and text meaning vectors, and the text meaning vector of each text meaning segment is obtained by inputting an attention mechanism AM algorithm through all vector groups corresponding to all keywords of the text meaning segment, obtaining a text meaning vector of the text meaning segment; the vector group is a vector group formed by word vectors of the keywords and disambiguation parameter vectors, the disambiguation parameter vector of each keyword is a vector formed by disambiguation parameters between the keyword and various effective words in the corpus, and the dimension of the disambiguation parameter vector is the same as that of the word vector of the keyword; word vectors of the keywords are formed by connecting co-occurrence matrixes of the keywords to symbolized normalized vectors of the co-occurrence matrixes. According to the method, a duplicate generation algorithm is counteracted, so that the AI cannot effectively reduce the weight of the article, and the reliability of the current article duplicate checking technology is improved.
Owner:HUAZHONG UNIV OF SCI & TECH

Text content structure plagiarism judgment method based on genetic map

The invention relates to the technical field of text plagiarism detection, and discloses a text content structure plagiarism judgment method based on a genetic atlas, which comprises the following steps: segmenting a text into plot exons and transition introns, analyzing and extracting structural features through a dependency syntax, decomposing into three-stage plot units, constructing a directed acyclic graph taking plot units as nodes and logic dependency relationships as edges, and allocating dynamic weights; structural and semantic similarities are respectively calculated through a graph neural network and a BERT model, narrative logic similarities are obtained through analysis, and comprehensive similarities are obtained through weighted fusion; performing segmented hash on the text gene map, performing chaining to form a four-layer association chain, recording structural change through an intelligent contract, and designing a zero-knowledge proof verification mechanism; setting a multi-level judgment threshold, executing suspected content and plot plagiarism judgment in combination with multi-dimensional similarity, and updating a model and a threshold strategy based on user feedback. According to the method and the device, accurate, safe and flexible text content structure plagiarism judgment requirements can be met.
Owner:BEIJING BANGCLE TECH CO LTD

Water and electricity science and technology document duplicate checking method and device based on knowledge graph

The invention relates to the technical field of computers, in particular to a hydropower science and technology document duplicate checking method and device based on a knowledge graph, and the method comprises the steps: obtaining historical science and technology document information; reading the information content of the historical science and technology document, and converting the historical science and technology document into a text file; performing data cleaning on the text file, and performing annotation processing to obtain a preprocessed text; carrying out analysis processing based on the structure of the preprocessed text, and carrying out structure disassembly by utilizing a regular expression to obtain a disassembled text; performing keyword extraction and content summarization on the disassembled text by utilizing a natural language processing algorithm and a large language model to obtain a knowledge graph data set; and inputting the knowledge graph data set into a knowledge graph integration tool to obtain the document duplicate checking knowledge graph. An information extraction prompt project for a large model is constructed, so that the efficiency and accuracy of information extraction can be improved, and deeper content understanding and decision support can be provided for users.
Owner:HUANENG CLEAN ENERGY RES INST +2

Intelligent comparison and analysis method and system for test answer codes

The application provides an examination answer code similarity intelligent comparison and analysis method and system, relates to the field of educational informatization technology, and comprises the following steps: determining an optimal insertion position through self-adaption, collecting an execution path sequence and a memory access mode to construct a dynamic behavior characteristic vector; extracting a static structure characteristic and fusing the dynamic characteristic; adopting a characteristic coding network and an adversarial discrimination network comprising three sub-networks to construct a similarity measurement model; constructing a multi-level adversarial sample library based on heuristic rules to perform adversarial training; and finally identifying and positioning a suspected plagiarism code segment. The application can effectively identify various plagiarism deformation strategies and improve the accuracy of code plagiarism detection.
Owner:ATA ONLINE (BEIJING) EDUCATION TECH LTD

A method and device for detecting paper plagiarism based on similarity

This application belongs to the field of computer technology, and discloses a method and device for detecting paper plagiarism based on similarity. The method includes: obtaining the paper to be detected and cleaning it to obtain the target paper; obtaining the historical paper database and screening to obtain multiple papers to be compared; calculating the digital fingerprints and word frequency vectors of the target paper and each paper to be compared; calculating the fingerprint similarity between the target paper and each paper to be compared based on the digital fingerprints; calculating the word frequency similarity between the target paper and each paper to be compared based on the word frequency vectors; and obtaining the plagiarism detection result according to the fingerprint similarity and the word frequency similarity. This application can reduce the amount of data to be processed in similarity calculation, improve the detection efficiency while ensuring the calculation accuracy.
Owner:GUANGZHOU KEAO INFORMATION TECH CO LTD

Local sequence comparison method and device for detecting continuous similar line number of file

PendingCN121957674Areduce complexitysimple resultReverse engineeringSoftware reuseLocal sequence alignmentRound complexity
The invention discloses a local sequence alignment algorithm and device for detecting the number of continuous similar rows of a file, and belongs to the technical field of code clone detection. The method comprises the following steps: carrying out preprocessing of blank removal, blank line removal and annotation removal on two files to obtain a line sequence; an improved Smith-Waterman algorithm is utilized, a'row 'is taken as a unit, difflib.Sequence Matcher > = 0.8 is taken as a similar threshold to construct a score matrix, and all continuous matching fragments with the length > = minlines are extracted through backtracking; and outputting starting and stopping line numbers and line-by-line contrast contents of each segment in a descending order according to the similar line number. According to the method, a language parser and time complexity O (n.m) are not needed, continuous similar regions of Type-1 to Type-3 clone can be positioned at one time, the result is visual, parameters can be configured, and the method is suitable for repeated code cleaning, plagiarism detection and reuse rate measurement of million row-level codes.
Owner:HAINAN CHEYOUJIA INFORMATION TECH CO LTD

Graphical User Interface for Uploading Academic Plagiarism Detection Files to Electronic Devices

1. Name of this design: Graphical User Interface for Uploading Academic Plagiarism Detection Files for Electronic Devices. 2. Purpose of this design: An electronic device. 3. The key design feature of this product is its graphical user interface. 4. The image or photograph that best illustrates the design's key features: the front view. 5. The electronic device of this design is a conventional design, and the rear view, left view, right view, top view and bottom view are omitted. 6. Purpose of the graphical user interface: Used for the interface display and human-computer interaction of academic plagiarism check file uploads. 7. Human-computer interaction method of graphical user interface: The main view displays the graphical user interface for uploading academic plagiarism check files. The upper right corner of the main view interface displays the "Historical Reports" button and the "Login / Register" button, which are used to view the history of academic plagiarism checking and the user's registration and login, respectively. The bottom of the interface displays a "Drag and drop or click to upload" button. Users can drag and drop the files that need to be checked for academic plagiarism to the "Drag and drop or click to upload" button, or click / touch the button to select the files that need to be checked for academic plagiarism and upload them. In each view, the "Historical Reports" button, "Login / Register" button, and "Drag and Drop or Click Upload" button in the main view are used for human-computer interaction and to implement functions. 8. Other situations requiring explanation: Other explanation: "X" in the interface represents a variable area containing text, numbers, or symbols.
Owner:WANGYIYOUDAO INFORMATION TECH BEIJING CO LTD

Method and system for detecting plagiarism in thesis

According to an aspect of the present disclosure, there is provided a thesis plagiarism detection method performed by a computing system. The thesis plagiarism detection method may comprise acquiring figure data for a target thesis, the figure data including images and text, acquiring first feature data for the figure data by applying the figure data to a first machine learning model, determining whether a thesis associated with second feature data having a similarity above a predetermined threshold with the acquired first feature data is found and determining the target thesis as a plagiarized thesis when it is determined that the thesis associated with the second feature data is found.
Owner:KOREA INST OF SCI & TECH INFORMATION

A game model plagiarism detection method based on three-dimensional model geometric features

The application discloses a kind of game model plagiarism detection methods based on three-dimensional model geometric features, comprising the following steps: S1, input processing: obtain original three-dimensional game model and suspected infringement three-dimensional model, as input data;S2, point cloud preprocessing: two three-dimensional models are respectively subjected to approximate surface uniform sampling, and model data is sampled into point cloud form;S3, feature extraction: the two point clouds after preprocessing are input into the twin neural network of shared weight, and the multi-level feature representation of two point clouds is respectively extracted using feature extractor.The application relates to the technical field of copyright protection of three-dimensional model, and the beneficial effects of the application are efficient three-dimensional model similarity detection: the application improves the detection efficiency and accuracy of three-dimensional game model similarity by introducing twin network architecture and improving its structure, combined with deep learning method to extract point cloud features.
Owner:YANBIAN UNIV

System for generating software code right authentication certificate, verification system and method

The invention discloses a system for generating a software code right authentication certificate and a verification system and method, and the system comprises an AI semantic analysis module which is used for analyzing a target code through an AI code large model, and generating a high-dimensional semantic embedding vector; the semantic fingerprint generation module is used for hashing the vector to obtain a binary semantic fingerprint; the extraction module is used for extracting a plurality of structural features from the abstract syntax tree; the structure fingerprint generation module is used for hashing the plurality of structure features to obtain a structure fingerprint; the core package construction module is used for constructing a code DNA core package; the signature module is used for encrypting the hash value of the core packet by using a private key; the trusted timestamp acquisition module is used for carrying out overall hash to obtain a timestamp authentication hash value, and obtaining a trusted timestamp certificate according to the timestamp authentication hash value; and the code DNA core package, the developer signature and the trusted timestamp certificate jointly serve as a right authentication certificate. According to the method, the accuracy of code plagiarism detection can be improved, and reliable copyright protection is provided.
Owner:BEIJING UNITED TRUST TECH SERVICE CO LTD

Content data detection method, detection device, electronic equipment, medium and product

The application discloses a content data detection method, a detection device, an electronic device, a medium and a product. The method comprises the following steps: obtaining to-be-detected content data; performing feature extraction on the to-be-detected content data to obtain feature data of the to-be-detected content data; obtaining target reference sample data in a preset reference sample data set according to the feature data of the to-be-detected content data; and determining that the to-be-detected content data is a plagiarism sample in a case where the similarity between the feature data of the to-be-detected content data and the feature data of the target reference sample data is greater than a preset similarity threshold. The method quickly screens the target reference sample data from the preset reference sample data set, then compares and analyzes the feature data of the to-be-detected content data and the target reference sample data, and finally determines whether there is a plagiarism sample, thereby realizing multi-level detection logic from macroscopic retrieval of the preset reference sample data set to microscopic comparison of feature data details, and improving the plagiarism detection precision.
Owner:郭颖