Duplicate checking method and system for products in various stages of software research and development
Through the integrated plagiarism check method, combined with abstract syntax trees, code intention labels and cross-modal attention mechanism, the limitations of plagiarism checking in the existing technology are solved, and efficient cross-domain and cross-modal plagiarism checking services are achieved.
Patent Information
- Application Number
- CN202510298018.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
AI Technical Summary
The existing plagiarism check technology is limited to a specific field or type and cannot provide cross-domain and cross-modal plagiarism check services, resulting in significant bottlenecks in the effectiveness and efficiency of plagiarism checks.
An integrated plagiarism check method is adopted, including code plagiarism check, document plagiarism check and functional plagiarism check. By analyzing software code, an abstract syntax tree with hash index is constructed, code intention tags are extracted, and text, images and table comparisons are combined with a cross-modal attention mechanism to achieve efficient plagiarism checks for various types of content.
It provides a unified plagiarism checking platform, which can handle various types of plagiarism checking tasks such as code, documents and functions at the same time, significantly improving the accuracy and efficiency of plagiarism checking and supporting cross-scenario and multi-modal plagiarism checking tasks.
Smart Images

Figure CN120144166A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of software engineering, and particularly relates to a duplicate checking method and system for products in each stage of software development. Background Art
[0002] With the rapid development of information technology, various duplicate checking technologies have been widely applied in multiple fields, especially in scenarios such as academic research, software development, and document management. However, existing duplicate checking technologies are usually limited to a specific field or type, and fail to provide cross-field and cross-modal duplicate checking services, resulting in significant bottlenecks in duplicate checking effects and efficiency. There are also some obvious disadvantages and limitations in current duplicate checking service platforms, which are specifically as follows:
[0003] 1. Limitations of code duplicate checking technology: Existing code duplicate checking technologies mainly rely on technologies such as text string matching, hash algorithms, or static code analysis. Although these methods can identify some similar codes, in the face of code refactoring, naming changes, format adjustments, etc., the accuracy and effect of duplicate checking are often significantly reduced, and the similarity between codes cannot be comprehensively identified.
[0004] 2. Singularity of document duplicate checking technology: Since documents may contain various types of content such as text, images, and tables, existing document duplicate checking technologies often can only perform duplicate checking on text content, ignoring the duplicate checking of non-text content such as images and tables. This makes it impossible to effectively identify the similarity in some documents containing similar images or tables, thus affecting the comprehensiveness and accuracy of duplicate checking.
[0005] 3. Singularity of duplicate checking platform services: Currently, many duplicate checking technologies are independent of each other and cannot handle multiple different types of duplicate checking tasks such as code, documents, and functions at the same time. Users often need to rely on multiple different duplicate checking tools, which not only increases the operation complexity but also reduces the duplicate checking efficiency. Summary of the Invention
[0006] The purpose of the present invention is to provide an integrated duplicate checking solution for code, document, and function duplicate checking tasks involved in the software product development process. This solution covers all stages of software development, can perform duplicate checking services efficiently and accurately, solves the limitation problems of duplicate checking of different types of content in the prior art, and significantly improves the accuracy and efficiency of different duplicate checking types.
[0007] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0008] A duplicate checking method for products in each stage of software development includes the following steps:
[0009] 1) Code duplication detection: Parse software code to construct an abstract syntax tree with a hash index; Extract code intent tags based on code semantics and map them to nodes through the hash index; Calculate code similarity through code intent tag matching or code structure comparison for code duplication detection.
[0010] 2) Document duplication detection: Perform text comparison, image comparison, and table comparison according to the cross-modal attention mechanism to detect document products in the software R & D stage.
[0011] 3) Function duplication detection: Generate function descriptions based on code intent tags, calculate the cosine similarity of the function descriptions in text and semantics with those in the function library, perform text-level matching and semantic-level matching to achieve function duplication detection.
[0012] Furthermore, the steps of constructing an abstract syntax tree with a hash index in step 1) include: Use a code parsing tool to parse the input code, construct an abstract syntax tree according to the syntax structure of the code; Extract key node information from the abstract syntax tree, calculate the hash value, associate the node and the hash value and store them in a hash table to establish a node hash index.
[0013] Furthermore, the steps of extracting code intent tags based on code semantics in step 1) include: Analyze the abstract syntax tree, use a deep learning model for code semantic parsing, extract the core function description of the code and generate code intent tags.
[0014] Furthermore, the steps of calculating code similarity through code intent tag matching in step 1) include: Quickly locate the code intent tags to be compared through the hash index; Calculate the similarity of intent tags of different code segments through BERT embedding and cosine similarity; Perform function category matching through KNN or TF-IDF methods to determine similar codes.
[0015] Furthermore, the steps of calculating code similarity through code structure comparison in step 1) include: Use a graph neural network to calculate the logical similarity of the code, and identify the code logical similarity across files and functions through control flow graph and data flow analysis.
[0016] Furthermore, the steps of text comparison in step 2) include: Use the difflib sequence comparison algorithm to calculate the longest common subsequence of the document to determine the surface similarity of the document; Extract the feature vectors of the document, and calculate the differences between the feature vectors using the Manhattan distance and cosine similarity to determine the content similarity of the document.
[0017] Further, the steps of image comparison in step 2) include: extracting the images in the document; generating the hash values of the images using a hash algorithm and calculating the image similarity based on the hash values; converting the images into grayscale images and calculating the pixel differences between different images to obtain the image similarity; extracting the key feature points in the images using the ORB algorithm and calculating the similarity of the key feature points between different images to obtain the image similarity.
[0018] Further, the steps of table comparison in step 2) include: extracting the table structure in the document and identifying the text and numerical information in the table; calculating the similarity of the text content in the table using the Levenshtein algorithm; directly comparing the numerical content in the table according to the set precision.
[0019] Further, the steps of generating the function description based on the code intention tags in step 3) include: integrating the code intention tags of multiple code snippets to generate a complete function description; removing stop words from the generated function description, performing part-of-speech tagging, and identifying elements to generate a standardized function description.
[0020] Further, step 3) also includes an extension and optimization step, specifically including: supporting cross-language function duplicate checking, deep learning rule fusion, template semantic dual matching, and data management level optimization.
[0021] A duplicate checking system for products in each stage of software development, used to implement the above duplicate checking method, includes:
[0022] Presentation layer: providing a user interface, including a Web end and an API port for submitting duplicate checking task requests;
[0023] Interface layer: implementing the interaction between the user's duplicate checking task request and the duplicate checking service module;
[0024] Gateway layer: responsible for the load balancing and scheduling of requests, and intelligently distributing requests according to the duplicate checking task type, where the duplicate checking task type includes code duplicate checking, document duplicate checking, and function duplicate checking;
[0025] Application layer: including a duplicate checking service module, responsible for code duplicate checking, document duplicate checking, and function duplicate checking, and used to process duplicate checking tasks;
[0026] Parallel computing layer: supporting parallel task scheduling and execution;
[0027] Data layer: responsible for storing duplicate checking task data, results, and related metadata;
[0028] Resource layer: responsible for managing computing and storage resources and supporting dynamic expansion.
[0029] The beneficial effects obtained by the present invention are as follows:
[0030] 1. The present invention provides a unified duplicate checking platform that can simultaneously handle various types of duplicate checking tasks such as code, documents, and functions, simplifies the complex process that requires multiple independent tools in traditional duplicate checking tools, improves the operation efficiency and user experience; at the same time, it supports flexible configuration options, and can adjust duplicate checking strategies and parameters according to different application scenarios (such as academic duplicate checking, software development, document management, etc.), providing customized duplicate checking solutions to meet the needs of multi-modal duplicate checking tasks.
[0031] 2. The present invention supports multi-dimensional duplicate checking of documents. Specialized similarity detection algorithms are designed for text, image, and table contents respectively, which can handle various changes of different data types, ensuring the comprehensiveness and accuracy of duplicate checking results. At the same time, combined with deep learning and efficient algorithm optimization, the performance and accuracy of the duplicate checking service are improved, reducing repetitive labor and copyright risks, and significantly improving work efficiency.
[0032] 3. The present invention adopts a duplicate checking method based on code intention tags, which can deeply capture the similarities at the code structure and logic levels, and is especially suitable for handling changes such as code refactoring, function renaming, and format adjustment, improving the duplicate checking accuracy and efficiency.
[0033] 4. The present invention proposes an innovative function duplicate checking method. Based on code intention tags, it infers the code function implementation, combines text similarity and function semantic matching, provides higher-precision function duplicate checking results, and ensures the correct classification of function modules with different expressions.
[0034] 5. The present invention designs a parallel duplicate checking technology architecture, which can efficiently process large-scale data sets, improve the duplicate checking speed, and at the same time ensure the accuracy. It is especially suitable for the duplicate checking needs of large-scale documents and code libraries. Combined with an intelligent optimization mechanism, it automatically adjusts resource allocation and strategies, further improving the processing efficiency and system performance.
[0035] 6. The present invention has high scalability and high availability, can flexibly adjust resources and functions according to business needs, supports high concurrency and large-scale task processing, and ensures the stable operation of the duplicate checking service.
[0036] 7. The present invention simplifies complex duplicate checking tasks through the concept of tooling, optimizes resource allocation, improves the degree of intelligence and automation in the duplicate checking process, and further improves work efficiency and system performance.
[0037] 8. The present invention provides comprehensive duplicate checking services for different fields, supports cross-scenario and multi-modal duplicate checking tasks, solves the deficiencies of existing technologies in duplicate checking accuracy and efficiency, and has broad application prospects. Brief Description of the Drawings
[0038] Figure 1It is a flowchart of the duplicate checking method for products in each stage of software development in the embodiments of the present invention.
[0039] Figure 2 It is an architecture diagram of the duplicate checking system for products in each stage of software development in the embodiments of the present invention. Detailed implementation manners
[0040] To make the technical features, advantages or technical effects in the above technical solutions of the present invention more obvious and understandable, the following is a detailed description through embodiments.
[0041] The embodiments of the present invention provide a duplicate checking method for products in each stage of software development, and its processing flow is as Figure 1 shown, including processing means for code duplicate checking, document duplicate checking and function duplicate checking, which are specifically described as follows:
[0042] I. Code duplicate checking
[0043] To ensure the high efficiency and high precision of code duplicate checking, this part integrates a variety of technical means, from surface text similarity to deep structure analysis, and then to intelligent optimization mechanisms, realizing a comprehensive and high-performance code duplicate checking function. The specific steps are as follows:
[0044] 1. Abstract syntax tree construction:
[0045] During the code duplicate checking process, first use a professional code parsing tool (such as the ast module in Python or JDT in Java) to parse the input code. The parsing tool reads the code line by line, disassembles it into a series of syntax units according to the syntax rules, such as expressions, statements, function definitions, etc., and constructs an abstract syntax tree (AST) based on the hierarchical relationship and logical structure between these units. During the construction of the AST, the tool will automatically ignore non-functional elements such as variable names, spaces, and comments, and focus on extracting the core syntax structure of the code, so as to ensure that the generated AST can accurately reflect the logical relationship and functional structure of the code. This lays a solid foundation for subsequent code analysis and duplicate checking work.
[0046] 2. Abstract syntax tree hash index construction:
[0047] After the AST is established, extract the key node information such as packages, files, functions, etc. in it. For each node, design a hash function, take its key information (such as package name, file name, function name, and the structural features of the node, etc.) as input, and calculate the corresponding hash value. Then, associate each AST node with its corresponding hash value and store it in the hash table. In this way, during the subsequent duplicate checking process, the AST nodes and related code segments to be compared can be quickly located through the hash index, thus greatly improving the efficiency of searching and retrieval.
[0048] 3. Code intention label extraction:
[0049] 1) Code parsing and syntax conversion: Based on the constructed AST, deeply analyze the code structure, and use code block splitting technology to disassemble the code into logical units, thereby further improving the accuracy of analysis.
[0050] 2) Code intention label generation: Use a deep learning model (such as a CodeBERT variant) for code semantic parsing; extract the core function description of the code through a code summary generation model and convert it into a code intention label; at the same time, map these labels to the corresponding code segments in combination with a hash index.
[0051] 4. Code similarity calculation:
[0052] 1) Based on code intention label matching: Use the hash index constructed by the AST to quickly locate the code intention labels of the packages, files, and functions to be compared. Calculate the similarity of intention labels of different code segments through BERT embedding and cosine similarity, and use methods such as KNN and TF-IDF for function category matching. If the code intention labels of two packages are completely dissimilar, the detailed comparison of their internal files and functions can be directly skipped, and the fast location and duplicate checking of label intentions can be used to reduce unnecessary calculations.
[0053] 2) Code structure comparison: Use a graph neural network (GNN) to calculate the logical similarity of the code, and combine the control flow graph (CFG) and data flow analysis (DFA) to identify the code logical similarity across files and functions.
[0054] 3) Intelligent optimization: Dynamically select a matching strategy according to the complexity of the code segment. For short code segments, use intention label matching; for complex code, use a comparison method that combines AST and control dependence graph (CDG). At the same time, distributed duplicate checking is supported to improve the processing ability of large-scale code libraries.
[0055] II. Document duplicate checking
[0056] For document products such as requirement specifications and test outlines generated in each stage of software development, cross-modal attention mechanisms can be combined to perform duplicate checking through the following text comparison, image comparison, and table comparison.
[0057] 1. Text comparison:
[0058] 1) Classical sequence alignment algorithm: Use the difflib sequence alignment algorithm in the Python standard library to quickly compare the document text and accurately identify the differences between strings. This algorithm effectively processes the surface similarity of the text by calculating the longest common subsequence (LCS) of two text sequences.
[0059] 2) Feature vector similarity calculation: To measure the similarity of document content in depth, the feature vectors of document texts are further extracted, and the Manhattan distance and cosine similarity are used to quantitatively calculate the differences between vectors. The Manhattan distance and cosine similarity can intuitively reflect the degree of difference in the content distribution of document texts, providing more accurate data support for the duplicate check results.
[0060] 2. Image comparison:
[0061] 1) Image extraction: Use pdfplumber to extract the image content in PDF files and python-docx to extract the image content in DOC / DOCX files to ensure the integrity of the image quality and make it available for subsequent processing. Each extracted image will be converted to a standard format (such as PNG or JPEG) for subsequent processing.
[0062] 2) Hashing algorithm: Use the hashing algorithm to generate the hash value of the image, convert the image into a string of numbers, and quickly calculate the similarity between different images.
[0063] 3) Grayscale image algorithm: Convert the image to a grayscale image and obtain the similarity value by calculating the pixel differences between different images.
[0064] 4) Feature point similarity: To measure the similarity of image content in depth, the ORB algorithm is used to extract the key feature points in the image. The ORB algorithm first uses the FAST corner detector to locate the key points in the image and generates a unique feature representation for each key point through the rotation-invariant BRIEF descriptor. Then, by comparing the set of feature points extracted from the images, the similarity between them is calculated to identify the commonalities and differences between the images. The advantage of the ORB algorithm is that it is highly robust to rotation, scaling, and illumination changes, etc., and can effectively capture the key information of the image, thus providing efficient and accurate comparison results for image duplicate checking.
[0065] 3. Table comparison:
[0066] 1) Table extraction: Use pdfplumber to extract the row and column structure of the table from PDF documents and identify the text and numerical information in the table. Pdfplumber can effectively identify the tables embedded in PDF files and convert the tables into structured data by parsing the cell content for subsequent processing. For the tables in Word documents, use python-docx to extract the table structure and cell content and convert them into a data format suitable for comparison, including information such as text, numerical values, and formats, and uniformly process the table structure.
[0067] 2) Text comparison: The Levenshtein algorithm is used to compare the text content in the table cells, calculate the minimum edit distance of insertion, deletion, and replacement operations, so as to quantify the similarity of the text content. For the numerical content in the table, the values are directly compared by setting the precision requirements to handle possible decimal point differences or format changes in the table.
[0068] III. Function duplication checking
[0069] 1. Function inference based on code intention tags:
[0070] 1) Obtaining code intention tags: Reuse the code intention tags extracted in the code duplication checking stage, which already contain the core function description information of the code.
[0071] 2) Integrating function information: Integrate the code intention tags of multiple related code snippets to sort out the complete function description information, so as to comprehensively reflect the functions implemented by the code.
[0072] 2. Processing function intention tags:
[0073] Apply natural language processing techniques to the integrated function description, remove stop words and perform part-of-speech tagging; combine dependency syntax analysis to identify elements such as function verbs, operation objects, inputs and outputs, and form a standardized description.
[0074] 3. Similarity calculation based on semantic matching:
[0075] Use cosine similarity to calculate the matching degree between the input function description and the existing function library; perform text-level matching and semantic-level matching on the input function description to further improve the accuracy of duplication checking.
[0076] 4. Expansion and optimization:
[0077] 1) Cross-language function duplication checking: Support cross-language function duplication checking, convert function descriptions in different languages through a translation API, and calculate cross-language semantic similarity in combination with a multi-language embedding model to expand the scope of application of the technology.
[0078] 2) Deep learning rule fusion: In the process of function extraction and classification, combine the semantic features extracted by the model with domain rules to improve the robustness and applicability of the classification results, especially suitable for the processing of function modules in sub-sectors.
[0079] 3) Template and semantic double matching: Combine semantic embedding and template matching strategies to optimize the coverage and accuracy of function label similarity calculation. Semantic matching focuses on deep semantic associations, while template matching quickly identifies standardized function descriptions, and the two complement and cooperate with each other.
[0080] 4) Optimization of the data management hierarchy: Optimize the storage structure through gerund parsing, simplify the data management process, and improve the duplicate checking efficiency through an efficient indexing mechanism to support large-scale function duplicate checking tasks.
[0081] An embodiment of the present invention also provides a duplicate checking system for products at each stage of software development. The backend is developed based on Python, supporting parallel computing and containerized deployment. The system architecture is as Figure 2 shown, including multiple levels, and each level works together to achieve efficient duplicate checking processing.
[0082] Presentation layer: Provides a user interface, supporting the Web side and API ports. Users can submit duplicate checking tasks through these interfaces.
[0083] Interface layer: Realizes the interaction between user requests and the duplicate checking service module, supports multiple input formats, and provides parameter parsing and verification functions.
[0084] Gateway layer: Responsible for load balancing and scheduling of requests, ensuring the high availability of the system, and intelligently distributing requests according to task types (such as text, code, pictures, functions).
[0085] Application layer: Contains the core duplicate checking service module, using deep learning and advanced algorithms (such as SVM, grayscale image, hash algorithm, etc.) to process duplicate checking tasks.
[0086] Parallel computing layer: Supports parallel task scheduling and execution, improves the duplicate checking efficiency through parallel computing, and is especially suitable for processing large-scale tasks.
[0087] Data layer: Responsible for storing duplicate checking task data, results, and related metadata, and adopting efficient databases (such as MySQL) and caching technologies (such as Redis) to improve query speed.
[0088] Resource layer: Manages computing and storage resources, supports dynamic expansion to cope with growing processing requirements.
[0089] This system adopts a parallel duplicate checking technology architecture, which can efficiently process large-scale data sets, ensuring the high speed and high accuracy of duplicate checking tasks. Through asynchronous task scheduling and task distribution technologies, this system can provide excellent performance when processing massive code and document duplicate checking tasks.
[0090] Although the present invention has been disclosed above in embodiments, it is not intended to limit the present invention. Appropriate modifications or equivalent replacements made by those of ordinary skill in the art to the technical solutions of the present invention shall all be covered within the protection scope of the present invention. The protection scope of the present invention shall be defined by the claims.
Claims
1. A method for checking for duplicate content in products at all stages of software development, characterized in that: The following steps are involved: 1) Code duplication check: parse the software code and build an abstract syntax tree with hash index; Extract code intent labels based on code semantics and map them to nodes through hash indexes; calculate code similarity and check for code duplication through code intent label matching or code structure comparison; 2) Document duplication check: Based on the cross-modal attention mechanism, text comparison, image comparison and table comparison are performed to check the duplication of document products in the software development stage; 3) Function duplication check: Generate function description based on code intent label, calculate the text and semantic cosine similarity between the function description and the function description in the function library, perform text level matching and semantic level matching to realize function duplication check.
2. The duplicate checking method according to claim 1, characterized in that: The step of constructing an abstract syntax tree with a hash index in step 1) includes: using a code parsing tool to parse the input code and construct an abstract syntax tree according to the grammatical structure of the code; extracting key node information in the abstract syntax tree, calculating the hash value, associating the node and the hash value and storing them in a hash table, and establishing a node hash index.
3. The duplicate checking method according to claim 1, characterized in that: The step of extracting code intent labels based on code semantics in step 1) includes: analyzing the abstract syntax tree, performing code semantic parsing using a deep learning model, extracting the core functional description of the code and generating code intent labels.
4. The duplicate checking method according to claim 1, characterized in that: The step of calculating code similarity by matching code intent labels in step 1) includes: quickly locating the code intent labels to be compared by hash index; calculating the intent label similarity of different code snippets by BERT embedding and cosine similarity; and matching functional categories by KNN or TF-IDF method to determine similar codes.
5. The duplicate checking method according to claim 1, characterized in that: The step of calculating code similarity through code structure comparison in step 1) includes: using a graph neural network to calculate the logical similarity of the code, and identifying the logical similarity of the code across files and functions through control flow graph and data flow analysis.
6. The duplicate checking method according to claim 1, characterized in that: The text comparison step in step 2) includes: using the difflib sequence alignment algorithm to calculate the longest common subsequence of the documents to determine the surface similarity of the documents; extracting the feature vectors of the documents, using Manhattan distance and cosine similarity to calculate the difference between the feature vectors, and determining the content similarity of the documents.
7. The duplicate checking method according to claim 1, characterized in that: The image comparison step in step 2) includes: extracting images from the document; using a hash algorithm to generate a hash value of the image, and calculating the image similarity based on the hash value; converting the image into a grayscale image, and calculating the pixel difference between different images to obtain the image similarity; using the ORB algorithm to extract key feature points in the image, and calculating the similarity of the key feature points between different images to obtain the image similarity.
8. The duplicate checking method according to claim 1, characterized in that: The table comparison step in step 2) includes: Extract the table structure in the document and identify the text and numerical information in the table; use the Levenshtein algorithm to calculate the similarity of the text content in the table; and directly compare the numerical content in the table according to the set accuracy.
9. The duplicate checking method according to claim 1, characterized in that: The step of generating a functional description based on the code intent label in step 3) includes: integrating the code intent labels of multiple code snippets to generate a complete functional description; removing stop words from the generated functional description, performing part-of-speech tagging, and identifying elements to generate a standardized functional description.
10. A duplicate checking system for products at all stages of software development, used to implement the duplicate checking method described in any one of claims 1 to 9, characterized in that: include: Presentation layer: provides a user interface, including a web terminal and API port for submitting duplicate checking task requests; Interface layer: realizes the interaction between the user's duplicate checking task request and the duplicate checking service module; Gateway layer: responsible for load balancing and scheduling of requests, and intelligently distributes requests according to the type of duplicate checking task, which includes code duplicate checking, document duplicate checking, and function duplicate checking; Application layer: Contains the duplicate checking service module, which is responsible for checking code duplicates, document duplicates, and function duplicates, and is used to handle duplicate checking tasks; Parallel computing layer: supports parallel task scheduling and execution; Data layer: responsible for storing duplicate checking task data, results and related metadata; Resource layer: responsible for managing computing and storage resources and supporting dynamic expansion.
Citation Information
Cited By
Multi-mode-based duplicate checking method and device and document duplicate checking system
CN121074932A