A method for checking duplicates by using a picture classification and similar information system scheme

By combining image classification and similarity calculation methods with deep learning and feature matching algorithms, the problem of insufficient image similarity analysis in existing technologies is solved, enabling efficient and accurate identification of plagiarism in design schemes, and has broad application prospects.

CN116543205BActive Publication Date: 2026-03-31JIANGSU CENTURY INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-28
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing plagiarism detection technologies are mainly based on text or code similarity comparisons, lacking effective analysis of image features, resulting in insufficient image similarity comparisons.

Method used

Image classification and similarity calculation methods are adopted. Deep learning image classification and feature matching algorithms are used, combined with the Faiss vector database for image feature vector retrieval, and similarity thresholds are used to determine whether there is plagiarism in the design scheme.

Benefits of technology

It enables comprehensive similarity analysis of non-textual information in design schemes, improves the accuracy and efficiency of plagiarism detection, can identify plagiarism behavior, and has broad application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116543205B_ABST
    Figure CN116543205B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of information system scheme plagiarism checking method using picture classification and similar information system scheme, comprising the following steps: step 1: reading relevant information system scheme, obtaining the corresponding picture in scheme, the picture in information system scheme is classified and identified, and classification storage is carried out;Step 2: feature extraction is carried out to the picture in each classification, and the image is converted into feature vector by feature extraction network model algorithm;Step 3: using vector retrieval technology Faiss vector database, retrieval is carried out, the similarity of the feature vector of the design scheme and the feature vector of the picture in each folder is obtained, and the picture with the highest similarity is found;Step 4: according to the similarity threshold value, judge whether the picture in scheme exists plagiarism;Step 5: the corresponding weight of different picture classification is multiplied by the number of repeated pictures to obtain the similar value of the picture of this kind, and the similar values of different classification pictures are summarized, and whether the entire scheme exists plagiarism is judged according to the similarity threshold value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a computational analysis method, specifically a method for deduplicating information system schemes using image classification and similarity, belonging to the field of image recognition technology. Background Technology

[0002] Plagiarism detection technology has become an important component of information management, especially in the field of information systems where the need is even more urgent. Existing plagiarism detection technologies mainly rely on comparing the similarity of text or code, but because image features are more complex, image-based plagiarism detection has limitations. Therefore, a new method is needed to address this issue. Summary of the Invention

[0003] This invention addresses the problems existing in the prior art by providing a method for detecting plagiarism in information system solutions using image classification and similarity. This technical solution applies image analysis and vector retrieval techniques to the plagiarism detection of information system solutions. This method can more comprehensively consider the similarity of non-textual information in the design solution and can better identify plagiarism.

[0004] To achieve the above objectives, the technical solution of the present invention is as follows: a method for deduplicating information system schemes using image classification and similarity, characterized in that the method includes the following steps:

[0005] Step 1: Read the relevant information system plan, obtain the corresponding images in the plan, classify and identify the images in the information system plan, and store them in categories;

[0006] Step 2: Extract features from the images in each category, and convert the images into feature vectors using a feature extraction network model algorithm;

[0007] Step 3: Use the Faiss vector database for vector retrieval to search and obtain the similarity between the feature vector of the design scheme and the feature vector of the image in each folder, and find the image with the highest similarity.

[0008] Step 4: Determine whether the images in the solution are plagiarized based on the similarity threshold;

[0009] Step 5: Multiply the corresponding weights of different image categories by the number of duplicate images to obtain the similarity value of that category of images. Summarize the similarity values ​​of images from different categories and determine whether the entire scheme is plagiarized based on the similarity threshold.

[0010] As an improvement of the present invention, step 1: the images in the information system scheme can be classified using a deep learning-based image classification algorithm, such as a convolutional neural network (CNN), specifically divided into functional architecture diagram, system structure diagram, system deployment diagram, technical architecture diagram, basic function interface image, and special function interface image.

[0011] As an improvement of the present invention, step 2: adopt a similarity calculation method based on feature matching, such as the SIFT algorithm or the ORB algorithm.

[0012] As an improvement of the present invention, step 3: the Faiss vector database is searched. When the image size is less than 10,000 images, a plane index is used: IndexFlatL2. The distance between the queried data and all data in the index is calculated to obtain the L2 distance (Euclidean distance) between them.

[0013] As an improvement of the present invention, step 5: multiply the corresponding weight of different image categories by the number of duplicate images to obtain the duplicate similarity value of that category of images, and summarize the similarity values ​​of images of different categories. Among them, the functional architecture diagram, technical architecture diagram, and system deployment diagram are not used as the basis for duplicate judgment because there are similarities in the actual situation, and a reminder is given; the system structure diagram is a single diagram, and the weight is set higher accordingly; the system interface related images need to be further judged because they use a unified framework. They are divided into basic functional interfaces, including login interface, personnel and organization management interface, permission management interface, etc., with correspondingly lower weights, and dedicated functional interfaces have special characteristics and correspondingly higher weights.

[0014] Compared to existing technologies, this invention has the following advantages: the method for detecting plagiarism in information system solutions using image classification and similarity provides a more comprehensive consideration of the similarity of non-textual information in the design scheme, enabling better identification of plagiarism and significantly improving the efficiency and accuracy of plagiarism detection. The method provided by this invention can not only be effectively applied to the field of information systems but also extended to other fields such as art design, industrial design, and architectural design. This method also has broad application prospects, applicable to various intellectual property fields, image retrieval fields, etc., and has positive significance for protecting intellectual property rights, improving work efficiency, and safeguarding public interests.

[0015] The method for deduplication of information system solutions using image classification and similarity provided by this invention has the following technical advantages:

[0016] 1) It takes into account the similarity of non-textual information in the design scheme more comprehensively. Compared with traditional text similarity comparison methods, the method of this invention can classify and calculate the similarity of non-textual information such as images in the design scheme, which can more accurately identify plagiarism and thus improve the accuracy of plagiarism detection;

[0017] 2) Improved efficiency in plagiarism detection. This invention uses a deep learning-based image classification algorithm for image classification, which can automatically classify images in design schemes, greatly reducing the workload of manual classification; at the same time, it uses a feature matching-based similarity calculation method, which has a fast calculation speed and can quickly perform plagiarism detection, thus improving the efficiency of plagiarism detection.

[0018] 3) It has broad application prospects. The method provided by this invention can not only be applied to the field of information system design, but also extended to other fields, such as art design, industrial design, and architectural design. This method also has broad application prospects, and can be applied to various intellectual property fields, image retrieval fields, etc., which has positive significance for protecting intellectual property rights, improving work efficiency, and safeguarding public interests.

[0019] 4) Strong robustness. The image classification algorithm and similarity calculation method used in this invention can overcome the problems of noise, deformation, and distortion in the design schemes, and have strong robustness, enabling more accurate identification of design schemes with high similarity;

[0020] In summary, the method for detecting duplicate information system design schemes using image classification and similarity provided by this invention has advantages such as high efficiency, accuracy, robustness, and wide applicability, and has good practicality and promotional value. Attached Figure Description

[0021] Figure 1 This is a schematic diagram of the overall steps in Embodiment 1 of the present invention. Detailed Implementation

[0022] To enhance understanding of the present invention, the embodiments will be described in detail below with reference to the accompanying drawings.

[0023] Example 1: See Figure 1 A method for deduplication of information system solutions using image classification and similarity, the method comprising the following steps:

[0024] Step 1: Read the relevant information system plan, obtain the corresponding images in the plan, classify and identify the images in the information system plan, and store them in categories;

[0025] Step 2: Extract features from the images in each category, and convert the images into feature vectors using a feature extraction network model algorithm;

[0026] Step 3: Use the Faiss vector database for vector retrieval to search and obtain the similarity between the feature vector of the design scheme and the feature vector of the image in each folder, and find the image with the highest similarity.

[0027] Step 4: Determine whether the images in the solution are plagiarized based on the similarity threshold;

[0028] Step 5: Multiply the corresponding weights of different image categories by the number of duplicate images to obtain the similarity value of that category of images. Summarize the similarity values ​​of images from different categories and determine whether the entire scheme is plagiarized based on the similarity threshold.

[0029] The specific process is as follows: Step 1: The images in the information system solution can be classified using a deep learning-based image classification algorithm, such as Convolutional Neural Network (CNN). Specifically, the classification is divided into functional architecture diagram, system structure diagram, system deployment diagram, technical architecture diagram, basic function interface images, and special function interface images.

[0030] Step 2: Use a feature-matching-based similarity calculation method, such as the SIFT algorithm or the ORB algorithm.

[0031] Step 3: Retrieve from the Faiss vector database. When the number of images is less than 10,000, use a flat index: IndexFlatL2. For the queried data, calculate the distance between it and all data in the index to obtain the L2 distance (Euclidean distance) between them.

[0032] Step 5: Multiply the corresponding weight of different image categories by the number of duplicate images to obtain the similarity value of that category. Summarize the similarity values ​​of images from different categories. Among them, functional architecture diagrams, technical architecture diagrams, and system deployment diagrams are not used as the basis for duplicate judgment because there are similarities in actual situations, but a reminder will be given; the system structure diagram is a single diagram, so the weight is set higher accordingly; system interface related images need to be further judged because they use a unified framework. They are divided into basic function interfaces, including login interface, personnel and organization management interface, and permission management interface, with correspondingly lower weights, and dedicated function interfaces with special characteristics, with correspondingly higher weights.

[0033] Example 2: See Figure 1 As shown, a method for deduplicating information system solutions using image classification and similarity includes the following steps:

[0034] Step 1.1: Read the information system solution document "Design Scheme for Archives Management System" and obtain the corresponding images;

[0035] Step 1.2: It is determined that the "Records Management System Design Scheme" document contains images;

[0036] Deployment 1.3 to perform image classification and detection: Obtain functional architecture Figure 1 Zhang, System Structure Figure 1 Zhang, System Deployment Figure 1 Zhang, Technical Architecture Figure 1Zhang, 5 images of basic function interfaces, 25 images of special function interfaces, and 5 other images (flowcharts, data flow diagrams, etc.);

[0037] Step 2: Extract features from the images in each category, and convert the images into feature vectors using a feature extraction network model algorithm;

[0038] Step 3: Search historical data to find the image with the highest similarity, and obtain 1 similar image corresponding to the functional architecture diagram, 1 similar image corresponding to the system structure diagram, 1 similar image corresponding to the system deployment diagram, 1 similar image corresponding to the technical architecture diagram, 5 similar images corresponding to the basic function interface image, 5 similar images corresponding to the special function interface image, and 1 similar image corresponding to other (flowchart, data flow diagram, etc.).

[0039] Step 4: Based on a similarity threshold of 85%, filter to obtain 1 similar image corresponding to the functional architecture diagram, 1 similar image corresponding to the system structure diagram, 1 similar image corresponding to the system deployment diagram, 1 similar image corresponding to the technical architecture diagram, 5 similar images corresponding to the basic function interface image, 3 similar images corresponding to the special function interface image, and 1 similar image corresponding to other (flowcharts, data flow diagrams, etc.).

[0040] Step 5.1: Calculate the weights of the results. The system structure diagram score = 1 (image) * 0.5 (weight), the basic function interface images = 5 (images) * 0.01 (weight), the special function interface images = 3 (images) * 0.3 (weight), and the other images = 5 (images) * 0.01 (weight).

[0041] Step 5.2: Summarize and calculate the results. The repetition value of the scheme = 0.5 + 0.05 + 0.9 + 0.01, which gives 1.46.

[0042] Step 5.3: Since the result 1.46 is greater than the threshold 1, the information system solution document "Design Scheme of Archive Management System" is duplicated with the historical database.

[0043] It should be noted that the above embodiments are not intended to limit the scope of protection of the present invention. Equivalent transformations or substitutions made based on the above technical solutions all fall within the scope of protection of the claims of the present invention.

Claims

1. A method for duplicate detection using a picture classification and similar pair information system scheme, characterized in that, The method comprises the following steps: Step 1: read the relevant information system scheme, obtain the corresponding pictures in the scheme, classify and store the pictures in the information system scheme; Step 2: feature extraction is performed on the pictures in each category, and the image is converted into a feature vector through a feature extraction network model algorithm; Step 3: use the vector retrieval technology Faiss vector database to retrieve, obtain the similarity of the design scheme feature vector and the picture feature vector in each folder, and find the picture with the highest similarity; Step 4: judge whether there is plagiarism in the picture in the scheme according to the first similarity threshold; Step 5: multiply the corresponding weight of different picture categories by the number of repeated pictures to obtain the repetition similarity value of this type of picture, and summarize the similarity values of different categories of pictures, and judge whether there is plagiarism in the whole scheme according to the second similarity threshold.

2. The method of claim 1, wherein the method further comprises: Step 1: the pictures in the information system scheme are classified by using an image classification algorithm based on deep learning, which is specifically classified into functional architecture diagram, system structure diagram, system deployment diagram, technical architecture diagram, basic function interface picture and special function interface picture.

3. The method of claim 1, wherein the method further comprises: Step 2: adopt a similarity calculation method based on feature matching, and adopt SIFT algorithm or ORB algorithm.

4. The method of claim 1, wherein the method further comprises: Step 3: the Faiss vector database is used for retrieval, when the image size is below 10,000 pictures, the flat index IndexFlatL2 is used, the distance between the query data and all data in the index is calculated to obtain the L2 distance between them.

5. The method of claim 1, wherein the method further comprises: Step 5: multiply the corresponding weight of different picture categories by the number of repeated pictures to obtain the repetition similarity value of this type of picture, and summarize the similarity values of different categories of pictures, wherein, the functional architecture diagram, the technical architecture diagram and the system deployment diagram are not used as the basis for repetition judgment due to the existence of similar situations in actual situation, and a prompt is given; the system structure diagram is one picture, and the weight is set relatively high; the system interface related pictures need to be further judged due to the existence of uniform framework, and are divided into basic function interface including login interface, personnel organization management interface and permission management interface with relatively low weight, and special function interface with relatively high weight.

Citation Information

Patent Citations

  • Science and technology project duplicate checking method for carrying out big data matching calculation based on image partitioning

    CN110929069A

  • Medical paper image duplicate checking system based on deep learning

    CN115171117A