A Data Deduplication Method and System Based on the DBSCAN Algorithm with Tolerant Clustering Bias

By classifying users using the DBSCAN algorithm, which is tolerant of clustering bias, the security issues of data uploaded by users from the same organization are resolved, and security and privacy protection are achieved in the data deduplication process.

CN115994133BActive Publication Date: 2026-07-17SHANDONG ZHENGZHONG COMP NETWORK TECH CONSULTING +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANDONG ZHENGZHONG COMP NETWORK TECH CONSULTING
Filing Date
2022-12-09
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing encrypted data deduplication schemes struggle to effectively differentiate data popularity when faced with user uploads from the same organization, leading to reduced security and increased risk of internal data breaches.

Method used

The DBSCAN algorithm, which is tolerant of clustering bias, is used to classify users. During the classification process, a certain degree of bias is tolerated. Newly added sample points are either assigned to already clustered classes or treated as noise points. Different encryption methods are adopted in combination with popularity judgment to reduce the risk of internal data leakage.

Benefits of technology

By using the DBSCAN algorithm, which tolerates clustering bias, the popularity of data can be effectively distinguished. Appropriate encryption methods are employed to reduce the risk of internal data leakage and ensure the confidentiality of non-popular data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115994133B_ABST
    Figure CN115994133B_ABST
Patent Text Reader

Abstract

This invention relates to the field of data deduplication technology, and provides a data deduplication method and system based on the DBSCAN algorithm with tolerance for clustering bias. The method includes: obtaining tags for files to be uploaded by a user, and determining the popularity of the files based on the tags; if the popularity of the files is less than a set value, obtaining semantically secure encrypted ciphertext, and classifying users according to user attributes using the DBSCAN algorithm with tolerance for clustering bias, updating the file's popularity contribution based on the classification results; if the popularity of the files is greater than the set value, obtaining converged encrypted ciphertext; wherein, the DBSCAN algorithm with tolerance for clustering bias assigns newly added users to already clustered classes or treats them as noise points until the number of biased points exceeds a threshold, at which point DBSCAN clustering is re-performed. This reduces the risk of internal data leakage.
Need to check novelty before this filing date? Find Prior Art