De-duplication retrieval method, device and equipment for file uploading

CN122152775APending Publication Date: 2026-06-05CHINA MOBILE JIUTIAN ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MOBILE JIUTIAN ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
Filing Date
2026-02-03
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing file upload deduplication technologies face difficulties in cross-type file identification and retrieval efficiency, resulting in low semantic recognition accuracy, low retrieval efficiency, and wasted storage resources, failing to meet practical application needs.

Method used

By obtaining the semantic vector and hash signature of the file to be uploaded, searching using a target vector database, and combining feature vector comparison to determine duplicate files, a multi-head attention mechanism and reinforcement learning algorithm are used to optimize the judgment threshold, and a composite index of MinHash and IVF_HNSW is constructed for efficient filtering.

Benefits of technology

It improves the accuracy and retrieval efficiency of cross-type file deduplication, reduces redundant data storage, and enhances system performance and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122152775A_ABST
    Figure CN122152775A_ABST
Patent Text Reader

Abstract

The application provides a deduplication retrieval method, device and equipment for file uploading, and belongs to the field of data processing. The method comprises the following steps: obtaining a semantic vector and a hash signature of a to-be-uploaded file; performing retrieval in a target vector database according to the semantic vector of the to-be-uploaded file and the hash signature of the to-be-uploaded file to obtain a similar file, wherein the target vector database stores semantic vectors and hash signatures of uploaded files; obtaining a first feature vector of the to-be-uploaded file and a second feature vector of the similar file; determining whether the to-be-uploaded file is a duplicate file according to the first feature vector and the second feature vector; and uploading the to-be-uploaded file if it is not a duplicate file. The deduplication retrieval efficiency and accuracy during file uploading can be improved, and redundant data storage can be reduced.
Need to check novelty before this filing date? Find Prior Art