Layered storage and verification system and method for unstructured data lake

Through the unstructured data lake hierarchical storage and verification system, problems such as low retrieval efficiency, high storage cost, and poor data consistency in unstructured data management are solved, and efficient indexing, intelligent hierarchical storage and automated verification are realized, improving data storage efficiency and security.

CN120296224AInactive Publication Date: 2025-07-11SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510779813.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology has problems such as low retrieval efficiency, high storage cost, poor data consistency and insufficient scalability in unstructured data management, which is difficult to meet the needs of enterprises for high-performance and low-cost storage.

Method used

The unstructured data lake hierarchical storage and verification system is adopted, including the application layer, the data lake interface layer, the metadata management layer, the unstructured data storage layer and the verification service layer. Through unified access interface, dynamic routing, multi-dimensional retrieval, and automated verification mechanisms, the hot and cold layered storage and data consistency management are realized.

Benefits of technology

It realizes efficient indexing, intelligent layered storage and automated verification of unstructured data lakes, improves data storage efficiency, access performance and security reliability, and meets the needs of different users and business scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296224A_ABST
    Figure CN120296224A_ABST
Patent Text Reader

Abstract

The invention provides an unstructured data lake hierarchical storage and verification system and method, and relates to the technical field of data storage and verification. The system comprises an application layer, a data lake interface layer, a metadata management layer, an unstructured data storage layer and a verification service layer, the application layer realizes data access and operation by calling services provided by the data lake interface layer; the data lake interface layer provides a unified access interface to realize unified input and output of structured and unstructured data; the metadata management layer is used for storing metadata information of user files and supporting multi-dimensional retrieval; the unstructured data storage layer is used for storing user files transmitted through the data lake interface layer, supporting various storage media and storage schemes, and automatically allocating the storage media based on a preset rule; and the verification service layer provides a data verification service. According to the method, automatic scheduling, unified access and multi-dimensional verification management of cold and hot hierarchical storage in the unstructured data lake are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data storage and verification, and in particular to a hierarchical storage and verification system and method for unstructured data lakes. Background Art

[0002] With the deepening of digital transformation, the proportion of unstructured data (such as pictures, videos, logs, documents, etc.) in enterprise data lakes has been continuously rising, exceeding 80% of the total data volume. These data are diverse in form and huge in scale, bringing unprecedented management challenges to enterprises. Traditional data management technologies are mainly designed for structured data and cannot effectively cope with the complexity of unstructured data. For example, unstructured data lacks a unified indexing mechanism, resulting in low retrieval efficiency. It often relies on full-scale scanning or manual marking, which is not only time-consuming and laborious but also difficult to meet the requirements of real-time analysis. In addition, the storage cost issue of unstructured data has become increasingly prominent. Traditional solutions usually adopt a single storage medium (such as fully using high-performance NAS), resulting in a large amount of expensive resources being occupied by low-frequency accessed data, causing waste of storage costs.

[0003] In terms of data consistency, there are also obvious deficiencies in the existing technologies. Problems such as path failures and checksum mismatches are prone to occur between metadata and storage entities, and the traditional manual verification mechanism is inefficient and difficult to detect data corruption or loss in real time. Especially in a large-scale data lake environment, the consistency and integrity of data are difficult to guarantee, increasing the risks of business interruption and data leakage. In addition, with the explosive growth of data volume, the scalability of traditional file systems also faces severe challenges. A single storage architecture is difficult to support the elastic expansion of EB-level unstructured data and cannot meet the enterprise's needs for high-performance and low-cost storage.

[0004] In summary, the existing technologies have problems such as low retrieval efficiency, high storage cost, poor data consistency, and insufficient scalability in the management of unstructured data. There is an urgent need for an innovative solution that can organically combine structured data and unstructured data to achieve efficient indexing, intelligent hierarchical storage, and automated verification, thereby providing enterprises with a high-performance, low-cost, and strongly consistent data lake infrastructure. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a hierarchical storage and verification system and method for unstructured data lakes to achieve hierarchical storage and verification of unstructured data lakes in view of the above-mentioned deficiencies of the existing technologies.

[0006] To solve the above technical problems, the technical solutions adopted by the present invention are as follows: On the one hand, the present invention provides an unstructured data lake hierarchical storage and verification system, including an application layer, a data lake interface layer, a metadata management layer, an unstructured data storage layer, and a verification service layer; wherein: The application layer interacts with the data lake interface layer to achieve data access and operations by calling the services provided by the data lake interface layer; The data lake interface layer provides a unified access interface, shields the underlying storage differences upward, and is compatible with object storage and file systems downward to achieve the unified input and output of structured and unstructured data; The metadata management layer is used to store metadata information such as the ID, storage path, checksum, storage type, access frequency, and business tags of user files, and supports multi-dimensional retrieval by tags, attributes, and time ranges; The unstructured data storage layer is used to store user files transmitted through the data lake interface layer, supports multiple storage media and storage schemes, and automatically allocates storage media based on preset rules; The verification service layer provides data verification services to ensure the integrity and accuracy of data.

[0007] Furthermore, the application layer is responsible for hosting BI visualization tools and various business system APIs, and uniformly receiving and forwarding user data query, upload, download, and management requests.

[0008] Furthermore, the data lake interface layer provides a unified access interface by encapsulating RESTful API and NFS / bucket protocols. The NFS client mounts the remote directory to the user's local file system, enabling the user to access remote files as if operating local files; the bucket encapsulates the S3 protocol API to provide the read and write capabilities of object storage.

[0009] Furthermore, the metadata management layer includes a structured metadata database and a metadata index engine; the structured metadata database is used to store the metadata information of files, recording key information such as the ID, storage path, checksum, storage type, access frequency, and business tags of files; the metadata index engine is used to retrieve unstructured data and supports fast query functions by tags and file attributes.

[0010] Furthermore, the unstructured data storage layer includes an NFS hot storage cluster and object storage bucket cold storage, and automatically allocates storage media based on preset rules through a dynamic routing module.

[0011] Furthermore, the verification service layer consists of a timed verification engine and an event-driven verification engine. The timed verification engine scans the full-scale storage layer at regular intervals and compares the records in the metadata database to verify the existence, checksum, and storage compliance of files, ensuring file integrity and storage standardization. The event-driven verification engine is triggered in real time based on NFS inotify events or object storage event notifications to promptly detect problems during file changes. The dual engines cooperate to ensure the existence, integrity, and compliance of files.

[0012] On the other hand, the present invention also provides a method for hierarchical storage and verification of unstructured data lakes, including three processes: data writing, data reading, and data verification. Among them, the data writing process includes: The user initiates a file writing request through the application layer, and the data lake interface layer receives the file sent by the user through the unified access interface. The metadata management layer extracts the file meta-information, and the verification service layer calculates and records the checksum to ensure the validity of the uploaded file. The dynamic routing module determines the storage target according to the file size and historical access frequency rules: if the file size is greater than the set value or the access frequency is less than the set frequency, it is written into the object storage bucket; otherwise, it is mounted to the NFS hot storage. After the storage is completed, update each field of the structured metadata database in the metadata management layer and synchronously push it to the metadata indexing engine. The data reading process includes: The business system initiates a file reading request, and the reading request is transferred to the data lake interface layer through the application layer. The metadata management layer retrieves and returns the storage medium and path of the target file according to the reading request of the data lake interface layer. The data lake interface layer automatically selects NFS mounting or object storage SDK call, reads the file from the corresponding storage layer and returns it to the business system. The data verification process includes: Timed verification: Scan NFS and object storage at the set period, and compare the actual checksum of each file with the records in the metadata database in full scale. If the two do not match, trigger an alarm and enter the automatic repair or re-storage strategy. Event-driven verification: When a file is created, modified, or deleted on NFS or object storage, the verification process is triggered in real time through inotify or storage event notifications to quickly locate and handle exceptions.

[0013] The beneficial effects produced by adopting the above technical solutions are as follows: A hierarchical storage and verification system and method for an unstructured data lake provided by the present invention realizes the automated scheduling, unified access, and multi-dimensional verification management of cold and hot hierarchical storage in the unstructured data lake, greatly improving the storage efficiency, access performance, and security reliability of data, and can meet the requirements of different users and business scenarios. Description of the Drawings

[0014] Figure 1 It is a structural block diagram of a hierarchical storage and verification system for an unstructured data lake provided by an embodiment of the present invention; Figure 2 It is a data writing flow chart provided by an embodiment of the present invention; Figure 3 It is a dynamic routing decision flow chart provided by an embodiment of the present invention. Detailed Embodiments

[0015] The following combines the drawings and embodiments to further describe in detail the specific embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0016] In the prior art, many enterprises use a single storage architecture (such as fully using high-performance NAS or fully using object storage) to manage unstructured data. Although this solution simplifies storage management, its disadvantages are obvious.

[0017] Serious resource waste: Fully using high-performance NAS will cause data with low-frequency access (such as historical archive files) to occupy a large amount of expensive resources, driving up storage costs; while fully using object storage may not be able to meet the low-latency requirements of high-frequency access data, affecting business performance.

[0018] Unable to balance performance and cost: A single storage strategy cannot dynamically adjust the storage location according to the data access pattern, resulting in unreasonable resource allocation.

[0019] At the same time, many enterprises use independent storage systems (such as the separation of NFS and object storage) to meet different business needs, but there are the following disadvantages:

[0020] High management complexity: Different storage systems use different interfaces and protocols, and the business system needs to interface separately, increasing the complexity of development and operation and maintenance.

[0021] Traditional methods rely on a manual verification mechanism to regularly check the consistency of metadata and storage entities to ensure data integrity, but there are the following disadvantages:

[0022] Low efficiency and easy to miss: Manual inspection cannot detect data corruption or loss in real time, increasing the risk of business interruption and data leakage.

[0023] Some existing solutions add tags or attribute information to unstructured data through static metadata indexing technology to improve retrieval efficiency, but there are the following disadvantages: Lack of dynamic routing ability: Static indexing cannot automatically optimize the storage location according to file attributes (such as size, access frequency), resulting in high-frequency access data that may be stored on low-performance media, while low-frequency access data occupies high-performance resources.

[0024] Limited retrieval efficiency: In the face of a large amount of unstructured data, static indexing still relies on full-scale scanning or manual marking, making it difficult to meet the needs of real-time analysis.

[0025] It can be seen that the existing technologies have the following deficiencies in the management of unstructured data: Low retrieval efficiency of unstructured data: Unstructured data lacks a unified indexing mechanism and relies on full-scale scanning or manual marking, resulting in time-consuming and laborious retrieval and making it difficult to meet the needs of real-time analysis. For example, retrieving historical images in a medical imaging system may take several minutes or even longer.

[0026] Waste of storage resources: Traditional single storage media (such as full use of high-performance NAS) causes low-frequency access data (such as historical archived files) to occupy expensive resources, driving up storage costs. For example, video streaming companies store a large number of low-frequency access historical video files, resulting in waste of resources.

[0027] Insufficient data consistency and security: Problems such as path failure and checksum mismatch are likely to occur between metadata and storage entities, and the manual verification mechanism is inefficient, making it difficult to detect data corruption or loss in real time. For example, the corruption or loss of transaction log files in the financial industry may trigger compliance risks.

[0028] Complex hybrid storage management: The interfaces and protocols of multiple storage media (such as NFS, object storage) are different, and business systems need to interface separately, with high development complexity and difficulty in achieving unified governance. For example, when manufacturing industries manage real-time production logs and historical production data at the same time, the separate storage systems lead to complex management.

[0029] In this embodiment, a hierarchical storage and verification system for an unstructured data lake, as Figure 1 shown, includes an application layer, a data lake interface layer, a metadata management layer, an unstructured data storage layer, and a verification service layer; among them: The application layer interacts with the data lake interface layer and realizes data access and operations by calling the services provided by the data lake interface layer; In this embodiment, the application layer provides a user interface (UI) or an application programming interface (API), allowing users or applications to access and operate on the data stored in the data lake; it supports basic operations such as data query, retrieval, upload, download, and deletion. It provides advanced functions such as data analysis, processing, and visualization to meet the needs of different users and business scenarios.

[0030] In this embodiment, the application layer is responsible for hosting BI visualization tools and various business system APIs, and uniformly receiving and forwarding users' query, upload, download, and management requests.

[0031] The data lake interface layer provides a unified access interface, shielding the underlying storage differences upwards and being compatible with object storage and file systems downwards, realizing the unified input and output of structured and unstructured data; In this embodiment, the data lake interface layer provides a unified data access interface, encapsulating the complexity of the underlying storage system, so that the upper-layer applications do not need to care about the underlying storage details. It supports multiple data formats and protocols, such as HTTP, RESTful API, etc., in order to integrate with different types of applications, realize data routing, load balancing, and cache management, and improve the performance and reliability of data access. The data lake interface layer interacts with the application layer, receiving and processing data access requests from the upper-layer applications. It interacts with the metadata management layer to obtain metadata information to support the precise positioning and access of data. It interacts with the unstructured data storage layer to perform data read and write operations.

[0032] In this embodiment, the data lake interface layer provides a unified access interface by encapsulating RESTful API and NFS / bucket protocols (such as S3). The NFS client mounts the remote directory to the user's local file system. For example, a certain directory is mounted to the user's local / data_lake / hot path, enabling the user to access the remote file as if operating on a local file; the bucket SDK encapsulates the S3 protocol API, providing the read and write capabilities of object storage, facilitating the operation of the object storage service without directly interacting with the complex S3 protocol, thus simplifying the development process.

[0033] The metadata management layer is used to store metadata information such as the ID, storage path, checksum (MD5 / SHA256), storage type, access frequency, and business tags of user files, and supports multi-dimensional retrieval by tags, attributes, and time range; In this embodiment, the metadata management layer manages the metadata of unstructured data, including information such as the attributes, location, size, and access permissions of the data, provides query, update, and deletion services for the metadata, supports the rapid positioning and access of the data, realizes the indexing and caching of the metadata, and improves the performance of metadata access. The metadata management layer interacts with the data lake interface layer to respond to the query and update requests for the metadata from the interface layer. It interacts with the unstructured data storage layer to obtain and update the metadata information in the storage layer. It interacts with the verification service layer to provide metadata for the data verification process.

[0034] In this embodiment, the metadata management layer includes a structured metadata database (using MySQL) and a metadata index engine (using Elasticsearch); the structured metadata database is used to store the metadata information of files, recording key information such as the ID, storage path, checksum (such as MD5 or SHA256), storage type (such as NFS or storage bucket), access frequency, and business tags of the files. These information provide basic support for the management and positioning of files; the metadata index engine focuses on the rapid retrieval of unstructured data. By supporting the rapid query function according to tags and file attributes, users can efficiently find the required content from a large amount of data, thus improving the retrieval efficiency and data utilization value of the entire data management system.

[0035] The unstructured data storage layer is used to store user files transmitted through the data lake interface layer, supports multiple storage media and storage schemes, and automatically allocates storage media based on preset rules; In this embodiment, the unstructured data storage layer is responsible for the actual storage of unstructured data, supports multiple storage media and storage schemes, such as distributed file systems, object storage, etc. It provides basic storage services such as reading, writing, copying, migrating, and deleting data. It realizes the distributed storage and fault tolerance mechanism of the data to ensure the high availability and reliability of the data. It interacts with the data lake interface layer to respond to the read and write requests for the data from the interface layer. The unstructured data storage layer interacts with the metadata management layer to obtain and update the metadata information in the storage layer. It interacts with the verification service layer to provide data blocks for the data verification process.

[0036] In this embodiment, the unstructured data storage layer includes an NFS hot storage cluster (low latency, high IOPS, suitable for small file and high-frequency access scenarios) and an object storage bucket cold storage (large capacity, low cost, suitable for large file and low-frequency access scenarios), and automatically allocates storage media based on preset rules (such as file size > 100MB or access frequency < 1 time / week) through a dynamic routing module (an automated program component), as Figure 2 shown.

[0037] The verification service layer provides data verification services to ensure data integrity and accuracy, and implements functions of regular verification, online verification, and repair of data, so as to timely detect and repair data corruption problems. It supports a variety of verification algorithms and strategies to meet the requirements of different data and application scenarios. The verification service layer interacts with the metadata management layer to obtain metadata information to support the precise positioning and verification of data. It also interacts with the unstructured data storage layer to obtain data blocks for verification operations. During the verification process, if data corruption problems are found, the upper-layer application or administrator can be notified through the data lake interface layer for repair.

[0038] In this embodiment, the verification service layer consists of a timed verification engine and an event-driven verification engine. The timed verification engine scans the full-scale storage layer on a daily / weekly / monthly cycle and compares the records in the metadata database to verify the existence, checksum, and storage compliance of files, ensuring file integrity and storage standardization; the event-driven verification engine is triggered in real time based on NFS inotify events or object storage event notifications to timely detect possible problems during the file change process. The dual mechanisms cooperate to ensure the existence, integrity, and compliance of files.

[0039] In this embodiment, a method for hierarchical storage and verification of unstructured data lakes includes three processes: data writing, data reading, and data verification; among them, the data writing process is as Figure 3 shown, including: Users initiate a file writing request through the application layer, and the data lake interface layer receives the file sent by the user through a unified access interface; The metadata management layer extracts the file metadata information, and the verification service layer calculates and records the checksum to ensure the validity of the uploaded file; The dynamic routing module determines the storage target according to the file size and historical access frequency rules: if the file > 100 MB or the access frequency < 1 time / week, it is written to the object storage bucket, otherwise it is mounted to the NFS hot storage; In this embodiment, after the user initiates a file writing request, first judge whether the file size exceeds 100MB. If it exceeds, it is stored in the object storage bucket; if it does not exceed, further judge whether the file access frequency is less than 1 time / week. If it is less, it is stored in the object storage bucket, otherwise it is stored in the NAS. No matter which medium it is stored in, the storage location field in the metadata database needs to be updated.

[0040] After the storage is completed, update the storage type, path, and other fields in the structured metadata database MySQL in the metadata management layer, and synchronously push them to the metadata index engine Elasticsearch index.

[0041] The data reading process includes: The business system initiates a file reading request, and the reading request is transferred to the data lake interface layer through the application layer; The metadata management layer retrieves and returns the storage medium and path of the target file according to the reading request of the data lake interface layer; The data lake interface layer automatically selects NFS mounting or object storage SDK call, reads the file from the corresponding storage layer and returns it to the business system.

[0042] The data verification process includes: Timed verification: The system scans NFS and object storage at the set period, compares the actual checksum of each file with the records in the metadata database in full volume. If the two do not match, an alarm is triggered and an automatic repair or re - storage strategy is entered; Event - driven verification: When a file is created, modified or deleted on NFS or object storage, the verification process is triggered in real - time through inotify or storage event notification to quickly locate and handle exceptions.

[0043] Finally, it should be noted that: The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: They can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope defined by the claims of the present invention.

Claims

1. A hierarchical storage and verification system for unstructured data lakes, characterized in that: It includes application layer, data lake interface layer, metadata management layer, unstructured data storage layer and verification service layer; among which: The application layer interacts with the data lake interface layer to access and operate data by calling services provided by the data lake interface layer; The data lake interface layer provides a unified access interface, shields the underlying storage differences upward, is compatible with object storage and file systems downward, and realizes unified input and output of structured and unstructured data; The metadata management layer is used to store metadata information such as user file ID, storage path, checksum, storage type, access frequency, and business tags, and supports multi-dimensional retrieval by tag, attribute, and time range; The unstructured data storage layer is used to store user files transmitted through the data lake interface layer, supports multiple storage media and storage solutions, and automatically allocates storage media based on preset rules; The verification service layer provides data verification services to ensure the integrity and accuracy of data.

2. The hierarchical storage and verification system for unstructured data lake according to claim 1, wherein: The application layer is responsible for carrying BI visualization tools and various business system APIs, and uniformly receiving and forwarding users' data query, upload, download and management requests.

3. The hierarchical storage and verification system for unstructured data lake according to claim 1, wherein: The data lake interface layer provides a unified access interface by encapsulating the RESTful API and the NFS / bucket protocol. The NFS client mounts the remote directory to the user's local file system, allowing the user to access the remote file like a local file. The bucket encapsulates the S3 protocol API to provide object storage read and write capabilities.

4. A hierarchical storage and verification system for unstructured data lakes according to claim 1, wherein: The metadata management layer includes a structured metadata database and a metadata index engine; The structured metadata database is used to store metadata information of files, and record key information such as file ID, storage path, checksum, storage type, access frequency, and service tag; The metadata indexing engine is used to retrieve unstructured data and supports fast query functions based on tags and file attributes.

5. A hierarchical storage and verification system for unstructured data lakes according to claim 1, characterized in that: The unstructured data storage layer includes NFS hot storage clusters and object storage bucket cold storage, and automatically allocates storage media based on preset rules through a dynamic routing module.

6. The hierarchical storage and verification system for unstructured data lake according to claim 1, characterized in that: The verification service layer consists of a timed verification engine and an event-driven verification engine. The timed verification engine periodically scans the entire storage layer and compares the metadata database records to verify the existence, checksum, and storage compliance of the file to ensure file integrity and storage compliance. The event-driven verification engine is triggered in real time based on NFS inotify events or object storage event notifications to promptly detect problems that arise during file changes. The dual engines work together to ensure the existence, integrity, and compliance of files.

7. A method for hierarchical storage and verification of unstructured data lakes, implemented based on the unstructured data lake hierarchical storage and verification system described in claim 1, characterized in that: It includes three processes: data writing, data reading and data verification; The data writing process includes: The user initiates a file write request through the application layer, and the data lake interface layer receives the file sent by the user through the unified access interface; The metadata management layer extracts file metadata and the verification service layer calculates and records the checksum to ensure the validity of the uploaded file; The dynamic routing module determines the storage destination according to the rules of file size and historical access frequency: if the file size is greater than the set value or the access frequency is less than the set frequency, it is written to the object storage bucket; otherwise, it is mounted to the NFS hot storage; After storage is completed, update each field of the structured metadata database in the metadata management layer and synchronously push it to the metadata index engine; The data reading process includes: The business system initiates a file reading request, and the reading request is transferred to the data lake interface layer through the application layer; The metadata management layer retrieves and returns the storage medium and path of the target file according to the reading request of the data lake interface layer; The data lake interface layer automatically selects NFS mounting or object storage SDK call, reads the file from the corresponding storage layer and returns it to the business system; The data verification process includes: Timed verification: Scan NFS and object storage at the set period, compare the actual checksum of each file with the record in the metadata database in full quantity. If the two do not match, trigger an alarm and enter the automatic repair or re-storage strategy; Event-driven verification: When a file is created, modified or deleted on NFS or object storage, the verification process is triggered in real time through inotify or storage event notification to quickly locate and handle exceptions.

Citation Information

Patent Citations

  • Data frame system

    CN106933555A

  • Distributed storage management system for library mass data

    CN110019521A

  • Meteorological data file directory service generation method and system based on virtual data lake

    CN115484276A

  • Online cloud storage platform based on hadoop

    CN206932239U

Cited By

  • Cross-application Excel data import processing method and system based on cloud storage

    CN121012829A

  • Database design system and method of fan integrated design platform

    CN121326882A