Security mass data hierarchical distributed storage management system and method
Through the security massive data hierarchical distributed storage management system, the problems of inaccurate classification and insufficient data access path optimization are solved, efficient data storage and access are achieved, and the system is scalable and stable.
Patent Information
- Application Number
- CN202510515428.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The lack of dual verification mechanism of traditional security systems leads to inaccurate classification, insufficient optimization of data access paths, low overall efficiency, and difficult to provide a clear classification framework, affecting the efficiency and quality of data processing and retrieval.
A large security data hierarchical distributed storage management system is adopted, including data acquisition module, type analysis module, framework construction module and data layout module. Through the preparatory type-experiment case dual verification model, a directional acyclic graph structure is built, and the data storage layout is optimized using the cosine similarity algorithm to reasonably allocate storage tasks.
It achieves the improvement of classification accuracy, eliminates loop dependence, optimizes data access paths, improves data storage and access efficiency, ensures the scalability and stability of the system, and rationally allocates storage resources, avoids the problems of excessive use or imbalance of a single storage node.
Smart Images

Figure CN120335728A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of management technologies, and particularly to a hierarchical distributed storage management system and method for a large amount of security data. Background Art
[0002] In today's society, security assurance is of utmost importance. With the rapid development of technologies, security data has become the core element of the entire security system. The traditional security system has accumulated a large amount of data during long-term operation, covering aspects such as video surveillance recordings, personnel and vehicle access records, and equipment operation status. The application scenarios of the security system have also become increasingly rich, involving multiple fields such as national defense, public security, commerce, and civil use;
[0003] Due to the lack of a dual-verification mechanism, the traditional system has inaccurate classification, insufficient optimization of data access paths, and low overall efficiency. It is difficult to provide a clear classification framework, which affects the efficiency and quality of data processing and retrieval. Therefore, we propose a hierarchical distributed storage management system and method for a large amount of security data. Summary of the Invention
[0004] The purpose of the present invention is to provide a hierarchical distributed storage management system and method for a large amount of security data.
[0005] To achieve the above purpose, the present invention provides the following technical solution: A hierarchical distributed storage management system for a large amount of security data, the distributed storage management system includes;
[0006] A data acquisition module, which is used to acquire different security data, obtain security cases, and at the same time obtain the permission to use cloud storage;
[0007] A type analysis module, which analyzes the source types, data property types, and data usage types of security data to obtain a distribution type library. At the same time, a type editing unit is constructed. The user has the permission to edit the distribution type library through the type editing unit. When the user edits the distribution type library, a preliminary type will be generated according to the user's editing information, and then case extraction will be performed from the security cases to obtain experimental cases. It is analyzed whether the data in the experimental cases can be classified into the preliminary type. When it cannot be classified into the preliminary type, the preliminary type information will be deleted. When it can be classified into the preliminary type, the preliminary type will be recorded in the distribution type library;
[0008] A framework construction module, which establishes storage nodes in the cloud storage according to the distribution type library, and then constructs a directed acyclic graph structure, where the vertices represent storage nodes, the directed edges represent the data flow logic, and there is no cyclic dependence path in the graph, that is, there is no path that starts from a certain vertex, passes through a series of vertices along the directed edge, and then returns to the starting vertex;
[0009] The data analysis module retrieves the distribution types recorded in the distribution type library, analyzes the data related to the distribution types in the security data, and intercepts it to obtain sub-security data;
[0010] The data layout module analyzes the correlation between sub-security data through the cosine similarity algorithm according to the data flow and processing logic, and allocates the data with a correlation ≥ 0.8 to adjacent storage nodes.
[0011] As a further solution of the present invention: after obtaining the distribution type library in the type analysis module, extract the cases in the security cases with a similarity ≥ 90% to the distribution types of the security data, respectively extract the values therein, analyze the same information between the security data and different security cases, calculate the overlap index between the security data and different security cases according to the same information, and perform sorting processing on the overlap index, wherein the sorting method is descending order.
[0012] As a further solution of the present invention: when calculating the similarity index in the type analysis module, let the value in the security data be A Z Let the value regarding the distribution type in different security cases be S Z Let the overlap index be X X Let the number of extracted values be L:
[0013]
[0014] Calculate the overlap index through the above formula.
[0015] As a further solution of the present invention: when obtaining the experimental cases in the type analysis module, analyze the number of security cases, and extract max(10, 10%P) security cases as experimental cases, where P is the total number of cases.
[0016] As a further solution of the present invention: when reasonably allocating sub-security data to different storage nodes in the data layout module, analyze the storage capacity and read-write performance of different storage nodes, calculate the priority storage index of different storage nodes according to the storage capacity and read-write performance of the storage nodes, perform sorting processing on the priority storage index, and grant different priority usage permissions to different storage nodes in the order of the priority storage index.
[0017] As a further solution of the present invention: when calculating the priority storage index in the data layout module, let the storage capacity value of different storage nodes be C Z Let the read-write performance value of different storage nodes be D X Let the priority storage index of different storage nodes be Y X :
[0018] Y X= D X ·C Z
[0019] Calculate the priority storage indices of different storage nodes through the above formula.
[0020] As a further solution of the present invention: when different storage nodes in the data layout module are granted different priority usage permissions, dynamic capacity thresholds will be set for the storage capacities of different storage nodes. When the used capacity of a single storage node reaches 80% of its total capacity, writing data to this node will be suspended. If more than 75% of the total number of nodes reach the capacity threshold, the system will automatically trigger an expansion mechanism.
[0021] As a further solution of the present invention: when the data layout module distributes sub-security data to different storage nodes, it dynamically calculates the number of shards D according to the size of the sub-security data and the minimum shard capacity threshold of the storage node, where D ≥ 2 and D ≤ the number of currently available storage nodes, and then writes different shards to different storage nodes simultaneously to make the writing operations parallel, and the user has the permission to edit the number of shards into which the sub-security data is divided.
[0022] In addition, a hierarchical distributed storage management method for security mass data is also provided, including the following steps:
[0023] S100. Collect different security data, obtain the permission to use cloud storage, analyze the source types, data property types, and data usage types of the security data, and obtain a distribution type library;
[0024] S200. Generate a preliminary type according to the user's editing information, then extract cases from security cases to obtain experimental cases, and analyze whether the data in the experimental cases can be classified into the preliminary type. When it cannot be classified into the preliminary type, the preliminary type information will be deleted. When it can be classified into the preliminary type, the preliminary type will be recorded in the distribution type library;
[0025] S300. Construct a directed acyclic graph structure, where the directed acyclic graph structure means that the edges in the graph have directions, each edge points from one vertex to another vertex, and at the same time, there are no cycles in the graph, that is, there is no path that starts from a certain vertex, passes through a series of vertices along the directed edges, and then returns to the starting vertex;
[0026] S400. Analyze the data about the distribution type in the security data and intercept it to obtain sub-security data;
[0027] S500. Analyze the correlation between sub-security data through the cosine similarity algorithm, and allocate the data with a correlation ≥ 0.8 to adjacent storage nodes.
[0028] By adopting the above technical solution, compared with the prior art, the beneficial effects of the present invention are:
[0029] 1. The present invention improves classification accuracy, eliminates loop dependencies, and improves data access path optimization efficiency by preparing a category-experimental case dual verification model. The category analysis module helps the system to accurately classify massive data, provides a clear framework for subsequent data processing and retrieval, and ensures the accuracy and effectiveness of the distribution category library. The framework construction module provides a clear hierarchy and data flow for data storage, improves the efficiency of data storage and access, and makes the system have good scalability. The data layout module can optimize the data storage layout and improve the data access speed and system performance.
[0030] 2. The present invention provides objective historical data support for subsequent preliminary category verification through the category analysis module, effectively processes the similarity calculation of numerical data, avoids evaluation bias caused by dimensional differences, helps the system to classify and analyze new data based on historical experience, reduces the workload of case extraction and analysis, improves system operation efficiency, prevents the problem of being unable to accurately verify preliminary categories due to too few extracted cases, and ensures the scientific nature of preliminary category verification;
[0031] 3. The present invention can reasonably allocate storage tasks through the data layout module, give priority to the use of storage nodes with better performance, improve the overall efficiency of data storage and access, optimize the utilization of storage resources, enable the system to better cope with the storage needs of massive data, make the evaluation of storage nodes more scientific and accurate, help the system to more reasonably allocate data storage tasks, effectively avoid the problem of excessive use or unbalanced storage of a single storage node, improve the utilization rate of storage resources and the stability of the system, and can be adjusted according to actual needs and the situation of the storage nodes to achieve the best storage effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 A schematic diagram of a system flow in an embodiment of the present invention;
[0033] Figure 2 Schematic diagram of the method steps in an embodiment of the present invention. DETAILED DESCRIPTION
[0034] The specific embodiments of the present invention will be further described below in conjunction with the accompanying drawings. It should be noted that the description of these embodiments is used to help understand the present invention, but does not constitute a limitation of the present invention.
[0035] In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0036] Please refer to the attachedFigure 1 - Attachment Figure 2 For a hierarchical distributed storage management system for a large amount of security data in the present invention, the distributed storage management system includes;
[0037] A data acquisition module, configured to acquire different security data, obtain security cases, and at the same time obtain the permission to use cloud storage;
[0038] A type analysis module, which analyzes the source types, data property types, and data usage types of security data to obtain a distribution type library. At the same time, a type editing unit is constructed. Users have the permission to edit the distribution type library through the type editing unit. When the user edits the distribution type library, preliminary types will be generated according to the user's editing information, and then case extraction will be performed from the security cases to obtain experimental cases. Analyze whether the data in the experimental cases can be classified into the preliminary types. When it cannot be classified into the preliminary types, the preliminary type information will be deleted. When it can be classified into the preliminary types, the preliminary types will be recorded in the distribution type library;
[0039] A framework construction module, which establishes storage nodes in the cloud storage according to the distribution type library, and then constructs a directed acyclic graph structure, where the vertices represent storage nodes, the directed edges represent the data flow logic, and there is no cyclic dependence path in the graph, that is, there is no path that starts from a certain vertex, passes through a series of vertices along the directed edges, and then returns to the starting vertex;
[0040] A data analysis module, which obtains the distribution types recorded in the distribution type library, analyzes the data in the security data regarding the distribution types, and intercepts it to obtain sub-security data;
[0041] A data layout module, and according to the data flow and processing logic, analyzes the correlation between sub-security data through the cosine similarity algorithm, and allocates the data with a correlation ≥ 0.8 to adjacent storage nodes;
[0042] For example, data that is often accessed together is allocated to adjacent nodes to reduce the search time when reading data.
[0043] In an embodiment of the present invention: After the distribution type library in the type analysis module is obtained, cases with a similarity ≥ 90% to the distribution types of the security data in the security cases are extracted, and the numerical values therein are respectively extracted. Analyze the same information between the security data and different security cases, and calculate the overlap index between the security data and different security cases according to the same information, and perform sorting processing on the overlap index, where the sorting processing method is descending order.
[0044] In an embodiment of the present invention: When calculating the similarity index in the type analysis module, let the numerical value in the security data be A Z , let the numerical value regarding the distribution type in different security cases be SZ , let the overlap index be X X , let the number of extracted values be L:
[0045]
[0046] Calculate the overlap index through the above formula.
[0047] In one embodiment of the present invention: when the experimental cases in the type analysis module are obtained, the number of security cases will be analyzed, and max(10, 10%P) security cases will be extracted as experimental cases, where P is the total number of cases.
[0048] In one embodiment of the present invention: when the data layout module reasonably distributes the sub-security data to different storage nodes, it will analyze the storage capacity and read-write performance of different storage nodes, calculate the priority storage index of different storage nodes according to the storage capacity and read-write performance of the storage nodes, and sort the priority storage indexes, and grant different priority usage permissions to different storage nodes in the order of the priority storage indexes.
[0049] In one embodiment of the present invention: when calculating the priority storage index in the data layout module, let the storage capacity value of different storage nodes be C Z , let the read-write performance value of different storage nodes be D X , let the priority storage index of different storage nodes be Y X :
[0050] Y X = D X ·C Z
[0051] Calculate the priority storage index of different storage nodes through the above formula.
[0052] In one embodiment of the present invention: when different storage nodes are granted different priority usage permissions in the data layout module, a dynamic capacity threshold will be set for the storage capacity of different storage nodes. When the used capacity of a single storage node reaches 80% of its total capacity, writing data to this node will be suspended. If more than 75% of the total number of nodes reach the capacity threshold, the system will automatically trigger the expansion mechanism.
[0053] In one embodiment of the present invention: when the data layout module distributes the sub-security data to different storage nodes, the number of shards D is dynamically calculated according to the size of the sub-security data and the minimum shard capacity threshold of the storage nodes, where D≥2 and D≤the number of current available storage nodes, and then different shards are written to different storage nodes simultaneously to make the writing operations parallel, and the user has the permission to edit the number of shards into which the sub-security data is divided.
[0054] Example 1. Please refer to the appendix Figure 1 - appendix Figure 2 to collect different security data, obtain the permission to use cloud storage, analyze the source types, data property types, and data usage types of the security data to obtain a distribution type library, generate a preliminary type based on the user's editing information, then extract cases from security cases to obtain experimental cases, analyze whether the data in the experimental cases can be classified into the preliminary type. When it cannot be classified into the preliminary type, the preliminary type information will be deleted. When it can be classified into the preliminary type, the preliminary type will be recorded in the distribution type library, construct a directed acyclic graph structure, where the vertices represent storage nodes, the directed edges represent the data flow logic, and there is no cyclic dependency path in the graph, that is, there is no path that starts from a certain vertex, passes through a series of vertices along the directed edges, and then returns to the starting vertex. Analyze the data on the distribution type in the security data and intercept it to obtain sub-security data, analyze the relevance between the sub-security data through the cosine similarity algorithm, and allocate the data with a relevance ≥ 0.8 to adjacent storage nodes.
[0055] Example 2. Please refer to the appendix Figure 1 - appendix Figure 2 to analyze the source types, data property types, and data usage types of the security data to obtain a distribution type library, generate a preliminary type based on the user's editing information, then extract cases from security cases to obtain experimental cases, analyze whether the data in the experimental cases can be classified into the preliminary type. When it cannot be classified into the preliminary type, the preliminary type information will be deleted. When it can be classified into the preliminary type, the preliminary type will be recorded in the distribution type library. Extract the cases in the security cases with a similarity ≥ 90% to the distribution type of the security data, analyze the same information between the security data and different security cases, calculate the overlap index between the security data and different security cases based on the same information, and perform a sorting process on the overlap index. Analyze the number of security cases and extract max(10, 10%P) security cases as experimental cases, where P is the total number of cases.
[0056] Example 3. Please refer to the appendix Figure 1 - appendix Figure 2, determine the directed edge relationship between nodes. According to the real-time load rate, remaining capacity, and read / write performance scores of the storage nodes, allocate the sub-security data to different storage nodes according to weights. Analyze the storage capacity and read / write performance of different storage nodes, calculate the priority storage index of different storage nodes based on the storage capacity and read / write performance of the storage nodes, sort the priority storage index, and grant different priority usage permissions to different storage nodes in the order of the priority storage index. Set a dynamic capacity threshold for the storage capacity of different storage nodes. When the used capacity of a single storage node reaches 80% of its total capacity, suspend writing data to this node. If more than 75% of the total number of nodes reach the capacity threshold, the system automatically triggers the capacity expansion mechanism and migrates and allocates the data of high-load nodes. Dynamically calculate the number of shards D according to the size of the sub-security data and the minimum shard capacity threshold of the storage nodes, where D≥2 and D≤the current number of available storage nodes, and then write different shards to different storage nodes simultaneously to make the write operation parallel.
[0057] Specifically, in the data layout module, set the capacity limit boundary of the storage nodes (for example, each node can store at most 20GB). When the number of storage nodes reaching the capacity limit boundary exceeds 75% of the total number of nodes, the system automatically increases the capacity limit boundary value of all nodes and migrates and allocates the data of high-load nodes.
[0058] Working principle:
[0059] First, the source types, data nature types and data usage types of security data are analyzed to obtain a distribution category library. Preliminary categories are generated according to the user's editing information. Then, cases are extracted from security cases to obtain experimental cases. It is analyzed whether the data in the experimental cases can be summarized as preliminary categories. When they cannot be summarized as preliminary categories, the preliminary category information will be deleted. When they can be summarized as preliminary categories, the preliminary categories will be recorded in the distribution category library. The cases with a distribution category similarity of ≥90% with the security data in the security cases are extracted. The same information between security data and different security cases is analyzed, and the overlapping index between security data and different security cases is calculated based on the same information. The overlapping index is sorted and processed. The number of security cases is analyzed, and max (10, 10% P) security cases are extracted as experimental cases, where P is the total number of cases. The directed edge relationship between nodes is judged, and the real-time load rate, remaining capacity and read-write performance of the storage node are used to determine the relationship. It can score, distribute sub-security data to different storage nodes according to weights, analyze the storage capacity and read-write performance of different storage nodes, and calculate the priority storage index of different storage nodes according to the storage capacity and read-write performance of the storage nodes, and sort the priority storage index. Different storage nodes are granted different priority usage permissions in order of the priority storage index, and dynamic capacity thresholds are set for the storage capacity of different storage nodes. When the used capacity of a single storage node reaches 80% of its total capacity, writing data to the node is suspended. If more than 75% of the total number of nodes reaches the capacity threshold, the system automatically triggers the capacity expansion mechanism and migrates and allocates the data of high-load nodes. The number of shards D is dynamically calculated according to the size of the sub-security data and the minimum shard capacity threshold of the storage node, where D≥2 and D≤the number of currently available storage nodes. Then, different shards are written to different storage nodes at the same time, so that the write operation is carried out in parallel. At this point, the entire workflow ends.
[0060] Although the present invention is disclosed as above in terms of preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art may make possible changes and modifications without departing from the spirit and scope of the present invention. Therefore, any modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention shall fall within the scope of protection defined by the claims of the present invention.
Claims
1. A hierarchical distributed storage management system for a large amount of security data, characterized in that: The distributed storage management system includes; A data acquisition module, which is used to acquire different security data, obtain security cases, and at the same time obtain the permission to use cloud storage; A type analysis module, which analyzes the source types, data property types, and data usage types of security data to obtain a distribution type library. At the same time, a type editing unit is constructed. Users have the permission to edit the distribution type library through the type editing unit. When the user edits the distribution type library, a preliminary type will be generated according to the user's editing information, and then case extraction will be performed from the security cases to obtain experimental cases. It analyzes whether the data in the experimental cases can be summarized into the preliminary type. When it cannot be summarized into the preliminary type, the preliminary type information will be deleted. When it can be summarized into the preliminary type, the preliminary type will be recorded in the distribution type library; A framework construction module, which establishes storage nodes in cloud storage according to the distribution type library, and then constructs a directed acyclic graph structure, where the vertices represent storage nodes, the directed edges represent the data flow logic, and there is no cyclic dependency path in the graph, that is, there is no path that starts from a certain vertex, passes through a series of vertices along the directed edge, and then returns to the starting vertex; A data analysis module, which obtains the distribution types recorded in the distribution type library, analyzes the data in the security data regarding the distribution types, and intercepts it to obtain sub-security data; A data layout module, and according to the data flow and processing logic, analyzes the correlation between sub-security data through the cosine similarity algorithm, and allocates the data with a correlation ≥ 0.8 to adjacent storage nodes.
2. The hierarchical distributed storage management system for security mass data according to claim 1, wherein: After the distribution type library in the type analysis module is obtained, cases with a similarity ≥ 90% to the distribution types of security data in the security cases are extracted, and the values therein are respectively extracted. The same information between the security data and different security cases is analyzed, and the overlap index between the security data and different security cases is calculated according to the same information, and the overlap index is sorted, and the sorting method is in descending order.
3. The security mass data hierarchical distributed storage management system according to claim 2, characterized in that: When calculating the similarity index in the type analysis module, let the value in the security data be A Z , let the value regarding the distribution type in different security cases be S Z , let the overlap index be X X , let the number of extracted values be L: The overlap index is calculated through the above formula.
4. The hierarchical distributed storage management system for security mass data according to claim 3, characterized in that: When the experimental cases in the type analysis module are obtained, it will analyze the number of security cases and extract max(10, 10%P) security cases as experimental cases, where P is the total number of cases.
5. An anti-theft security mass data hierarchical distributed storage management system according to claim 1, characterized in that: When the data layout module reasonably distributes sub-security data to different storage nodes, it will analyze the storage capacity and read-write performance of different storage nodes, calculate the priority storage index of different storage nodes according to the storage capacity and read-write performance of the storage nodes, and sort the priority storage index. Different priority usage permissions are granted to different storage nodes in the order of the priority storage index.
6. The hierarchical distributed storage management system for security mass data according to claim 5, wherein: When preferentially storing exponents in the data layout module during calculation, let the storage capacity values of different storage nodes be C Z , let the read / write performance values of different storage nodes be D X , let the preferential storage exponents of different storage nodes be Y X : Y X = D X · C Z The priority storage index of different storage nodes is calculated through the above formula.
7. An anti-theft security mass data hierarchical distributed storage management system according to claim 6, characterized in that: When different storage nodes in the data layout module are granted different priority usage permissions, a dynamic capacity threshold will be set for the storage capacity of different storage nodes. When the used capacity of a single storage node reaches 80% of its total capacity, writing data to this node will be suspended. If more than 75% of the total nodes reach the capacity threshold, the system will automatically trigger an expansion mechanism.
8. The security mass data hierarchical distributed storage management system according to claim 1, characterized in that: When the data layout module distributes the sub-security data to different storage nodes, it dynamically calculates the number of shards D according to the size of the sub-security data and the minimum shard capacity threshold of the storage nodes, where D ≥ 2 and D ≤ the number of currently available storage nodes. Then, it writes different shards to different storage nodes simultaneously to make the write operations parallel. The user has the permission to edit the number of shards into which the sub-security data is divided.
9. A management method for a security mass data hierarchical distributed storage management system according to any one of claims 1-8, characterized in that, It includes the following steps: S100. Collect different security data, obtain the permission to use cloud storage, analyze the source types, data property types, and data usage types of the security data, and obtain a distribution type library. S200. Generate a preliminary type according to the user's editing information, then extract cases from security cases to obtain experimental cases, and analyze whether the data in the experimental cases can be classified into the preliminary type. When it cannot be classified into the preliminary type, the preliminary type information will be deleted. When it can be classified into the preliminary type, the preliminary type will be recorded in the distribution type library. S300. Construct a directed acyclic graph structure, where the vertices represent storage nodes, the directed edges represent the data flow logic, and there are no circular dependency paths in the graph, that is, there is no path that starts from a certain vertex, passes through a series of vertices along the directed edges, and then returns to the starting vertex. S400. Analyze the data about the distribution type in the security data and intercept it to obtain sub-security data. S500. Analyze the correlation between sub-security data through the cosine similarity algorithm, and allocate the data with a correlation ≥ 0.8 to adjacent storage nodes.