Data storage management classification system based on cloud computing
By designing a data storage management classification system in a cloud computing environment, dynamically selecting the optimal storage path and performing load balancing, the problem of unreasonable data storage path selection in the existing technology is solved, and the stability and reliability of the system are improved.
Patent Information
- Application Number
- CN202510059174.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the selection of data storage paths mostly relies on static rules or single parameters, and the adaptability of data characteristics and storage paths is not fully considered, and dynamic monitoring and optimization of path loads are also lacked.
A data storage management classification system based on cloud computing is designed, including data acquisition module, preprocessing module, feature extraction module, classification module and storage module. By calculating the adaptability of the classification value and storage path, dynamically selecting the optimal storage path, and data migration is carried out when the load exceeds the threshold to ensure path load balancing.
It improves the accuracy and rationality of storage path allocation, avoids storage path overload or resource waste, significantly improves the stability and reliability of the storage system, and supports dynamic adjustment of the classification value range of storage paths according to the data distribution.
Smart Images

Figure CN120029538A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data management, and in particular to a data storage management classification system based on cloud computing. Background Art
[0002] With the rapid development of cloud computing and big data technology, the amount of data has shown explosive growth. The storage and management of massive data has gradually become a core challenge in the field of modern information technology. In the existing technology, data storage management is usually optimized in the following ways: First, through static storage path allocation rules, data is allocated to pre-defined storage paths according to type or format; second, a load balancing strategy is adopted to monitor the resource occupancy of storage nodes (such as disk space or access frequency) to dynamically adjust the distribution of stored data; third, the hot and cold data tiered storage method is used to store high-frequency access data on high-performance media, and low-frequency data on low-cost devices.
[0003] After searching, a Chinese patent (publication number: CN118331948A) discloses an intelligent data management system and method using artificial intelligence and big data technology. The patent includes a data management platform, which is communicatively connected to a collection module, a screening module, a classification module and a storage module: the collection module is used to process and analyze the data uploaded by the sender; the screening module is used to receive the data transmitted by the collection module and perform screening processing: by calculating the regression distance between the position of the data on the feature layer of different classification dimensions and the center of different classification dimensions, it is determined whether the regression distance exceeds the preset distance value. If so, the data meets the classification requirements, and the data is transmitted to the classification module, the first classification score of the data is calculated based on the regression distance, and the first classification score is multiplied by the classification confidence predicted by the classification branch to obtain a second classification score, and the data is classified according to the second classification score.
[0004] In the prior art, the selection of storage paths mostly relies on static rules or single parameters, without fully considering the adaptability of data characteristics and storage paths, and also lacks dynamic monitoring and optimization of path loads. Therefore, the present invention proposes a data storage management classification system based on cloud computing. Summary of the invention
[0005] The purpose of the present invention is to provide a data storage management classification system based on cloud computing to solve the problems mentioned in the above background technology.
[0006] The present invention can be implemented through the following technical solutions: A data storage management classification system based on cloud computing, including a data acquisition module, a preprocessing module, a feature extraction module, a classification module and a storage module;
[0007] The data acquisition module is used to collect data from various data sources, including structured data sources, unstructured data sources and real-time data streams;
[0008] The preprocessing module is used to preprocess the collected data to improve the data quality. The preprocessing includes:
[0009] Data cleaning, deleting or correcting missing values, outliers and redundant data;
[0010] Data conversion, standardizing or formatting diverse raw data;
[0011] Noise reduction processing: filtering and smoothing signal data or time series data to reduce noise interference;
[0012] Data integration: aligning and fusing data from multiple sources, such as synchronizing by timestamp or merging data based on associations;
[0013] Feature standardization: normalize or standardize numerical features to eliminate the impact of dimensional differences on classification results;
[0014] The feature extraction module is responsible for extracting key features of data to provide a basis for subsequent data classification, including a sorting unit and a feature matching unit;
[0015] The feature matching unit performs a hierarchical operation on the preset key features and classifies the data extraction range from large to small to obtain an extraction framework;
[0016] The sorting unit matches the data features layer by layer based on the extraction framework, completes the classification at each level, and marks the corresponding features of the data, and finally generates a set of marks for each data;
[0017] The classification module quickly identifies the tag set of each data based on the preset algorithm and rules in the cloud computing environment, and calculates each tag based on the level of the extraction framework to obtain the corresponding classification value;
[0018] The storage module stores the data in a corresponding location of the cloud computing storage resource according to the classification value.
[0019] A further technical improvement of the present invention is that: the sorting unit adopts a tree-like hierarchical classification structure, which sets a root node, an intermediate node and a leaf node of the tree-like classification;
[0020] Among them, the root node is used to preliminarily divide the data to form large-category data subsets;
[0021] The intermediate nodes further extract modality-related intermediate features based on the output of the first layer;
[0022] Leaf nodes are refined into key features related to specific tasks based on intermediate features.
[0023] A further technical improvement of the present invention is that the calculation of the data classification value by the recognition unit includes the following steps:
[0024] Q1, extract the tag set S of each data = {s 1 ,s 2 ,...s n ,}, where each s i is the feature matching result of the i-th layer;
[0025] Q2. Calculate the matching value of the tag set of each layer of the extraction frame, defined as M i , represents the matching degree of all features in the i-th layer;
[0026] Q3. Set weights W for different levels i , the weight of the first layer is low, and the weight of the layers gradually increases;
[0027] Q4. Through the formula Calculate the classification value of the corresponding data;
[0028] In the formula, M i represents the feature matching value of the i-th layer, which is the weighted average of all feature matching results of this layer; mi represents the total number of features contained in the i-th layer; R ij is the matching result of the jth feature of the i-th layer, with a value range of [0, 1], indicating the matching degree between the current data and the feature, 1 indicates a complete match, and 0 indicates a complete mismatch; P ij is the priority weight of the jth feature in the i-th layer, reflecting the importance of the feature in this layer, and the value range is a positive number;
[0029] W i is the weight of the i-th layer, which is determined by the preset layer weight and satisfies
[0030] A further technical improvement of the present invention is that the recognition unit normalizes the classification value so that its range is fixed between [0, 1] or [0, 100].
[0031] A further technical improvement of the present invention is that: the storage module is provided with a synonym library, in which the synonym library stores the synonymous features of each feature, that is, for each feature f i , create the corresponding synonym set Z i ={f i1 ,f i2 ,...f im}, add possible synonymous expressions;
[0032] The extraction framework introduces a synonym library as a core reference in the feature classification rules at each level to ensure that all feature matching operations rely on the synonym library;
[0033] When the extraction framework is matched with the data, the data is traversed through the synonym library, and each feature d in each data i ;
[0034] If the data feature d i ∈Synonym set S i , then the match is successful and mapped to the standard feature f i ;
[0035] Finally, the tag set of the data is s = {Z 1 ,Z 2 ,...Z i}.
[0036] A further technical improvement of the present invention is that the synonym library is provided with a dynamic update mechanism, which sets corresponding level permissions and corresponding level weight modification permissions for different levels of the extraction framework, allowing the corresponding permission level to add synonyms of the corresponding level or modify the existing mapping relationship.
[0037] A further technical improvement of the present invention is that the storage module dynamically modifies the storage path of the data based on the classification value of the data, including the following steps:
[0038] S1. Map the classification value V to different storage path ranges and assign an initial load value L to each storage path. i , indicating the resource usage of the current storage path;
[0039] S2. Calculate the compatibility between the classification value V and the storage path. The formula used is:
[0040]
[0041] In the formula, F i is the fitness of path i, L i is the load value of path i, ε is a constant used to place the denominator at zero when the load value is 0;
[0042] H i is the classification value V and the path range center C i The matching degree is expressed as H i =max(1-|VC i ,0);
[0043] S3, according to the calculated fitness F i, select the storage path with the highest adaptability, and after storing the data in the selected path, update the load value of the path;
[0044] S4. Regularly check the load values of all paths. When the load of a path exceeds the preset load threshold, trigger data migration.
[0045] S5. Generate a range index table of corresponding classification values for each path. When searching for data, query the index table according to the classification value V to quickly locate the path.
[0046] A further technical improvement of the present invention is that in S4, data migration includes:
[0047] G1. Recalculate the classification value V and fitness F for the high load path i , migrate some data to other paths with higher adaptability;
[0048] G2. Dynamically adjust the path range and center according to data distribution:
[0049] If the amount of data on a path is too large, split the range;
[0050] If the amount of data on a certain path is small, merge the ranges.
[0051] A further technical improvement of the present invention is that the implementation steps of G1 include:
[0052] J1. Get the load value L of all paths i , and select some data with lower classification value V from the high-load path for migration;
[0053] J2. Determine the amount of data that needs to be migrated based on the load difference of the path;
[0054] J3. After excluding the current path, the remaining paths are sorted according to the fitness F i Sort and select the fitness F i The highest path is used as the target path;
[0055] J4, moving the selected part of the data from the high-load path to the target path, and updating the load values of the high-load path and the target path;
[0056] If there are too many data in the classification value V range of a certain path, the range is further refined;
[0057] Compared with the prior art, the present invention has the following beneficial effects:
[0058] The present invention calculates the classification value and the storage path fitness, comprehensively considers the matching degree between the data characteristics and the storage path, and the load of the current path, dynamically selects the optimal storage path, improves the accuracy and rationality of storage path allocation, and can timely adjust the storage strategy when the load exceeds the threshold, migrate the data to the path with higher fitness, avoid storage path overload or resource waste, and significantly improve the stability and reliability of the storage system;
[0059] The present invention also supports dynamic adjustment of the classification value range of the storage path according to data distribution, realizing flexible optimization of storage rules. This dynamic adjustment mechanism can adapt to changes in data characteristic distribution at different time stages, enhancing the scalability and adaptability of the storage system. At the same time, while fully utilizing storage resources, the present invention reduces the impact of hot data on system performance through load balancing and data migration strategies, while reducing the risk of idle storage resources, and realizing overall performance improvement of the storage system. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to facilitate understanding by those skilled in the art, the present invention is further described below with reference to the accompanying drawings.
[0061] Figure 1 It is a system block diagram of the present invention. DETAILED DESCRIPTION
[0062] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the specific implementation methods, structures, features and effects of the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments.
[0063] Example 1
[0064] See also Figure 1 As shown, the present invention provides a data storage management classification system based on cloud computing, including a data acquisition module, a preprocessing module, a feature extraction module, a classification module and a storage module;
[0065] The data collection module is used to collect data from various data sources, including structured data sources, unstructured data sources and real-time data streams;
[0066] The preprocessing module is used to preprocess the collected data to improve the data quality. The preprocessing includes:
[0067] Data cleaning, deleting or correcting missing values, outliers and redundant data;
[0068] Data conversion, standardizing or formatting diverse raw data;
[0069] Noise reduction processing: filtering and smoothing signal data or time series data to reduce noise interference;
[0070] Data integration, aligning and fusing multi-source data, for example, synchronizing according to timestamps or merging data based on correlation relationships;
[0071] Feature standardization, normalizing or standardizing numerical features to eliminate the impact of dimensional differences on classification results;
[0072] The feature extraction module is responsible for extracting key features of the data, providing a basis for subsequent data classification, including a sorting unit and a feature matching unit;
[0073] The feature matching unit performs a hierarchical operation on the preset key features, classifying the extraction scope of the data from large to small to obtain an extraction framework;
[0074] The extraction framework includes:
[0075] The first layer: Global data feature classification;
[0076] The processing methods include: Classifying according to data types: such as images, texts, audios, time series, and structured data;
[0077] Coarse classification based on domain knowledge: For example, financial data is divided into transaction data and account data; medical data is divided into imaging data and laboratory reports;
[0078] Classification based on data formats: such as JSON, CSV, and video files;
[0079] The second layer: Specific modality feature classification;
[0080] The processing methods include:
[0081] For image data, extract global texture and color histogram features, and use shallow features of a pre-trained deep network;
[0082] For text data:
[0083] Extract keywords based on term frequency-inverse document frequency (TF-IDF), and generate semantic vectors using Word2Vec or BERT;
[0084] For time series data:
[0085] Extract global trends, periodicity, and seasonality features, and use difference analysis to remove long-term trends and retain local fluctuations;
[0086] For audio data:
[0087] Extract time-domain features and extract frequency-domain features;
[0088] The third layer: Fine-grained key feature classification;
[0089] The processing methods include:
[0090] Image data: detect local key points and edge features, extract high-level features of deep networks, and generate semantic feature vectors;
[0091] Text data: Identify contextual semantic features (such as named entities, syntactic structures) and extract keywords or phrases related to specific tasks;
[0092] Time series data: Identify abnormal points or mutation points and extract local window features (such as short-term moving average and local fluctuation amplitude);
[0093] Audio data: Identify key frequency components and time domain events, and generate feature descriptions of event sequences (such as starting point, peak time);
[0094] The sorting unit matches the data with the extraction framework and identifies the features that meet the different levels of the extraction framework. The sorting unit classifies the data level by level and marks the features corresponding to the data to obtain a set of marks for each data.
[0095] Based on preset algorithms and rules, the classification module quickly identifies the tag sets of each data in the cloud computing environment, and calculates each tag based on the hierarchy of the extraction framework to obtain the corresponding classification value;
[0096] The storage module stores the data in a corresponding location of the cloud computing storage resource according to the classification value.
[0097] The calculation of the classification value of the data by the recognition unit includes the following steps:
[0098] Q1, extract the tag set S of each data = {s 1 ,s 2 ,...s n ,}, where each s i is the feature matching result of the i-th layer;
[0099] Q2. Calculate the matching value of the tag set of each layer of the extraction frame, defined as M i , represents the matching degree of all features in the i-th layer;
[0100] Q3. Set weights W for different levels i , the weight of the first layer is low, and the weight of the layers gradually increases;
[0101] Q4. Through the formula Calculate the classification value of the corresponding data;
[0102] In the formula, M irepresents the feature matching value of the i-th layer, which is the weighted average of all feature matching results of this layer; mi represents the total number of features contained in the i-th layer; R ij is the matching result of the jth feature of the i-th layer, with a value range of [0, 1], indicating the matching degree between the current data and the feature, 1 indicates a complete match, and 0 indicates a complete mismatch; P ij is the priority weight of the jth feature in the i-th layer, reflecting the importance of the feature in this layer, and the value range is a positive number;
[0103] W i is the weight of the i-th layer, which is determined by the preset layer weight and satisfies
[0104] In this embodiment, the extraction framework includes three layers of features, and the matching results are as follows:
[0105] First layer: R 11 =0.7, R 12 =0.6, weight P 11 =0.5, P 12 =0.5, weight W 1 =0.2, then
[0106]
[0107] Second layer: R 21 =0.9, R 22 =0.8, R 23 =0.7, weight P 21 =0.4, P 22 =0.3, P 23 =0.3, weight W 2 =0.3, then
[0108] Third layer: R 31 =1.0, R 32 =0.9, weight P 31 =0.6, P 32 =0.4, weight W 3 =0.5, then
[0109]
[0110] final,
[0111] The recognition unit normalizes the classification value so that its range is fixed between [0, 1] or [0, 100];
[0112] In this embodiment, the classification values are normalized to a percentage system, and the result is V=84.4.
[0113] Example 2
[0114] A data storage management classification system based on cloud computing, comprising a data acquisition module, a preprocessing module, a feature extraction module, a classification module and a storage module;
[0115] The data collection module is used to collect data from various data sources, including structured data sources, unstructured data sources and real-time data streams;
[0116] The preprocessing module is used to preprocess the collected data to improve the data quality. The preprocessing includes:
[0117] Data cleaning, deleting or correcting missing values, outliers and redundant data;
[0118] Data conversion, standardizing or formatting diverse raw data;
[0119] Noise reduction processing: filtering and smoothing signal data or time series data to reduce noise interference;
[0120] Data integration: aligning and fusing data from multiple sources, such as synchronizing by timestamp or merging data based on associations;
[0121] Feature standardization: normalize or standardize numerical features to eliminate the impact of dimensional differences on classification results;
[0122] The feature extraction module is responsible for extracting key features of the data and providing a basis for subsequent data classification, including sorting units and feature matching units;
[0123] The feature matching unit performs a hierarchical operation on the preset key features and classifies the data extraction range from large to small to obtain an extraction framework;
[0124] The extraction framework includes:
[0125] First layer: global data feature classification;
[0126] Processing methods include: classification according to data type: such as images, text, audio, time series, structured data;
[0127] Rough classification based on domain knowledge: For example, financial data is divided into transaction data and account data; medical data is divided into imaging data and laboratory reports;
[0128] Classification based on data format: such as JSON, CSV, video files;
[0129] The second layer: classification of specific modality features;
[0130] The treatment methods include:
[0131] Image data, extract global texture and color histogram features, and use shallow features of pre-trained deep networks;
[0132] Text data:
[0133] Extract keywords based on word frequency statistics (TF-IDF) and generate semantic vectors using Word2Vec or BERT;
[0134] Time Series Data:
[0135] Extract global trends, cyclical and seasonal characteristics, use differential analysis to remove long-term trends and retain local fluctuations;
[0136] Audio Data:
[0137] Extract time domain features, extract frequency domain features;
[0138] The third layer: fine-grained key feature classification;
[0139] The treatment methods include:
[0140] Image data: detect local key points and edge features, extract high-level features of deep networks, and generate semantic feature vectors;
[0141] Text data: Identify contextual semantic features (such as named entities, syntactic structures) and extract keywords or phrases related to specific tasks;
[0142] Time series data: Identify abnormal points or mutation points and extract local window features (such as short-term moving average and local fluctuation amplitude);
[0143] Audio data: Identify key frequency components and time domain events, and generate feature descriptions of event sequences (such as starting point, peak time);
[0144] The sorting unit adopts a tree-like hierarchical classification structure, which sets the root node, intermediate node and leaf node of the tree-like classification;
[0145] The root node is used to preliminarily divide the data into large data subsets. The classification basis includes:
[0146] Modality classification: distinguishing between images, text, audio, time series, structured data, etc.;
[0147] Field classification: such as classification of characteristics in fields such as finance, medical care, and industry;
[0148] Format classification: based on the storage format of data;
[0149] The intermediate nodes further extract the intermediate features related to the modality from the output of the first layer. The classification basis includes:
[0150] Image data: extract global characteristics such as color histogram and texture features;
[0151] Text data: extract word frequency (TF-IDF) and word embedding vectors;
[0152] Time series data: extract trends, periodicity, and volatility characteristics;
[0153] Audio data: extract MFCC features or frequency domain features;
[0154] Leaf nodes are refined into key features related to specific tasks based on intermediate features. The classification basis includes:
[0155] Extract fine-grained features based on local characteristics of the data:
[0156] Image: edge detection, key point extraction, semantic feature representation;
[0157] Text: syntactic analysis, named entity recognition;
[0158] Time series: local mutation point detection, short-term fluctuation characteristics;
[0159] Audio: time-frequency event analysis;
[0160] The sorting unit matches the data with the extraction framework and identifies the features that meet the different levels of the extraction framework. The sorting unit classifies the data level by level and marks the features corresponding to the data to obtain a set of marks for each data.
[0161] The tree structure in the sorting unit mainly realizes the classification organization of data, and gradually refines the data classification from the root node to the leaf node. The extraction framework of each layer extracts the key features required by the layer according to specific rules or algorithms to support the branch decision of the tree structure;
[0162] Based on preset algorithms and rules, the classification module quickly identifies the tag sets of each data in the cloud computing environment, and calculates each tag based on the hierarchy of the extraction framework to obtain the corresponding classification value;
[0163] The storage module stores the data in a corresponding location of the cloud computing storage resource according to the classification value;
[0164] The calculation of the classification value of the data by the recognition unit includes the following steps:
[0165] Q1, extract the tag set S of each data = {s 1 ,s 2 ,...s n ,}, where each s i is the feature matching result of the i-th layer;
[0166] Q2. Calculate the matching value of the tag set of each layer of the extraction frame, defined as M i , represents the matching degree of all features in the i-th layer;
[0167] Q3. Set weights W for different levels i , the weight of the first layer is low, and the weight of the layers gradually increases;
[0168] Q4. Through the formula Calculate the classification value of the corresponding data;
[0169] Compared with the first embodiment, the storage module is provided with a synonym library, in which the synonym features of each feature are stored, that is, for each feature f i , create the corresponding synonym set Z i ={f i1 ,f i2 ,...f im}, add possible synonymous expressions;
[0170] The extraction framework introduces a synonym library as a core reference in the feature classification rules at each level to ensure that all feature matching operations rely on the synonym library;
[0171] When the extraction framework matches the data, the data is traversed through the synonym library, and each feature d in each data i ;
[0172] If the data feature d i ∈Synonym set S i , then the match is successful and mapped to the standard feature f i ;
[0173] Finally, the tag set of the data is s = {Z 1 ,Z 2 ,...Z i};
[0174] The synonym library is equipped with a dynamic update mechanism, which sets the modification permissions of the corresponding level weights for different levels of the extraction framework by setting the permissions of the corresponding level, allowing the corresponding permission level to add synonyms of the corresponding level or modify the existing mapping relationship;
[0175] The storage module dynamically modifies the storage path of the data based on the classification value of the data, including the following steps:
[0176] S1. Map the classification value V to different storage path ranges and assign an initial load value L to each storage path. i , indicating the resource usage of the current storage path;
[0177] In this embodiment, when V∈[0, 0.3], it is a cold storage path;
[0178] When V∈(0.3, 0.7], it is a conventional storage path;
[0179] When V∈(0.7, 1.0], it is a heat storage path;
[0180] S2. Calculate the compatibility between the classification value V and the storage path. The formula used is:
[0181]
[0182] In the formula, F i is the fitness of path i, L i is the load value of path i, ε is a constant used to place the denominator at zero when the load value is 0;
[0183] H i is the classification value V and the path range center C i The matching degree is expressed as H i =max(1-|VC i ,0);
[0184] S3, according to the calculated fitness F i , select the storage path with the highest adaptability, and after storing the data in the selected path, update the load value of the path;
[0185] S4. Regularly check the load values of all paths. When the load of a path exceeds the preset load threshold, trigger data migration.
[0186] Data migration includes:
[0187] G1. Recalculate the classification value V and fitness F for the high load path i , migrate some data to other paths with higher adaptability, including:
[0188] J1. Get the load value L of all paths i , and select some data with lower classification value V from the high-load path for migration;
[0189] J2. Determine the amount of data that needs to be migrated based on the load difference of the path;
[0190] J3. After excluding the current path, the remaining paths are sorted according to the fitness F i Sort and select the fitness F i The highest path is used as the target path;
[0191] J4, moving the selected part of the data from the high-load path to the target path, and updating the load values of the high-load path and the target path;
[0192] If there are too many data in the classification value V range of a certain path, the range is further refined;
[0193] G2. Dynamically adjust the path range and center according to data distribution:
[0194] In this embodiment, if the amount of data in V∈(0.6, 0.9] exceeds the preset quantity threshold, it is split into V∈(0.6, 0.75] and V∈(0.75, 0.9];
[0195] S5. Generate a range index table of corresponding classification values for each path. When searching for data, query the index table according to the classification value V to quickly locate the path.
[0196] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technical personnel in this field can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A data storage management classification system based on cloud computing, comprising a data acquisition module, a preprocessing module, a feature extraction module, a classification module and a storage module, characterized in that: The data acquisition module is used to collect data from various data sources; The preprocessing module is used to preprocess the collected data; The feature extraction module includes a sorting unit and a feature matching unit, which are responsible for extracting key features of the data; The feature matching unit is responsible for classifying the data according to the preset key features and gradually narrowing the data extraction scope to obtain the extraction framework; The sorting unit matches the data features layer by layer based on the extraction framework, completes the classification at each level, and marks the corresponding features of the data, and finally generates a set of marks for each data; The classification module quickly identifies the tag set of each data based on the preset algorithm and rules in the cloud computing environment, and calculates each tag based on the level of the extraction framework to obtain the corresponding classification value; The storage module dynamically selects the optimal location in the cloud computing storage resource to store data according to the classification value.
2. The data storage management classification system based on cloud computing according to claim 1, characterized in that: The sorting unit adopts a tree-like hierarchical classification structure, which sets the root node, intermediate node and leaf node of the tree-like classification; Among them, the root node is used to preliminarily divide the data to form large-category data subsets; The intermediate nodes further extract modality-related intermediate features based on the output of the first layer; Leaf nodes are refined into key features related to specific tasks based on intermediate features.
3. The data storage management classification system based on cloud computing according to claim 1, characterized in that: The calculation of the data classification value by the recognition unit includes the following steps: Q1, extract the tag set S of each data = {s1, s2, ...s n ,}, where each s i is the feature matching result of the i-th layer; Q2. Calculate the matching value of the tag set of each layer of the extraction frame, defined as M i , represents the matching degree of all features in the i-th layer; Q3. Set weights W for different levels i ; Q4. Through the formula Calculate the classification value of the corresponding data; In the formula, M i represents the feature matching value of the i-th layer; mi represents the total number of features contained in the i-th layer; R ij is the matching result of the jth feature of the i-th layer; P ij is the priority weight of the jth feature in the i-th layer; W i is the weight of the i-th layer, which is determined by the preset layer weight and satisfies 4. The data storage management classification system based on cloud computing according to claim 3, characterized in that: The recognition unit normalizes the classification value so that its range is fixed between [0, 1] or [0, 100].
5. The data storage management classification system based on cloud computing according to claim 1, characterized in that: The storage module is provided with a synonym library, in which the synonym features of each feature are stored, that is, for each feature f i , create the corresponding synonym set Z i ={f i1 ,f i2 ,...f im }, add synonymous expressions to it; The extraction framework introduces a synonym library as a core reference in the feature classification rules at each level to ensure that all feature matching operations rely on the synonym library; When the extraction framework is matched with the data, the data is traversed through the synonym library, and each feature d in each data i ; If the data feature d i ∈Synonym set S i , then the match is successful and mapped to the standard feature f i ; Finally, the tag set of the data is s = {Z1, Z2, ... Z i }.
6. The data storage management classification system based on cloud computing according to claim 5, characterized in that: The synonym library is provided with a dynamic update mechanism, which sets corresponding level permissions and corresponding level weight modification permissions for different levels of the extraction framework, allowing the corresponding permission level to add synonyms of the corresponding level or modify the existing mapping relationship.
7. The data storage management classification system based on cloud computing according to claim 5, characterized in that: The storage module dynamically modifies the storage path of the data based on the classification value of the data, including the following steps: S1. Map the classification value V to different storage path ranges and assign an initial load value L to each storage path. i ; S2. Calculate the compatibility between the classification value V and the storage path. The formula used is: In the formula, F i is the fitness of path i, L i is the load value of path i, ε is a constant used to place the denominator at zero when the load value is 0; H i is the classification value V and the path range center C i The matching degree is expressed as H i =max(1-|VC i |,0); S3, according to the calculated fitness F i , select the storage path with the highest adaptability, and after storing the data in the selected path, update the load value of the path; S4. Regularly check the load values of all paths. When the load of a path exceeds a preset load threshold, trigger data migration. S5. Generate a range index table of corresponding classification values for each path. When searching for data, query the index table according to the classification value V to quickly locate the path.
8. The data storage management classification system based on cloud computing according to claim 7, characterized in that: In S4, data migration includes: G1. Recalculate the classification value V and fitness F for the high load path i , migrate some data to other paths with higher adaptability; G2. Dynamically adjust the path range and center according to data distribution: If the amount of data on a path is too large, split the range; If the amount of data on a certain path is small, merge the ranges.
9. The data storage management classification system based on cloud computing according to claim 8, characterized in that: The implementation steps of G1 include: J1. Get the load value L of all paths i , and select some data with lower classification value V from the high-load path for migration; J2. Determine the amount of data that needs to be migrated based on the load difference of the path; J3. After excluding the current path, the remaining paths are sorted according to the fitness F i Sort and select the fitness F i The highest path is used as the target path; J4. Move the selected part of the data from the high-load path to the target path, and update the load values of the high-load path and the target path.
Citation Information
Patent Citations
Intelligent data management system and method applying artificial intelligence and big data technology
CN118331948A