Object storage intelligent layering method, system and equipment based on decision tree
Through the intelligent hierarchical method of object storage based on decision tree, the data storage pool is dynamically adjusted, and the simple problems of data hierarchical delay, migration loss and hierarchical model in the existing technology are solved, achieving more efficient and accurate data hierarchy and improved access performance.
Patent Information
- Application Number
- CN202510128916.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-13
AI Technical Summary
The existing data hierarchical storage technology has problems such as hierarchical delay, data migration loss and hierarchical model, and cannot respond to changes in data heat, affect system performance, and be difficult to meet access needs in different time periods.
The intelligent hierarchical method of object storage based on decision tree is adopted. By collecting and preprocessing object data, setting access time windows and storage tags, extracting and dimensionality reduction feature data, building and deploying decision tree models, classifying and storing predictions of uploaded object data, and dynamically adjusting object data in the storage pool.
It realizes saving data migration losses, improving access performance, improving data hierarchy accuracy and realizing dynamic data hierarchy, which can better adapt to access needs in different time periods and optimize the utilization of storage resources.
Smart Images

Figure CN119989173A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data storage, and in particular to a decision tree-based object storage intelligent tiering method, system and device. Background Art
[0002] With the rapid development of information technology, the demand for data storage and management is growing. In order to improve the efficiency and performance of storage systems, data tiered storage technology has emerged. The shortcomings of existing technologies and their practical applications are mainly reflected in the following aspects: Hot data identification: Identify hot data by analyzing existing access information. Although this method can initially distinguish the popularity of data, it has a certain delay and needs to be adjusted according to the access frequency over a period of time. It cannot respond to changes in data popularity in real time.
[0003] Data migration: Migrate the identified hot data between different storage media to achieve hierarchical data storage. However, the data migration process will not only cause the loss of storage medium space, but also increase the time cost of data migration, affecting the overall performance of the system.
[0004] Classified storage: Classified storage is performed based on the known heat values of object storage, which usually only implements the simplest hot-cold tiered model. This method cannot dynamically migrate data based on more detailed time characteristics (such as whether it is daytime or working day), and it is difficult to meet the access requirements of different time periods, limiting the further improvement of data access performance.
[0005] In summary, the existing technology has the following problems in data tiered storage: 1. Tier delay, which needs to be adjusted according to the access frequency over a period of time and cannot respond to changes in data popularity in real time; 2. Data migration loss, the data migration process will cause loss of storage medium space and migration time, affecting the overall performance of the system; 3. The tiered model is simple and can only implement the simplest hot and cold tiered model. It is impossible to dynamically migrate data according to more detailed time characteristics, making it difficult to meet access requirements in different time periods, limiting the further improvement of data access performance.
[0006] Therefore, a new technical solution is needed to overcome the shortcomings of existing technologies, achieve more efficient and accurate data tiered storage, and improve data access performance. Summary of the invention
[0007] In view of this, in order to overcome the deficiencies of the prior art, the present application aims to provide a decision tree-based object storage intelligent tiering method, system and device.
[0008] According to a first aspect of the present application, a decision tree-based object storage intelligent tiering method is provided, the method comprising: Collect object data, pre-process the collected object data, and set access time windows and storage tags for the obtained standardized object data; Perform feature extraction on the stored standardized data, perform dimensionality reduction processing on the extracted features, and obtain feature data; Build a training model, use feature data to train and test the training model, and select a decision tree model from the training model for production deployment; The deployed decision tree model is used to perform classification storage prediction on the uploaded object data, and the object data is stored in the corresponding storage pool according to the classification storage prediction results.
[0009] Optionally, in the decision tree-based object storage intelligent tiering method of the present application, object data includes object metadata, access log data and business scenario data, object metadata includes object size, file type, creation time, modification time, user and storage layer data, access log data includes access timestamp, user ID, access type and access result data, and business scenario data includes scenario event data and scenario activity data.
[0010] Optionally, in the decision tree-based object storage intelligent tiering method of the present application, the collected object data is preprocessed, including: deleting duplicate object data, deleting erroneous object data, and deleting object data missing key features, wherein the key features include file type, creation time, and access timestamp.
[0011] Optionally, in the object storage intelligent tiering method based on a decision tree of the present application, setting an access time window and a storage label for the obtained standardized object data includes: Setting a continuous first time period, a second time period and a third time period, and sequentially setting corresponding first access times, second access times and third access times for the first time period, the second time period and the third time period; When the first access number is greater than 0 and not greater than a set threshold, a low-frequency object tag is set for the corresponding object data, and the low-frequency object tag indicates that the object data is stored in a low-frequency storage pool; When the first access number is greater than a set threshold, a standard object tag is set for the corresponding object data, the standard object tag indicating that the object data is stored in a standard storage pool; When the second access count is greater than 0, an archiving object tag is set for the corresponding object data, and the archiving object tag indicates that the object data is stored in the archiving storage pool; When the third access number is greater than 0, a deep archive tag is set for the corresponding object data, and the deep archive tag indicates that the object data is stored in the deep archive storage pool.
[0012] Optionally, in the decision tree-based object storage intelligent stratification method of the present application, feature extraction is performed on the stored standardized data, and the extracted features are subjected to dimensionality reduction processing to obtain feature data, including: extracting basic features and user features of the object data from the standardized data, and locating and acquiring feature data from the extracted basic features and user features through linear discriminant.
[0013] Optionally, in the object storage intelligent tiering method based on a decision tree in the present application, a training model is constructed, feature data is used to train and test the training model, and a decision tree model is selected from the training model for production deployment, including: Divide the feature data into training set and test set through data segmentation; Construct a training model, and use the training set to train the constructed training model; Use the test set to test the trained model and calculate the accuracy of the classification storage of the training model; The training model with the highest classification storage accuracy is deployed to the production environment as a decision tree model.
[0014] Optionally, in the decision tree-based object storage intelligent tiering method of the present application, a test set is used to test the trained training model to calculate the accuracy of the classification storage of the training model, including: using the test set as the input of the training model, outputting the predicted storage label of the object data in the test set, comparing the output predicted storage label with the storage label of the object data, and calculating the accuracy of the classification storage of the training model based on the matching degree between the predicted storage label and the storage label.
[0015] Optionally, in the object storage intelligent tiering method based on a decision tree of the present application, a deployed decision tree model is used to perform classified storage prediction on the uploaded object data, and the object data is stored in a corresponding storage pool according to the classified storage prediction result, including: The decision tree model is used regularly to reclassify and predict the object data in the storage pool, and the storage of the object data in the storage pool is adjusted according to the re-obtained classification storage prediction results.
[0016] According to a second aspect of the present application, a decision tree-based object storage intelligent tiering system is provided, the system comprising a tiering service end, the tiering service end comprising: An object data acquisition module, which acquires object data and pre-processes the acquired object data; A time window and storage label setting module, used for setting an access time window and a storage label for the obtained standardized object data; The feature data extraction and dimensionality reduction module is used to extract features from the stored standardized data, perform dimensionality reduction processing on the extracted features, and obtain feature data; The decision tree model building and deployment module is used to build a training model, use feature data to train and test the training model, and select a decision tree model from the training model for production deployment; The object data classification prediction storage module is used to use the deployed decision tree model to perform classification storage prediction on the uploaded object data, and store the object data in the corresponding storage pool according to the classification storage prediction results.
[0017] According to a third aspect of the present application, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect of the present application when executing the program.
[0018] The object storage intelligent tiering method, system and device based on decision tree of the present application have the following beneficial technical effects: 1. Save data migration loss: By predicting the future access frequency of the object, the data is directly allocated to different storage media layers, avoiding the extra time and space loss in the data migration process and improving the efficiency of the storage system.
[0019] 2. Improve access performance: The optimal storage media layer is recommended for each object based on the prediction results, so that users can obtain the required data more quickly when accessing object storage, thereby improving overall access performance.
[0020] 3. Improve the accuracy of data stratification: Models trained using a large amount of historical data can effectively eliminate occasional hot spot data, making data stratification more accurate and avoiding inaccurate stratification caused by data fluctuations.
[0021] 4. Realize dynamic data stratification: The ability of dynamic data stratification is realized, and hot and cold stratification can be performed according to different time characteristics (such as whether it is daytime or working day), so that the storage system can better adapt to the access needs of different time periods and further optimize the utilization of storage resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0023] Figure 1 This is an example diagram of the architecture of an object storage intelligent tiering system based on a decision tree according to an embodiment of the present application; Figure 2This is an example diagram of the architecture of a tiered server of an object storage intelligent tiered system based on a decision tree according to an embodiment of the present application; Figure 3 A flowchart of a method for intelligent tiering of object storage based on a decision tree according to an embodiment of the present application; Figure 4 This is an example diagram of performing data feature analysis according to a decision tree-based object storage intelligent tiering method according to an embodiment of the present application; Figure 5 This is an example diagram of model training for an object storage intelligent tiering method based on a decision tree according to this embodiment; Figure 6 A schematic diagram of the structure of the device provided in this application. DETAILED DESCRIPTION
[0024] The embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0025] It should be noted that the following embodiments and features in the embodiments may be combined with each other in the absence of conflict; and, based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in the field without making any creative work are within the scope of protection of the present disclosure.
[0026] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein may be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on the present disclosure, it should be understood by those skilled in the art that an aspect described herein may be implemented independently of any other aspect, and two or more of these aspects may be combined in various ways. For example, any number of aspects described herein may be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein may be used to implement this device and / or practice this method.
[0027] Figure 1 FIG. 1 is an example diagram of an architecture of an object storage intelligent tiering system based on a decision tree according to an embodiment of the present application. Figure 1 As shown, the system may include a layered server 101, a communication network 102 and / or one or more layered clients 103. Figure 1 The example shown is a plurality of hierarchical clients 103 .
[0028] The tiering server 101 can be any appropriate server for storing information, data, programs and / or any other suitable type of content. In some embodiments, the tiering server 101 can perform appropriate functions. For example, in some embodiments, the tiering server 101 can be used to perform intelligent tiering of object storage based on a decision tree. As an optional example, in some embodiments, the tiering server 101 can be used to implement intelligent tiering of object storage by building a decision tree model. For example, the tiering server 101 can be used to: collect object data, pre-process the collected object data, set access time windows and storage labels for the obtained standardized object data; perform feature extraction on the stored standardized data, perform dimensionality reduction processing on the extracted features, and obtain feature data; build a training model, use feature data to train and test the training model, and select a decision tree model from the training model for production deployment; use the deployed decision tree model to classify and store the uploaded object data, and store the object data in the corresponding storage pool according to the classification storage prediction results.
[0029] Figure 2 FIG. 1 is an example diagram of the architecture of a tiered server of an object storage intelligent tiered system based on a decision tree according to an embodiment of the present application, such as Figure 2 As shown, the layered service end of this embodiment includes: An object data acquisition module, which acquires object data and pre-processes the acquired object data; A time window and storage label setting module, used for setting an access time window and a storage label for the obtained standardized object data; The feature data extraction and dimensionality reduction module is used to extract features from the stored standardized data, perform dimensionality reduction processing on the extracted features, and obtain feature data; The decision tree model building and deployment module is used to build a training model, use feature data to train and test the training model, and select a decision tree model from the training model for production deployment; The object data classification prediction storage module is used to use the deployed decision tree model to perform classification storage prediction on the uploaded object data, and store the object data in the corresponding storage pool according to the classification storage prediction results.
[0030] As another example, in some embodiments, the tiering service end 101 may send the decision tree-based object storage intelligent tiering method to the tiering client 103 for user use based on the request of the tiering client 103 .
[0031] As an optional example, in some embodiments, the tiering client 103 is used to provide a visual tiering interface, which is used to receive a user's selection input operation for performing intelligent tiering of object storage based on a decision tree, and to, in response to the selection input operation, obtain a tiering interface corresponding to the option selected by the selection input operation from the tiering server 101 and display the tiering interface, wherein the tiering interface at least displays information on performing intelligent tiering of object storage based on a decision tree and operation options for the information on performing intelligent tiering of object storage based on a decision tree.
[0032] In some embodiments, the communication network 102 can be any suitable combination of one or more wired and / or wireless networks. For example, the communication network 102 can include any one or more of the following: the Internet, an intranet, a wide area network (WAN), a local area network (LAN), a wireless network, a digital subscriber line (DSL) network, a frame relay network, an asynchronous transfer mode (ATM) network, a virtual private network (VPN), and / or any other suitable communication network. The layered client 103 can be connected to the communication network 102 via one or more communication links (e.g., communication link 104), and the communication network 102 can be linked to the layered server 101 via one or more communication links (e.g., communication link 105). The communication link can be any communication link suitable for transmitting data between the layered client 103 and the layered server 101, such as a network link, a dial-up link, a wireless link, a hard-wired link, any other suitable communication link, or any suitable combination of such links.
[0033] The tiering client 103 may include any one or more clients that present an interface related to intelligent tiering of object storage based on a decision tree in an appropriate form for use and operation by a user. In some embodiments, the tiering client 103 may include any suitable type of device. For example, in some embodiments, the tiering client 103 may include a mobile device, a tablet computer, a laptop computer, a desktop computer, and / or any other suitable type of client device.
[0034] Although the layered service end 101 is illustrated as one device, in some embodiments, any suitable number of devices may be used to perform the functions performed by the layered service end 101. For example, in some embodiments, multiple devices may be used to implement the functions performed by the layered service end 101. Alternatively, the functions of the layered service end 101 may be implemented using a cloud service.
[0035] Based on the above system, an embodiment of the present application provides a cloud platform interface design walkthrough method based on a rule engine, which is described below through the following embodiments.
[0036] Figure 3The following is a flowchart of a method for intelligent tiering of object storage based on a decision tree according to an embodiment of the present application. The method for intelligent tiering of object storage based on a decision tree in this embodiment can be executed on a tiering service end, and the method for intelligent tiering of object storage based on a decision tree includes the following steps: Step S201: collect object data, pre-process the collected object data, and set an access time window and a storage tag for the obtained standardized object data.
[0037] In this embodiment, object data includes object metadata, access log data and business scenario data. Object metadata includes object size, file type, creation time, modification time, user and storage layer data. Access log data includes access timestamp, user ID, access type and access result data. Business scenario data includes scenario event data and scenario activity data. Among them, access type includes read access and write access, and access result includes access success and access failure.
[0038] As an optional example, in this embodiment, object data preprocessing includes: deleting duplicate object data, deleting erroneous object data, and deleting object data with missing key features, where the key features include file type, creation time, and access timestamp.
[0039] After obtaining the standardized object data, this embodiment sets the time window length according to the access characteristics of the object data.
[0040] As an optional example, this embodiment sets a continuous first time period, a second time period, and a third time period, and sets corresponding first access times, second access times, and third access times for the first time period, the second time period, and the third time period in sequence. For example, the number of accesses within 0-7 days is set to t1, the number of accesses within 7-30 days is set to t2, and the number of accesses within 30-90 days is set to t3.
[0041] This embodiment divides different storage pools according to the hardware resources of the distributed storage system, including a low-frequency storage pool, a standard storage pool, an archive storage pool, and a deep archive storage pool, and determines the segmentation value in combination with the actual access frequency of the object data storage.
[0042] As an optional example, when the first access count is greater than 0 and not greater than the set threshold, a low-frequency object label is set for the corresponding object data, and this low-frequency object label indicates that the object data is stored in the low-frequency storage pool; when the first access count is greater than the set threshold, a standard object label is set for the corresponding object data, and this standard object label indicates that the object data is stored in the standard storage pool; when the second access count is greater than 0, an archived object label is set for the corresponding object data, and this archived object label indicates that the object data is stored in the archived storage pool; when the third access count is greater than 0, a deep-archived label is set for the corresponding object data, and this deep-archived label indicates that the object data is stored in the deep-archived storage pool. For example, when 0 < t1 < 10, a low-frequency object label is set for the corresponding object data, and this low-frequency object label indicates that the object data is stored in the low-frequency storage pool; when t1 > 10, a standard object label is set for the corresponding object data, and this standard object label indicates that the object data is stored in the standard storage pool; when t2 > 0, an archived object label is set for the corresponding object data, and this archived object label indicates that the object data is stored in the archived storage pool; when t3 > 0, a deep-archived label is set for the corresponding object data. In practical applications, the label settings for all object data can be carried out in a sequential manner.
[0043] Step S202: Extract features from the stored standardized data, perform dimensionality reduction on the extracted features, and obtain feature data.
[0044] In this embodiment, the basic features and user features of the object data are extracted from the standardized data. As an optional example, the following basic features of the object data are extracted in this application: 1. Object size, logarithmic transformation is performed to reduce the impact of large files; 2. File type, which is mapped to an integer value using label encoding, such as text = 1, video = 2, picture = 3, audio = 4. 3. Creation time and last modification time, the absolute values of these two features cannot be introduced into the model for feature analysis, and time characteristics can be constructed based on them, such as whether it is on a weekday, holiday, night, etc. for extraction; 4. Bucket ID to which the object belongs, object upload method (ordinary upload, segmented upload, append upload, overwrite upload). Secondly, the following user features of the object data are extracted in this embodiment: 1. Directly perform label encoding on the user ID; 2. Process advanced features of user information, such as the department to which the user belongs, user gender, age, whether the user is a cloud user, whether the user is the main user, whether the user is an anonymous user, whether the quota is enabled, the quota amount, etc.
[0045] When there are too many features and they are highly correlated, it is necessary to perform dimensionality reduction processing on the feature data. As an optional example, this embodiment locates and obtains feature data from the extracted basic features and user features through linear discriminant. For example, this embodiment can use LDA linear discriminant analysis to find the feature subset that best distinguishes objects of different categories. The feature data is input into the LDA model constructed by Python's scikit-learn and pandas libraries to obtain a set of optimal feature data. LDA (Latent Dirichlet Allocation) is a generative probability model that is mainly used for topic modeling tasks in natural language processing. It can discover potential topic structures from a large number of documents and effectively mine meaningful topic information by representing documents as topic probability distributions. scikit-learn is an open source Python machine learning library that provides simple and effective data mining and data analysis tools. scikit-learn is built on scientific computing libraries such as NumPy, SciPy, and matplotlib, and is widely used in machine learning tasks such as data preprocessing, model selection, evaluation, and deployment. pandas is an open source Python data analysis library that provides high-performance, easy-to-use data structures and data analysis tools. Pandas is built on the NumPy library and is widely used in data analysis tasks such as data cleaning, data exploration, data processing, and data visualization. Figure 4 This is an example diagram of performing data feature analysis according to a decision tree-based object storage intelligent tiering method according to an embodiment of the present application.
[0046] Step S203: construct a training model, use feature data to train and test the training model, and select a decision tree model from the training model for production deployment.
[0047] As an optional example, in this embodiment, the feature data is divided into a training set and a test set by data segmentation; a training model is constructed, and the constructed training model is trained using the training set; the trained training model is tested using the test set, and the classification storage accuracy of the training model is calculated; the training model with the highest classification storage accuracy is deployed to the production environment as a decision tree model. It should be noted that in this embodiment, the test set is used as the input of the training model, the predicted storage label of the object data in the test set is output, the output predicted storage label is compared with the storage label of the object data, and the classification storage accuracy of the training model is calculated based on the matching degree between the predicted storage label and the storage label.
[0048] For example, this embodiment first needs to divide the training set and the test set, using 70%-80% of the feature data as the training set and the remaining feature data as the test set. Use the train_test_split tool to randomly divide the data to ensure that the two parts of the data are similarly distributed. Then, use the scikit-learn library of python to build a training model. After initializing the training model, build a pipeline to combine data preprocessing and model training, define possible hyperparameter values, and automatically traverse the hyperparameter space through cross-validation and GridSearchCV to find the best combination. GridSearchCV is a tool for hyperparameter tuning in the Scikit-learn library. It traverses a given parameter grid and uses cross-validation to evaluate the model performance of each parameter combination, and finally returns the best parameter combination and the corresponding model. In this embodiment, train_test_split is a function in the scikit-learn library, which is used to randomly divide the data set into a training set and a test set. It is a common step in machine learning and is used to evaluate the performance of the model and ensure that the model has good generalization ability on unseen data. Figure 5 This is an example diagram of model training according to a decision tree-based object storage intelligent tiering method of this embodiment.
[0049] In practical applications, this embodiment can also provide a more detailed view through a confusion matrix to show the classification between different categories. In this embodiment, the F1 score can be used to comprehensively consider the precision and recall rate to improve the applicability to unbalanced data sets. This embodiment can also use the AUC-ROC curve to measure the model's ability to distinguish in binary or multi-classification problems, thereby achieving the evaluation of the training model. In this embodiment, the training model is applied to the test set, the accuracy of the classification storage is calculated, the training model with the best indicators and the highest classification accuracy is selected for saving, and the saved training model is deployed to the production environment. In this embodiment, the F1 score (F1 Score) is an indicator used in machine learning to evaluate the performance of the classification model, and is particularly suitable for processing unbalanced data sets. The F1 score is the harmonic mean of the precision (Precision) and the recall rate (Recall), which can comprehensively consider the accuracy and recall of the model. The AUC-ROC curve is an important tool for evaluating the performance of binary classifiers in machine learning. The ROC curve (Receiver Operating Characteristic Curve) shows the performance of the classifier by plotting the True Positive Rate (TPR) and False Positive Rate (FPR) at different decision thresholds. The AUC (Area Under the Curve) is the area under the ROC curve, which is used to quantify the performance of the classifier under all possible classification thresholds.
[0050] Step S204: using the deployed decision tree model to perform classification storage prediction on the uploaded object data, and storing the object data in a corresponding storage pool according to the classification storage prediction result.
[0051] The decision tree model is used regularly to reclassify and predict the object data in the storage pool, and the storage of the object data in the storage pool is adjusted according to the re-obtained classification storage prediction results.
[0052] This embodiment constructs different storage pools. As an optional example, a standard storage pool is created using all SSD (solid state drive) hard disks or high-performance HDD (mechanical drive) hard disks to provide the fastest read and write speeds. A low-frequency storage pool is created by mixing a small number of SSDs and more economical SATA HDD (Serial ATA Hard Disk Drive) hard disks, which has a lower cost but maintains a faster retrieval speed; an archive storage pool is created using all SATA HDD hard disks, which is suitable for long-term preservation and rarely accessed data; a tape library or an optical disk library provides a large-capacity, low-cost long-term storage solution, which is suitable for rarely accessed data. When object data is uploaded, its metadata is input into the decision tree model, the prediction results are output in real time, and the object data is stored in the corresponding storage pool according to the prediction results. During the operation of the storage system, the access mode of the existing object data is re-evaluated regularly, and its storage layer is adjusted according to the latest prediction demerits. In practical applications, this embodiment can also monitor and track the deviation between the actual access situation and the prediction result in real time, and trigger an alarm when the deviation exceeds the set threshold range. At the same time, feedback from operation and maintenance personnel or end users is collected to improve the algorithm of the decision tree model and the strategy of feature engineering.
[0053] The object storage intelligent tiering method based on a decision tree according to the embodiment of the present application has the following beneficial technical effects: 1. Save data migration loss: By predicting the future access frequency of the object, the data is directly allocated to different storage media layers, avoiding the extra time and space loss in the data migration process and improving the efficiency of the storage system.
[0054] 2. Improve access performance: The optimal storage media layer is recommended for each object based on the prediction results, so that users can obtain the required data more quickly when accessing object storage, thereby improving overall access performance.
[0055] 3. Improve the accuracy of data stratification: Models trained using a large amount of historical data can effectively eliminate occasional hot spot data, making data stratification more accurate and avoiding inaccurate stratification caused by data fluctuations.
[0056] 4. Realize dynamic data stratification: The ability of dynamic data stratification is realized, and hot and cold stratification can be performed according to different time characteristics (such as whether it is daytime or working day), so that the storage system can better adapt to the access needs of different time periods and further optimize the utilization of storage resources.
[0057] like Figure 6As shown, the present application also provides a device, including a processor 310, a communication interface 320, a memory 330 for storing a processor executable computer program, and a communication bus 340. The processor 310, the communication interface 320, and the memory 330 communicate with each other through the communication bus 340. The processor 310 implements the above-mentioned decision tree-based object storage intelligent tiering method by running the executable computer program.
[0058] Among them, the computer program in the memory 330 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0059] The system embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, i.e., they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected based on actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art may understand and implement it without creative effort.
[0060] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiment.
[0061] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A decision tree-based object storage intelligent tiering method, characterized in that: The method comprises: Collect object data, pre-process the collected object data, and set access time windows and storage tags for the obtained standardized object data; Perform feature extraction on the stored standardized data, perform dimensionality reduction processing on the extracted features, and obtain feature data; Build a training model, use feature data to train and test the training model, and select a decision tree model from the training model for production deployment; The deployed decision tree model is used to perform classification storage prediction on the uploaded object data, and the object data is stored in the corresponding storage pool according to the classification storage prediction results.
2. The object storage intelligent tiering method based on decision tree according to claim 1, characterized in that: Object data includes object metadata, access log data and business scenario data. Object metadata includes object size, file type, creation time, modification time, user and storage layer data. Access log data includes access timestamp, user ID, access type and access result data. Business scenario data includes scenario event data and scenario activity data.
3. The object storage intelligent tiering method based on decision tree according to claim 1, characterized in that: The collected object data is preprocessed, including: deleting duplicate object data, deleting erroneous object data, and deleting object data with missing key features, wherein the key features include file type, creation time, and access timestamp.
4. The object storage intelligent tiering method based on decision tree according to claim 1, characterized in that: Set access time windows and storage tags for acquired standardized object data, including: Setting a continuous first time period, a second time period and a third time period, and sequentially setting corresponding first access times, second access times and third access times for the first time period, the second time period and the third time period; When the first access number is greater than 0 and not greater than a set threshold, a low-frequency object tag is set for the corresponding object data, and the low-frequency object tag indicates that the object data is stored in a low-frequency storage pool; When the first access number is greater than a set threshold, a standard object tag is set for the corresponding object data, the standard object tag indicating that the object data is stored in a standard storage pool; When the second access count is greater than 0, an archiving object tag is set for the corresponding object data, and the archiving object tag indicates that the object data is stored in the archiving storage pool; When the third access number is greater than 0, a deep archive tag is set for the corresponding object data, and the deep archive tag indicates that the object data is stored in the deep archive storage pool.
5. The object storage intelligent tiering method based on decision tree according to claim 1, characterized in that: Feature extraction is performed on the stored standardized data, and dimension reduction processing is performed on the extracted features to obtain feature data, including: extracting basic features and user features of object data from the standardized data, and locating and obtaining feature data from the extracted basic features and user features through linear discrimination.
6. The object storage intelligent tiering method based on decision tree according to claim 1, characterized in that: Build a training model, use feature data to train and test the training model, and select a decision tree model from the training model for production deployment, including: Divide the feature data into training set and test set through data segmentation; Construct a training model, and use the training set to train the constructed training model; Use the test set to test the trained model and calculate the accuracy of the classification storage of the training model; The training model with the highest classification storage accuracy is deployed to the production environment as a decision tree model.
7. The object storage intelligent tiering method based on decision tree according to claim 1, characterized in that: The trained training model is tested using a test set to calculate the accuracy of classification storage of the training model, including: taking the test set as the input of the training model, outputting the predicted storage label of the object data in the test set, comparing the output predicted storage label with the storage label of the object data, and calculating the accuracy of classification storage of the training model according to the matching degree between the predicted storage label and the storage label.
8. The object storage intelligent tiering method based on decision tree according to claim 1, characterized in that: The deployed decision tree model is used to classify and predict the uploaded object data for storage, and the object data is stored in the corresponding storage pool according to the classification and storage prediction results, including: The decision tree model is used regularly to reclassify and predict the object data in the storage pool, and the storage of the object data in the storage pool is adjusted according to the re-obtained classification storage prediction results.
9. An object storage intelligent tiering system based on a decision tree, characterized in that: The system includes a layered service end, and the layered service end includes: An object data acquisition module, which acquires object data and pre-processes the acquired object data; A time window and storage label setting module, used for setting an access time window and a storage label for the obtained standardized object data; The feature data extraction and dimensionality reduction module is used to extract features from the stored standardized data, perform dimensionality reduction processing on the extracted features, and obtain feature data; The decision tree model building and deployment module is used to build a training model, use feature data to train and test the training model, and select a decision tree model from the training model for production deployment; The object data classification prediction storage module is used to use the deployed decision tree model to perform classification storage prediction on the uploaded object data, and store the object data in the corresponding storage pool according to the classification storage prediction results.
10. A computer device, characterized in that: The computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the method according to any one of claims 1 to 8 when executing the program.
Citation Information
Cited By
Data hierarchical storage method and device based on access characterization, equipment and medium
CN120743183A
Unstructured storage hierarchical strategy optimization method based on machine learning
CN121116940A
File storage method and device and related equipment
CN121578954A