Data life cycle management method and device based on large model, equipment and medium
Through the data lifecycle management method based on the big model, the target value prediction model predicts data value trends and uncertainties, and the data management is automated, solving the problems of resource waste and inefficiency in traditional methods, and achieving efficient and accurate data management.
Patent Information
- Application Number
- CN202510519422.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-25
AI Technical Summary
Traditional data lifecycle management methods rely on simple rules and manual operations, resulting in inefficient redundant data storage, resource waste and management, making it difficult to make quick and reasonable data archiving, deleting or migration decisions.
The data life cycle management method based on the big model is adopted, and the data life cycle management is automatically completed by collecting multi-dimensional feature information, using the target value prediction model to predict data value trends and uncertainty quantitative indicators, and the target management decision is determined based on the value measurement type, so as to automatically complete the data life cycle management.
It realizes efficient and accurate data management, significantly improves the intelligence level of data management, rationalizes storage decisions, and reduces resource waste and manual errors.
Smart Images

Figure CN120371827A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a data life cycle management method, device, equipment and medium based on a large model. Background Art
[0002] In today's digital age, data is being generated and accumulated at an unprecedented rate, and the amount of data is growing exponentially. Enterprises and organizations are facing increasingly severe data management challenges. As an important part of data management, data life cycle management plays a crucial role in ensuring the availability, integrity and security of data, as well as optimizing the use of storage resources.
[0003] In terms of data storage, traditional methods usually rely on simple rules and experience to make storage decisions. This leads to redundant data being randomly stored on expensive storage media, occupying a large amount of storage resources, and some useless data being retained in the storage system for a long time, resulting in a waste of storage space. When storage resources are scarce, it is difficult to quickly and reasonably make decisions on data archiving, deletion or migration, thus affecting the overall performance and efficiency of the storage system.
[0004] In terms of the execution of storage decisions, traditional methods often rely on manual operations, which are inefficient and error-prone. With the continuous increase in the amount of data and the increasing complexity of storage requirements, it has become increasingly difficult to manually formulate and execute storage decisions. Manual operations not only consume a large amount of time and energy, but also are prone to errors, resulting in data loss or waste of storage resources.
[0005] In summary, how to manage the data life cycle more reasonably and efficiently is an issue to be solved in this field. Summary of the Invention
[0006] In view of this, the purpose of the present invention is to provide a data life cycle management method, device, equipment and medium based on a large model, which can manage the data life cycle more reasonably and efficiently. The specific solutions are as follows:
[0007] In the first aspect, the present application discloses a data life cycle management method based on a large model, including:
[0008] Collect multi-dimensional feature information of target data to be managed in the current life cycle;
[0009] Use a target value prediction large model to predict the multi-dimensional feature information to obtain the value trend of the target data and a quantification index of the uncertainty of the value trend; wherein, the training data of the target value prediction large model includes the feature information of historical data and the actual value trend label;
[0010] Determine the target management decision generation logic of the target data according to the value measurement type of the value trend;
[0011] Obtain a target decision corresponding to the value trend and the uncertainty quantification index based on the target management decision generation logic, and perform life cycle management of the target data this time based on the target decision.
[0012] Optionally, obtaining the target value prediction large model includes:
[0013] Construct an initial value prediction large model;
[0014] Divide the training data into each first training data group; wherein, the first training data group includes single-dimensional feature information of historical data and actual value trend labels;
[0015] Successively use each of the first training data groups to train the initial value prediction large model to obtain a first trained large model;
[0016] Divide the training data into second training data groups; wherein, the second training data group includes multi-dimensional feature information of historical data and actual value trend labels;
[0017] Successively use each of the second training data groups to train the first trained large model, and adjust the influence weights between the single-dimensional feature information of the historical data and the actual value trend labels to obtain a second trained large model, and obtain a target value prediction large model based on the second trained large model.
[0018] Optionally, the obtaining the target value prediction large model based on the second trained large model includes:
[0019] Establish a target loss function including a data error term and a regularization term;
[0020] Use the target loss function to update the parameters of the second trained large model to obtain an updated large model that meets the preset low complexity condition;
[0021] Evaluate the contribution degree of each model parameter in the updated large model to obtain the contribution degree scores of each model parameter; wherein, each model parameter includes model weights, neurons, and convolution kernels in the updated large model;
[0022] Prune the updated large model according to the contribution degree scores of each model parameter to obtain a target value prediction large model.
[0023] Optionally, the determining the target management decision generation logic of the target data according to the value measurement type of the value trend includes:
[0024] If the value measurement type of the value trend is a quantitative value type, determine the preset threshold comparison decision generation logic as the target management decision generation logic of the target data;
[0025] Correspondingly, obtaining the target decision corresponding to the value trend and the uncertainty quantification index based on the target management decision generation logic includes:
[0026] Compare the value scores corresponding to each preset time node in the value trend with the value thresholds corresponding to each preset time node in the preset threshold comparison decision generation logic to obtain a first preliminary decision;
[0027] Adjust the first preliminary decision according to the magnitude relationship between the uncertainty quantification index and the uncertainty threshold in the preset threshold comparison decision generation logic to obtain a first target decision.
[0028] Optionally, determining the target management decision generation logic of the target data according to the value measurement type of the value trend includes:
[0029] If the value measurement type of the value trend is a hierarchical value type, determine the preset level - oriented decision generation logic as the target management decision generation logic of the target data;
[0030] Correspondingly, obtaining the target decision corresponding to the value trend and the uncertainty quantification index based on the target management decision generation logic includes:
[0031] Match the target value levels corresponding to each preset time node in the value trend with each preset value level corresponding to each preset time node in the preset level - oriented decision generation logic to obtain a second preliminary decision;
[0032] Adjust the second preliminary decision according to the magnitude relationship between the uncertainty quantification index and the uncertainty threshold in the preset level - oriented decision generation logic to obtain a second target decision.
[0033] Optionally, collecting the multi - dimensional feature information of the target data to be currently managed in the life cycle includes:
[0034] Obtain the target data to be currently managed in the life cycle; wherein, the target data is the data for the first time to be managed in the life cycle obtained based on a preset interface or the data that is not for the first time to be managed in the life cycle in the data storage carrier;
[0035] Collect multi-dimensional feature information of the target data; wherein, the multi-dimensional feature information includes any one or several of data creation time feature, recent access time feature, data size feature, data type feature, associated business process feature, user operation record feature, data update frequency feature, and business requirement feature.
[0036] Optionally, the obtaining of the target decision corresponding to the value trend and the uncertainty quantification index based on the target management decision generation logic includes:
[0037] Screen out the target decision corresponding to the value trend and the uncertainty quantification index from each preset decision of the target management decision generation logic; the types of each preset decision include data deletion type and data storage type, wherein the data storage type includes data archiving type, data migration type, and data hot storage type;
[0038] Correspondingly, the performing of the current life cycle management on the target data based on the target decision includes:
[0039] If the type of the target decision is the data archiving type, perform compression processing and encryption processing on the target data in sequence to obtain processed data, and store the processed data in a preset low-cost storage carrier;
[0040] If the type of the target decision is the data migration type, determine the next storage carrier level based on the current storage carrier level of the target data, and migrate the target data from the current storage carrier to a preset storage carrier corresponding to the next storage carrier level;
[0041] If the type of the target decision is the data hot storage type, save the target data in a preset high-cost storage carrier;
[0042] Generate index information of the target data in the corresponding storage carrier, so as to search for the target data in the corresponding storage carrier based on the index information.
[0043] In a second aspect, the present application discloses a data life cycle management device based on a large model, including:
[0044] A feature collection module, configured to collect multi-dimensional feature information of target data to be currently subjected to life cycle management;
[0045] A value prediction module, configured to use a target value prediction large model to predict the multi-dimensional feature information to obtain the value trend of the target data and the uncertainty quantification index of the value trend; wherein, the training data of the target value prediction large model includes each feature information of historical data and an actual value trend label;
[0046] A logic determination module, configured to determine the target management decision generation logic of the target data according to the value measurement type of the value trend;
[0047] A data management module, configured to obtain a target decision corresponding to the value trend and the uncertainty quantification index based on the target management decision generation logic, and perform the current life cycle management on the target data based on the target decision.
[0048] In a third aspect, the present application discloses an electronic device, including:
[0049] A memory, configured to store a computer program;
[0050] A processor, configured to execute the computer program to implement the steps of the foregoing data life cycle management method based on a large model.
[0051] In a fourth aspect, the present application discloses a computer-readable storage medium, configured to store a computer program; wherein, when the computer program is executed by a processor, the steps of the foregoing data life cycle management method based on a large model are implemented.
[0052] The beneficial effects of this application are as follows: This application collects multi-dimensional characteristic information of target data to be subjected to life cycle management currently; uses a target value prediction large model to predict the multi-dimensional characteristic information to obtain the value trend of the target data and a quantification index of the uncertainty of the value trend; wherein, the training data of the target value prediction large model includes each characteristic information of historical data and an actual value trend label; determines the generation logic of the target management decision for the target data according to the value measurement type of the value trend; obtains a target decision corresponding to the value trend and the uncertainty quantification index based on the generation logic of the target management decision, and performs the current life cycle management on the target data based on the target decision. Thus, it can be seen that this application collects multi-dimensional characteristic information of target data to be subjected to life cycle management currently, and the training data of the target value prediction large model includes each characteristic information of historical data and an actual value trend label, so the multi-dimensional characteristic information can be predicted by using the target value prediction large model, thereby obtaining the value trend of the target data and the quantification index of the uncertainty of the value trend, and further determining the generation logic of the target management decision for the target data according to the value measurement type of the value trend, and obtaining a target decision corresponding to the value trend and the uncertainty quantification index based on the generation logic of the target management decision. That is to say, in the process of generating the target decision for the target data, not only the value dimension of the data is considered, but also a reasonable decision generation logic is determined according to the value measurement type of the data, that is, this application determines a management decision generation logic that is more matched with the value measurement method, so that the generated target decision is more reasonable. Further, the current life cycle management of the target data is automatically completed based on the target decision, ensuring the efficiency and accuracy of data management and significantly improving the intelligent level of data management. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0054] Figure 1 It is a flowchart of a data life cycle management method based on a large model disclosed in this application;
[0055] Figure 2 It is a schematic diagram of a specific data life cycle management disclosed in this application;
[0056] Figure 3 It is a schematic structural diagram of a data life cycle management device based on a large model disclosed in this application;
[0057] Figure 4 A structural diagram of an electronic device disclosed in this application. Specific implementation manners
[0058] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0059] In today's digital age, data is being generated and accumulated at an unprecedented speed, the amount of data is growing exponentially, and enterprises and organizations are facing increasingly severe data management challenges. As an important part of data management, data life cycle management plays a crucial role in ensuring the availability, integrity, and security of data, as well as optimizing the use of storage resources.
[0060] In terms of data storage, traditional methods usually rely on simple rules and experience to make storage decisions. This results in redundant data being randomly stored on expensive storage media, occupying a large amount of storage resources, and some useless data being retained in the storage system for a long time, causing waste of storage space. When storage resources are scarce, it is difficult to make quick and reasonable decisions on data archiving, deletion, or migration, thus affecting the overall performance and efficiency of the storage system.
[0061] In terms of the execution of storage decisions, traditional methods often rely on manual operations, which are inefficient and error-prone. As the amount of data continues to increase and storage requirements become increasingly complex, it has become increasingly difficult to manually formulate and execute storage decisions. Manual operations not only consume a large amount of time and energy but are also prone to errors, resulting in data loss or waste of storage resources.
[0062] Therefore, this application correspondingly provides a data life cycle management solution based on a large model, which can manage the data life cycle more reasonably and efficiently.
[0063] See Figure 1 As shown, the embodiments of this application disclose a data life cycle management method based on a large model, including:
[0064] Step S11: Collect multi-dimensional feature information of the target data to be managed in the current life cycle.
[0065] In this embodiment, collecting the multi-dimensional feature information of the target data to be subject to life cycle management currently includes: obtaining the target data to be subject to life cycle management currently; where the target data is the data that is subject to life cycle management for the first time obtained based on a preset interface or the data that is not subject to life cycle management for the first time in the data storage carrier; collecting the multi-dimensional feature information of the target data; where the multi-dimensional feature information includes any one or several of the feature information such as data creation time feature, recent access time feature, data size feature, data type feature, associated business process feature, user operation record feature, data update frequency feature, and business requirement feature.
[0066] Based on a preset interface, obtain the target data to be subject to life cycle management currently. This target data can be the data that is subject to life cycle management for the first time or the data that is not subject to life cycle management for the first time in the data storage carrier. That is to say, in this embodiment, not only can life cycle management be performed on the unmanaged data, but also life cycle management can be performed again on the data that has been managed. In other words, the target data is the data that needs to be subject to life cycle management. Collect the multi-dimensional feature information of the target data. The multi-dimensional feature information includes any one or several of the feature information such as data creation time feature, recent access time feature, data size feature, data type feature, associated business process feature, user operation record feature, data update frequency feature, and business requirement feature.
[0067] Step S12: Use the target value prediction large model to predict the multi-dimensional feature information to obtain the value trend of the target data and the uncertainty quantification index of the value trend; where the training data of the target value prediction large model includes the feature information of historical data and the actual value trend label.
[0068] It should be noted that the multi-dimensional feature information can also be pre-converted in format to the input format acceptable to the large model. For example, convert the data creation time feature and the recent access time feature into time series features, encode the data type feature, vectorize the text data in the multi-dimensional feature information, etc. That is, use the target value prediction large model to predict the multi-dimensional feature information after format conversion to obtain the value trend of the target data and the uncertainty quantification index of the value trend.
[0069] In this embodiment, obtaining the target value prediction large model includes: constructing an initial value prediction large model; dividing the training data into each first training data group; where the first training data group includes the single-dimensional feature information of historical data and the actual value trend label; sequentially using each of the first training data groups to train the initial value prediction large model to obtain a first trained large model; dividing the training data into second training data groups; where the second training data group includes the multi-dimensional feature information of historical data and the actual value trend label; sequentially using each of the second training data groups to train the first trained large model, and adjusting the influence weight between each single-dimensional feature information of the historical data and the actual value trend label to obtain a second trained large model, and obtaining the target value prediction large model based on the second trained large model.
[0070] It can be understood that the target value prediction large model is obtained in advance, and the training data of the model and the initial value prediction large model need to be obtained separately.
[0071] The training data is the feature information of historical data and the actual value trend label. It should be noted that in the process of obtaining the training data, first, historical data that meets the requirements needs to be obtained. Specifically: strictly screening and cleaning the collected initial historical data, removing the noise data, duplicate data and error data therein. At the same time, in order to ensure the representativeness of the data, historical data is collected from different time periods, different business scenarios and different data sources. That is to say, initial historical data is collected from each time period, each business scenario, each data source, and the initial historical data is data-cleaned to obtain historical data that meets the requirements. Secondly, after collecting historical data that meets the requirements, the historical data needs to be labeled. Specifically, there are two aspects of data labeling. On the one hand, the actual value of the historical data is quantitatively labeled to obtain the actual value trend label. For example, continuous value scores or relative value levels are used to represent the value of the data at different time points. The relative value levels include high, medium, and low. On the other hand, the storage decision of the historical data is labeled. The storage decision types specifically include data deletion type and data storage type. Among them, the data storage type includes data archiving type, data migration type and data hot storage type.
[0072] An initial value prediction large model can be constructed using a hybrid model architecture, combining the advantages of Transformer and GNN (Graph Neural Network) to process multi-dimensional features and capture the complex correlation relationships between data.
[0073] Furthermore, in order to improve the model prediction accuracy in this embodiment, a two-stage training mode is adopted to train the model. First, the training data is divided into each first training data group. The first training data group includes the single-dimensional feature information of historical data and the actual value trend label. That is, the model is trained from a single-dimensional perspective first to obtain the first post-training large model, enabling the model to learn the relationship between the single-dimensional feature information and the actual value trend. In this way, even if the feature information dimension of the target data collected during the actual prediction process is small, value prediction can be achieved. Secondly, the training data is divided into the second training data group. The second training data group includes the multi-dimensional feature information of historical data and the actual value trend label. That is, the model is further trained from a multi-dimensional perspective. Since there may be mutual influences among the multi-dimensional feature information, which in turn affects the actual value, enabling the model to learn the relationship between the multi-dimensional feature information and the actual value trend can improve the model's prediction accuracy. That is to say, the influence weights between each single-dimensional feature information and the actual value trend label in the first post-training large model can be adjusted to obtain the second post-training large model, and the target value prediction large model is obtained based on the second post-training large model. In this way, the finally obtained target value prediction large model can use the multi-dimensional feature information for more accurate value prediction.
[0074] It can be understood that in the process of training the first post-training large model from the initial value prediction large model and the process of training the second post-training large model from the first post-training large model, the model parameters can be adjusted by methods such as the backpropagation algorithm and the gradient descent method to minimize the loss function and improve the model prediction accuracy.
[0075] In this embodiment, obtaining the target value prediction large model based on the second post-training large model includes: establishing a target loss function including a data error term and a regularization term; using the target loss function to update the parameters of the second post-training large model to obtain an updated large model that meets the preset low complexity condition; evaluating the contribution degrees of each model parameter in the updated large model to obtain the contribution degree scores of each model parameter; where each model parameter includes the model weights, neurons, and convolution kernels in the updated large model; pruning the updated large model according to the contribution degree scores of each model parameter to obtain the target value prediction large model.
[0076] After obtaining the second trained large model, although the accuracy of the second trained large model in predicting the value trend is relatively high, the complexity of the second trained large model is also very high. Therefore, in this embodiment, while ensuring the accuracy of model prediction, the model can also be simplified, thereby reducing the computing resources required for value prediction. An objective loss function including a data error term and a regularization term is established, and the objective loss function is used to update the parameters of the second trained large model to obtain an updated large model that meets the preset low-complexity condition, that is, reducing the complexity of the model; the contribution degrees of each model parameter in the updated large model are evaluated to obtain the contribution degree scores of each model parameter; wherein, each model parameter includes model weights, neurons, and convolution kernels in the updated large model; the updated large model is pruned according to the contribution degree scores of each model parameter to obtain the target value prediction large model, that is, retaining the model parameters with relatively high contribution degrees to value trend prediction in the model, thereby further simplifying the model.
[0077] Based on the parameter update and pruning processes in the process of obtaining the target value prediction large model from the second trained large model, methods such as grid search, random search, or Bayesian optimization can be used to find the best hyperparameter configuration, and techniques such as Dropout (random inactivation) and L2 regularization can be applied to reduce the model complexity and improve the model generalization ability. Additionally, the hyperparameters of the model can be adjusted through methods such as grid search or Bayesian optimization. That is to say, the objective loss function is used to update the parameters of the second trained large model. Specifically, it is based on any one or several optimization algorithms among the grid search algorithm, random search algorithm, Bayesian optimization algorithm, random inactivation algorithm, and L2 regularization algorithm, and the objective loss function is used to update the parameters of the second trained large model.
[0078] Step S13: Determine the target management decision generation logic of the target data according to the value measurement type of the value trend.
[0079] There can be two value measurement types for the generated value trend. One value measurement type is the quantitative value type, specifically referring to that the value trend is presented in the form of specific scores. For example, the model predicts that the value score of the target data will gradually decrease from 80 points to 60 points in the next month, indicating that the value of this data is gradually decreasing. The other value measurement type is the hierarchical value type, such as high, medium, and low. This presentation method is more intuitive and convenient for making decisions quickly.
[0080] In the first specific embodiment, the determining the target management decision generation logic of the target data according to the value measurement type of the value trend includes: if the value measurement type of the value trend is the quantitative value type, then determine the preset threshold comparison decision generation logic as the target management decision generation logic of the target data.
[0081] When the value measurement type of the value trend is the quantified value type, the preset threshold comparison decision generation logic is determined as the target management decision generation logic for the target data. That is to say, next, in the process of generating the target decision, it is necessary to generate the target decision based on the preset threshold comparison decision generation logic.
[0082] In the second specific embodiment, the determining of the target management decision generation logic for the target data according to the value measurement type of the value trend includes: if the value measurement type of the value trend is the hierarchical value type, the preset level-oriented decision generation logic is determined as the target management decision generation logic for the target data.
[0083] When the value measurement type of the value trend is the hierarchical value type, the preset level-oriented decision generation logic is determined as the target management decision generation logic for the target data. That is to say, next, in the process of generating the target decision, it is necessary to generate the target decision based on the preset level-oriented decision generation logic.
[0084] Step S14: Obtain the target decision corresponding to the value trend and the uncertainty quantification index based on the target management decision generation logic, and perform the life cycle management of the target data for this time based on the target decision.
[0085] In this embodiment, the obtaining of the target decision corresponding to the value trend and the uncertainty quantification index based on the target management decision generation logic includes: screening out the target decision corresponding to the value trend and the uncertainty quantification index from each preset decision of the target management decision generation logic; the types of each preset decision include the data deletion type and the data storage type, where the data storage type includes the data archiving type, the data migration type, and the data hot storage type.
[0086] From the preset decisions in the generation logic of the target management decision, screen out the target decisions corresponding to the value trend and the uncertainty quantification index. That is to say, determine the life cycle management method corresponding to the value trend and the uncertainty quantification index, and then complete the life cycle management of the target data. It can be understood that this embodiment includes multiple life cycle management methods, and the levels corresponding to different life cycle management methods are also different, that is, the levels corresponding to each preset decision are also different. In this way, different life cycle management methods can be implemented for the data, which is more flexible and targeted. Further, the types of each preset decision are mainly divided into two categories, namely the data deletion type and the data storage type. The life cycle management of the data deletion type is to delete the target data. Among them, the data storage type includes the data archiving type, the data migration type, and the data hot storage type. That is to say, even when performing the life cycle management of storing the target data, the specific methods are also different, including the data archiving type, the data migration type, and the data hot storage type.
[0087] In this embodiment, the life cycle management of the target data based on the target decision includes: if the type of the target decision is the data archiving type, sequentially perform compression processing and encryption processing on the target data to obtain processed data, and store the processed data in a preset low-cost storage carrier; if the type of the target decision is the data migration type, determine the next storage carrier level based on the current storage carrier level of the target data, and migrate the target data from the current storage carrier to a preset storage carrier corresponding to the next storage carrier level; if the type of the target decision is the data hot storage type, save the target data in a preset high-cost storage carrier; generate index information of the target data in the corresponding storage carrier, so as to search for the target data in the corresponding storage carrier based on the index information.
[0088] If the type of the target decision is the data archiving type, sequentially perform compression processing and encryption processing on the target data to obtain processed data, and store the processed data in a preset low-cost storage carrier. The life cycle management corresponding to the data archiving type includes processing such as sorting, compressing, and encrypting the data to reduce the occupancy of storage space and improve the security of the data. At the same time, in order to facilitate subsequent query and retrieval, detailed index information will be established to record key information such as the archiving time and storage location of the data.
[0089] If the type of the target decision is data migration type, determine the next storage carrier level based on the current storage carrier level of the target data, and migrate the target data from the current storage carrier to a preset storage carrier corresponding to the next storage carrier level. Data migration refers to migrating the target data from the current storage carrier to the next storage carrier, and the current storage carrier level is different from the next storage carrier level. Specifically, data migration can be from a higher storage carrier level to a lower storage carrier level, or from a lower storage carrier level to a higher storage carrier level.
[0090] If the type of the target decision is data hot storage type, save the target data to a preset high-cost storage carrier; generate index information of the target data in the corresponding storage carrier, so as to search for the target data in the corresponding storage carrier based on the index information. When the value of the target data is relatively high, the type of the target decision is data hot storage type, and the target data is saved to a preset high-cost storage carrier, and index information of the target data in the corresponding storage carrier is generated, so as to search for the target data in the corresponding storage carrier based on the index information, improving the subsequent data search efficiency.
[0091] In the first specific embodiment, the obtaining of the target decision corresponding to the value trend and the uncertainty quantification index based on the target management decision generation logic includes: comparing the value scores corresponding to each preset time node in the value trend with the value thresholds corresponding to each preset time node in the preset threshold comparison decision generation logic to obtain a first preliminary decision; adjusting the first preliminary decision according to the magnitude relationship between the uncertainty quantification index and the uncertainty threshold in the preset threshold comparison decision generation logic to obtain a first target decision.
[0092] When the target management decision generation logic is the preset threshold comparison decision generation logic, compare the value scores corresponding to each preset time node in the value trend with the value thresholds corresponding to each preset time node in the preset threshold comparison decision generation logic to obtain a first preliminary decision. Specifically, it can also be to compare the value score with the corresponding value threshold under the preset critical time node. The value threshold can be set according to the specific actual situation. For example, in the preset time nodes, if the value scores are all higher than the value threshold, it indicates that the value of the target data is relatively high. Then the first preliminary decision indicates that the target data needs to be stored at a high level, such as the data hot storage type; adjust the first preliminary decision according to the magnitude relationship between the uncertainty quantification index and the uncertainty threshold in the preset threshold comparison decision generation logic to obtain a first target decision. Among them, the uncertainty quantification index can specifically be the confidence level. For example, if the uncertainty quantification index is 90% and is greater than the uncertainty threshold of 75%, it means that the value scores corresponding to each preset time node in the value trend are very reliable. Also, because the data hot storage type is already the highest data life cycle management, the first preliminary decision can be directly determined as the first target decision. Another example is that if the uncertainty quantification index is 50% and is less than the uncertainty threshold of 75%, it means that the value scores corresponding to each preset time node in the value trend are less reliable. If the highest data life cycle management is adopted, it may lead to waste of storage resources. Therefore, the first preliminary decision can be adjusted. For example, a first target decision of the data archiving type can be obtained. It can be understood that there can be multiple value thresholds and uncertainty thresholds in the preset threshold comparison decision generation logic, so that more detailed storage decisions can be obtained.
[0093] In the second specific embodiment, obtaining the target decision corresponding to the value trend and the uncertainty quantification index based on the target management decision generation logic includes: matching the target value levels corresponding to each preset time node in the value trend with the preset value levels corresponding to each preset time node in the preset level-oriented decision generation logic to obtain a second preliminary decision; adjusting the second preliminary decision according to the magnitude relationship between the uncertainty quantification index and the uncertainty threshold in the preset level-oriented decision generation logic to obtain a second target decision.
[0094] When the target management decision generation logic is the preset level-oriented decision generation logic, match the target value levels corresponding to each preset time node in the value trend with each preset value level corresponding to each preset time node in the preset level-oriented decision generation logic to obtain a second preliminary decision. Among them, the preset value levels in the preset level-oriented decision generation logic mainly include high, medium, and low. For example, at a preset time node, the matching of the target value level and the preset value level is low. Therefore, the type of the second preliminary decision can be the data deletion type; adjust the second preliminary decision according to the magnitude relationship between the uncertainty quantification index and the uncertainty threshold in the preset level-oriented decision generation logic to obtain a second target decision. The uncertainty quantification index is used to characterize the reliability of the value trend predicted by the model. The uncertainty quantification index of the value trend can specifically be the confidence level. Then, the higher the confidence level, the more reliable the value trend is, and there is no need to adjust the second preliminary decision. On the contrary, if the confidence level is lower, it indicates that the second preliminary decision needs to be adjusted. Then, if the data life cycle management level corresponding to the second preliminary decision is relatively high, at this time, it is necessary to adjust it in the direction of a lower data life cycle management level. It should be noted that when adjusting in the direction of a lower data life cycle management level, the data deletion type will not be selected, and the target data will still be stored, but the carrier on which the target data is stored may have a lower cost. If the data life cycle management level corresponding to the second preliminary decision is relatively low, at this time, it is necessary to adjust it in the direction of a higher data life cycle management level; it can be understood that there can also be multiple uncertainty thresholds in the preset level-oriented decision generation logic.
[0095] The beneficial effects of this application are as follows: This application collects multi-dimensional feature information of target data to be subject to lifecycle management currently; uses a target value prediction large model to predict the multi-dimensional feature information to obtain the value trend of the target data and the uncertainty quantification index of the value trend; wherein, the training data of the target value prediction large model includes the feature information of historical data and the actual value trend label; determines the target management decision generation logic of the target data according to the value measurement type of the value trend; obtains a target decision corresponding to the value trend and the uncertainty quantification index based on the target management decision generation logic, and performs the current lifecycle management of the target data based on the target decision. It can be seen that this application collects multi-dimensional feature information of target data to be subject to lifecycle management currently, and the training data of the target value prediction large model includes the feature information of historical data and the actual value trend label, so the multi-dimensional feature information can be predicted by the target value prediction large model to obtain the value trend of the target data and the uncertainty quantification index of the value trend, and further determine the target management decision generation logic of the target data according to the value measurement type of the value trend, and obtain a target decision corresponding to the value trend and the uncertainty quantification index based on the target management decision generation logic. That is to say, in the process of generating the target decision of the target data, not only the value dimension of the data is considered, but also a reasonable decision generation logic is determined according to the value measurement type of the data, that is, this application determines a management decision generation logic that is more matched with the value measurement method, so that the generated target decision is more reasonable. Further, the current lifecycle management of the target data is automatically completed based on the target decision, ensuring the efficiency and accuracy of data management and significantly improving the intelligent level of data management.
[0096] The following takes Figure 2Taking a specific data lifecycle management schematic diagram shown as an example, the present application is explained accordingly. In the data collection and processing part, business data is collected and processed, which may include data cleaning, data conversion, data classification and data labeling. Next, in the large model part, the model is trained and tuned using the collected data, and then the target value prediction large model can be used to complete the decision generation part and the lifecycle management part. The decision generation part and the lifecycle management part can be specifically: when it is predicted that the data value will continue to decrease in the future and is lower than a certain threshold, an automatic archiving operation is triggered; when the data value is extremely low and has not been used for a long time, a deletion operation is performed; when storage resources are tight and the value of some data is low, the data is migrated to a lower-cost storage medium. These storage decisions will be dynamically adjusted according to different business scenarios and storage resource conditions to achieve the purpose of resource optimization. From the data value dimension, when it is predicted that the data value will continue to decrease in the future and is lower than a certain threshold, the corresponding storage operation needs to be triggered. The setting of this threshold is not fixed, but will be based on different business scenarios. The data value is dynamically adjusted according to the type of data. When it is predicted that the data value will continue to decrease in the future and is lower than a certain threshold, the automatic archiving operation is triggered, including sorting, compressing and encrypting the data to reduce the storage space occupied and improve the security of the data. At the same time, in order to facilitate subsequent query and retrieval, detailed index information will be established to record key information such as the archiving time and storage location of the data. When the data value is extremely low and has not been used for a long time, the deletion operation is performed. Before deleting the data, the system will conduct strict review and confirmation. At the same time, in order to prevent accidental deletion, a certain deletion log will be retained to record the deleted data information and deletion time so that it can be traced and restored when needed. When storage resources are tight and the value of some data is low, the data will be migrated to a storage medium with lower cost. The data migration process needs to ensure the integrity and consistency of the data while minimizing the impact on the business. During the migration process, incremental migration, parallel migration and other technologies will be used to improve migration efficiency. At the same time, data verification will be performed to ensure that the migrated data is consistent with the original data.
[0097] See also Figure 3 As shown, the embodiment of the present application discloses a data lifecycle management device based on a large model, including:
[0098] A feature collection module 11 is used to collect multi-dimensional feature information of target data to be currently managed for lifecycle;
[0099] A value prediction module 12 is configured to use a target value prediction large model to predict the multi-dimensional feature information, so as to obtain the value trend of the target data and a quantification index of the uncertainty of the value trend; wherein, the training data of the target value prediction large model includes the feature information of historical data and actual value trend labels;
[0100] A logic determination module 13 is configured to determine the generation logic of the target management decision for the target data according to the value measurement type of the value trend;
[0101] A data management module 14 is configured to obtain a target decision corresponding to the value trend and the uncertainty quantification index based on the generation logic of the target management decision, and perform the current life cycle management on the target data based on the target decision.
[0102] The beneficial effects of this application are as follows: This application collects the multi-dimensional feature information of the target data to be subjected to life cycle management currently; uses a target value prediction large model to predict the multi-dimensional feature information, so as to obtain the value trend of the target data and a quantification index of the uncertainty of the value trend; wherein, the training data of the target value prediction large model includes the feature information of historical data and actual value trend labels; determines the generation logic of the target management decision for the target data according to the value measurement type of the value trend; obtains a target decision corresponding to the value trend and the uncertainty quantification index based on the generation logic of the target management decision, and performs the current life cycle management on the target data based on the target decision. It can be seen that this application collects the multi-dimensional feature information of the target data to be subjected to life cycle management currently, and the training data of the target value prediction large model includes the feature information of historical data and actual value trend labels, so the multi-dimensional feature information can be predicted by using the target value prediction large model, thereby obtaining the value trend of the target data and a quantification index of the uncertainty of the value trend, and further determining the generation logic of the target management decision for the target data according to the value measurement type of the value trend, and obtaining a target decision corresponding to the value trend and the uncertainty quantification index based on the generation logic of the target management decision. That is to say, in the process of generating the target decision for the target data, not only the value dimension of the data is considered, but also a reasonable decision generation logic is determined according to the value measurement type of the data, that is, this application determines a management decision generation logic that is more compatible with the value measurement method, so that the generated target decision is more reasonable. Further, the current life cycle management of the target data is automatically completed based on the target decision, ensuring the efficiency and accuracy of data management, and significantly improving the intelligent level of data management.
[0103] Furthermore, an embodiment of this application also provides an electronic device. Figure 4It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure shall not be regarded as any limitation on the scope of use of this application.
[0104] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of this application. Specifically, it may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the data life cycle management method based on a large model executed by an electronic device disclosed in any of the foregoing embodiments.
[0105] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device; the communication interface 24 can create a data transmission channel between the electronic device and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of this application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.
[0106] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 can be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 21 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for the rendering and drawing of the content to be displayed on the display screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0107] In addition, as a carrier for resource storage, the memory 22 can be a read-only memory, a random access memory, a magnetic disk, an optical disc, etc. The resources stored thereon include an operating system 221, a computer program 222, data 223, etc. The storage method can be temporary storage or permanent storage.
[0108] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device, so as to enable the processor 21 to perform operations and processing on the massive data 223 in the memory 22. It can be Windows, Unix, Linux, etc. In addition to the computer program that can be used to complete the data life cycle management method based on the large model executed by the electronic device disclosed in any of the foregoing embodiments, the computer program 222 can further include computer programs that can be used to complete other specific tasks. The data 223 can include not only the data transmitted by external devices received by the electronic device, but also the data collected by its own input / output interface 25, etc.
[0109] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the data life cycle management method based on the large model disclosed above. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.
[0110] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0111] Those skilled in the art may further realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application. The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable EPROM (Erasable Programmable Read Only Memory), electrically erasable programmable EEPROM (Electrically Erasable Programmable read only memory), registers, hard disk, removable disk, CD-ROM (Compact Disc Read-Only Memory), or any other form of storage medium well-known in the technical field.
[0112] Finally, it should also be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0113] The above has introduced in detail a data life cycle management method, apparatus, device, and medium provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A data life cycle management method based on a large model, characterized in that, Including: Collecting multi-dimensional feature information of target data to be subjected to lifecycle management currently; Using a target value prediction large model to predict the multi-dimensional feature information to obtain a value trend of the target data and a quantification index of the uncertainty of the value trend; wherein, training data of the target value prediction large model includes each feature information of historical data and an actual value trend label; Determining a target management decision generation logic of the target data according to a value measurement type of the value trend; Obtaining a target decision corresponding to the value trend and the uncertainty quantification index based on the target management decision generation logic, and performing this lifecycle management on the target data based on the target decision.
2. The data life cycle management method based on a large model according to claim 1, wherein, Obtaining the target value prediction large model includes: Constructing an initial value prediction large model; Dividing the training data into each first training data group; wherein, the first training data group includes single-dimensional feature information of historical data and an actual value trend label; Successively using each of the first training data groups to train the initial value prediction large model to obtain a first post-training large model; Dividing the training data into second training data groups; wherein, the second training data group includes multi-dimensional feature information of historical data and an actual value trend label; Successively using each of the second training data groups to train the first post-training large model, and adjusting an influence weight between each single-dimensional feature information of the historical data and the actual value trend label to obtain a second post-training large model, and obtaining a target value prediction large model based on the second post-training large model.
3. The data life cycle management method based on a large model according to claim 2, wherein The obtaining the target value prediction large model based on the second post-training large model includes: Establishing a target loss function including a data error term and a regularization term; Using the target loss function to update parameters of the second post-training large model to obtain an updated large model that meets a preset low complexity condition; Evaluating contribution degrees of each model parameter in the updated large model to obtain contribution degree scores of each model parameter; wherein, each model parameter includes a model weight, a neuron, and a convolution kernel in the updated large model; Pruning the updated large model according to the contribution degree scores of each model parameter to obtain a target value prediction large model.
4. The data life cycle management method based on a large model according to claim 1, characterized in that The determining the target management decision generation logic of the target data according to the value measurement type of the value trend includes: If the value measurement type of the value trend is a quantified value type, determining a preset threshold comparison decision generation logic as the target management decision generation logic of the target data; Correspondingly, the obtaining the target decision corresponding to the value trend and the uncertainty quantification index based on the target management decision generation logic includes: Comparing value scores corresponding to each preset time node in the value trend with value thresholds corresponding to each preset time node in the preset threshold comparison decision generation logic to obtain a first preliminary decision; Adjust the first preliminary decision according to the magnitude relationship between the uncertainty quantification index and the uncertainty threshold in the comparison decision generation logic of the preset threshold to obtain a first target decision.
5. The data life cycle management method based on a large model according to claim 1, characterized in that The generation logic of the target management decision for the target data determined according to the value measurement type of the value trend includes: If the value measurement type of the value trend is a hierarchical value type, determine the preset level-oriented decision generation logic as the generation logic of the target management decision for the target data; Correspondingly, obtaining the target decision corresponding to the value trend and the uncertainty quantification index based on the generation logic of the target management decision includes: Match the target value levels corresponding to the preset time nodes in the value trend with the preset value levels corresponding to the preset time nodes in the preset level-oriented decision generation logic to obtain a second preliminary decision; Adjust the second preliminary decision according to the magnitude relationship between the uncertainty quantification index and the uncertainty threshold in the preset level-oriented decision generation logic to obtain a second target decision.
6. The data life cycle management method based on a large model according to claim 1, wherein Collecting the multi-dimensional feature information of the target data to be subjected to life cycle management currently includes: Obtain the target data to be subjected to life cycle management currently; wherein, the target data is the data for the first life cycle management obtained based on a preset interface or the data that is not for the first life cycle management in the data storage carrier; Collect the multi-dimensional feature information of the target data; wherein, the multi-dimensional feature information includes any one or several of the data creation time feature, the most recent access time feature, the data size feature, the data type feature, the associated business process feature, the user operation record feature, the data update frequency feature, and the business requirement feature.
7. The data life cycle management method based on a large model according to any one of claims 1 to 6, characterized in that Obtaining the target decision corresponding to the value trend and the uncertainty quantification index based on the generation logic of the target management decision includes: Screen out the target decision corresponding to the value trend and the uncertainty quantification index from each preset decision in the generation logic of the target management decision; the types of each preset decision include a data deletion type and a data storage type, wherein the data storage type includes a data archiving type, a data migration type, and a data hot storage type; Correspondingly, performing the current life cycle management on the target data based on the target decision includes: If the type of the target decision is the data archiving type, perform compression processing and encryption processing on the target data in sequence to obtain processed data, and store the processed data in a preset low-cost storage carrier; If the type of the target decision is the data migration type, determine the next storage carrier level based on the current storage carrier level of the target data, and migrate the target data from the current storage carrier to a preset storage carrier corresponding to the next storage carrier level; If the type of the target decision is the data hot storage type, save the target data in a preset high-cost storage carrier; Generate index information of the target data in the corresponding storage medium, so as to search for the target data in the corresponding storage medium based on the index information.
8. A data life cycle management device based on a large model, characterized in that, It includes: A feature collection module, configured to collect multi-dimensional feature information of the target data to be subject to life cycle management currently; A value prediction module, configured to use a target value prediction large model to predict the multi-dimensional feature information to obtain the value trend of the target data and a quantization index of the uncertainty of the value trend; wherein, the training data of the target value prediction large model includes the feature information of historical data and the actual value trend label; A logic determination module, configured to determine the generation logic of the target management decision of the target data according to the value measurement type of the value trend; A data management module, configured to obtain a target decision corresponding to the value trend and the uncertainty quantization index based on the target management decision generation logic, and perform the current life cycle management on the target data based on the target decision.
9. An electronic device, characterized in that, It includes: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the steps of the data life cycle management method based on a large model according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, For storing a computer program; wherein, when the computer program is executed by a processor, the steps of the data life cycle management method based on a large model according to any one of claims 1 to 7 are implemented.