Method and system for intelligent operation of full life cycle of data assets
By generating initial distributed storage strategies, predicting data retrieval probabilities, and performing real-time optimization, the problem of unreasonable data storage layout was solved, data access efficiency and rational allocation of storage resources were improved, storage strategies were ensured to adapt to changes in business needs, and the operational efficiency and value of data assets were enhanced.
Patent Information
- Application Number
- CN202510739346.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2045-06-04
AI Technical Summary
Existing data asset operation methods make it difficult to generate reasonable initial distributed storage strategies, resulting in unreasonable data storage layouts, affecting data access speed and storage costs, and making it impossible to predict data usage in different business scenarios. Consequently, storage strategies cannot adapt to changes in business needs in a timely manner, reducing the operational efficiency and value of data assets.
An initial distributed storage strategy is generated based on data asset inventory vectors and key identifier vectors for directory retrieval. The probability of joint calls for each type of joint data in different business scenarios is predicted. The initial strategy is improved based on the importance level of the business scenario, and the storage strategy is optimized in real time based on the actual call records in the latest cycle.
It enables data storage strategies to proactively match business needs, improve data access efficiency, rationally allocate storage resources, ensure that storage strategies are adjusted in a timely manner as business changes occur, and continuously improve the operational efficiency and value of data assets.
Smart Images

Figure CN120811900B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data asset operation, in particular to a full life cycle intelligent operation method and system of data assets. BACKGROUND
[0002] In the digital era, data has become a crucial asset for enterprises, like the "blood" of enterprise operation, containing tremendous value. With the rapid development of information technology and the increasing complexity of enterprise business, the scale of data assets is growing explosively, and their types are becoming increasingly diversified, covering structured, semi-structured and unstructured data, etc. The full life cycle intelligent operation of data assets is of great significance to enterprises. From the generation, storage, use to the final archiving or deletion of data, effective management and operation can help enterprises fully tap the value of data, provide strong support for decision-making, and improve the competitiveness of enterprises. Through intelligent operation, enterprises can achieve rational storage of data assets, improve data access efficiency, and reduce storage costs; they can also accurately predict data usage according to business needs, optimize data allocation and application, and make data better serve various business scenarios. In the context of the continuous development of big data, artificial intelligence and other technologies, the full life cycle intelligent operation method and system of data assets have broad application prospects and are expected to become a key driving force for enterprise digital transformation, promoting sustainable development of enterprises in the digital economic tide.
[0003] However, the existing data asset operation method is difficult to generate a reasonable initial distributed storage strategy, resulting in unreasonable data storage layout, affecting data access speed and storage cost. And because it is impossible to predict the usage of data in different business scenarios, it is difficult to plan the allocation of data resources in advance. Subsequent improvement of the initial generated storage layout is also difficult, making it difficult to form an actual storage layout that meets the actual business needs of enterprises. In addition, it is also unable to cope with the dynamic changes in data usage, and cannot optimize the actual storage layout of data assets in real time, resulting in storage strategies that cannot adapt to changes in business needs in a timely manner, reducing the operation efficiency and value of data assets.
[0004] Therefore, the present application proposes a full life cycle intelligent operation method and system of data assets. SUMMARY
[0005] The application provides a full life cycle intelligent operation method and system of data assets, which generates an initial distributed storage strategy based on a data asset inventory vector and a directory search key identification vector, can preliminarily plan data asset storage layout, effectively integrates various data assets of an enterprise, and lays a foundation for subsequent operation. The joint call probability of each kind of joint data in different business scenarios is predicted through the same vector, which helps the enterprise to understand the data use trend in advance, so that the data storage strategy is more forward-looking and better matches the business demand. In combination with the importance level of the business scenario and the joint call probability, the initial strategy is improved, and the actual distributed storage strategy obtained can more reasonably allocate storage resources according to the business priority and data call possibility, and improve the data access efficiency. Finally, the actual distributed storage strategy is optimized in real time according to the actual call records in the latest period, so that the storage strategy is adjusted in time with the change of business, the data asset operation efficiency is continuously improved, and the value of data assets is maximized.
[0006] The application provides a full life cycle intelligent operation method of data assets, comprising:
[0007] Based on the data asset inventory vector and the directory search key identification vector of all data asset blocks of an enterprise, an initial distributed storage strategy of all data asset blocks of the enterprise is generated;
[0008] Based on the data asset inventory vector and the directory search key identification vector of all data asset blocks of an enterprise, the joint call probability of each kind of joint data of all data asset blocks of the enterprise in each business scenario is predicted;
[0009] Based on the importance level of each business scenario and the joint call probability of each kind of joint data in each business scenario, the initial distributed storage strategy is improved to obtain an actual distributed storage strategy;
[0010] Based on the actual call records of all data asset blocks in the latest period, the actual distributed storage strategy is optimized in real time to obtain a latest distributed storage strategy.
[0011] Optionally, the acquisition process of the data asset inventory vector and the directory search key identification vector of all data asset blocks of an enterprise comprises:
[0012] The multi-dimensional asset classification attributes and multi-dimensional right information of each data asset block of the enterprise are determined, and the data asset inventory vector of each data asset block is generated based on the multi-dimensional asset classification attributes and multi-dimensional right information of each data asset block;
[0013] The target search key identification of each data asset block is generated based on the multi-dimensional search basis of each asset data block of the enterprise.
[0014] Optionally, based on the data asset inventory vector and the directory search key identification vector of all data asset blocks of the enterprise, an initial distributed storage strategy of all data asset blocks of the enterprise is generated, including:
[0015] Based on the data asset inventory vector and the directory search key identification vector of all data asset blocks of the enterprise, all data asset blocks of the enterprise are divided to obtain all first effective division clusters and all second effective division clusters.
[0016] Based on the value evaluation value of all data asset blocks, an average value evaluation value of each first effective division cluster and an average value evaluation value of each second effective division cluster are determined.
[0017] According to the average value evaluation value of all first effective division clusters from large to small, all first effective division clusters are sorted to obtain the sorting value of all first effective division clusters, and at the same time, according to the value evaluation value of all data asset blocks in each first effective division cluster from large to small, all data asset blocks in each first effective division cluster are sorted to obtain the sorting value of each data asset block in each first effective division cluster, and based on the sorting value of each data asset block in each first effective division cluster and the sorting value of the corresponding first effective division cluster, a first division position two-dimensional array of each data asset block is generated.
[0018] According to the average value evaluation value of all second effective division clusters from large to small, all second effective division clusters are sorted to obtain the sorting value of all second effective division clusters, and at the same time, according to the value evaluation value of all data asset blocks in each second effective division cluster from large to small, all data asset blocks in each second effective division cluster are sorted to obtain the sorting value of each data asset block in each second effective division cluster, and based on the sorting value of each data asset block in each second effective division cluster and the sorting value of the corresponding second effective division cluster, a second division position two-dimensional array of each data asset block is generated.
[0019] Based on the first division position two-dimensional array and the second division position two-dimensional array of each data asset block, a two-dimensional division position pointing vector of each data asset block is determined, and based on the two-dimensional division position pointing vector of all data asset blocks, an initial distributed storage strategy of all data asset blocks of the enterprise is generated.
[0020] Optionally, based on the data asset inventory vector and the directory search key identification vector of all data asset blocks of the enterprise, all data asset blocks of the enterprise are divided to obtain all first effective division clusters and all second effective division clusters, including:
[0021] The data asset blocks of the enterprise are randomly divided to obtain a plurality of first data asset block clusters, and the similarity of the data asset inventory vectors of each two data asset blocks in each first data asset block cluster is calculated, the average of the similarity of the data asset inventory vectors of each two first data asset block clusters in each first data asset block cluster is taken as the intra-cluster cohesion of each first data asset block cluster, the average of the intra-cluster cohesion of all first data asset block clusters obtained in the current division process is taken as the division performance value of the current division process, and it is judged whether the division performance value of the current division process is greater than a preset division performance threshold value; if yes, all first data asset block clusters in the current division process are taken as all first effective division clusters; otherwise, the data asset blocks of the enterprise are continuously randomly divided until the division performance value of the latest division process is greater than the preset division performance threshold value, and then all newly obtained second data asset block clusters are taken as second effective division clusters.
[0022] The data asset blocks of the enterprise are randomly divided to obtain a plurality of second data asset block clusters, and the similarity of the directory search key identification vectors of each two data asset blocks in each second data asset block cluster is calculated, the average of the similarity of the directory search key identification vectors of each two second data asset block clusters in each second data asset block cluster is taken as the intra-cluster cohesion of each second data asset block cluster, the average of the intra-cluster cohesion of all second data asset block clusters obtained in the current division process is taken as the division performance value of the current division process, and it is judged whether the division performance value of the current division process is greater than a preset division performance threshold value; if yes, all second data asset block clusters in the current division process are taken as all second effective division clusters; otherwise, the data asset blocks of the enterprise are continuously randomly divided until the division performance value of the latest division process is greater than the preset division performance threshold value, and then all newly obtained second data asset block clusters are taken as second effective division clusters.
[0023] Optionally, a two-dimensional division position pointing vector of each data asset block is determined based on the first division position two-dimensional array and the second division position two-dimensional array of each data asset block, and an initial distributed storage strategy of all data asset blocks of the enterprise is generated based on the two-dimensional division position pointing vectors of all data asset blocks, including:
[0024] The first division position point and the second division position point of each data asset block in a preset coordinate system are respectively determined based on the first division position two-dimensional array and the second division position two-dimensional array of each data asset block;
[0025] The vector from the first division position point to the second division position point of each data asset block in the preset coordinate system is taken as the two-dimensional division position pointing vector of each data asset block;
[0026] Based on the similarity between the two-dimensional division position pointing vectors of each two data asset blocks, an initial distributed storage network containing all storage nodes of all data asset blocks is constructed as an initial distributed storage strategy of all data asset blocks of the enterprise.
[0027] Optionally, based on the data asset inventory vector and the directory retrieval key identification vector of all data asset blocks of the enterprise, the joint calling probability of each joint data of all data asset blocks of the enterprise in each business scenario is predicted, including:
[0028] The data asset inventory calling range vector and the target retrieval key identification calling range vector in each business scenario are determined.
[0029] The ratio of the total number of elements in the data asset inventory vector of each data asset block that meet the value range of the same element position in the data asset inventory calling range vector in each business scenario to the total number of all elements in the data asset inventory vector is taken as the first calling probability of each data asset block in the corresponding business scenario.
[0030] The ratio of the total number of elements in the target retrieval key identification vector of each data asset block that meet the value range of the same element position in the target retrieval key identification calling range vector in each business scenario to the total number of all elements in the target retrieval key identification vector is taken as the second calling probability of each data asset block in the corresponding business scenario.
[0031] Based on the first calling probability and the second calling probability of each data asset block in each business scenario, the joint calling probability of each joint data of all data asset blocks of the enterprise in each business scenario is predicted.
[0032] Optionally, based on the first calling probability and the second calling probability of each data asset block in each business scenario, the joint calling probability of each joint data of all data asset blocks of the enterprise in each business scenario is predicted, including:
[0033] The mean of the first calling probability and the second calling probability of each data asset block in each business scenario is taken as the comprehensive calling probability of each data asset block in each business scenario.
[0034] The product of the comprehensive calling probabilities of all data asset blocks contained in each joint data of all data asset blocks of the enterprise in the same business scenario is taken as the joint calling probability of the corresponding joint data in the corresponding business scenario.
[0035] Optionally, based on the importance level of each business scenario and the joint call probability of each joint data in each business scenario, the initial distributed storage strategy is improved to obtain an actual distributed storage strategy, including:
[0036] Based on the importance level of all kinds of business scenarios and the joint call probability of all kinds of joint data involved by each data asset block in all kinds of business scenarios, the total call probability of each data asset block is obtained;
[0037] Based on the similarity between the total call probabilities of all pairs of data asset blocks, the initial distributed storage strategy is improved to obtain the actual distributed storage strategy.
[0038] Optionally, based on the actual call records of all data asset blocks in the latest period, the actual distributed storage strategy is optimized in real time to obtain a latest distributed storage strategy, including:
[0039] Based on the actual call records of all data asset blocks in the latest period, the actual call frequency of each data asset block in the latest period is determined;
[0040] Based on the similarity between the actual call frequencies of all pairs of data asset blocks in the latest period, the actual distributed storage strategy is optimized in real time to obtain the latest distributed storage strategy.
[0041] The present application provides a kind of data asset's whole life cycle intelligent operation system, including:
[0042] Initial storage strategy generation module, for based on the data asset inventory vector and directory search key identification vector of all data asset blocks of enterprise, generate the initial distributed storage strategy of all data asset blocks of enterprise;
[0043] Joint call probability determination module, for based on the data asset inventory vector and directory search key identification vector of all data asset blocks of enterprise, predict the joint call probability of each joint data of all data asset blocks of enterprise in each business scenario;
[0044] Initial storage strategy improvement module, for based on the importance level of each business scenario and the joint call probability of each joint data in each business scenario, the initial distributed storage strategy is improved to obtain an actual distributed storage strategy;
[0045] Storage strategy real-time optimization module, for based on the actual call records of all data asset blocks in the latest period, the actual distributed storage strategy is optimized in real time to obtain a latest distributed storage strategy.
[0046] The beneficial effects generated by the present application relative to the prior art are: based on the data asset inventory vector and the directory retrieval key identification vector, an initial distributed storage strategy is generated, which can preliminarily plan the data asset storage layout, effectively integrate various types of data assets of the enterprise, and lay a foundation for subsequent operation. By the same vector, the joint call probability of each kind of joint data in different business scenarios is predicted, which helps the enterprise to understand the data usage trend in advance, so that the data storage strategy is more forward-looking and better matches the business needs. In combination with the importance level of the business scenario and the joint call probability, the initial strategy is improved, and the actual distributed storage strategy obtained can more reasonably allocate storage resources according to the business priority and data call probability, and improve the data access efficiency. Finally, the actual distributed storage strategy is optimized in real time according to the actual call record in the latest period, so as to ensure that the storage strategy is adjusted in time with the change of business, continuously improves the operation efficiency of data assets, and maximizes the value of data assets.
[0047] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the present application. The objects and other advantages of the present application can be realized and attained by the structure particularly pointed out in the specification.
[0048] The technical solutions of the present application will be further described in detail below with the help of the accompanying drawings and examples. BRIEF DESCRIPTION OF DRAWINGS
[0049] The accompanying drawings are included to provide a further understanding of the present application, and constitute a part of the specification, and are used to explain the present application together with embodiments of the present application, and do not constitute a limitation on the present application. In the drawings:
[0050] Figure 1 The flow chart of the full life cycle intelligent operation method of data assets in the embodiments of the present application. DETAILED DESCRIPTION
[0051] The preferred embodiments of the present application will be described below in conjunction with the accompanying drawings, and it should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and do not limit the present application.
[0052] As shown in Figure 1 The present application provides an embodiment of a full life cycle intelligent operation method of data assets, which comprises:
[0053] Based on the data asset inventory vector and the directory retrieval key identification vector of all data asset blocks of the enterprise, an initial distributed storage strategy of all data asset blocks of the enterprise is generated;
[0054] The data asset inventory vector and the directory search key identification vector of all data asset blocks of the enterprise are used to retrieve a joint call probability of each joint data of all data asset blocks of the enterprise in each business scenario;
[0055] Based on the importance level of each business scenario and the joint call probability of each joint data in each business scenario, the initial distributed storage strategy is improved to obtain an actual distributed storage strategy;
[0056] Based on the actual call records of all data asset blocks in the latest period, the actual distributed storage strategy is optimized in real time to obtain a latest distributed storage strategy.
[0057] In this embodiment, all data asset blocks of the enterprise refer to the entire set of data assets owned by the enterprise, which are divided into data asset blocks according to certain rules. These data asset blocks cover various types of data generated and collected by the enterprise in the operation process, and are the basic units of data asset operation and management, such as customer information data block, sales data block, etc.
[0058] In this embodiment, the initial distributed storage strategy of all data asset blocks of the enterprise: a preliminary storage strategy generated based on the data asset inventory vector and the directory search key identification vector of all data asset blocks of the enterprise. By dividing and sorting the data asset blocks, the relative position relationship of each data asset block in the storage network is determined, and an initial distributed storage network is constructed to preliminarily plan the storage layout of data assets, providing a basic framework for data storage, aiming to effectively integrate data assets and provide basic support for subsequent storage optimization and data access.
[0059] In this embodiment, the joint call probability of each joint data of all data asset blocks of the enterprise in each business scenario: a value reflecting the possibility of calling different joint data of the enterprise in various business scenarios, which is predicted by using the data asset inventory vector and the directory search key identification vector. This probability helps the enterprise to understand the usage trend of data in different business scenarios in advance, so as to more reasonably plan data storage and allocate resources.
[0060] In this embodiment, business scenario refers to different business situations or business activity categories involved in the operation process of the enterprise. For example, the sales business scenario covers customer negotiation, order processing, etc.; the financial business scenario involves budgeting, financial statement generation, etc. Each business scenario has different requirements for data assets. By analyzing the joint call probability of data in different business scenarios, combined with the importance level of the scenario, the data asset storage strategy can be optimized to make data better serve various businesses of the enterprise.
[0061] In this embodiment, the actual distributed storage strategy: combined with the importance level of each business scenario and the joint call probability of each joint data in each business scenario, the storage strategy obtained after improving the initial distributed storage strategy. By calculating the total call probability of each data asset block, and adjusting the initial storage strategy according to the similarity between the total call probability of each data asset block, the storage strategy can allocate storage resources more reasonably according to business priority and data call probability, improve data access efficiency, and better meet the actual business needs of enterprises.
[0062] In this embodiment, the actual call record of all data asset blocks in the latest period: in the latest period set, the detailed record of the actual call of all data asset blocks of the enterprise. These records contain the number of times and time of each data asset block being called, which directly reflects the actual use of data assets in the period, provides real data basis for real-time optimization of the actual distributed storage strategy, so that the storage strategy can adapt to the dynamic changes of business data use.
[0063] In this embodiment, the latest distributed storage strategy: based on the actual call record of all data asset blocks in the latest period, the storage strategy obtained after real-time optimization of the actual distributed storage strategy. By determining the actual call frequency of each data asset block in the latest period, adjusting the actual distributed storage strategy according to the similarity between the actual call frequency of each data asset block, ensuring that the storage strategy can be updated in time according to the change of data use, continuously improving the operation efficiency of data assets, and maximizing the value of data assets.
[0064] In an alternative embodiment, the process of obtaining the data asset inventory vector and the directory retrieval key identification vector of all data asset blocks of the enterprise includes:
[0065] Determine the multi-dimensional asset classification attribute and multi-dimensional right information of each data asset block of the enterprise, and generate the data asset inventory vector of each data asset block based on the multi-dimensional asset classification attribute and multi-dimensional right information of each data asset block;
[0066] Based on the multi-dimensional retrieval basis of each asset data block of the enterprise, generate the target retrieval key identification of each data asset block.
[0067] In this embodiment, the multi-dimensional asset classification attribute and multi-dimensional right information of each data asset block:
[0068] Multi-dimensional asset classification attributes: Refers to the characteristic description of classifying data asset blocks from multiple dimensions. These dimensions can include the source of data (such as internal system generation, external purchase, etc.), the type of data (structured, semi-structured or unstructured), the business field to which the data belongs (sales, finance, human resources, etc.), the sensitivity of the data (public, internal use, confidential, etc.), etc. By classifying data asset blocks from multiple dimensions, the characteristics and uses of data assets can be better understood.
[0069] Multi-dimensional right confirmation information: Refers to the right attribution and related limitation information of data asset blocks in multiple aspects. For example, the ownership of the data (belongs to the enterprise itself, partners or other subjects), the scope of the right to use (which departments and personnel can use, the purpose of use is limited, etc.), the authorization period of the data (long-term effective, specific time period, etc.), and the compliance-related information of the data (whether it complies with industry regulations, privacy policies, etc.). Clear multi-dimensional right confirmation information is crucial for the legal and compliant use of data assets.
[0070] In this embodiment, the data asset inventory vector of each data asset block is generated based on the multi-dimensional asset classification attributes and multi-dimensional right confirmation information of each data asset block: The multi-dimensional asset classification attributes and multi-dimensional right confirmation information of each data asset block are converted into a vector form of data asset inventory vector through certain algorithms or rules. Each element in the vector may correspond to a certain attribute or quantized representation of right confirmation information in a certain dimension. For example, assume that the asset classification attribute dimension has data type (structured = 1, semi-structured = 2, unstructured = 3), business field (sales = 1, finance = 2, etc.), and right confirmation information has ownership (enterprise itself = 1, partners = 2), etc. Then for a structured data asset block belonging to the sales field and owned by the enterprise itself, its data asset inventory vector may be [1, 1, 1, …], and in this way, complex classification and right confirmation information is integrated into a vector form for analysis and calculation, providing a data basis for subsequent generation of initial distributed storage strategy and other operations.
[0071] In this embodiment, the multi-dimensional retrieval basis of each asset data block: The relevant basis for retrieving data asset blocks described from multiple angles. This may include keywords of information contained in data asset blocks, time range of data creation or update, data associated business processes, data format, etc. For example, in the sales data asset block, the multi-dimensional retrieval basis may be customer name, sales time interval, sales product category, etc., which provides multiple dimensional clues for quickly and accurately locating and querying data asset blocks, helping to improve the efficiency and accuracy of data retrieval.
[0072] In this embodiment, the multi-dimensional search based on each asset data block of the enterprise generates a target search key identification for each data asset block: according to the multi-dimensional search basis of each data asset block, a key identification used to identify the data asset block for search is extracted by a specific algorithm or rule. This target search key identification can be a code, a string or other forms of identifier, which condenses the key information in the multi-dimensional search basis, so that the corresponding data can be quickly located according to this key identification when searching for the data asset block. For example, the first letter of the customer name in the sales data asset block, the specific code of the sales time and the code of the sales product category are combined into a string as the target search key identification, when the relevant sales data needs to be searched, the corresponding asset data block can be quickly found by this key identification.
[0073] In an alternative embodiment, the data asset inventory vector and the directory search key identification vector of all data asset blocks of the enterprise are used to generate an initial distributed storage strategy for all data asset blocks of the enterprise, including:
[0074] The data asset inventory vector and the directory search key identification vector of all data asset blocks of the enterprise are used to divide all data asset blocks to obtain all first effective division clusters and all second effective division clusters;
[0075] The average value evaluation value of each first effective division cluster and the average value evaluation value of each second effective division cluster are determined based on the value evaluation value of all data asset blocks, that is, the average value of the value evaluation value of all data asset blocks in the first effective division cluster or the second effective division cluster is taken as the average value evaluation value of the corresponding first effective division cluster or the second effective division cluster;
[0076] All first effective division clusters are sorted in descending order of the average value evaluation value to obtain the sorting value of all first effective division clusters, and all data asset blocks in each first effective division cluster are sorted in descending order of the value evaluation value to obtain the sorting value of each data asset block in each first effective division cluster. Based on the sorting value P1 of each data asset block in each first effective division cluster and the sorting value P2 of the corresponding first effective division cluster, a first division position two-dimensional array (P2, P1) of each data asset block is generated;
[0077] The second effective partition clusters are sorted in descending order according to the average value evaluation values of all the second effective partition clusters to obtain the sorting values of all the second effective partition clusters, and meanwhile, all the data asset blocks in each second effective partition cluster are sorted in descending order according to the value evaluation values of all the data asset blocks in each second effective partition cluster to obtain the sorting values of each data asset block in each second effective partition cluster, and a second partition position two-dimensional array (P4, P3) of each data asset block is generated based on the sorting value P3 of each data asset block in each second effective partition cluster and the sorting value P4 of the corresponding second effective partition cluster;
[0078] A two-dimensional partition position pointing vector (P4-P2, P3-P1) of each data asset block is determined based on the first partition position two-dimensional array and the second partition position two-dimensional array of each data asset block, and an initial distributed storage strategy of all the data asset blocks of the enterprise is generated based on the two-dimensional partition position pointing vectors of all the data asset blocks.
[0079] In this embodiment, the value evaluation value of each data asset block is a quantitative representation of the value of each data asset block of the enterprise. For example, the importance of the business field involved by the data asset block, if the data is related to the core business of the enterprise, the value evaluation value may be higher; the scarcity of the data, if the data asset block contains unique or difficult to obtain data, the value evaluation value will also be correspondingly improved; the quality of the data, such as accuracy, integrity, etc., high-quality data usually has a higher value evaluation value; the frequency of use of the data in the past business and the degree of influence on business decision-making, the data asset block which is frequently used and plays a key role in decision-making, the value evaluation value will be higher, by determining the performance of the data asset block in the above aspects and assigning values, the value evaluation value of the data asset block can be obtained.
[0080] In an alternative implementation, based on the data asset inventory vector and the directory retrieval key identification vector of all the data asset blocks of the enterprise, all the data asset blocks of the enterprise are partitioned to obtain all the first effective partition clusters and all the second effective partition clusters, including:
[0081] The enterprise's all data asset blocks are randomly divided to obtain a plurality of first data asset block clusters, and similarity of data asset inventory vectors of each two data asset blocks in each first data asset block cluster is calculated. The average of the similarity of the data asset inventory vectors of each two first data asset block clusters in each first data asset block cluster is taken as the cluster cohesion of each first data asset block cluster. The average of the cluster cohesion of all first data asset block clusters obtained in the current division process is taken as the division performance value of the current division process. Whether the division performance value of the current division process is greater than a preset division performance threshold is judged. If yes, all first data asset block clusters of the current division process are taken as all first effective division clusters. Otherwise, the enterprise's all data asset blocks are continuously randomly divided until the division performance value of the latest division process is greater than the preset division performance threshold. Then, all newly obtained second data asset block clusters are taken as second effective division clusters.
[0082] The enterprise's all data asset blocks are randomly divided to obtain a plurality of second data asset block clusters, and similarity of directory search key identification vectors of each two data asset blocks in each second data asset block cluster is calculated. The average of the similarity of the directory search key identification vectors of each two second data asset block clusters in each second data asset block cluster is taken as the cluster cohesion of each second data asset block cluster. The average of the cluster cohesion of all second data asset block clusters obtained in the current division process is taken as the division performance value of the current division process. Whether the division performance value of the current division process is greater than a preset division performance threshold is judged. If yes, all second data asset block clusters of the current division process are taken as all second effective division clusters. Otherwise, the enterprise's all data asset blocks are continuously randomly divided until the division performance value of the latest division process is greater than the preset division performance threshold. Then, all newly obtained second data asset block clusters are taken as second effective division clusters.
[0083] In this embodiment, the similarity between all vectors can be obtained by using the cosine similarity calculation method.
[0084] In this embodiment, the preset division performance threshold is a standard value set in advance, for example, can be 0.74, which is used to judge the division effect of the enterprise data asset blocks. When the data asset blocks are divided, the division performance value is calculated and compared with the threshold. If it is greater than it, the division is effective, and if it is less than it, the division is continued. In this way, the division of the data asset blocks is reasonable, and a better initial distributed storage strategy is generated.
[0085] In an alternative embodiment, a two-dimensional division position pointing vector of each data asset block is determined based on the first division position two-dimensional array and the second division position two-dimensional array of each data asset block, and an initial distributed storage strategy of all data asset blocks of the enterprise is generated based on the two-dimensional division position pointing vectors of all data asset blocks, including:
[0086] determining a first partition position point (P2, P1) and a second partition position point (P4, P3) of each data asset block in the preset coordinate system based on the first partition position two-dimensional array (P2, P1) and the second partition position two-dimensional array (P4, P3) of each data asset block;
[0087] taking the vector from the first partition position point to the second partition position point of each data asset block in the preset coordinate system as the two-dimensional partition position pointing vector (P4-P2, P3-P1) of each data asset block;
[0088] constructing an initial distributed storage network containing all storage nodes of all data asset blocks as an initial distributed storage strategy of all data asset blocks of the enterprise based on the similarity between the two-dimensional partition position pointing vectors of each two data asset blocks.
[0089] In this embodiment, the initial distributed storage network containing all storage nodes of all data asset blocks is constructed as an initial distributed storage strategy of all data asset blocks of the enterprise based on the similarity between the two-dimensional partition position pointing vectors of each two data asset blocks, including: first, obtaining the two-dimensional partition position pointing vector of each data asset block. Then, by comparing the similarity between the vectors of each two data asset blocks, the connection relationship between the storage nodes is determined. The two data asset blocks with high similarity correspond to storage nodes that are more closely connected in the network (such as being closer in physical location, or having faster network connection speed. In this way, when accessing related data later, the associated data can be obtained faster, improving data management efficiency). In this way, all storage nodes corresponding to all data asset blocks are constructed into an initial distributed storage network, which forms the initial distributed storage strategy of the data asset blocks of the enterprise, with the purpose of reasonably arranging data storage and improving data management efficiency.
[0090] In an alternative embodiment, based on the data asset inventory vector and the directory retrieval key identification vector of all data asset blocks of the enterprise, the joint calling probability of each joint data of all data asset blocks of the enterprise in each business scenario is predicted, including:
[0091] determining the data asset inventory calling range vector and the target retrieval key identification calling range vector in each business scenario;
[0092] taking the ratio of the total number of elements in the data asset inventory vector of each data asset block that meet the value range of the same element position in the data asset inventory calling range vector in each business scenario to the total number of all elements in the data asset inventory vector as the first calling probability of each data asset block in the corresponding business scenario;
[0093] The ratio of the total number of elements in the target search key identification calling range vector that meet the value range of the same element position in the target search key identification calling range vector to the total number of elements in the target search key identification vector is taken as the second calling probability of each data asset block under the corresponding business scenario.
[0094] Based on the first calling probability and the second calling probability of each data asset block under each business scenario, the joint calling probability of each joint data of all data asset blocks of the enterprise under each business scenario is predicted.
[0095] In this embodiment, for each business scenario, a specific vector corresponding thereto needs to be specified. The data asset inventory calling range vector specifies the value range of each element in the data asset inventory vector under the business scenario, which defines the range of data assets that can be used in this business scenario in terms of classification attributes and right information, etc. For example, under the marketing business scenario, the vector can specify that the data source needs to be related to market research, the data sensitivity degree is public or internal use, etc.
[0096] The target search key identification calling range vector specifies the value range of each element in the target search key identification vector under the business scenario, which limits the key identification features of the data asset block used for searching in this business scenario. For example, in the supply chain management business scenario, the target search key identification can be required to contain specific supplier codes, product category codes, etc. By determining the two vectors, a basis is provided for predicting the calling probability of data assets under various business scenarios.
[0097] In an alternative embodiment, in order to accurately predict the joint calling probability of each joint data of all data asset blocks of the enterprise under each business scenario based on the first calling probability and the second calling probability of each data asset block under each business scenario, the method comprises:
[0098] The mean of the first calling probability and the second calling probability of each data asset block under each business scenario is taken as the comprehensive calling probability of each data asset block under each business scenario;
[0099] The product of the comprehensive calling probabilities of all data asset blocks contained in each joint data of all data asset blocks of the enterprise under the same business scenario is taken as the joint calling probability of the corresponding joint data under the corresponding business scenario.
[0100] In an alternative embodiment, based on the importance level of each business scenario and the joint calling probability of each joint data under each business scenario, the initial distributed storage strategy is improved to obtain an actual distributed storage strategy, which comprises:
[0101] obtain the total calling probability of each data asset block based on the importance level of all kinds of business scenarios and the joint calling probability of all kinds of joint data involving each data asset block in all kinds of business scenarios;
[0102] improve the initial distributed storage strategy based on the similarity between the total calling probabilities of all pairs of data asset blocks, and obtain the actual distributed storage strategy.
[0103] In this embodiment, the importance level of all kinds of business scenarios: this is a quantitative classification of the relative importance of various business scenarios in the enterprise. Different business scenarios of an enterprise, such as production, sales, and research and development, have different degrees of influence on the operation and development of the enterprise. By setting the importance level, the priority of each business scenario can be clearly distinguished. For example, the importance level of core business scenarios (such as scenarios related to the production of the enterprise's main profit products) may be set to high, while the importance level of some auxiliary business scenarios (such as internal communication scenarios of non-core departments) is relatively low. This level is an important reference factor for subsequent adjustment of data storage strategy, which can make resources preferentially tilt towards important business scenarios.
[0104] In this embodiment, the total calling probability of each data asset block is obtained based on the importance level of all kinds of business scenarios and the joint calling probability of all kinds of joint data involving each data asset block in all kinds of business scenarios: the importance level of each business scenario and the joint calling probability of joint data of each data asset block in different business scenarios are comprehensively considered to calculate the total calling probability of each data asset block. For example, suppose the joint calling probability of joint data containing a certain data asset block q in business scenario A (importance level 3, full score 5) is 0.6, and the joint calling probability of joint data containing a certain data asset block q in business scenario B (importance level 2) is 0.4. When calculating the total calling probability, the calling probability will be weighted in combination with the importance of the business scenario, and the total calling probability of the data asset block may be obtained by a calculation method similar to "total calling probability = importance level of business scenario A ÷ 5 × joint calling probability in business scenario A + importance level of business scenario B ÷ 5 × joint calling probability in business scenario B". This total calling probability reflects the comprehensive calling possibility of the data asset block in the overall business.
[0105] In this embodiment, the initial distributed storage strategy is improved based on the similarity between the total call probabilities of all two-two data asset blocks, and an actual distributed storage strategy is obtained: after calculating the total call probability of all data asset blocks, the similarity of the total call probability of each two data asset blocks is compared. If the total call probability of two data asset blocks is high, it means that they are similar in use in the business, and should be closer in storage layout. According to this similarity relationship, the initial distributed storage strategy is adjusted, for example, data asset blocks with high similarity are arranged in adjacent storage nodes or storage areas, so as to obtain an actual distributed storage strategy that is more in line with the actual needs of the business, so as to improve the access efficiency of data and the reasonable allocation of storage resources.
[0106] In an alternative embodiment, the actual distributed storage strategy is optimized in real time based on the actual call records of all data asset blocks statistically in the latest period, and a latest distributed storage strategy is obtained, including:
[0107] Based on the actual call records of all data asset blocks statistically in the latest period, the actual call frequency of each data asset block in the latest period is determined.
[0108] The actual distributed storage strategy is optimized in real time based on the similarity between the actual call frequencies of all two-two data asset blocks in the latest period, and a latest distributed storage strategy is obtained.
[0109] In this embodiment, the actual call frequency of each data asset block in the latest period is determined based on the actual call records of all data asset blocks statistically in the latest period: the latest period can be a pre-set period of time, such as a week, a month, etc. In this period, the system records the actual call of all data asset blocks to form the actual call record. By counting the number of times each data asset block is called in the period and dividing by the length of the period (or the number of relevant time units), the actual call frequency of each data asset block in the latest period can be obtained. For example, in a one-month period, a data asset block is called 20 times, and the actual call frequency is 1 call per day, calculated based on 20 working days in a month. This actual call frequency reflects the recent use activity of each data asset block.
[0110] In this embodiment, the actual distributed storage strategy is optimized in real time based on the similarity between the actual call frequencies of all two-two data asset blocks in the latest period, and the latest distributed storage strategy is obtained: the similarity between the actual call frequencies of all two-two data asset blocks is calculated, and if the actual call frequencies of two data asset blocks are similar, it means that their use activity patterns are similar. Based on this similarity, the existing actual distributed storage strategy is adjusted in real time. For example, data asset blocks with high actual call frequency similarity are stored closer to each other, so that when data is accessed, data transmission time can be reduced and overall data access efficiency can be improved, so that the latest distributed storage strategy that conforms to the current business data usage situation is obtained, and the storage strategy can adapt to the dynamic changes of business data requirements.
[0111] The application also provides an embodiment of a full-life-cycle intelligent operation system of data assets, comprising:
[0112] An initial storage strategy generation module is configured to generate an initial distributed storage strategy of all data asset blocks of an enterprise based on the data asset inventory vectors and the directory retrieval key identification vectors of all data asset blocks of the enterprise;
[0113] A joint call probability determination module is configured to predict the joint call probability of each joint data of all data asset blocks of an enterprise in each business scenario based on the data asset inventory vectors and the directory retrieval key identification vectors of all data asset blocks of the enterprise;
[0114] An initial storage strategy improvement module is configured to improve the initial distributed storage strategy based on the importance level of each business scenario and the joint call probability of each joint data in each business scenario, and obtain an actual distributed storage strategy;
[0115] A storage strategy real-time optimization module is configured to optimize the actual distributed storage strategy in real time based on the actual call records of all data asset blocks in the latest period, and obtain a latest distributed storage strategy.
[0116] Based on the data asset inventory vector and the directory retrieval key identification vector, an initial distributed storage strategy is generated, which can preliminarily plan the data asset storage layout, effectively integrate various types of data assets of the enterprise, and lay a foundation for subsequent operation. By predicting the joint call probability of each joint data in different business scenarios through the same vector, the enterprise can anticipate data usage trends in advance, making the data storage strategy more forward-looking and better matching business needs. Combined with the importance level of the business scenario and the joint call probability, the initial strategy is improved to obtain an actual distributed storage strategy that can allocate storage resources more reasonably according to business priority and data call probability, improving data access efficiency. Finally, the actual distributed storage strategy is optimized in real time according to the actual call records in the latest period, ensuring that the storage strategy is adjusted in a timely manner as the business changes, continuously improving the efficiency of data asset operation, and maximizing the value of data assets.
[0117] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application belong to the scope of the present application and its equivalent technology, the present application also intends to include these modifications and variations.
Claims
1. A method for intelligent operation of data assets throughout their entire lifecycle, characterized in that: include: Based on the data asset inventory vector and catalog retrieval key identifier vector of all data asset blocks of the enterprise, an initial distributed storage strategy for all data asset blocks of the enterprise is generated. Among them, based on the data asset inventory vectors and catalog retrieval key identifier vectors of all data asset blocks of the enterprise, an initial distributed storage strategy for all data asset blocks of the enterprise is generated, including: Based on the data asset inventory vector and catalog retrieval key identifier vector of all data asset blocks of the enterprise, all data asset blocks of the enterprise are divided to obtain all first effective partition clusters and all second effective partition clusters. The average value of each first effective partition cluster and the average value of each second effective partition cluster are determined based on the value assessment of all data asset blocks. Sort all first effective partition clusters in descending order of their average value assessment value to obtain a sorting value for all first effective partition clusters. At the same time, sort all data asset blocks in each first effective partition cluster in descending order of their value assessment value to obtain a sorting value for each data asset block in each first effective partition cluster. Based on the sorting value of each data asset block in each first effective partition cluster and the sorting value of the corresponding first effective partition cluster, generate a two-dimensional array of the first partition positions for each data asset block. Sort all second effective partition clusters in descending order of their average value assessment value to obtain a sorting value for all second effective partition clusters. At the same time, sort all data asset blocks in each second effective partition cluster in descending order of their value assessment value to obtain a sorting value for each data asset block in each second effective partition cluster. Based on the sorting value of each data asset block in each second effective partition cluster and the sorting value of the corresponding second effective partition cluster, generate a two-dimensional array of second partition positions for each data asset block. Based on the first and second two-dimensional arrays of partition positions for each data asset block, a two-dimensional partition position pointing vector for each data asset block is determined. Based on the two-dimensional partition position pointing vectors of all data asset blocks, an initial distributed storage strategy for all data asset blocks of the enterprise is generated. Based on the data asset inventory vector and catalog retrieval key identifier vector of all data asset blocks of an enterprise, the probability of joint access of each type of joint data of all data asset blocks of an enterprise in each business scenario is predicted; The process of obtaining the data asset inventory vector and catalog retrieval key identifier vector for all data asset blocks of the enterprise includes: Determine the multidimensional asset classification attributes and multidimensional ownership information of each data asset block of the enterprise, and generate a data asset inventory vector for each data asset block based on the multidimensional asset classification attributes and multidimensional ownership information of each data asset block; Based on the multi-dimensional retrieval criteria of each asset data block of the enterprise, a target retrieval key identifier vector is generated for each data asset block; Specifically, based on the data asset inventory vectors and directory retrieval key identifier vectors of all data asset blocks of the enterprise, the probability of joint retrieval of each type of joint data in each business scenario of all data asset blocks of the enterprise is predicted, including: Determine the data asset inventory call range vector and the target retrieval key identifier call range vector for each business scenario; The ratio of the total number of elements in the data asset inventory vector of each data asset block that meet the value range of the same element position in the data asset inventory call range vector under each business scenario to the total number of all elements in the data asset inventory vector is taken as the first call probability of each data asset block under the corresponding business scenario. The ratio of the total number of elements in the target retrieval key identifier vector of each data asset block that meet the value range of the same element position in the target retrieval key identifier call range vector under each business scenario to the total number of all elements in the target retrieval key identifier vector is taken as the second call probability of each data asset block under the corresponding business scenario. Based on the first and second call probabilities of each data asset block in each business scenario, the joint call probability of each joint data of all data asset blocks of the enterprise in each business scenario is predicted; Based on the importance level of each business scenario and the probability of joint access to each type of federated data in each business scenario, the initial distributed storage strategy is improved to obtain the actual distributed storage strategy, including: Based on the importance level of all business scenarios and the joint call probability of all types of joint data involved in each data asset block under all business scenarios, the total call probability of each data asset block is obtained. The initial distributed storage strategy is improved based on the similarity between the total call probabilities of all pairs of data asset blocks to obtain the actual distributed storage strategy. Based on the actual usage records of all data asset blocks in the latest period, the actual distributed storage strategy is optimized in real time to obtain the latest distributed storage strategy, including: Based on the actual call records of all data asset blocks in the latest period, the actual call frequency of each data asset block in the latest period is determined. The actual distributed storage strategy is optimized in real time based on the similarity between the actual call frequencies of all pairs of data asset blocks in the latest period to obtain the latest distributed storage strategy.
2. The intelligent operation method for the entire lifecycle of data assets according to claim 1, characterized in that, Based on the data asset inventory vectors and catalog retrieval key identifier vectors of all data asset blocks of the enterprise, all data asset blocks of the enterprise are divided to obtain all first-effective partition clusters and all second-effective partition clusters, including: The system arbitrarily divides all of the enterprise's data asset blocks to obtain multiple first data asset block clusters. It calculates the similarity of the data asset inventory vectors of each pair of data asset blocks in each first data asset block cluster. The average similarity of the data asset inventory vectors of all pairs of first data asset block clusters in each first data asset block cluster is taken as the intra-cluster cluster degree of each first data asset block cluster. The average intra-cluster cluster degree of all first data asset block clusters obtained in the current division process is taken as the division performance value of the current division process. It is determined whether the division performance value of the current division process is greater than the preset division performance threshold. If so, all first data asset block clusters in the current division process are taken as all first effective division clusters. Otherwise, the system continues to arbitrarily divide all of the enterprise's data asset blocks until the division performance value of the latest division process is greater than the preset division performance threshold. Then, all the newly obtained first data asset block clusters are taken as first effective division clusters. The enterprise's data asset blocks are arbitrarily divided to obtain multiple second data asset block clusters. The similarity of the directory retrieval key identifier vectors of each pair of data asset blocks in each second data asset block cluster is calculated. The average similarity of the directory retrieval key identifier vectors of all pairwise second data asset block clusters in each second data asset block cluster is taken as the intra-cluster cluster degree of each second data asset block cluster. The average intra-cluster cluster degree of all second data asset block clusters obtained in the current division process is taken as the division performance value of the current division process. It is determined whether the division performance value of the current division process is greater than the preset division performance threshold. If so, all second data asset block clusters in the current division process are regarded as all second effective partition clusters. Otherwise, the enterprise's data asset blocks are arbitrarily divided until the division performance value of the latest division process is greater than the preset division performance threshold. Then, all the newly obtained second data asset block clusters are regarded as second effective partition clusters.
3. The intelligent operation method for the entire lifecycle of data assets according to claim 1, characterized in that, Based on the first and second two-dimensional arrays of partition positions for each data asset block, a two-dimensional partition position pointing vector for each data asset block is determined. Based on the two-dimensional partition position pointing vectors of all data asset blocks, an initial distributed storage strategy for all data asset blocks of the enterprise is generated, including: Based on the first and second two-dimensional arrays of the first and second division positions of each data asset block, the first and second division positions of each data asset block in the preset coordinate system are determined respectively. The vector pointing from the first division point to the second division point of each data asset block in the preset coordinate system is taken as the two-dimensional division point pointing vector of each data asset block. Based on the similarity between the two-dimensional partitioning position pointing vectors of every two data asset blocks, an initial distributed storage network containing all storage nodes of all data asset blocks is constructed as the initial distributed storage strategy for all data asset blocks of the enterprise.
4. The intelligent operation method for the entire lifecycle of data assets according to claim 1, characterized in that, Based on the first and second call probabilities of each data asset block in each business scenario, the joint call probability of each type of combined data across all data asset blocks of the enterprise in each business scenario is predicted, including: The average of the first and second call probabilities of each data asset block in each business scenario is taken as the comprehensive call probability of each data asset block in each business scenario. The product of the overall call probabilities of all data asset blocks contained in each type of federated data of an enterprise under the same business scenario is taken as the joint call probability of the corresponding type of federated data under the corresponding business scenario.
5. A smart operation system for the entire lifecycle of data assets, characterized in that: include: The initial storage strategy generation module is used to generate an initial distributed storage strategy for all data asset blocks of the enterprise, based on the data asset inventory vector and catalog retrieval key identifier vector of all data asset blocks. This strategy includes: Based on the data asset inventory vector and catalog retrieval key identifier vector of all data asset blocks of the enterprise, all data asset blocks of the enterprise are divided to obtain all first effective partition clusters and all second effective partition clusters. The average value of each first effective partition cluster and the average value of each second effective partition cluster are determined based on the value assessment of all data asset blocks. Sort all first effective partition clusters in descending order of their average value assessment value to obtain a sorting value for all first effective partition clusters. At the same time, sort all data asset blocks in each first effective partition cluster in descending order of their value assessment value to obtain a sorting value for each data asset block in each first effective partition cluster. Based on the sorting value of each data asset block in each first effective partition cluster and the sorting value of the corresponding first effective partition cluster, generate a two-dimensional array of the first partition positions for each data asset block. Sort all second effective partition clusters in descending order of their average value assessment value to obtain a sorting value for all second effective partition clusters. At the same time, sort all data asset blocks in each second effective partition cluster in descending order of their value assessment value to obtain a sorting value for each data asset block in each second effective partition cluster. Based on the sorting value of each data asset block in each second effective partition cluster and the sorting value of the corresponding second effective partition cluster, generate a two-dimensional array of second partition positions for each data asset block. Based on the first and second two-dimensional arrays of partition positions for each data asset block, a two-dimensional partition position pointing vector for each data asset block is determined. Based on the two-dimensional partition position pointing vectors of all data asset blocks, an initial distributed storage strategy for all data asset blocks of the enterprise is generated. The joint invocation probability determination module is used to predict the joint invocation probability of each type of joint data in each business scenario based on the data asset inventory vector and catalog retrieval key identifier vector of all data asset blocks of the enterprise. The process of obtaining the data asset inventory vector and catalog retrieval key identifier vector for all data asset blocks of the enterprise includes: Determine the multidimensional asset classification attributes and multidimensional ownership information of each data asset block of the enterprise, and generate a data asset inventory vector for each data asset block based on the multidimensional asset classification attributes and multidimensional ownership information of each data asset block; Based on the multi-dimensional retrieval criteria of each asset data block of the enterprise, a target retrieval key identifier vector is generated for each data asset block; Specifically, based on the data asset inventory vectors and directory retrieval key identifier vectors of all data asset blocks of the enterprise, the probability of joint retrieval of each type of joint data in each business scenario of all data asset blocks of the enterprise is predicted, including: Determine the data asset inventory call range vector and the target retrieval key identifier call range vector for each business scenario; The ratio of the total number of elements in the data asset inventory vector of each data asset block that meet the value range of the same element position in the data asset inventory call range vector under each business scenario to the total number of all elements in the data asset inventory vector is taken as the first call probability of each data asset block under the corresponding business scenario. The ratio of the total number of elements in the target retrieval key identifier vector of each data asset block that meet the value range of the same element position in the target retrieval key identifier call range vector under each business scenario to the total number of all elements in the target retrieval key identifier vector is taken as the second call probability of each data asset block under the corresponding business scenario. Based on the first and second call probabilities of each data asset block in each business scenario, the joint call probability of each joint data of all data asset blocks of the enterprise in each business scenario is predicted; The initial storage strategy improvement module is used to improve the initial distributed storage strategy based on the importance level of each business scenario and the probability of joint access to each type of federated data in each business scenario, to obtain the actual distributed storage strategy, including: Based on the importance level of all business scenarios and the joint call probability of all types of joint data involved in each data asset block under all business scenarios, the total call probability of each data asset block is obtained. The initial distributed storage strategy is improved based on the similarity between the total call probabilities of all pairs of data asset blocks to obtain the actual distributed storage strategy. The real-time storage strategy optimization module is used to optimize the actual distributed storage strategy in real time based on the actual call records of all data asset blocks in the latest period, and obtain the latest distributed storage strategy, including: Based on the actual call records of all data asset blocks in the latest period, the actual call frequency of each data asset block in the latest period is determined. The actual distributed storage strategy is optimized in real time based on the similarity between the actual call frequencies of all pairs of data asset blocks in the latest period to obtain the latest distributed storage strategy.
Citation Information
Patent Citations
Data asset segmentation verification method and equipment
CN113868601A
Distributed storage cluster resource energy consumption optimization method and device and electronic equipment
CN116301585A