Blockchain-based enterprise data asset management method and system
By analyzing the storage and interaction of enterprise data in the sidechain, and adjusting the synchronization strategy in conjunction with changes in sidechain load, the sidechain expansion is optimized, solving the problem of poor performance of existing sidechain expansion strategies and achieving efficient enterprise data management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANDONG FUTURE DATA TECHNOLOGY CO LTD
- Filing Date
- 2025-06-16
- Publication Date
- 2026-05-08
AI Technical Summary
Existing methods employ a fixed sidechain extension strategy, which is ineffective for enterprise-level data asset management and fails to improve the efficiency and security of data management.
By acquiring enterprise asset data stored in each business module of each sidechain, analyzing the interaction information and data differences between business modules, determining the initial storage strategy, and adjusting the synchronization level in real time in conjunction with the load changes of the sidechain, the sidechain expansion strategy is optimized.
It reduces the complexity of enterprise data management, improves the response speed and efficiency of data management, reduces the pressure on the main chain, and achieves efficient collaborative management between the side chain and the main chain.
Smart Images

Figure CN120655162B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and specifically to a blockchain-based enterprise data asset management method and system. Background Technology
[0002] Blockchain is a distributed ledger technology that records data in a decentralized manner, ensuring data security, transparency, and immutability. The core objectives of enterprise data asset management are to ensure data security, transparency, and efficiency. Therefore, managing enterprise data assets based on blockchain enables enterprises to improve data credibility, security, and efficiency during the data management process.
[0003] With technological advancements, sidechain technology has emerged to facilitate the transfer of digital assets between different blockchains. It connects different blockchains to enable blockchain expansion. For enterprise-level data assets, which may involve large amounts of transaction and log data, the scale and real-time throughput of private blockchains (main chains) are limited, and directly expanding the main chain is costly. Therefore, developing a superior enterprise sidechain expansion strategy can provide enterprises with automated and intelligent data management solutions. However, existing methods using fixed sidechain expansion strategies are less effective for enterprise-level data asset management. Summary of the Invention
[0004] To address the technical problem that existing methods employing fixed sidechain expansion strategies are ineffective for enterprise-level data asset management, the present invention aims to provide a blockchain-based enterprise data asset management method and system. The specific technical solution adopted is as follows:
[0005] In a first aspect, the present invention provides a blockchain-based enterprise data asset management method, comprising:
[0006] Obtain enterprise asset data for each business module in each sidechain;
[0007] Based on the interaction information between each business module in each sidechain and each business module in other sidechains, and combined with the core differences between the data contained in different business modules, the sidechains are merged to determine the initial storage strategy of the blockchain structure.
[0008] Under the initial storage strategy, based on the memory usage and bandwidth information of each sidechain at each update time, the load changes of the corresponding sidechain are analyzed, and the synchronization degree of each sidechain at each update time is adjusted in real time in combination with the update status of the sidechain at historical update times, so as to determine the global storage strategy of the blockchain structure.
[0009] Preferably, the acquisition of enterprise asset data storage in each business module of each sidechain specifically includes:
[0010] Obtain the confidentiality tag value, access frequency, and update frequency for each type of enterprise data in the enterprise asset data;
[0011] Based on the confidentiality tag value of each type of enterprise data, the access frequency and update frequency of all enterprise data of the same type, the core degree index of each type of enterprise data is obtained;
[0012] Based on the differences in core indicators among different types of enterprise data, enterprise data is classified into different core levels;
[0013] Enterprise data belonging to the same business and the same core level are treated as the same business module. Different types of enterprise data contained in each business module are stored in a sidechain. Each sidechain includes at least one business module.
[0014] Preferably, the process of obtaining the core index for each type of enterprise data based on the confidentiality marker value for each type of enterprise data, the access frequency and update frequency of all enterprise data of the same type, specifically includes:
[0015] Based on the confidentiality tag value of each type of enterprise data and the average access frequency of all enterprise data of the same type, determine the first importance coefficient corresponding to each type of enterprise data;
[0016] The second importance coefficient for each type of enterprise data is determined based on the difference between the maximum update frequency of all types of enterprise data and the mean update frequency of all enterprise data for each type.
[0017] The coreness index for each type of enterprise data is obtained based on the first importance coefficient and the second importance coefficient.
[0018] Preferably, the step of classifying enterprise data into different core levels based on the differences in core indicators between different types of enterprise data specifically includes:
[0019] The core indicators of all types of enterprise data are arranged in a preset order to obtain a core sequence;
[0020] Based on the difference between every two adjacent core degree indicators in the core degree sequence, determine the core difference value corresponding to every two adjacent core degree indicators; based on the ratio between each core difference value and the mean of all core difference values, determine the grade classification reference value corresponding to each core difference value.
[0021] Following the order from largest to smallest, a preset number of reference values for level division are obtained as dividing points to divide the core degree sequence, thus determining that the enterprise data of each data segment belongs to a core level.
[0022] Preferably, the step of merging the sidechains based on the interaction information between each business module in each sidechain and each business module in other sidechains, combined with the core differences between the data contained in different business modules, to determine the initial storage strategy of the blockchain structure, specifically includes:
[0023] Based on the interaction frequency between any two different business modules and the repetition of interaction objects that interact with each business module, the degree of data dependency between any two different business modules can be obtained.
[0024] Based on the differences in core metrics between any two different business modules, and combined with the degree of data dependency, the degree of data merging between any two different business modules can be obtained.
[0025] The initial storage strategy for the blockchain structure is to merge and store two sidechains containing two different business modules whose data merging degree is greater than or equal to a preset merging threshold.
[0026] Preferably, the step of determining the data dependency between any two different business modules based on the interaction frequency between them and the repetition of interaction objects that interact with each business module specifically includes:
[0027] Select any two different business modules as the first business module and the second business module respectively;
[0028] All different business modules that have interactive operations with the first business module are collected to form a first interaction set; all different business modules that have interactive operations with the second business module are collected to form a second interaction set.
[0029] The degree of data dependency between the first and second business modules is determined based on the number of business modules contained in the intersection of the first and second interaction sets and the interaction frequency between the first and second business modules.
[0030] Preferably, the step of determining the degree of data merging between any two different business modules based on the differences in core metrics between them, combined with the degree of data dependency, specifically includes:
[0031] Based on the difference between the average coreity index of all types of enterprise data contained in the first business module and the average coreity index of all types of enterprise data contained in the second business module, similarity feature factors are determined; the normalized value of the product of the data dependence degree of the first business module and the second business module and the similarity feature factors is taken as the degree of data merging between the first business module and the second business module.
[0032] Preferably, the step of analyzing the load changes of the corresponding sidechain based on the memory usage and bandwidth information of each sidechain at each update moment, and adjusting the synchronization degree of each sidechain at each update moment in real time in conjunction with the update status of the sidechain at historical update moments to determine the global storage strategy of the blockchain structure specifically includes:
[0033] Obtain the memory usage and remaining bandwidth usage of each sidechain at each update time;
[0034] The load of the corresponding sidechain is analyzed based on the memory usage and remaining bandwidth usage. The dynamic load weight of each sidechain at each update time is obtained by combining the degree of change in the load of each sidechain at adjacent update times.
[0035] For any sidechain, the ratio of the dynamic load weight of each update moment to the adjacent previous historical update moment is used as an adjustment coefficient. The synchronization time interval of the historical update moment is adjusted using the adjustment coefficient to obtain the adjustment time interval of each update moment. The time interval is used to characterize the data synchronization frequency between the corresponding sidechain and the main chain.
[0036] Preferably, the step of analyzing the load of the corresponding sidechain based on the memory occupancy rate and remaining bandwidth occupancy rate, and combining the degree of load change of each sidechain at adjacent update times to obtain the dynamic load weight of each sidechain at each update time, specifically includes:
[0037] For any sidechain, the load coefficient of the sidechain at each update time is obtained based on the memory utilization and remaining bandwidth utilization of the sidechain at each update time. The memory utilization is positively correlated with the load coefficient, and the remaining bandwidth utilization is negatively correlated with the load coefficient.
[0038] By using the ratio of the load coefficient of the sidechain at each update time to that of the adjacent previous historical update time, the load coefficient at each update time is adjusted to obtain the dynamic load weight of the sidechain at each update time. The dynamic load weight is a normalized value.
[0039] Secondly, the present invention provides a blockchain-based enterprise data asset management system, which implements the steps of a blockchain-based enterprise data asset management method. The blockchain-based enterprise data asset management system specifically includes:
[0040] The data acquisition module is used to acquire enterprise asset data from each business module in each sidechain.
[0041] The initial strategy analysis module is used to merge sidechains based on the interaction information between each business module in each sidechain and each business module in other sidechains, combined with the core differences between the data contained in different business modules, and to determine the initial storage strategy of the blockchain structure.
[0042] The global strategy analysis module is used to analyze the load changes of each sidechain based on the memory usage and bandwidth information of each sidechain at each update time under the initial storage strategy. It also adjusts the synchronization degree of each sidechain at each update time in real time by combining the update status of the sidechain at historical update times, and determines the global storage strategy of the blockchain structure.
[0043] The embodiments of the present invention have at least the following beneficial effects:
[0044] This invention first identifies the business modules stored in different sidechains based on the initial data storage of enterprise assets, providing a data foundation for subsequent analysis of the interaction characteristics between these modules. Then, it analyzes the interaction information between different business modules and, combined with the core differences in the data contained within each module, deeply analyzes the differences and similarities in the importance of the data between them. This analysis is used to determine whether the data of each business module is suitable for merged storage management, thus determining the initial storage strategy and reducing the complexity of enterprise data management to some extent. Furthermore, based on the determined initial storage strategy, it deeply analyzes the load status of the sidechains by combining their memory usage and bandwidth information. Then, by combining historical update data, it adjusts the level of sidechain data synchronization at each moment in real time, adaptively determining the data synchronization strategy between the main chain and each sidechain. This adaptive sidechain expansion reduces the pressure on the main chain, effectively improving the response speed and management efficiency of enterprise data asset management. Attached Figure Description
[0045] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 This is a flowchart of the steps of a blockchain-based enterprise data asset management method provided by the present invention;
[0047] Figure 2 This is a flowchart of the sub-steps of step S100 provided by the present invention;
[0048] Figure 3 This is a flowchart of the sub-steps of step S200 provided by the present invention;
[0049] Figure 4 This is a flowchart of the sub-steps of step S300 provided by the present invention;
[0050] Figure 5 This is a schematic diagram of the structure of a blockchain-based enterprise data asset management system provided by the present invention. Detailed Implementation
[0051] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a blockchain-based enterprise data asset management method and system proposed in accordance with the present invention.
[0052] Before introducing the specific solutions provided in the embodiments of this application, some terms used in this application will be explained to facilitate understanding by those skilled in the art, and will not be used to limit the scope of this application.
[0053] (1) Main Chain: The main chain refers to the officially launched and independently operating blockchain network, also known as the mainnet or mother chain. It is the core chain of the entire blockchain network, carrying the network's basic functions and transaction records. Each block on the main chain contains a large number of transaction records, which are connected together in a certain order to form an immutable ledger. Every node on the main chain needs to participate in the block verification and accounting process, which is called the consensus mechanism.
[0054] (2) Sidechain: A sidechain is a chain that exists in parallel with the main chain and can interoperate with the main chain through smart contracts. Sidechains are completely independent of the main chain, but the two ledgers can "interoperate" and interact. Sidechains typically have higher processing speeds and lower fees, and different consensus mechanisms can be customized as needed. The security of a sidechain primarily depends on the interoperability between the smart contracts and the main chain.
[0055] In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments may be combined in any suitable form.
[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0057] The following description, in conjunction with the accompanying drawings, details a specific solution for a blockchain-based enterprise data asset management method and system provided by this invention.
[0058] Please see Figure 1 The diagram illustrates a flowchart of a blockchain-based enterprise data asset management method according to an embodiment of the present invention. The method includes the following steps:
[0059] Step S100: Obtain enterprise asset data storage in each business module of each sidechain.
[0060] A typical blockchain consists of a series of chronologically ordered "blocks." Each block contains a certain number of transaction records or data and is linked to the previous block using cryptographic technology, forming a chain-like structure, such as the main chain. Every node in the network can participate in verifying and recording transactions, and each node has a complete copy of the ledger, ensuring the data is open and transparent.
[0061] In enterprise data management, the sheer volume and high frequency of data access can lead to a decrease in the efficiency of a limited-scale main chain. Therefore, this embodiment introduces sidechains to extend the main chain. A sidechain can be viewed as a new blockchain attached to the main chain, typically smaller in scale. This means that instead of delegating all computations and tasks to the main chain, a "sidechain" is added to customize its services. The main chain and sidechain communicate through two-way anchoring, enabling data synchronization between them.
[0062] Therefore, to improve the operating efficiency of the main chain, this embodiment stores the most critical data on the main chain, while storing less critical, less frequently accessed data on the side chains. This reduces the transaction load on the main chain and improves processing speed.
[0063] Based on this, the company first acquires relevant data assets from its database, such as financial data, transaction data, work documents, and announcements. Data belonging to the same business is stored in the same business module. Then, based on the importance of different business modules, they are stored in specific sidechains. This ensures that more important company data is stored in a more secure blockchain. The number of sidechains and the number of business modules stored within them are determined by company professionals to initially formulate a scaling strategy.
[0064] At this point, a preliminary expansion strategy for sidechains has been established. To further improve data management efficiency, further analysis of merging and synchronization operations can be conducted on the preliminary expansion strategy.
[0065] Step S200: Based on the interaction information between each business module in each sidechain and each business module in other sidechains, and combined with the core differences between the data contained in different business modules, the sidechains are merged to determine the initial storage strategy of the blockchain structure.
[0066] The main purpose of this embodiment is to optimize the sidechain expansion strategy. If the enterprise's business is large, the initial sidechain expansion strategy will lead to too many sidechains attached to the main chain, which will cause management difficulties and increase the complexity of cross-chain operations. Therefore, it is necessary to merge sidechains containing different business modules to optimize the efficiency of sidechain distributed management.
[0067] Based on this, we first analyze the frequency of interaction between each business module in each sidechain and each business module in other sidechains to determine the degree of dependency between two different business modules. Business modules with higher interaction frequency have a greater degree of dependency.
[0068] Then, based on the results of the dependency analysis, we analyze the differences in the coreness of various types of enterprise data contained in the two different business modules. This can quantitatively characterize the necessity of merging the sidechains corresponding to the two different business modules, that is, the degree of data merging.
[0069] Finally, when the data merging degree of two different business modules meets the threshold requirement, the enterprise data contained in the two corresponding sidechains can be merged and stored to improve the efficiency of sidechain distributed management.
[0070] Step S300: Under the initial storage strategy, based on the memory usage and bandwidth information of each sidechain at each update time, analyze the load changes of the corresponding sidechain, and adjust the synchronization degree of each sidechain at each update time in real time in combination with the update status of the sidechain at historical update times to determine the global storage strategy of the blockchain structure.
[0071] Based on the initial storage strategy, considering that the size of sidechains is generally limited, when multiple sidechains handle different data requests, if some sidechains have excessive load while others have low load, it may lead to bottlenecks and affect the overall system performance. Therefore, it is necessary to set up load balancing mechanisms for different sidechains to ensure stability.
[0072] Since sidechain architecture enhances the scalability and performance of the main chain, data synchronization between the main chain and sidechains is essential. Through data synchronization, the main chain ensures that every data change on the sidechain is correctly recorded on the main chain, preventing data loss, duplicate transactions, or incorrect asset states. Therefore, the data synchronization frequency between the main chain and sidechains needs to be adjusted based on dynamic load, giving sidechains with lower loads a higher synchronization frequency and those with higher loads a lower synchronization frequency, thereby improving the collaborative efficiency between multiple sidechains and the main chain. In this way, based on the initial storage strategy, optimizing the data synchronization frequency between the sidechains and the main chain further optimizes the sidechain expansion strategy, resulting in a global storage strategy for a blockchain structure with superior final management performance.
[0073] In some embodiments, the main purpose of step S100 is to transfer some non-core types of enterprise data to the sidechain to improve the efficiency of the mainchain. Therefore, this step requires classifying the different data asset types in the enterprise according to their importance and determining the corresponding storage location based on the business modules with different levels of importance, so as to clarify the enterprise asset data storage in each business module of each sidechain. Specifically, as... Figure 2 As shown, step S100 can be implemented by steps S101 to S104.
[0074] Step S101: Obtain the confidentiality tag value, access frequency, and update frequency of each type of enterprise data in the enterprise asset data.
[0075] Enterprise asset data includes different types of enterprise data, such as financial data, transaction information data, work documents, announcements, logistics data, sales order data, inventory management data, etc. These different types of enterprise data can be collected from the enterprise's internal database.
[0076] First, for different types of enterprise data, the level of sensitivity determines the degree of importance. Generally, the leakage of highly sensitive enterprise data, such as financial statements and contracts, can have a more serious impact on the enterprise. Therefore, in this embodiment, a confidentiality flag value is obtained for each type of enterprise data. The magnitude of this value represents the confidentiality level and importance of the corresponding enterprise data.
[0077] As a specific example, in this embodiment, the enterprise's professional operations and maintenance personnel mark the confidentiality level of each type of enterprise data to obtain the corresponding confidentiality mark value. The confidentiality mark value can be 1, 2, 3, 4 and 5. The larger the value, the higher the confidentiality level of the corresponding type of enterprise data, that is, the higher the importance level.
[0078] In enterprise data management, update frequency is a core indicator for measuring data dynamism. By quantifying data values in a simple and intuitive way, data activity can be quickly determined, providing a basis for storage optimization and resource allocation. In this embodiment, the number of data update operations for each type of enterprise data is collected in the enterprise database to quantify the update frequency of the corresponding data type.
[0079] Specifically, the total number of data update operations for each type of enterprise data within a certain period is obtained, and the ratio of the total number of update operations to the length of the time range is used as the update frequency for each type of enterprise data. The length of the time range can be set to 7 days; in other embodiments, the implementer can set this value according to the specific implementation scenario.
[0080] Furthermore, the access frequency of enterprise data represents the degree of interaction between the corresponding data, and the more frequently an enterprise accesses the data, the more likely it is to involve the core of the enterprise. In this embodiment, the access frequency of each type of enterprise data is quantified by collecting the number of times each type of enterprise data is accessed in the enterprise database.
[0081] Specifically, the total number of times each type of enterprise data is accessed within a certain period of time is obtained, and the ratio of the total number of accesses to the length of the time period is taken as the access frequency of each type of enterprise data.
[0082] Thus, the update frequency and access frequency of each type of enterprise data can characterize the degree to which the corresponding enterprise data belongs to the core of the enterprise's business, providing a data foundation for analyzing the coreness of enterprise data.
[0083] Step S102: Based on the confidentiality tag value of each type of enterprise data, the access frequency and update frequency of all enterprise data of the same type, obtain the core index of each type of enterprise data.
[0084] The first step is to determine the first importance coefficient for each type of enterprise data based on the confidentiality tag value for each type of enterprise data and the average access frequency of all enterprise data of the same type.
[0085] The second step is to determine the second importance coefficient for each type of enterprise data based on the difference between the maximum update frequency of all types of enterprise data and the mean update frequency of all enterprise data for each type.
[0086] The third step is to obtain the core index of each type of enterprise data based on the first importance coefficient and the second importance coefficient.
[0087] As a concrete example, taking any type of enterprise data as an illustration, the formula for calculating the core index of type a enterprise data can be expressed as:
[0088]
[0089] in, Let 'a' represent the core index of type a enterprise data, where 'a' represents type a enterprise data, a = 1, 2, 3, ..., N, and N represents the total number of enterprise data types. This represents the confidentiality marker value for type a of enterprise data. This represents the average access frequency of all enterprise data of type a. This represents the average update frequency of all enterprise data of type a. This represents the maximum update frequency for all types of enterprise data. This represents a preset hyperparameter, with a value range of (0, 0.1). In this embodiment, the value is 0.01, which is intended to prevent the denominator from being 0 and affecting the data analysis results.
[0090] The most important coefficient is the core value. The higher the value, the more frequently the enterprise data is accessed and the higher the degree of confidentiality required. This indicates that the enterprise data of this type is more important and more likely to belong to the core part of the enterprise's business. The corresponding core value is higher.
[0091] The second most important coefficient represents the relative update status of enterprise data of type a. The smaller the value, the closer the update frequency of enterprise data of this type is to the maximum update status, indicating that the data of this type of enterprise is updated more frequently. Combined with the larger value of the first most important coefficient, the more likely the data represented is to belong to the core part.
[0092] Thus, by following the same method, we can obtain the core index for each type of enterprise data. This index represents the extent to which each type of enterprise data belongs to the core of the enterprise and reflects the importance of each type of enterprise data. It is understandable that the higher the core index value, the higher the confidentiality and storage security requirements for the type of enterprise data.
[0093] Step S103: Based on the differences in core indicators between different types of enterprise data, classify enterprise data into different core levels.
[0094] The purpose of this step is to classify different types of enterprise data by coreity to determine the appropriate storage location.
[0095] The first step is to arrange the core indicators of all types of enterprise data in a preset order to obtain a core indicator sequence. In this embodiment, the core indicators of all types of enterprise data are arranged in descending order to obtain the core indicator sequence.
[0096] The second step is to determine the core difference value corresponding to each pair of adjacent core degree indicators based on the difference between them in the core degree sequence; and to determine the grade classification reference value corresponding to each core difference value based on the ratio between each core difference value and the mean of all core difference values.
[0097] In this embodiment, for the core degree sequence, the absolute value of the difference between the i-th core degree indicator and the (i+1)-th core degree indicator is calculated to obtain the core difference value corresponding to the i-th adjacent core degree indicator, i=1,2,3,...,N, where N represents the total number of enterprise data types and also reflects the total number of core indicators in the core degree sequence. Then, N-1 core difference values can be calculated, and the ratio between each core difference value and the mean of all core difference values is used as the level classification reference value corresponding to each core difference value.
[0098] It is understandable that this embodiment uses the ratio of the core difference value to the corresponding mean value to reflect the relative difference of each core difference value, so as to avoid the excessive difference between two adjacent core degree indicators affecting the segmentation result.
[0099] The third step is to obtain a preset number of reference values for level division as dividing points in descending order, and then divide the core degree sequence to obtain the enterprise data of each data segment belonging to a core level.
[0100] When the reference values for level classification are arranged in descending order, the changes in adjacent coreity indicators can be seen more intuitively, and an appropriate number of divisions can be selected as needed. For example, in this embodiment, all types of enterprise data are divided into three levels, so only two division points need to be selected. Therefore, the preset number in this embodiment is 2. Two level classification reference values are selected in descending order. The positions of the two adjacent coreity indicators corresponding to the level classification reference values are used as the division points. Data segmentation processing of the coreity sequence can be performed to obtain three data segments. Enterprise data of the same type corresponding to the coreity indicators contained in each data segment can be regarded as the same core level.
[0101] Furthermore, after the core sequence is tiered, the first data segment corresponding to the enterprise data type represents the high core level. Enterprise data at this level has the highest level of confidentiality and is accessed frequently, requiring the highest level of security and privacy protection. The second data segment corresponds to the auxiliary level, where enterprise data has a medium level of confidentiality and a medium frequency of access and updates, requiring security and privacy protection that is slightly lower than the core level. The third data segment corresponds to the non-core level, where enterprise data has a lower level of confidentiality and is accessed and updated less frequently, and can be stored on a less secure blockchain.
[0102] Step S104: Enterprise data belonging to the same business and the same core level are treated as the same business module. Different types of enterprise data contained in each business module are stored in a sidechain. A sidechain includes at least one business module.
[0103] In this embodiment, a fixed number of sidechains is used, and the initial expansion strategy sets up a corresponding sidechain for each business. Specifically, all types of enterprise data at the same core level contained in each business are stored in one sidechain. That is, one sidechain stores one business module, and one business module stores all types of enterprise data contained in one business that belong to the same core level.
[0104] Understandably, through examples, customer management, transaction information, and sales order data belong to sales and marketing operations; supplier information, inventory management, and logistics data belong to supply chain management operations; and billing, payment, budget, and financial reporting data belong to financial accounting operations. It can be seen that different companies have different businesses that include different types of corporate data, depending on the specific implementation scenario.
[0105] In some embodiments, such as Figure 3 The step S200 shown can be implemented by steps S201 to S203.
[0106] Step S201: Based on the interaction frequency between any two different business modules and the repetition of interaction objects that have interactive operations with each business module, the data dependency degree between any two different business modules is obtained.
[0107] Because sidechains store various types of enterprise data, and these different types of data correspond to different business needs—for example, some enterprise data is used for high-frequency trading or low-latency applications, while other enterprise data only needs to be stored and queried—if there are many interactions between two different business modules, storing them on different sidechains may lead to frequent cross-chain operations, thereby increasing management difficulty and reducing data management efficiency.
[0108] Therefore, by analyzing the details of the interactions between different business modules, we can quantify the degree of operational and data dependencies between each pair of business modules. The stronger the dependency between two different business modules, the greater the help in improving data management efficiency through sidechain merging.
[0109] Based on this feature, we take any two different business modules as an example to analyze the dependency characteristics. Specifically, we obtain any two different business modules as the first business module and the second business module, respectively.
[0110] The first step is to identify all different business modules that interact with the first business module, forming a first interaction set; and to identify all different business modules that interact with the second business module, forming a second interaction set. It's understandable that these interactions between different business modules manifest as data requests, service calls, or transaction coordination, etc. For example, the sales business module might call the inventory query interface of the supply chain business module; the finance business module might synchronize transaction settlement results with the sales business module, which can be directly obtained from the company's management logs.
[0111] The second step is to determine the data dependency between the first and second business modules based on the number of business modules contained in the intersection of the first and second interaction sets and the interaction frequency between the first and second business modules.
[0112] Specifically, the total number of different business modules contained in the intersection of the first and second interaction sets is denoted as the feature count. The feature count represents the overlap of business modules that have interactive operations between two different business modules. The larger the value, the higher the overlap of the interactive operations between the two different business modules, the greater the business similarity, and the higher the management efficiency of storing these two business modules in the same blockchain in the daily management work of an enterprise.
[0113] Furthermore, this embodiment obtains the ratio of the number of interactions between the first business module and the second business module within a preset historical time period to the length of the historical time period, as the interaction frequency between the first and second business modules. The preset historical time period is the 7 days prior to the current moment, which can be set by the implementer according to the specific implementation scenario. By analyzing the frequency of interaction between two different business modules within a certain historical time period, it can reflect the need for collaborative management between different business modules in the company's daily management. The more frequent the interaction, the greater the necessity for collaborative management, that is, the greater the dependence between the two.
[0114] The product of the interaction frequency between the first and second business modules and the number of features is used as the data dependency degree between the two business modules. The data dependency degree comprehensively characterizes the extent of dependence between the two different business modules in enterprise data management based on the results of feature analysis from both aspects. A higher value indicates a greater dependency of the data contained in the two business modules on enterprise data management, making them more suitable for unified storage. In other words, merging the sidechains containing these two business modules can effectively improve enterprise data management efficiency.
[0115] Step S202: Based on the differences in core metrics between any two different business modules, and in conjunction with the degree of data dependency, the degree of data merging between any two different business modules is obtained.
[0116] Firstly, the degree of data dependency reflects the efficiency of merging data management between two different business modules from the perspective of the interaction information between them. Furthermore, considering the core nature and importance of the enterprise data contained in the two different business modules, the efficiency of merging data management can be better reflected by combining data characteristic attributes.
[0117] Specifically, based on the difference between the average coreity index of all types of enterprise data contained in the first business module and the average coreity index of all types of enterprise data contained in the second business module, similarity feature factors are determined; the normalized value of the product of the data dependence degree of the first business module and the second business module and the similarity feature factors is taken as the degree of data merging between the first business module and the second business module.
[0118] As a concrete example, if we take business module A as the first business module and business module B as the second business module, then the formula for calculating the degree of data merging between the first and second business modules can be expressed as:
[0119]
[0120] in, This indicates the degree of data merging between the first and second business modules. A represents business module A, which is the first business module, and B represents business module B, which is the second business module. This indicates the degree of data dependency between the first business module and the second business module. This represents the average coreity index, which is the mean of all types of enterprise data contained in the first business module. This represents the average coreity index, which is the mean of all types of enterprise data contained in the second business module. This represents a preset hyperparameter, with a value range of (0, 0.1). In this embodiment, the value is 0.01 to prevent a denominator of 0 from affecting the data analysis results. This is the normalization function.
[0121] The similarity feature factor represents the result of negatively correlated processing of the differences between the average coreity indicators corresponding to the first and second business modules. This reflects the difference in average coreity metrics between the first and second business modules, and the ratio. The greater the difference between 1 and 1, the greater the difference between their average coreity indices, and the smaller the value of the corresponding similarity feature factor, which means the smaller the similarity between them.
[0122] The more similar the core of two business modules, the more similar the importance of their corresponding data types. In this case, the greater the dependency between the two business modules, the better the effect of reducing management complexity after classifying and merging the sidechains of the two business modules, and the greater the value of the corresponding data merging degree.
[0123] In some embodiments, the similarity feature factor can also be expressed as ,in This represents the difference between the average core index corresponding to the first business module and the second business module. exp represents an exponential function with the natural constant e as the base. In other words, this function is used to negatively correlate the difference between the average core index corresponding to the first business module and the second business module to obtain similarity feature factors.
[0124] Step S203: Merge and store the two sidechains containing two different business modules whose data merging degree is greater than or equal to the preset merging threshold to obtain the initial storage strategy of the blockchain structure.
[0125] The degree of data merging between two different business modules represents the efficiency of data management between the two business modules. The higher the value of the degree of data merging, the more effectively the complexity of data management can be reduced and the efficiency of data management can be improved by merging and classifying the sidechains where the two different business modules are located. Therefore, the enterprise data of the sidechains where the two different business modules are located can be merged and stored.
[0126] Specifically, in this embodiment, the merging threshold is set to 0.8, which can be set by the implementer according to the specific implementation scenario. When the degree of data merging between two different business modules is greater than the merging threshold, it indicates that the necessity of merging and classifying management is greater. After merging and classifying storage, the data management efficiency can be effectively improved. That is, at this time, all data contained in the sidechains of the two different business modules that meet the threshold need to be merged and stored, which means they can be stored in the same sidechain.
[0127] It should be understood that priority is given to merging the two sidechains containing the business modules with the highest data merging degree values. Once a sidechain has been merged, it is no longer compared with business modules in other sidechains. Thus, by traversing all the business modules contained in all sidechains, the initial merging of sidechains is achieved, resulting in the initial storage strategy for the blockchain structure. It should be noted that the initial storage strategy, based on a fixed number of sidechain storage strategies, selects those parts that can be merged and stored based on similarities in the interaction and core information between business modules in the sidechains. This reduces the complexity of data management and improves its efficiency, resulting in the initial storage strategy for the blockchain structure.
[0128] In some embodiments, such as Figure 4 As shown, step S300 can be implemented by steps S301 to S303.
[0129] Step S301: Obtain the memory usage and remaining bandwidth usage of each sidechain at each update time.
[0130] Under normal circumstances, the size of sidechains is also limited. When multiple sidechains handle different data management requests, if some sidechains are overloaded while others are underloaded, it may lead to bottlenecks and affect the overall performance of data management. Therefore, it is necessary to analyze and adjust the load balancing of different sidechains to ensure the stability of data management.
[0131] Firstly, considering that different sidechains are independent of each other and have their own memory and bandwidth, we can analyze the load of the sidechains at different times, set a dynamic weight value for each sidechain, so that it can adjust the weight according to the current load, and then allocate different data requests to different sidechains to balance the load.
[0132] Secondly, the sidechain architecture improves the scalability and performance of the main chain. Data synchronization between the main chain and sidechains is essential. Through data synchronization, the main chain can ensure that every data change on the sidechain is correctly recorded on the main chain, preventing data loss, duplicate transactions, or incorrect asset states. Therefore, the frequency of data synchronization between the main chain and sidechains also needs to be adjusted according to the dynamic load.
[0133] Based on these two characteristics, to provide a data foundation for subsequent analysis, it is first necessary to obtain the memory usage and remaining bandwidth usage of each sidechain at each update time. The remaining bandwidth usage can be the difference between the total bandwidth and the real-time bandwidth usage. Specifically, in order to analyze the memory and bandwidth usage of the sidechains in the initial storage strategy in real time, their storage information is monitored at each update time.
[0134] It is understandable that the time interval between adjacent update moments is the same, and implementers can set it according to the specific implementation scenario. This embodiment uses the current update moment as an example for illustration. The methods for obtaining memory utilization and bandwidth remaining utilization are well-known technologies, such as directly obtaining them through system management logs, and will not be described in detail here. At the same time, considering that the unit of measurement may affect the data analysis results, the memory utilization and bandwidth remaining utilization obtained in this embodiment are both standardized data.
[0135] Step S302: Analyze the load of the corresponding sidechain based on the memory occupancy rate and the remaining bandwidth occupancy rate, and combine the degree of change in the load of each sidechain at adjacent update times to obtain the dynamic load weight of each sidechain at each update time.
[0136] Firstly, the load of each sidechain is quantitatively characterized by analyzing its memory and bandwidth usage in two steps. Firstly, for any sidechain, the load coefficient is obtained at each update time based on its memory usage and remaining bandwidth usage. The memory usage is positively correlated with the load coefficient, while the remaining bandwidth usage is negatively correlated with the load coefficient.
[0137] As a concrete example, the load factor of any sidechain at the current update moment can be expressed as: ,in This represents the load factor of the sidechain at the current update time, where t represents the current update time. This represents the memory usage of the sidechain at the current update moment, and the remaining bandwidth usage of the sidechain at the current update moment. This represents a preset hyperparameter, with a value range of (0, 0.1). In this embodiment, the value is 0.01, which is intended to prevent the denominator from being 0 and affecting the data calculation result.
[0138] The lower the memory usage and the higher the remaining bandwidth usage of a sidechain at the current update moment, the lower its load coefficient, indicating a lower load on the sidechain. A higher load coefficient for a sidechain corresponds to a higher load, and vice versa. The combination of high memory usage and low remaining bandwidth directly reflects a high load state.
[0139] The second step is to adjust the load coefficient of the sidechain at each update time by using the ratio of the load coefficient of the sidechain at each update time to the load coefficient of the adjacent previous historical update time, so as to obtain the dynamic load weight of the sidechain at each update time. The dynamic load weight is a normalized value.
[0140] If the sidechain's load is higher at the current update time compared to the previous update time, and the sidechain's load coefficient is also higher at the current update time, it indicates that the sidechain has frequent data requests and data storage at this time, and the corresponding load weight is higher. Conversely, if the load is lower, the load weight is also lower.
[0141] Specifically, for any sidechain, the dynamic load weight at the current moment can be expressed as: ,in, This indicates the dynamic load weight of the sidechain at the current update moment. This represents the load factor of the sidechain at the current update moment. This represents the load factor of the sidechain at the previous historical update time adjacent to the current update time. This is the normalization function.
[0142] It should be noted that dynamic load weight analysis is not performed for the current update time when historical update times are unavailable. Understandably, the dynamic load weights of all sidechains at the same update time can be obtained, and load is allocated based on these weights using a round-robin mechanism. A higher weight corresponds to a sidechain receiving more and more frequent data requests, while a lower weight reduces data requests, thus ensuring efficient system operation and scalability under high load.
[0143] Step S303: Based on the dynamic load weight of each update time and the adjacent previous historical update time, adjust the synchronization frequency of the sidechain and the main chain at each historical update time to obtain the adjustment time interval for each update time.
[0144] Considering that the data synchronization frequency between the main chain and side chains also needs to be adjusted according to the dynamic load, side chains with lower loads have higher synchronization frequencies, while side chains with higher loads have lower synchronization frequencies, in order to improve the coordination efficiency between multiple side chains and the main chain.
[0145] Specifically, obtaining the time interval for data synchronization between each sidechain and the main chain needs to be determined based on the specific implementation scenario. Taking any sidechain as an example, the ratio of the dynamic load weight of each update moment to the dynamic load weight of the adjacent previous historical update moment is used as an adjustment coefficient. The synchronization time interval of the historical update moments is adjusted using the adjustment coefficient to obtain the adjusted time interval for each update moment. The time interval is used to characterize the data synchronization frequency between the corresponding sidechain and the main chain.
[0146] More specifically, taking the current update moment as an example, the adjustment time interval of the sidechain at the current update moment can be expressed as: ,in, This indicates the adjustment time interval of the sidechain at the current update moment. This represents the synchronization time interval between the previous historical update time adjacent to the current update time of the sidechain. This indicates the dynamic load weight of the sidechain at the current update moment. This represents the dynamic load weight of the sidechain at the previous historical update time adjacent to the current update time. This indicates rounding up to the nearest integer.
[0147] The shorter the time interval for information synchronization between the sidechain and the main chain, the faster the synchronization frequency of the corresponding sidechain. The synchronization frequency of the corresponding sidechain is adjusted in real time according to different load update times to reduce system latency. It should be noted that if relevant parameters at historical moments cannot be obtained, no adjustment analysis will be performed for the current operation. It can be understood that when the adjacent previous historical moment has been adjusted, the obtained synchronization time interval is also the actual parameter after adjustment. Therefore, this embodiment describes it as a synchronization time interval. Implementers can analyze it according to specific implementation scenarios to obtain the actual synchronization degree between the sidechain and the main chain at historical update times.
[0148] If the dynamic load weight at the current update time The dynamic load weight is higher than that of the previous time step. This indicates that the load has increased, and the corresponding synchronization time interval should be increased, meaning the adjustment time interval should be set larger to reduce the synchronization frequency and minimize the impact of synchronization on the performance of the high-load sidechain. If the dynamic load weight at the current update time is lower than the dynamic load weight at the previous update time, it means the load has decreased, and the corresponding synchronization time interval should be reduced, meaning the adjustment time interval should be set smaller to increase the synchronization frequency and effectively ensure data consistency.
[0149] This mechanism improves the throughput and stability of multi-sidechain systems by dynamically adjusting data flow distribution and synchronization strategies based on real-time awareness of sidechain resource status, while reducing pressure on the main chain. By combining main chain and sidechain strategies, this approach enables distributed management of enterprise data assets. The sidechain locks data or transactions on the first blockchain (main chain), then creates and modifies the corresponding data or transactions on the second blockchain (sidechain). It then provides cryptographic proof of the correct locking on the first blockchain to facilitate interaction and synchronization, thus resolving the issues of high system response latency and low efficiency associated with a single main chain.
[0150] Meanwhile, due to the independence of sidechains from the main chain, the design can be more flexible. Sidechains can use consensus mechanisms, data storage methods or encryption algorithms different from those of the main chain to optimize performance, but this will also increase maintenance costs. Therefore, enterprises can determine the scale of sidechains according to their actual situation.
[0151] In some embodiments, such as Figure 5 As shown, this embodiment provides a blockchain-based enterprise data asset management system. This system implements the steps of a blockchain-based enterprise data asset management method. Specifically, the blockchain-based enterprise data asset management system includes:
[0152] The data acquisition module is used to acquire enterprise asset data from each business module in each sidechain.
[0153] The initial strategy analysis module is used to merge sidechains based on the interaction information between each business module in each sidechain and each business module in other sidechains, combined with the core differences between the data contained in different business modules, and to determine the initial storage strategy of the blockchain structure.
[0154] The global strategy analysis module is used to analyze the load changes of each sidechain based on the memory usage and bandwidth information of each sidechain at each update time under the initial storage strategy. It also adjusts the synchronization degree of each sidechain at each update time in real time by combining the update status of the sidechain at historical update times, and determines the global storage strategy of the blockchain structure.
[0155] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A blockchain-based enterprise data asset management method, characterized in that, The method includes the following steps: Obtain enterprise asset data for each business module in each sidechain; Based on the interaction information between each business module in each sidechain and each business module in other sidechains, and combined with the core differences between the data contained in different business modules, the sidechains are merged to determine the initial storage strategy of the blockchain structure. This includes: obtaining the confidentiality marker value, access frequency, and update frequency of each type of enterprise data in the enterprise asset data; obtaining the coreity index of each type of enterprise data based on the confidentiality marker value of each type of enterprise data, the access frequency and update frequency of all enterprise data of the same type; obtaining the data dependency degree between any two different business modules based on the interaction frequency between any two different business modules and the repetition of the interaction objects that interact with each business module; obtaining the data merging degree between any two different business modules based on the differences in the coreity index between any two different business modules and the data dependency degree; and merging and storing the two sidechains containing the two different business modules whose data merging degree is greater than or equal to a preset merging threshold to obtain the initial storage strategy of the blockchain structure. Under the initial storage strategy, based on the memory usage and bandwidth information of each sidechain at each update time, the load changes of the corresponding sidechain are analyzed, and the synchronization degree of each sidechain at each update time is adjusted in real time in combination with the update status of the sidechain at historical update times, so as to determine the global storage strategy of the blockchain structure.
2. The blockchain-based enterprise data asset management method according to claim 1, characterized in that, The acquisition of enterprise asset data storage in each business module of each sidechain specifically includes: Based on the differences in core indicators among different types of enterprise data, enterprise data is classified into different core levels; Enterprise data belonging to the same business and the same core level are treated as the same business module. Different types of enterprise data contained in each business module are stored in a sidechain. Each sidechain includes at least one business module.
3. The blockchain-based enterprise data asset management method according to claim 2, characterized in that, The core index for each type of enterprise data is obtained based on the confidentiality marker value for each type of enterprise data, the access frequency and update frequency of all enterprise data of the same type, and specifically includes: Based on the confidentiality tag value of each type of enterprise data and the average access frequency of all enterprise data of the same type, determine the first importance coefficient corresponding to each type of enterprise data; The second importance coefficient for each type of enterprise data is determined based on the difference between the maximum update frequency of all types of enterprise data and the mean update frequency of all enterprise data for each type. The coreness index for each type of enterprise data is obtained based on the first importance coefficient and the second importance coefficient.
4. The blockchain-based enterprise data asset management method according to claim 2, characterized in that, The method of classifying enterprise data into different core levels based on the differences in core indicators between different types of enterprise data specifically includes: The core indicators of all types of enterprise data are arranged in a preset order to obtain a core sequence; Based on the difference between every two adjacent core degree indicators in the core degree sequence, determine the core difference value corresponding to every two adjacent core degree indicators; based on the ratio between each core difference value and the mean of all core difference values, determine the grade classification reference value corresponding to each core difference value. Following the order from largest to smallest, a preset number of reference values for level division are obtained as dividing points to divide the core degree sequence, thus determining that the enterprise data of each data segment belongs to a core level.
5. A blockchain-based enterprise data asset management method according to claim 1, characterized in that, The process of determining the data dependency between any two different business modules based on the interaction frequency between them and the repetition of interaction objects that interact with each business module specifically includes: Select any two different business modules as the first business module and the second business module respectively; All different business modules that have interactive operations with the first business module are collected to form a first interaction set; all different business modules that have interactive operations with the second business module are collected to form a second interaction set. The data dependency between the first and second business modules is determined based on the number of business modules contained in the intersection of the first and second interaction sets and the interaction frequency between the first and second business modules.
6. A blockchain-based enterprise data asset management method according to claim 5, characterized in that, The process of determining the degree of data merging between any two different business modules based on the differences in core metrics and the degree of data dependency specifically includes: Based on the difference between the average coreity index of all types of enterprise data contained in the first business module and the average coreity index of all types of enterprise data contained in the second business module, similarity feature factors are determined; the normalized value of the product of the data dependency degree of the first business module and the similarity feature factors is taken as the degree of data merging between the first business module and the second business module.
7. A blockchain-based enterprise data asset management method according to claim 1, characterized in that, The process involves analyzing the load changes of each sidechain based on its memory usage and bandwidth information at each update moment, and adjusting the synchronization level of each sidechain at each update moment in real time, in conjunction with the historical update status of the sidechains, to determine the global storage strategy for the blockchain structure. Specifically, this includes: Obtain the memory usage and remaining bandwidth usage of each sidechain at each update time; The load of the corresponding sidechain is analyzed based on the memory utilization and remaining bandwidth utilization. The dynamic load weight of each sidechain at each update time is obtained by combining the degree of change in the load of each sidechain at adjacent update times. For any sidechain, the ratio of the dynamic load weight of each update moment to the adjacent previous historical update moment is used as an adjustment coefficient. The synchronization time interval of the historical update moment is adjusted using the adjustment coefficient to obtain the adjustment time interval of each update moment. The time interval is used to characterize the data synchronization frequency between the corresponding sidechain and the main chain.
8. A blockchain-based enterprise data asset management method according to claim 7, characterized in that, The step of analyzing the load status of the corresponding sidechain based on the memory occupancy rate and remaining bandwidth occupancy rate, and combining the degree of load change of each sidechain at adjacent update times, yields the dynamic load weight of each sidechain at each update time, specifically including: For any sidechain, the load coefficient of the sidechain at each update time is obtained based on the memory utilization and remaining bandwidth utilization of the sidechain at each update time. The memory utilization is positively correlated with the load coefficient, and the remaining bandwidth utilization is negatively correlated with the load coefficient. By using the ratio of the load coefficient of the sidechain at each update time to that of the adjacent previous historical update time, the load coefficient at each update time is adjusted to obtain the dynamic load weight of the sidechain at each update time. The dynamic load weight is a normalized value.
9. A blockchain-based enterprise data asset management system, characterized in that, This system is used to implement the steps of a blockchain-based enterprise data asset management method as described in any one of claims 1-8, wherein the blockchain-based enterprise data asset management system specifically includes: The data acquisition module is used to acquire enterprise asset data from each business module in each sidechain. The initial strategy analysis module is used to merge sidechains based on the interaction information between each business module in each sidechain and each business module in other sidechains, combined with the core differences between the data contained in different business modules, and to determine the initial storage strategy of the blockchain structure. The global strategy analysis module is used to analyze the load changes of each sidechain based on the memory usage and bandwidth information of each sidechain at each update time under the initial storage strategy. It also adjusts the synchronization degree of each sidechain at each update time in real time by combining the update status of the sidechain at historical update times, and determines the global storage strategy of the blockchain structure.
Citation Information
Patent Citations
Data storage optimization method based on block chain
CN119293049A
Data asset operation technology platform, operation method, storage medium and electronic equipment
CN119809643A