A Big Data-Based Data Asset Evaluation Method and System

By building a graph network of data assets and using GraphSAGE model and XGBoost model, the insufficient capture of dynamic relationship changes in data asset value evaluation is solved, and a more comprehensive data asset evaluation is achieved, which improves the credibility and practicality of the evaluation results.

CN119006165BActive Publication Date: 2025-07-04CHINA NAT INST OF STANDARDIZATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410936211.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-12
Publication Date
2025-07-04
Estimated Expiration
2044-07-12

AI Technical Summary

Technical Problem

The existing technology lacks effective capture of dynamic changes in relationships between data assets, resulting in insufficient comprehensiveness and accuracy of the evaluation of data assets value.

Method used

By collecting multi-source data and storing it in a central database and uploading it to the blockchain for data rights confirmation and recording, a graph network of data assets is built, a graph network is captured, the time changes of data relationships in the graph network are extracted, and the graph data model is constructed using Neo4j and Elasticsearch, and the data asset value evaluation is evaluated in combination with the GraphSAGE model and the XGBoost model, and access control is implemented.

Benefits of technology

It realizes accurate analysis of the dynamic relationship between data assets, identifies key nodes and important relationships, improves the depth and breadth of data analysis, provides a more dynamic and comprehensive evaluation perspective, and significantly improves the credibility and practicality of the evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119006165B_ABST
    Figure CN119006165B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for data asset evaluation based on big data, which relates to the technical field of data processing. It includes collecting multi-source data, storing it in a central database, and uploading it to a blockchain for data right confirmation and recording; constructing a graph network of data assets based on the central database, capturing the temporal changes of data relationships in the graph network, and extracting node features from the graph network; identifying factors affecting the value of data assets based on the extracted node features, and constructing an asset evaluation model to evaluate the value of data assets. By collecting multi-source data to construct a graph network of data assets, extracting node features of the graph network to identify factors affecting the value of data assets, and constructing an asset evaluation model for asset evaluation, the present invention can accurately analyze the dynamic relationships between data assets, identify key nodes and important relationships, improve the depth and breadth of data analysis, provide a more dynamic and comprehensive evaluation perspective, and significantly improve the credibility and practicality of evaluation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a method and system for data asset evaluation based on big data. Background Art

[0002] In the context of the rapid development of modern information technology, big data has become the core asset to support corporate strategic decision-making and operational optimization. The method for data asset evaluation is an important research direction in the fields of information science and data management, and the key lies in how to accurately and efficiently evaluate and utilize the value of these data assets. With the improvement of computing power and the progress of data acquisition technology, the integration and analysis of multi-source data have become possible, but at the same time, new challenges have emerged. Existing methods often lack the effective capture of the dynamic changes in the relationships between data assets, are insufficient in identifying the key nodes and important relationships of data assets, and there is still room for further improvement in the comprehensiveness and accuracy of the value evaluation of data assets. Therefore, a more dynamic and in-depth method for data asset evaluation is needed. Summary of the Invention

[0003] The purpose of the present invention is to provide a method and system for data asset evaluation based on big data to solve the above problems.

[0004] The present invention is achieved through the following technical solutions:

[0005] A method and system for data asset evaluation based on big data, which includes:

[0006] Collect multi-source data, store it in a central database, and upload it to the blockchain for data right confirmation and recording;

[0007] Construct a graph network of data assets based on the central database, capture the temporal changes in the data relationships in the graph network, and extract node features from the graph network;

[0008] Identify the factors affecting the value of data assets based on the extracted node features, and construct an asset evaluation model to evaluate the value of data assets;

[0009] Store the data generated during the data asset evaluation process and implement access control.

[0010] As a preferred solution of the method for data asset evaluation based on big data of the present invention, where: the step of collecting multi-source data, storing it in a central database, and uploading it to the blockchain for data right confirmation and recording refers to collecting data from each data source, performing preprocessing, and then storing it in the central database, writing a smart contract using the Solidity language, deploying the smart contract on the Ethereum platform, and storing the data on the blockchain.

[0011] As a preferred solution of the big data-based data asset evaluation method of the present invention, wherein: constructing a graph network of data assets based on a central database means installing a Neo4j graph database and Elasticsearch, constructing a graph data model, and using the transaction relationship between two nodes as an edge to connect relevant data asset nodes and user nodes;

[0012] Adding multi-dimensional attributes of data assets to the basic model and setting the weight of the edge according to the transaction frequency and amount of data assets;

[0013] Defining the centrality of data assets using the data asset influence score formula based on the interaction intensity and frequency between data assets:

[0014]

[0015] In the formula, S(v) represents the influence score of the asset node, w uv is the relationship weight between nodes u and v, t uv is the most recent interaction time, f is the time decay function, and d uv is the interaction frequency between nodes u and v;

[0016] Calculating the influence score of each data asset node and using the influence score as the influence score attribute of the node;

[0017] Establishing a real-time data monitoring system to monitor the dynamic changes of data assets in real time and pushing the updates to Neo4j, and using the exponential decay function to adjust the data asset value attribute of the data asset node in real time;

[0018] Using a bulk import tool to import data assets and transaction records into Neo4j, and updating the risk attribute of the data asset node based on market fluctuations and historical data asset using a risk assessment formula;

[0019] When it is detected that the initial data asset value and risk of the data asset change, automatically update the data asset value attribute and risk attribute of the data asset node in the graph network.

[0020] As a preferred solution of the big data-based data asset evaluation method of the present invention, wherein: capturing the time change of data relationships in the graph network and extracting node features from the graph network means adding a time attribute to each edge and defining a time window, performing time series analysis on the data asset relationships, calculating the statistical indicators of each data asset node in different time windows, using Cypher queries in Neo4j to generate statistical indicators, and adding the statistical indicators as attributes to the corresponding data asset nodes;

[0021] Obtain all nodes and corresponding edges in the graph network, configure the GraphSAGE model parameters, learn node embeddings by aggregating the neighbor information of each data asset node, and perform iterative updates;

[0022] Update the model weights through backpropagation and define the loss function. When the loss of the GraphSAGE model no longer decreases significantly during consecutive iterative processes, stop the iteration and output the model parameters to update the GraphSAGE model;

[0023] After the training is completed, extract the features of nodes and edges from the GraphSAGE model.

[0024] As a preferred solution of the data asset evaluation method based on big data according to the present invention, wherein: the factor that affects the value of the data asset identified based on the extracted node features refers to defining a list of assumed factors that affect the initial value of the data asset based on business knowledge and the node features extracted from the graph network, using the DoWhy library to draw a causal graph, determining that the transaction frequency of the data asset is the cause variable, and the initial value of the data asset is the result variable;

[0025] Use the regression analysis method to estimate the impact of factors on the initial value of the data asset for causal inference. The results of causal inference include the coefficients of each factor. Check the p-value of each factor and set a judgment threshold Q. Compare the p-value of each factor with the judgment threshold Q to judge the impact of the factor on the result variable. If the p-value of the factor is less than Q, the factor is statistically significant, indicating that the factor affects the result variable, otherwise it does not;

[0026] Build an XGBoost model to quantify the contribution of each factor to the initial value of the data asset. Use the features extracted from the graph network to train the XGBoost model. Use the cross-validation method to verify the performance of the XGBoost model. If the model performance does not improve, stop the training and obtain the trained XGBoost model;

[0027] Adjust the parameters of the XGBoost model based on the results of causal inference, and use the factors screened out from the causal inference as input variables to identify the factors that have the greatest impact on the initial value of the data asset.

[0028] As a preferred solution of the data asset evaluation method based on big data according to the present invention, wherein: building an asset evaluation model to evaluate the value of the data asset refers to building a linear regression model as the asset evaluation model, taking the initial value of the data asset as the dependent variable and the most influential factor as the independent variable, using the least squares method to estimate the model parameters, and using historical influencing factors and historical data asset value data as training data to input into the asset evaluation model for iterative training;

[0029] Define a loss function and an Adam optimizer to iteratively optimize the model parameters. When the loss of the asset evaluation model no longer significantly decreases during consecutive iterations, stop the iteration and output the updated model parameters to update the asset evaluation model.

[0030] Input the data to be evaluated into the asset evaluation model to obtain the value evaluation result of the data asset.

[0031] As a preferred solution of the data asset evaluation method based on big data according to the present invention, wherein: storing the data generated during the data asset evaluation process and implementing access control means that after verifying the data asset value evaluation result through a smart contract and recording it on the blockchain, synchronize the result to the central database for storage, collect the data generated during the data asset evaluation process, store it in the central database and implement access control, allowing authorized users to query and access the data.

[0032] Another object of the present invention is to provide a data asset evaluation system based on big data, which includes,

[0033] A data acquisition module, used to determine and configure the data source types to be integrated, collect data for preprocessing, conduct blockchain confirmation and recording, and store the data in the central database;

[0034] A graph network construction module, used to construct a graph database based on the central database, update the node attributes in real time according to the changes of the node data, and extract the node features in the graph network;

[0035] A causal inference module, used to draw a causal graph using the DoWhy library and conduct causal inference using a regression analysis method, and use the XGBoost model to quantify the contribution degree of each factor;

[0036] An asset evaluation module, used to construct an asset evaluation model based on the maximum influencing factor to evaluate the value of the data asset, verify the data asset value evaluation result through a smart contract, and store the data generated during the evaluation process in the central database.

[0037] A computer device, including: a memory and a processor; the memory stores a computer program, and when the processor executes the computer program, it implements the steps of data asset evaluation based on big data.

[0038] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of data asset evaluation based on big data.

[0039] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0040] The present invention constructs a graph network of data assets by collecting multi-source data, extracts the node features of the graph network to identify the factors affecting the value of data assets, and constructs an asset evaluation model for asset evaluation. It can accurately analyze the dynamic relationships between data assets, identify key nodes and important relationships, improve the depth and breadth of data analysis, provide a more dynamic and comprehensive evaluation perspective, and significantly enhance the credibility and practicality of the evaluation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of this application, and do not limit the embodiments of the present invention. In the drawings:

[0042] Figure 1 is a schematic flowchart of a data asset evaluation method based on big data;

[0043] Figure 2 is an implementation schematic diagram of graph network construction and update;

[0044] Figure 3 is a schematic structural diagram of a data asset evaluation method based on big data. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the embodiments and the drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and do not limit the present invention. It should be noted that the present invention has been in the actual R & D and use stage.

[0046] Example 1, as Figure 1 and Figure 2 shown, is the first embodiment of the present invention. This embodiment provides a data asset evaluation method based on big data. The data asset evaluation method based on big data includes:

[0047] S1. Collect multi-source data, store it in the central database, and upload it to the blockchain for data rights confirmation and recording;

[0048] Specifically, collecting multi-source data, storing it in the central database, and uploading it to the blockchain for data rights confirmation and recording means determining the types of data sources to be integrated. For each data source, configure the NiFi data flow group, set the data flow error handling logic, and set the processors in NiFi. Use SQL-like queries to remove invalid, blank, and duplicate data records from the data collected from each data source, standardize the data, and then store the data in the central database.

[0049] Calculate the hash value of the data, write a smart contract in Solidity, deploy the smart contract on the Ethereum platform, call the functions of the smart contract by writing a client application, and store the hash value, the current block timestamp, and the data source address on the blockchain.

[0050] By collecting and integrating multi-source data, the comprehensiveness and diversity of the data are ensured. This integration of multi-source data not only provides a richer data foundation but also improves the depth and breadth of data analysis. Storing data in a central database enables centralized management and efficient access to data, supporting large-scale data processing and analysis. Uploading to the blockchain for data rights confirmation and recording ensures the immutability and high credibility of the data, solving the problems of data tampering and loss existing in traditional data storage methods. By configuring the NiFi data flow group, automated data flow management from various data sources to the central database can be achieved. This automated management not only improves the efficiency of data processing but also reduces manual intervention and the error rate during data processing. Setting data flow error handling logic and processors ensures the stability and reliability of the data processing process, solving various abnormal situations that may occur during data processing. By calculating the hash value of the data, a unique identifier can be generated for each data item to ensure the integrity and consistency of the data. Writing a smart contract in Solidity and deploying the smart contract on the Ethereum platform, storing the hash value, block timestamp, and data source address on the blockchain through the smart contract to achieve data rights confirmation and recording. This method not only improves the security and credibility of data storage but also realizes the automation and transparency of data processing.

[0051] S2. Build a graph network of data assets based on the central database, capture the temporal changes of data relationships in the graph network, and extract node features from the graph network;

[0052] Specifically, building a graph network of data assets based on the central database means installing the Neo4j graph database, starting the Neo4j instance and performing preliminary configuration, including setting authentication information, configuring network interfaces and ports, installing Elasticsearch, and setting the cluster configuration and indexing strategy for integration with Neo4j, setting the node name, cluster name, and data path of Elasticsearch, and synchronizing the data of the two through the Neo4j-to-Elasticsearch plugin;

[0053] Build a graph data model, use the Cypher language to create data asset nodes, user nodes, transaction relationships, and their corresponding attributes in Neo4j, and use the transaction relationship between the two nodes as an edge to connect the relevant data assets and user nodes;

[0054] Add multi-dimensional attributes of data assets to the basic model, including static attributes (ID, type, and initial value) and dynamic attributes (market value, risk level) of data assets, and set the weight of the edge according to the transaction frequency and amount of data assets;

[0055] Define the centrality of data assets using the data asset influence score formula based on the interaction intensity and frequency between data assets:

[0056]

[0057] In the formula, S(v) represents the influence score of the asset node, w uv is the relationship weight between nodes u and v, t uv is the most recent interaction time, f is the time decay function, d uv is the interaction frequency between nodes u and v. The higher the frequency, the greater the weight. t uv is the time difference of the most recent interaction, and λ is the decay coefficient, indicating the impact of time on the interaction weight;

[0058] Calculate the influence score of each data asset node and use the influence score as the influence score attribute of the node;

[0059] Establish a real-time data monitoring system to monitor the dynamic changes of data assets in real time and push the updates to Neo4j. Use the exponential decay function to adjust the data asset value attribute of the data asset node in real time:

[0060] V t =V0·e -δt ,

[0061] In the formula, V t is the data asset value at time t, V0 is the initial data asset value, and δ is the decay constant, which is obtained through historical data regression analysis and represents the decay rate of the data asset value over time;

[0062] When the initial data asset value of the data asset is monitored to change, automatically update the data asset value attribute of the data asset node in the graph network;

[0063] Use the batch import tool to import data assets and transaction records into Neo4j, and update the risk attribute of the data asset node based on market fluctuations and historical data assets using the risk assessment formula:

[0064] R(v)=α·σ(v)+β·μ(v),

[0065] Wherein, R(v) is the risk score of data asset v, σ(v) is the standard deviation of the return of data asset v, representing the volatility of the return, μ(v) is the expected return of the data asset, and α and β are parameters for adjusting risk perception, which are determined by expert opinion or historical data regression analysis;

[0066] Wherein, r i is the return of the data asset in the i-th period, ε(v) is the average return of the data asset, and N is the total number of periods;

[0067] When the risk change of the data asset is monitored, the risk attributes of the data asset nodes in the graph network are automatically updated.

[0068] By installing the Neo4j graph database and performing preliminary configuration, including setting authentication information, configuring network interfaces and ports, the efficient storage and secure access of data assets are ensured. Integrating Elasticsearch and setting cluster configuration and indexing policies make the search and query of data more efficient. Using the Neo4j-to-Elasticsearch plugin to synchronize the data of both enhances the data management ability and query performance of the system. By using the Cypher language to create data asset nodes, user nodes and transaction relationships in Neo4j, and defining their respective attributes, the complex relationships between data assets can be intuitively represented. Taking the transaction relationship as the edge to connect the relevant data assets and user nodes constructs a complete graph data model, which can more accurately reflect the data interaction situation in the actual business. Adding multi-dimensional attributes of data assets, including static attributes and dynamic attributes, to the basic model makes the description of data assets more comprehensive and fine-grained. Setting the weight of the edge according to the transaction frequency and amount of data assets enhances the expression ability of the graph model, enabling it to reflect more complex interaction patterns and value evaluation. By using the data asset influence score formula, the interaction frequency and time factors are comprehensively considered. The influence of each data asset in the graph network can be quantified, and the dynamic influence of data assets can be more accurately reflected. By establishing a real-time data monitoring system, the changes of data assets can be dynamically monitored, and the exponential decay function is used to adjust the initial value of data asset nodes in real time, which can timely reflect market changes and improve the real-time and accuracy of data asset value evaluation. By using the risk assessment formula, the risk attributes of data assets can be dynamically evaluated and updated, and the risk level of data assets can be accurately evaluated, helping enterprises to better manage and control risks in the decision-making process.

[0069] Furthermore, capture the temporal changes in the data relationships in the graph network, extract node feature values from the graph network, add temporal attributes to each edge, and define a time window. Conduct time series analysis on the data asset relationships, calculate the statistical metrics of each data asset node in different time windows, including total transaction amount, average transaction amount, and maximum and minimum transaction amounts. Use Cypher queries in Neo4j to generate the statistical metrics and add the statistical metrics as attributes to the corresponding data asset nodes;

[0070] Obtain all nodes and corresponding edges in the graph network, configure the GraphSAGE model parameters, and learn node embeddings by aggregating the neighbor information of each data asset node and perform iterative updates;

[0071] Update the model weights through backpropagation and define a loss function. Stop the iteration and output the model parameters to update the GraphSAGE model when the loss of the GraphSAGE model no longer decreases significantly during consecutive iteration processes;

[0072] After training is completed, extract the features of nodes and edges from the GraphSAGE model.

[0073] By adding time attributes to each edge, the time point of each data interaction can be accurately recorded, laying a foundation for subsequent time series analysis. Defining time windows can divide continuous time data into several fixed-length intervals, facilitating data statistics and analysis within these intervals. This can capture the changes in data asset relationships over time and improve the understanding of the dynamics of data relationships. Through time series analysis of data asset relationships, the change patterns of data relationships can be identified, and statistical metrics such as total transaction volume, average transaction amount, and maximum and minimum transaction amounts can be calculated for each data asset node at different time windows. These statistical metrics can reflect the activity and importance of data assets, providing valuable information for subsequent data analysis and decision-making. By using Cypher queries, data can be efficiently extracted from the graph database and statistical metrics can be calculated, and these metrics can be added as attributes to the corresponding data asset nodes. This method can enhance the representational ability of the graph network, making it contain not only structural information but also rich attribute information, facilitating subsequent graph analysis tasks. By configuring the GraphSAGE model parameters to aggregate the neighbor information of each data asset node to learn node embeddings, the structural information and features of the node and its neighbors can be captured. The GraphSAGE model updates the node embeddings iteratively, enabling efficient representation learning in large-scale graph networks and enhancing the representational ability and analysis effect of the graph network. Through the backpropagation algorithm, the weights of the GraphSAGE model can be updated according to the gradient of the loss function to minimize the prediction error. Defining a reasonable loss function ensures that the loss of the model gradually decreases during continuous iterations until it reaches a convergent state. This can guarantee the effectiveness and stability of model training and improve the quality of node embeddings. By extracting the features of nodes and edges from the trained GraphSAGE model, a high-quality graph representation can be obtained, capturing the structural information and features of the node and its neighbors.

[0074] S3. Identify the factors affecting the value of data assets based on the extracted node features, and construct an asset evaluation model to evaluate the value of data assets;

[0075] Specifically, identifying the factors affecting the value of data assets based on the extracted node features means defining a list of hypothesized factors affecting the initial value of data assets based on business knowledge and the node features extracted from the graph network, including but not limited to factors such as transaction frequency, market volatility, and risk score. Use the DoWhy library to draw a causal graph, determining that the transaction frequency of the data asset is the cause variable and the initial value of the data asset is the result variable;

[0076] Use the regression analysis method to estimate the impact of factors on the initial value of data assets for causal inference. The results of causal inference include the coefficients of each factor. Check the p-value of each factor and set a judgment threshold Q. Compare the p-value of each factor with the judgment threshold Q to determine the impact of the factor on the result variable. If the p-value of the factor is less than Q, the factor is statistically significant, indicating that the factor affects the result variable; otherwise, it does not.

[0077] In statistics, the p-value is a probability value that represents the probability of observing the statistical result (or a more extreme result) under the condition that the null hypothesis is true. The null hypothesis usually assumes that the factor has no effect on the result (i.e., the coefficient is zero). If the p-value is below the pre-set threshold, the null hypothesis is rejected, and it is considered that the factor has a significant effect on the result.

[0078] Build an XGBoost model to quantify the contribution of each factor to the initial value of data assets. Use the features extracted from the graph network to train the XGBoost model. Use the cross-validation method to verify the performance of the XGBoost model. Stop training if the model performance does not improve, and obtain the trained XGBoost model.

[0079] Adjust the parameters of the XGBoost model based on the results of causal inference. Use the factors selected from causal inference as input variables to identify the factors that have the greatest impact on the initial value of data assets.

[0080] The judgment threshold Q is set through actual application scenarios and the experience of professionals. By defining a list of hypothesized factors that affect the initial value of data assets based on business knowledge and node features extracted from the graph network, including factors such as transaction frequency, market volatility, and risk scores, key factors affecting the value of data assets can be systematically identified and analyzed. Using the DoWhy library to draw a causal graph can clearly display and understand the causal relationships between these factors. By determining that the transaction frequency of data assets is the cause variable and the initial value of data assets is the result variable, a scientific and intuitive analysis method is provided. By estimating the impact of each factor on the initial value of data assets through regression analysis methods and conducting causal inference, the specific impact of each factor can be quantified. The results of causal inference include the coefficients of each factor, and the statistical significance of the factors is judged by checking the p-values. If the p-value of a factor is less than the set judgment threshold Q, the factor is statistically significant, indicating that the factor has a significant impact on the result variable. This method improves the credibility and scientific nature of the analysis results. By constructing an XGBoost model to quantify the contribution of each factor to the initial value of data assets, using the features extracted from the graph network to train the model, and using the cross-validation method to verify the model performance, the accuracy and stability of the model are ensured. If the model performance does not improve, the training is stopped, and the optimized XGBoost model is obtained. Based on the results of causal inference, the parameters of the XGBoost model are adjusted to make the model more accurately identify and quantify the factors that have the greatest impact on the initial value of data assets.

[0081] Furthermore, constructing an asset evaluation model to evaluate the value of data assets means constructing a linear regression model as the asset evaluation model, taking the initial value of data assets as the dependent variable and the most influential factor as the independent variable, using the least squares method to estimate the model parameters, and using historical influencing factors and historical data asset value data as training data to input into the asset evaluation model for iterative training;

[0082] Define a loss function and an Adam optimizer to iteratively optimize the model parameters. When the loss of the asset evaluation model no longer significantly decreases during continuous iteration, stop the iteration, output the updated model parameters, and update the asset evaluation model;

[0083] Input the data to be evaluated into the asset evaluation model to obtain the value evaluation result of the data assets.

[0084] By constructing a linear regression model, the linear relationship between the value of data assets and their influencing factors can be effectively captured. Taking the initial value of the data assets as the dependent variable and the most influential factor as the independent variable can simplify the complexity of the model while retaining the main explanatory ability for the asset value. Using the least squares method to estimate the model parameters ensures that the model can minimize the prediction error, improve the accuracy and stability of the prediction. Utilizing historical influencing factors and historical data asset value data as training data can make full use of past actual data and improve the prediction ability of the model. The iterative training process can continuously optimize the model parameters to make them more suitable for the current data characteristics and market environment. Defining a reasonable loss function can effectively measure the error between the predicted value and the actual value of the model, providing a clear goal for model optimization. Using the Adam optimizer for parameter optimization, due to its high computational efficiency and ability to adaptively adjust the learning rate, makes the model training process more stable and fast. When, in the continuous iterative process, the loss of the asset evaluation model no longer decreases significantly, stop the iteration to ensure that the model reaches the best state. Inputting the data to be evaluated into the trained asset evaluation model can quickly obtain the value evaluation result of the data assets. The model can give an accurate value prediction according to the features in the input data, combined with the trained parameters, improving the efficiency and accuracy of the evaluation work.

[0085] S4. Store the data generated during the data asset evaluation process and implement access control;

[0086] Specifically, storing the data generated during the data asset evaluation process and implementing access control means that after verifying the data asset evaluation result through a smart contract and recording it on the blockchain, the result is synchronously stored in the central database, collecting the data generated during the data asset evaluation process and storing it in the central database and implementing access control, allowing authorized users to query and access the data.

[0087] The intelligent contract is used to verify the data asset value assessment result and record it on the blockchain, ensuring the transparency and immutability of the assessment result. The automatic execution feature of the intelligent contract reduces manual intervention, lowering the operation risk and error rate. At the same time, the decentralized and distributed storage features of the blockchain further enhance the data security and credibility. The assessment result verified by the intelligent contract is synchronously stored in the central database, achieving centralized management of data. The central database provides efficient data query and access capabilities, ensuring that the assessment result can be quickly retrieved and used. This centralized storage method also facilitates subsequent data processing and analysis, improving the overall efficiency of data management. During the data asset evaluation process, a large amount of intermediate data and result data will be generated. By collecting this data and storing it in the central database, the evaluation process can be completely recorded, providing detailed data tracking and auditing capabilities. This method not only improves data integrity but also facilitates the backtracking and analysis of the evaluation process, further enhancing the credibility and transparency of the evaluation.

[0088] Embodiment 2, as Figure 3 shown, is the second embodiment of the present invention. This embodiment is different from the previous one and provides a data asset evaluation system based on big data, which includes

[0089] a data collection module, which is used to determine and configure the data source types to be integrated, collect data for preprocessing, conduct blockchain confirmation of rights and record, and store the data in the central database;

[0090] a graph network construction module, which is used to construct a graph database based on the central database, update node attributes in real time according to the changes of node data, and extract node features in the graph network;

[0091] a causal inference module, which is used to draw a causal graph using the DoWhy library and conduct causal inference using the regression analysis method, and use the XGBoost model to quantify the contribution degree of each factor;

[0092] an asset evaluation module, which is used to construct an asset evaluation model based on the most influential factor to evaluate the value of data assets, verify the data asset value evaluation result through the intelligent contract, and store the data generated during the evaluation process in the central database.

[0093] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0094] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a predefined sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device.

[0095] More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connection parts (electronic devices) having one or more wirings, portable computer disk cartridges (magnetic devices), random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memories), fiber optic devices, and portable compact disc read-only memories (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.

[0096] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0097] The specific embodiments described above further elaborate on the object, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A data asset evaluation method based on big data Comprising, characterized in that: Collect multi-source data, store it in a central database, and upload it to the blockchain for data right confirmation and recording; Build a graph network of data assets based on the central database, capture the temporal changes of data relationships in the graph network, and extract node features from the graph network; Identify factors affecting the value of data assets based on the extracted node features, and build an asset evaluation model to evaluate the value of data assets; Store the data generated during the data asset evaluation process and implement access control; The step of collecting multi-source data, storing it in a central database, and uploading it to the blockchain for data right confirmation and recording means collecting data from each data source, preprocessing it, storing it in the central database, writing a smart contract using the Solidity language, deploying the smart contract on the Ethereum platform, and storing the data on the blockchain; The step of building a graph network of data assets based on the central database means installing the Neo4j graph database and Elasticsearch, building a graph data model, and using the transaction relationship between two nodes as an edge to connect relevant data asset nodes and user nodes; On this basis, add multi-dimensional attributes of data assets and set the weight of the edge according to the transaction frequency and amount of data assets; Define the centrality of data assets using the data asset influence score formula based on the interaction intensity and frequency between data assets: Where, S(v) represents the influence score of the asset node, and w uv is the relationship weight between nodes u and v, and t uv is the most recent interaction time, f is the time decay function, and d uv is the interaction frequency between nodes u and v; Calculate the influence score of each data asset node and use the influence score as the influence score attribute of the node; Establish a real-time data monitoring system to monitor the dynamic changes of data assets in real time and push the updates to Neo4j, and use an exponential decay function to adjust the data asset value attribute of the data asset node in real time; Use a bulk import tool to import data assets and transaction records into Neo4j, and update the risk attribute of the data asset node based on market fluctuations and historical data asset using a risk assessment formula; When it is detected that the initial data asset value and risk of the data asset change, automatically update the data asset value attribute and risk attribute of the data asset node in the graph network.

2. The data asset evaluation method based on big data according to claim 1, characterized in that: The step of capturing the temporal changes of data relationships in the graph network and extracting node features from the graph network means adding a time attribute to each edge and defining a time window, performing time series analysis on the data asset relationships, calculating the statistical metrics of each data asset node in different time windows, using Cypher queries in Neo4j to generate the statistical metrics, and adding the statistical metrics as attributes to the corresponding data asset nodes; Obtain all nodes and corresponding edges in the graph network, configure the GraphSAGE model parameters, learn node embeddings by aggregating the neighbor information of each data asset node, and perform iterative updates; Update the model weights through backpropagation and define a loss function. When the loss of the GraphSAGE model no longer decreases significantly during consecutive iterations, stop the iteration, output the model parameters, and update the GraphSAGE model; After training is completed, extract the features of nodes and edges from the GraphSAGE model.

3. The data asset evaluation method based on big data according to claim 2, wherein: The factors that affect the value of data assets identified based on the extracted node features refer to defining a list of hypothesized factors that affect the initial value of data assets based on business knowledge and the node features extracted from the graph network, using the DoWhy library to draw a causal graph, determining the transaction frequency of the data asset as the cause variable, and the initial value of the data asset as the result variable; Using the regression analysis method to estimate the impact of factors on the initial value of data assets and conduct causal inference. The results of causal inference include the coefficients of each factor. Check the p-value of each factor and set a judgment threshold Q. Compare the p-value of each factor with the judgment threshold Q to determine the impact of the factor on the result variable. If the p-value of the factor is less than Q, then the factor is statistically significant, indicating that the factor affects the result variable, otherwise it does not; Construct an XGBoost model to quantify the contribution of each factor to the initial value of data assets. Use the features extracted from the graph network to train the XGBoost model, and use the cross-validation method to verify the performance of the XGBoost model. If the model performance does not improve, stop training and obtain the trained XGBoost model; Adjust the parameters of the XGBoost model based on the results of causal inference, and use the factors selected from the causal inference as input variables to identify the factors that have the greatest impact on the initial value of data assets.

4. The data asset evaluation method based on big data according to claim 3, characterized in that: The construction of the asset evaluation model to evaluate the value of data assets refers to constructing a linear regression model as the asset evaluation model, taking the initial value of the data asset as the dependent variable and the most influential factor as the independent variable, using the least squares method to estimate the model parameters, and using the historical influencing factors and historical data asset value data as training data to input into the asset evaluation model for iterative training; Define a loss function and an Adam optimizer to perform iterative optimization of the model parameters. When the loss of the asset evaluation model no longer decreases significantly during continuous iteration, stop the iteration, output the model parameters, and update the asset evaluation model; Input the data to be evaluated into the asset evaluation model to obtain the value evaluation result of the data asset.

5. A method for data asset evaluation based on big data according to claim 4, characterized in that: The storage of the data generated during the data asset evaluation process and the implementation of access control refer to synchronously storing the result into the central database after verifying the data asset value evaluation result through a smart contract and recording it on the blockchain, collecting the data generated during the data asset evaluation process and storing it in the central database and implementing access control, allowing authorized users to query and access the data.

6. A big data-based data asset evaluation system based on the big data-based data asset evaluation method according to any one of claims 1-5, characterized in that: Including, A data collection module, used to determine and configure the types of data sources to be integrated, collect data for preprocessing and conduct blockchain confirmation and recording, and store the data in the central database; A graph network construction module, used to construct a graph database based on the central database, update the node attributes in real time according to the changes in the node data, and extract the node features in the graph network; A causal inference module, used to draw a causal graph using the DoWhy library and conduct causal inference using the regression analysis method, and use the XGBoost model to quantify the contribution degree of each factor; An asset evaluation module, which is used to build an asset evaluation model based on the maximum influencing factors to evaluate the value of data assets, verify the data asset value evaluation result through a smart contract, and store the data generated during the evaluation process in a central database.

7. A computer device, comprising: A memory and a processor; the memory stores a computer program, characterized in that: when the processor executes the computer program, the steps of the data asset evaluation method based on big data according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, the steps of the data asset evaluation method based on big data according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Asset value evaluation method and device, model training method and device and readable storage medium

    CN116194911A

  • Virtual resource prediction release amount determination and value evaluation model training method

    CN116308734A