Public data resource asset evaluation system and method
By constructing server clusters and econometric models, the objectivity and efficiency issues of public data resource valuation have been resolved, achieving automated and scientific data asset valuation, applicable to data asset management and business activities.
Patent Information
- Application Number
- CN202511351532.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2026-01-02
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies lack a method that can objectively and efficiently assess the overall value of public data resources from a macro perspective. Traditional methods are costly, subjective in their results, and fail to reflect the potential revenue potential of the data.
An evaluation system based on a server cluster is constructed. Through data collection, econometric models and regression analysis, combined with a quality assessment engine, the system automatically calculates the output elasticity coefficient and net value of data elements, and uses distributed storage and econometric models to assess data value.
It enables objective and automated assessment of the value of public data resources. The assessment results are highly scientific and applicable to commercial activities such as data asset entry and pledge financing, supporting the optimization of data openness strategies and industrial policies.
Smart Images

Figure CN121258554A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of data governance and information technology, and in particular to a system and method for assessing the value of regional public data resource assets based on a server cluster architecture and using econometric models and big data processing technology. Specifically, it is a public data resource asset assessment system and method. Background Technology
[0002] With the advent of the digital economy era, data has been established as the fifth major factor of production. Public data resources are a crucial national strategic resource. As the marketization of data elements accelerates, scientifically and accurately assessing the asset value of public data resources has become a core prerequisite for their asset management, authorized operation, fiscal accounting, and financing. Governments at all levels have accumulated massive amounts of public data resources; how to scientifically and rationally assess the value of these assets has become a key prerequisite for their asset management, authorized operation, fiscal planning, and financing.
[0003] Currently, traditional intangible asset valuation methods face significant challenges when applied to data assets:
[0004] Cost-based approach: This approach assesses value by calculating the historical costs of data collection, storage, and management. Its drawback is that costs and value are severely mismatched; data requiring significant investment may have low value, while high-value data may have low acquisition costs, failing to reflect the potential profitability of data as a factor of production.
[0005] Market approach: This method assesses data by finding similar data assets trading in the market. Its drawbacks are: there are very few comparable public market transactions for public data resources, and their non-exclusivity and reusability make it difficult to establish standardized market prices. Furthermore, the data trading market is currently immature, thus this method has very poor applicability.
[0006] The income approach assesses value by predicting and discounting the future economic benefits of data assets. Its drawbacks include: the extremely wide range of applications for public data (G2B, G2C, G2G); the indirect nature of its benefits, manifested through improvements in social efficiency; the lag and synergistic nature of its benefits; and the difficulty in accurately predicting and attributing them to a single dataset, leading to strong subjectivity in prediction and large fluctuations in assessment results.
[0007] Therefore, existing technologies lack a method that can objectively and efficiently assess the overall value of public data resources from a macro perspective. Summary of the Invention
[0008] In view of the deficiencies of the prior art, the present application provides a public data resource asset evaluation system and method, which realizes accurate and efficient evaluation of the total value of regional public data assets through the construction of a dedicated server cluster and an evaluation engine, automation, quantification, and high reliability.
[0009] To achieve the above object, the present application is realized by the following technical solutions:
[0010] A public data resource asset evaluation system, characterized in that it comprises:
[0011] a data collection server, a distributed storage cluster, an evaluation engine server, and a management terminal.
[0012] The data collection server is configured with a data resource directory interface adapter for automatically collecting data resource directory quantity, API call quantity, and data storage quantity from heterogeneous government affairs platforms, and calculating the standardized quantitative value of data element input D based on the above.
[0013] The evaluation engine server is built-in with an econometrics analysis module for regression analysis based on a production function model to measure the output elasticity coefficient of data elements.
[0014] The regression analysis performed by the evaluation engine server uses a panel data fixed effect model and uses robust standard error or Newey-West standard error to overcome heteroscedasticity and serial correlation problems.
[0015] The system further comprises a data quality evaluation engine for calculating the weighted score of the integrity, accuracy, timeliness, and consistency indicators of the data, and using a Sigmoid function to normalize to obtain a quality adjustment coefficient K.
[0016] The system further comprises a time series prediction module for predicting the medium and long-term value of data assets based on the future trend of historical output elasticity coefficient (γ) and macroeconomic prediction data.
[0017] Preferably, the data collection server is configured with a web crawler and an API interface calling module for automatically collecting macroeconomic data, production factor data, and data governance cost data of various government departments, statistical bureaus, and public service units to form an original data set.
[0018] The distributed storage cluster is used for storing and managing the original data set and intermediate processing results, and is in communication connection with the data collection server and the evaluation engine server.
[0019] The evaluation engine server is built-in with an evaluation algorithm module for calling data from the storage cluster, performing value calculation based on a production function model, and outputting evaluation results.
[0020] The management terminal provides a human-computer interaction interface for configuring evaluation parameters, triggering evaluation tasks, and visualizing and displaying evaluation results.
[0021] Preferably, the time series prediction module is used to realize dynamic valuation of data assets. The working principle of this module is:
[0022] Historical trend analysis: based on historical data of the past 5 years, the ARIMA (autoregressive integrated moving average) model is used to predict the future trend of the output elasticity coefficient γ of each industry.
[0023] Macro prediction integration: access to macroeconomic prediction data released by authoritative agencies such as IMF and World Bank as the prediction input of future GDP Y of each industry.
[0024] Value prediction calculation: substitute the predicted γ value and Y value into the value calculation formula V_future=
[0025] Y_future*γ_future, to get the predicted value of data assets in the next 3-5 years.
[0026] Report generation: automatically generate the "Data Asset Value Prediction Report" to show the trend chart of future value, key inflection points and main driving factors, and provide decision support for strategic investment and risk management.
[0027] Preferably, the quantity of the data element input D i is quantified by the data collection server through the following steps:
[0028] Step 1, multi-source data collection, through API interface automatically docking data resource directory platform, data sharing exchange platform and cloud platform monitoring system of government departments at all levels, collecting the following original indicators;
[0029] Data resource directory quantity
[0030] Total number of data tables and total number of data fields
[0031] Daily API interface call frequency times / day
[0032] Daily API interface data return volume GB / day
[0033] Total amount of structured data storage TB
[0034] Step 2, standardization processing, the above heterogeneous indicators are standardized calculated by dimensionless and weighted summation, the calculation formula is:
[0035] D i = ω1*directory quantity + ω2*total number of data tables + ω3*API call volume + ω4*data storage capacity
[0036] The weight coefficients (ω1, ω2, ω3, ω4) are determined by principal component analysis (PCA) on historical data to objectively reflect the contribution of each index to the overall data.
[0037] Step 3, unit conversion, convert the final calculated D i value into standard measurement unit of 10,000 standard units for model use.
[0038] Preferably, the quality adjustment coefficient K is automatically calculated by the data quality evaluation engine, including the following steps:
[0039] Step 1, index quantification. Real-time monitoring and scoring of the four core dimensions of data assets;
[0040] Completeness (Completeness Score): Sc = (number of non-empty fields / total number of fields)
[0041] Accuracy (Accuracy Score): Sa = 1 - (number of data verification error records / total number of records)
[0042] Timeliness (Timeliness Score): St = e^(-λ*Δt), where Δt is the data update time interval, and λ is the industry decay coefficient
[0043] Consistency (Consistency Score): Sco = (number of records meeting embedded logic rules / total number of records)
[0044] Step 2, comprehensive scoring. Calculate the comprehensive quality score using the weight determined based on the analytic hierarchy process (AHP):
[0045] Score = Wc*Sc + Wa*Sa + Wt*St + Wco*Sco
[0046] The weight coefficients (Wc, Wa, Wt, Wco) are determined by the analytic hierarchy process (AHP) on historical data to objectively reflect the contribution of each index.
[0047] The weight determination steps based on the analytic hierarchy process (AHP) include:
[0048] Construct hierarchy: decompose the decision goal (data quality) into multiple dimensions (completeness, accuracy, etc.).
[0049] Construct judgment matrix: invite multiple experts to compare each dimension pairwise and use 1-9 scale method to judge their relative importance.
[0050] Calculate weight: calculate the eigenvector of the judgment matrix to obtain the preliminary weight of each dimension.
[0051] Consistency check: Calculate the consistency ratio CR, if CR < 0.1, it is considered that the judgment logic is consistent, and the weight is effective; otherwise, it needs to be compared again.
[0052] Step 3, normalization output. Normalize Score to the interval (0, 1] through the Sigmoid function to get the final quality adjustment coefficient K, K = 1 / (1+e^(-Score))
[0053] The quality evaluation engine is automatically run once a week to update the coefficient K and record the version number, ensuring the timeliness and accuracy of the value evaluation.
[0054] A public data resource asset evaluation method, characterized in that it comprises the following steps:
[0055] S1: Data collection and preprocessing
[0056] The data collection server regularly crawls or obtains the original data of the gross domestic product Y, capital investment K, labor input L and data governance cost C of each industry according to the preconfigured data source address list through the API interface; after cleaning, format standardization and missing value filling of the obtained heterogeneous data, it is stored in the distributed storage cluster.
[0057] S2: Output elasticity coefficient calculation
[0058] The evaluation engine server reads the historical panel data from the storage cluster, calls the built-in econometric analysis module, and performs multiple linear regression analysis based on the Cobb-Douglas production function model to fit the formula ln(Y) = ln(A) + αln(K) + βln(L) + γln(D) and solve the output elasticity coefficient γi of each industry data element. Where A is the total factor productivity, K is the capital investment, L is the labor input, and D is the data element input.
[0059] S3: Asset value calculation
[0060] The evaluation engine server reads the current economic data, calculates the contribution value of data in each industry according to the formula V i = Y i × γi, and sums up the total contribution value of regional public data resources V = ΣV i , Y i is the economic output of industry i.
[0061] S4: Calculation and output of asset net value
[0062] The evaluation engine server reads the total cost of data governance C, calls the data quality evaluation engine to obtain the data quality adjustment coefficient K, calculates the net asset value NAV according to the net asset value NAV=(V*K)-C, and finally writes the evaluation result NAV into the storage cluster and pushes it to the management terminal for display.
[0063] This set of economic model is creatively applied to the new and specific technical field of "public data asset value evaluation" for the first time. For the first time, "data" is explicitly put into the production function model as a core production factor D, together with K and L. Data D is introduced into the model as a core variable, an operational measurement scheme for data elements is defined, a complete evaluation system is constructed, and the long-standing technical problems of data assets are solved. It is applied to the new scene of data element marketization, and has produced significant technical effects:
[0064] The objective and automated evaluation of the value of data assets replaces the traditional method with strong subjectivity and high cost. The evaluation results can be used for data asset listing, pledge financing, transaction pricing and other specific business activities.
[0065] Preferably, the S1 collects the following data of the area to be evaluated in the evaluation period:
[0066] The economic output values of each sub-industry, such as industrial added value, financial industry added value, and total retail sales of social consumer goods, constitute the vector Y=(Y1, Y2,..., Y n )。
[0067] The traditional production factor input data of each industry mainly includes:
[0068] Capital input K: The total social fixed asset investment or industry depreciation can be used as a proxy variable.
[0069] Labor input L: The number of industry employees or total wages can be used as a proxy variable.
[0070] The total life cycle governance cost C of the regional public data resources, including hardware investment, operation and maintenance cost, and labor cost.
[0071] The data element D i input of each industry,
[0072] The quantification of the data element input (D i ) is automatically completed by the data collection server through the following steps:
[0073] Step 1, multi-source data collection, automatically connects the data resource directory platform, data sharing exchange platform and cloud platform monitoring system of government departments at all levels through API interface, and collects the following original indicators:
[0074] Number of data resource catalog
[0075] Total number of data tables and data fields
[0076] Daily API interface call frequency (times / day)
[0077] Daily API interface data return volume (GB / day)
[0078] Total amount of structured data storage (TB)
[0079] Step 2, standardization, the above heterogeneous indicators are standardized by dimensionless and weighted summation. The calculation formula is:
[0080] D i = ω1*(catalog quantity) + ω2*(total number of data tables) + ω3*(API call volume) + ω4*(data storage volume)
[0081] Wherein, the weight coefficient (ω1, ω2, ω3, ω4) is determined by principal component analysis (PCA) on historical data to objectively reflect the contribution of each index to the overall data input.
[0082] Step 3, unit conversion, the final calculated D i value is uniformly converted into standard measurement units such as ten thousand standard units for model use.
[0083] Preferably, S2 constructs a regression model in the form of Cobb-Douglas production function for each industry i:
[0084] Y i = A*K i ^α*L i ^β*D i ^γ
[0085] Wherein:
[0086] Y i is the economic output of industry i.
[0087] K i ,L i are the capital and labor inputs of industry i, respectively.
[0088] D i is the data element input of industry i, which can use data application scale or digital economy development index as a proxy variable.
[0089] γ is the data element output elasticity coefficient to be solved, which has the economic meaning that when the data input increases by 1%, the average output of the industry increases by γ%.
[0090] Using the historical panel data collected in step S1, the above formula is taken logarithm and converted into a linear model through multiple linear regression analysis by econometric software such as Stata, SPSS, and the elasticity coefficients γ1, γ2,..., γn of each industry can be estimated. n .
[0091] The econometric analysis module built in the evaluation engine server adopts
[0092] a panel data fixed effect model, and three key technical optimizations are made according to the characteristics of data elements:
[0093] Heteroscedasticity processing: robust standard errors (Robust Standard Errors) are forced to be used in regression calculation, effectively overcoming the problem of heteroscedasticity commonly existing in economic data, and ensuring the effectiveness of parameter estimation.
[0094] Serial correlation test: after regression, Durbin-Watson test is automatically performed, if serial correlation is detected, Newey-West standard error is used for correction to ensure the accuracy of statistical inference.
[0095] Multiple collinearity diagnosis: before regression, the variance inflation factor (VIF) of each explanatory variable (lnK, lnL, lnD) is calculated, if VIF>10, the "ridge regression" algorithm is started to replace the ordinary least squares method to eliminate the interference of multiple collinearity on parameter estimation.
[0096] The module finally outputs a complete report containing coefficient estimates, t-statistics, p-values and confidence intervals, and automatically filters out significant coefficients with p-value less than 0.05 for value calculation.
[0097] Preferably, the S3 calculates the value of data elements created in the industry i according to the output elasticity coefficient obtained in step S2: V i = Y i * γi
[0098] Where: γi represents the contribution of data to the output of industry i, so V i is the value of data in the industry.
[0099] The total macro value of regional public data resources is obtained by adding up the data values of each industry:
[0100] V = ΣV i
[0101] Preferably, the S4 introduces adjustment to more accurately reflect the net asset value:
[0102] NAV = (V*K)-C
[0103] wherein:
[0104] C is the total cost of data governance, deducted from the total value, representing net income.
[0105] K is the data quality adjustment coefficient (0 < K < 1), which is comprehensively evaluated according to the accuracy, timeliness, completeness, openness and other dimensions of the data, and the value is reduced or increased.
[0106] The quality adjustment coefficient K is automatically calculated by an independent data quality evaluation engine, and the core algorithm includes the following steps:
[0107] Step 1, index quantification. Real-time monitoring and scoring of the four core dimensions of data assets:
[0108] Completeness: Sc = (number of non-empty fields / total number of fields)
[0109] Accuracy: Sa = 1 - (number of data verification error records / total number of records)
[0110] Timeliness: St = e^(-λ*Δt), where Δt is the data update time interval, and λ is the industry decay coefficient
[0111] Consistency: Sco = (number of records meeting embedded logic rules / total number of records).
[0112] Step 2, comprehensive scoring. The comprehensive quality score is calculated using the weights determined based on the analytic hierarchy process (AHP):
[0113] Score = Wc*Sc + Wa*Sa + Wt*St + Wco*Sco
[0114] Step 3, normalized output. The Score is normalized to the interval (0, 1] by the Sigmoid function to obtain the final quality adjustment coefficient K.
[0115] K = 1 / (1 + e^(-Score))
[0116] The quality evaluation engine automatically runs once a week, updates the coefficient K, and records the version number, ensuring the timeliness and accuracy of the value evaluation.
[0117] The present application provides a public data resource asset evaluation system and method. It has the following beneficial effects:
[0118] 1) The present application uses standard production function model and regression analysis technology to measure data contribution weight, avoiding subjective speculation, making the evaluation result more scientific and credible; it is a top-down overall evaluation method, without enumerating specific data sets, suitable for macro and rapid value evaluation of massive public data assets.
[0119] 2) This invention not only considers the revenue generated by data, but also deducts its governance costs and introduces a quality adjustment factor, making the assessed net value closer to the actual economic value of the asset; the assessment results can clearly show the differences in the value contribution of data in different industries, providing precise data support for the government to optimize data opening strategies and industrial policies.
[0120] 3) Data collection and calculation are completed automatically through server clusters, avoiding subjective intervention and enabling rapid processing of massive amounts of data; based on rigorous econometric models, the evaluation process is traceable and verifiable, and the evaluation results are supported by economic theory; the distributed architecture design can support large-scale, multi-concurrency evaluation tasks and can be seamlessly integrated with other systems such as blockchain authentication platforms and data trading platforms through APIs. Attached Figure Description
[0121] Figure 1 This is a diagram of the main framework of the present invention;
[0122] Figure 2 This invention is based on a parameter interpretation table for the Cobb-Douglas production function model. Detailed Implementation
[0123] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0124] Please see the appendix Figure 1 and Figure 2 This invention provides a public data resource asset assessment system and method. By constructing a dedicated server cluster and assessment engine, it enables an automated, quantifiable, and highly reliable assessment of the total value of regional public data assets, achieving accurate and efficient evaluation.
[0125] To achieve the above objectives, the present invention provides the following technical solution:
[0126] A public data resource asset assessment system, characterized in that it includes:
[0127] Data acquisition server; distributed storage cluster; evaluation engine server; management terminal;
[0128] The data acquisition server is equipped with a data resource catalog interface adapter, which is used to automatically collect the number of data resource catalogs, API call volume and data storage volume from heterogeneous government affairs platforms, and calculate the standardized quantitative value of data element input D based on this.
[0129] The evaluation engine server is built-in econometrics analysis module, for regression analysis based on production function model, to measure the output elasticity coefficient of data elements.
[0130] The regression analysis performed by the evaluation engine server uses a panel data fixed effect model and uses robust standard errors or Newey-West standard errors to overcome heteroscedasticity and serial correlation problems.
[0131] The system also includes a data quality evaluation engine for calculating the weighted score of data integrity, accuracy, timeliness, consistency indicators, and using Sigmoid function normalization to obtain the quality adjustment coefficient K.
[0132] The system also includes a time series prediction module for predicting the medium and long-term value of data assets based on the future trend of historical output elasticity coefficient (γ) and macroeconomic prediction data.
[0133] Preferably, the data collection server is configured with a web crawler and API interface calling module for automatically collecting macroeconomic data, production factor data and data governance cost data from government departments, statistical bureaus and public service units to form an original data set.
[0134] The distributed storage cluster is used to store and manage the original data set and intermediate processing results, which is in communication connection with the data collection server and the evaluation engine server.
[0135] The evaluation engine server is built-in evaluation algorithm module, for calling data from the storage cluster, performing value calculation based on the production function model, and outputting the evaluation results.
[0136] The management terminal provides a human-computer interaction interface for configuring evaluation parameters, triggering evaluation tasks and visualizing evaluation results.
[0137] Preferably, the time series prediction module is used to realize the dynamic valuation of data assets. The working principle of this module is:
[0138] Historical trend analysis: based on the historical data of the past 5 years, the ARIMA autoregressive integrated moving average model is used to predict the future trend of the output elasticity coefficient γ of each industry.
[0139] Macro prediction integration: access to macroeconomic prediction data published by authoritative agencies such as IMF and World Bank as prediction input of future GDP Y of each industry.
[0140] Value prediction calculation: substitute the predicted γ value and Y value into the value calculation formula V_future=
[0141] Y_future*γ_future, to obtain the data asset value prediction value of the future 3-5 years.
[0142] Report generation: automatically generate a "data asset value prediction report", show the trend chart of future value, key inflection points and main driving factors, provide decision support for strategic investment and risk management.
[0143] Preferably, the data element input D i The quantification is automatically completed by the data acquisition server through the following steps:
[0144] Step 1, multi-source data acquisition, automatically docking the data resource directory platform, data sharing exchange platform and cloud platform monitoring system of government departments at all levels through API interface, collecting the following original indicators;
[0145] Data resource directory quantity
[0146] Total number of data tables and total number of data fields
[0147] Daily API interface call frequency times / day
[0148] Daily API interface data return volume GB / day
[0149] Total amount of structured data storage TB
[0150] Step 2, standardization processing, standardizing the above heterogeneous indicators by dimensionless and weighted summation, the calculation formula is:
[0151] D i = ω1*catalog quantity + ω2*total number of data tables + ω3*API call volume + ω4*data storage volume
[0152] Wherein, the weight coefficient (ω1, ω2, ω3, ω4) is determined by principal component analysis PCA on historical data to objectively reflect the contribution of each index to the overall data input;
[0153] Step 3, unit unification, the final calculated D i Value is uniformly converted into standard measurement unit of 10,000 standard units for model use.
[0154] Preferably, the quality adjustment coefficient K is automatically calculated by the data quality evaluation engine, including the following steps:
[0155] Step 1, index quantification. Real-time monitoring and scoring of the four core dimensions of data assets;
[0156] Completeness: Sc=(non-empty field number / total field number)
[0157] Accuracy: Sa = 1 - (number of data check error records / total number of records)
[0158] Timeliness: St = e^(-λ*Δt), where Δt is the data update time interval, λ is the industry decay coefficient
[0159] Consistency: Sco = (number of records that meet the embedded logic rules / total number of records)
[0160] Step 2, comprehensive score. The comprehensive quality score is calculated by using the weight determined based on the analytic hierarchy process (AHP);
[0161] Score = Wc*Sc + Wa*Sa + Wt*St + Wco*Sco
[0162] Step 3, normalized output. The Score is normalized to the interval (0, 1] by the Sigmoid function to obtain the final quality adjustment coefficient K, K = 1 / (1+e^(-Score))
[0163] The quality evaluation engine is automatically run once a week to update the coefficient K and record the version number, ensuring the timeliness and accuracy of the value evaluation.
[0164] A public data resource asset evaluation method, characterized by comprising the following steps:
[0165] S1: Data collection and preprocessing
[0166] The data collection server periodically crawls or obtains the original data of gross economic production Y, capital investment K, labor input L and data governance cost C of each industry according to the preconfigured data source address list through API interface; after cleaning, format standardization and missing value filling of the obtained heterogeneous data, the data is stored in the distributed storage cluster.
[0167] S2: Output elasticity coefficient calculation
[0168] The evaluation engine server reads the historical panel data from the storage cluster, calls the built-in econometric analysis module, and performs multiple linear regression analysis based on the Cobb-Douglas production function model to fit the formula ln(Y) = ln(A) + αln(K) + βln(L) + γln(D), and solve the output elasticity coefficient γi of each industry data element. Where A is the total factor productivity, K is the capital investment, L is the labor input, and D is the data element input.
[0169] S3: Asset value calculation
[0170] The evaluation engine server reads the current economic data and calculates the asset value V according to the formula V i = Y iThe contribution value of the calculated data of Xi in each industry is calculated, and the total contribution value V =∑V of the regional public data resource is obtained by summarizing i , Y i is the economic output of industry i.
[0171] S4: Net asset value calculation and output
[0172] The evaluation engine server reads the total cost C of data governance, calls the data quality evaluation engine to obtain the data quality adjustment coefficient K, calculates the net asset value NAV = (V x K) - C according to the net asset value (NAV), and finally writes the evaluation result NAV to the storage cluster and pushes it to the management terminal for display.
[0173] This set of economic models is creatively applied to the new and specific technical field of "public data asset value evaluation" for the first time, and for the first time, "data" is explicitly taken as the core production factor D, which is placed in the production function model together with K and L. Data D is introduced as a core variable into the model, an operational measurement scheme for data elements is defined, a complete evaluation system is constructed, and the long-standing technical problem of data assets is solved, applied to the new scene of data element marketization, and significant technical effects are produced:
[0174] Objective and automated evaluation of data asset value is realized, replacing the traditional method with strong subjectivity and high cost. The evaluation results can be used for data asset entry, pledge financing, transaction pricing, and other specific business activities.
[0175] Preferably, the S1 collects the following data of the region to be evaluated within the evaluation period:
[0176] The economic output values of each sub-industry, such as industrial added value, financial added value, and total retail sales of consumer goods, constitute the vector Y = (Y1, Y2,..., Y n ).
[0177] The traditional production factor input data of each industry mainly includes:
[0178] Capital input K: The total social fixed asset investment or industry depreciation can be used as a proxy variable.
[0179] Labor input L: The number of industry employees or total wages can be used as a proxy variable.
[0180] The total life cycle governance cost C of the regional public data resource, including hardware investment, operation and maintenance cost, and labor cost.
[0181] The data element D i input of each industry,
[0182] The data element input (Di ) is quantified, which is automatically completed by the data acquisition server through the following steps:
[0183] Step 1, multi-source data acquisition, automatically docking the data resource directory platform, data sharing exchange platform and cloud platform monitoring system of government departments at all levels through API interface, collecting the following original indicators:
[0184] Data resource directory quantity (pieces)
[0185] Total number of data tables and total number of data fields (pieces)
[0186] API interface daily call frequency (times / day)
[0187] API interface daily data return volume (GB / day)
[0188] Total amount of structured data storage (TB)
[0189] Step 2, standardization processing, the above heterogeneous indicators are standardized and calculated by dimensionless and weighted summation. The calculation formula is:
[0190] D i = ω1*(catalog quantity) + ω2*(total number of data tables) + ω3*(API call volume) + ω4*(data storage volume)
[0191] Wherein, the weight coefficient (ω1, ω2, ω3, ω4) is determined by principal component analysis (PCA) on historical data to objectively reflect the contribution of each indicator to the overall data input.
[0192] Step 3, unit conversion, the final calculated D i value is uniformly converted into standard measurement units such as ten thousand standard units for model use.
[0193] Preferably, the S2 constructs a regression model in the form of Cobb-Douglas production function for each industry i:
[0194] Y i = A*K i ^α*L i ^β*D i ^γ
[0195] Wherein:
[0196] Y i is the economic output of industry i.
[0197] K i ,L i are the capital and labor inputs of industry i, respectively.
[0198] D iThe data element input for the industry i is used as a proxy variable for the scale of the data application or the development index of the digital economy.
[0199] γ is the data element output elasticity coefficient to be solved, which has the economic meaning that when the data input increases by 1%, the output of the industry increases by γ% on average.
[0200] Using the historical panel data collected in step S1, the above formula is taken as a linear model after logarithmic transformation by using econometric software such as Stata, SPSS for multiple linear regression analysis, and the elasticity coefficients γ1, γ2,..., γn of various industries can be estimated. n .
[0201] The econometric analysis module built in the evaluation engine server adopts the core algorithm of
[0202] Panel data fixed effect model, and three key technical optimizations are made according to the characteristics of data elements:
[0203] Heteroscedasticity processing: robust standard errors (Robust Standard Errors) are forced to be used in regression calculation, which effectively overcomes the problem of heteroscedasticity commonly existing in economic data, and ensures the effectiveness of parameter estimation.
[0204] Serial correlation test: after regression, Durbin-Watson test is automatically performed, if serial correlation is detected, Newey-West standard error is used for correction to ensure the accuracy of statistical inference.
[0205] Multiple collinearity diagnosis: before regression, the variance inflation factor (VIF) of each explanatory variable (lnK, lnL, lnD) is calculated, if VIF> 10, the "ridge regression" algorithm is started to replace the ordinary least squares method to eliminate the interference of multiple collinearity on parameter estimation.
[0206] The module finally outputs a complete report containing coefficient estimates, t-statistics, p-values and confidence intervals, and automatically filters out significant coefficients with p-value less than 0.05 for value calculation.
[0207] Preferably, the S3 calculates the value of the data element created in the industry i according to the output elasticity coefficient obtained in step S2: V i = Y i * γi
[0208] Where: γi represents the contribution of data to the output of industry i, so V i is the value of data in the industry.
[0209] The total macro value of regional public data resources is obtained by adding up the data values of various industries:
[0210] V =∑V i
[0211] Preferably, S4 is more accurately reflect the net asset value, can be introduced to adjust:
[0212] NAV = (V * K) - C
[0213] Wherein:
[0214] C is the total cost of data governance, deducted from the total value, reflecting net income.
[0215] K is the data quality adjustment coefficient (0 < K < 1), according to the accuracy, timeliness, integrity, openness and other dimensions of the comprehensive evaluation, the value of the discount or value-added.
[0216] The quality adjustment coefficient K is automatically calculated by an independent data quality evaluation engine, and its core algorithm includes the following steps:
[0217] Step 1, index quantification. Four core dimensions of data assets are monitored and scored in real time:
[0218] Integrity: Sc = (non-empty field number / total field number)
[0219] Accuracy: Sa = 1-(data verification error record number / total record number)
[0220] Timeliness: St = e^(-λ*Δt), where Δt is the data update interval, λ is the industry attenuation coefficient
[0221] Consistency: Sco = (records that meet the embedded logic rules / total records).
[0222] Step 2, comprehensive score. The weight based on AHP is used to calculate the comprehensive quality score:
[0223] Score = Wc*Sc + Wa*Sa + Wt*St + Wco*Sco
[0224] Step 3, normalized output. The Score is normalized to the interval (0, 1] by Sigmoid function to get the final quality adjustment coefficient K.
[0225] K = 1 / (1 + e^(-Score))
[0226] The quality evaluation engine is automatically run once a week, update the coefficient K, and record the version number, to ensure the timeliness and accuracy of the value evaluation.
[0227] The evaluation engine server of this system can be a commercial server equipped with an Intel Xeon processor and 64GB of memory, and the distributed storage cluster can use Hadoop HDFS or cloud storage services. The system software is developed based on Java / Python and uses the Spark or TensorFlow framework for distributed regression calculations.
[0228] Example 1: Assessing the value of public data resources assets in a certain city in 2024.
[0229] The data collection server automatically retrieves economic data for various industries from the websites of the Municipal Bureau of Statistics, the Bureau of Industry and Information Technology, and the Big Data Bureau for 2020-2024.
[0230] The evaluation engine server reads the data, performs regression analysis, and obtains the financial industry data output elasticity coefficient γfinance = 0.12.
[0231] Calculate the contribution value of the 2024 data to the financial industry, Vfinance = 100 billion yuan × 0.12 = 12 billion yuan. Summarizing the value across all industries, we get V = 60 billion yuan.
[0232] Given the total governance cost C = 250 million yuan, and setting K = 0.9, the net asset value NAV is calculated to be (600 × 0.9) - 2.5 = 53.75 billion yuan.
[0233] The assessment result of 53.75 billion yuan was displayed on the large visual screen of the management terminal.
[0234] Example 2: Taking the assessment of the value of public data resource assets in a certain city in 2023 as an example.
[0235] S1: Obtain GDP data (Y1-Y10) for ten major industries including industry, finance, retail, transportation, and healthcare in 2023 from the city's statistics bureau. Obtain relevant data on fixed asset depreciation and employee wages from the finance bureau and human resources and social security bureau as proxy variables for K and L. Obtain the total annual cost C for data platform construction, operation and maintenance, and human resources from the big data bureau.
[0236] S2: Collect panel data for the above industries over the past 5 years (2018-2022). Using SPSS or Stata software, perform regression analysis for each industry, fit the production function, and obtain the output elasticity coefficients γ1-γ10 for each industry. For example, the calculated γfinance for the financial industry is 0.15, and the γtransport for the transportation industry is 0.08.
[0237] S3: Calculate the data value for each industry, such as VFinance = Y2023Finance * 0.15, VTransportation = Y2023Transportation * 0.08. Summarize the Vi values of all industries, including VFinance and VTransportation, to obtain the total contribution value V.
[0238] S4: Hire an expert committee or assess the city's data quality level as good according to the data open platform maturity model, corresponding to K=0.85. Calculate the net asset value NAV=(V*0.85)-C.
[0239] Generate an evaluation report, pointing out that financial, retail, and other industry data have the highest value, and suggest prioritizing opening; at the same time, data quality has room for improvement, and suggest strengthening governance.
[0240] Example 3: Provincial macro-strategic planning evaluation
[0241] Evaluation goal: Evaluate the overall value of public data resources in a certain province, and provide quantitative basis for formulating the "14th Five-Year" plan for the market development of data elements in the province.
[0242] Technical details implementation process:
[0243] S1: Data collection and preprocessing
[0244] Collect panel data of 15 categories of industries (such as digital economy core industries, intelligent manufacturing, wholesale and retail, smart tourism, financial technology, etc.) in 11 cities of the province from the provincial statistics bureau from 2018 to 2023, including:
[0245] Y: Annual added value of each industry (100 million yuan).
[0246] K: Fixed asset depreciation of each industry (100 million yuan).
[0247] L: Average wage of urban employees in each industry (100 million yuan).
[0248] Proxy variable of data input D: Use the proportion of digital economy core industry operating income to total industry income data from the provincial department of industry and information technology.
[0249] Governance cost C: Aggregate the annual budget of the provincial big data bureau and the data resource management bureau of each city for public data collection, governance, and platform operation, totaling about 8.5 billion yuan.
[0250] S2: Output elasticity coefficient calculation
[0251] Use Stata16.0 software to perform fixed-effects model regression analysis on the panel data of 15 industries for 6 years (a total of 15*6=90 sample points), controlling for regional and temporal heterogeneity.
[0252] Core regression equation (after taking the logarithm):
[0253] Where μi is the individual fixed effect (controls industry characteristics), λt Fixed effects for time (control common trends over years).
[0254] Example of regression results: The highest gamma for the core industry of digital economy was found to be 0.22; the gamma for traditional manufacturing was 0.09; and the gamma for smart tourism was 0.15.
[0255] S3 Asset Value Calculation
[0256] Based on 2023 data, calculate the data value of each industry and summarize. For example:
[0257] Digital economy core industry: V1 = industry added value * gamma = 12000 billion yuan * 0.22 = 2640 billion yuan
[0258] Traditional manufacturing: V2 = 18000 billion yuan * 0.09 = 1620 billion yuan
[0259] ...(other industry calculations omitted)
[0260] After summarizing, the total contribution value of public data resources in a certain province V = 2640 + 1620 +... ≈ 6500 billion yuan.
[0261] After review, the data quality of a certain province is high, and the coefficient K = 0.92.
[0262] Net asset value NAV = (6500 * 0.92) - 85 ≈ 5895 billion yuan.
[0263] Application effect:
[0264] The evaluation results show that the net asset value of public data resources in a certain province is nearly 6000 billion yuan, among which the data value contribution in the field of digital economy is the most outstanding.
[0265] This result provides key decision-making basis for the provincial government: firmly placing the development focus on promoting data opening and integration application in the core industry of digital economy, and setting corresponding data asset value-added targets in the planning.
[0266] Example 4: Precise assessment of specific fields in counties (taking the health data of a certain district as an example)
[0267] Evaluation goal: Precise assessment of the value of public data resources managed by the Health Bureau of a certain district (such as resident electronic health records, medical treatment data, public health data, etc.) in medical education research management, for department data asset table exploration.
[0268] Implementation process:
[0269] S1: Data collection and preprocessing
[0270] Subdivide the "health and health" field into 4 subfields:
[0271] Clinical diagnosis and treatment (Y1)
[0272] Public health (Y2)
[0273] Resident health services (Y3)
[0274] Health management and decision-making (Y4)
[0275] Proxy variables for output Y:
[0276] Y1: Annual medical service revenue of district hospitals (100 million yuan)
[0277] Y2: Estimated reduction in medical insurance expenditure due to effective chronic disease management (100 million yuan)
[0278] Y3: Annual financial allocation for resident health services (100 million yuan)
[0279] Y4: Administrative management cost saved through data-driven decision-making (100 million yuan)
[0280] Data input D: Using "number of times data is accessed / applied" as a proxy variable
[0281] Governance cost C: Annual hardware and software investment and human cost of the district health information center, approximately 0.15 billion yuan
[0282] S2: Estimation of output elasticity coefficient
[0283] Due to the fine granularity of district-level data, the production function model is supplemented and calibrated using expert scoring method (Delphi method) combined with analytic hierarchy process (AHP). Ten medical and health information experts are invited to estimate the contribution proportion γi of data in the output of the four sub-fields.
[0284] Calibrated results:
[0285] γ1 clinical = 0.15, data assists in clinical auxiliary diagnosis and rational drug use
[0286] γ2 public health = 0.40, data is the core of chronic disease management and infectious disease early warning
[0287] γ3 service = 0.30, data supports family doctors and appointment services
[0288] γ4 management = 0.25, data optimizes resource allocation and improves management efficiency
[0289] S3 & S4: Value calculation and summary
[0290] Calculate the value of data in each sub-field:
[0291] V1 = 10 billion yuan * 0.15 = 1.5 billion yuan
[0292] V2 = 0.8 billion yuan
[0293] V3 = 0.3 billion yuan
[0294] V4 = 0.125 billion yuan
[0295] The total value of the health data of the certain district of China V = 2.725 billion yuan.
[0296] Although embodiments of the present application have been shown and described, it is to be understood that various modifications, substitutions, replacements and changes can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalents.
Claims
1. A public data resource asset evaluation system, characterized in that, include: The system includes a data acquisition server; Distributed storage cluster; Evaluation engine server; Management terminal; The data acquisition server is equipped with a data resource catalog interface adapter, which is used to automatically collect the number of data resource catalogs, API call volume and data storage volume from the heterogeneous government affairs platform, and calculate the standardized quantitative value of data element input D based on this. The evaluation engine server has a built-in econometric analysis module, which is used to perform regression analysis based on the production function model and calculate the output elasticity coefficient of data elements. The regression analysis performed by the evaluation engine server uses a panel data fixed effects model and employs robust standard errors or Newey-West standard errors to overcome heteroscedasticity and serial correlation issues.
2. A public data resource asset valuation system according to claim 1, characterized in that, The system also includes a data quality assessment engine, which calculates weighted scores for data integrity, accuracy, timeliness, and consistency indicators, and normalizes them using the Sigmoid function to obtain a quality adjustment coefficient K.
3. A public data resource asset evaluation system according to claim 1, characterized in that, The system also includes a time series forecasting module, which is used to predict the medium- and long-term value of data assets based on the future trend of historical output elasticity coefficient γ and macroeconomic forecast data.
4. A public data resource asset evaluation system according to claim 1, characterized in that, The quantification of the data element input D is completed by the data acquisition server through the following steps: Step 1: Multi-source data collection. Automatically connect to the data resource catalog platform, data sharing and exchange platform, and cloud platform monitoring system of government departments at all levels through API interfaces to collect the following raw indicators; Number of data resource catalogs Total number of data tables and total number of data fields API interface average daily call count / API interface returns an average of GB / day Total TB of structured data storage Step 2: Standardization. The above heterogeneous indices are standardized by dimensionless conversion and weighted summation. The calculation formula is as follows: D = ω1 * number of directories + ω2 * total number of data tables + ω3 * number of API calls + ω4 * data storage size Among them, the weight coefficients (ω1, ω2, ω3, ω4) are determined by analyzing historical data using principal component analysis (PCA) to objectively reflect the contribution of each indicator to the overall data input. Step 3: Unit unification. The final calculated D value is converted into the standard unit of measurement, ten thousand standard units, for use in the model.
5. A public data resource asset evaluation system according to claim 2, characterized in that, The quality adjustment factor K is automatically calculated by the data quality assessment engine, including the following steps: Step 1: Quantify indicators and monitor and score the four core dimensions of data assets in real time. Completeness: Sc = (Number of non-null fields / Total number of fields) Accuracy: Sa = 1 - (Number of data verification errors / Total number of records) Timeliness: St=e^(-λ*Δt), where Δt is the data update time interval and λ is the industry decay coefficient. Consistency: Sco = (Number of records conforming to the embedded logic rules / Total number of records) Step 2: Comprehensive scoring. The overall quality score is calculated using weights determined by the Analytic Hierarchy Process (AHP). Score=Wc*Sc+Wa*Sa+Wt*St+Wco*Sco Among them, the weight coefficients (Wc, Wa, Wt, Wco) are determined by analyzing historical data using the analytic hierarchy process (AHP) to objectively reflect the contribution of each indicator. Step 3: Normalize the output. Normalize the Score to the (0,1] interval using the Sigmoid function to obtain the final quality adjustment coefficient K, K = 1 / (1 + e^(-Score)). The quality assessment engine runs automatically once a week, updates the coefficient K, and records the version number to ensure the timeliness and accuracy of the value assessment.
6. The method of using a public data resource asset appraisal system according to claim 1, characterized in that, Includes the following steps: S1: Data Acquisition and Preprocessing The data acquisition server periodically crawls or obtains raw data of GDP Y, capital input K, labor input L, and data governance cost C of various industries according to a pre-configured list of data source addresses; after cleaning, format standardization, and missing value filling of the acquired heterogeneous data, it stores it in the distributed storage cluster. S2: Calculation of output elasticity coefficient The evaluation engine server reads historical panel data from the storage cluster, calls the built-in econometric analysis module, performs multiple linear regression analysis based on the Cobb-Douglas production function model, fits the formula ln(Y)=ln(A)+αl n(K)+βl n(L)+γl n(D), and solves for the output elasticity coefficient γi of data elements in each industry. Where A is total factor productivity, K is capital input, and L is... D represents labor input, and D represents data element input. S3: Asset valuation. The evaluation engine server reads the current economic data and applies it according to formula V. i =Y i ×γi calculates the contribution value of data in various industries, and summarizes them to obtain the total contribution value V = ΣV of regional public data resources. i ;Y i For the economic output of industry i; S4: Net Asset Value Calculation and Output The evaluation engine server reads the total data governance cost C and calls the quality evaluation module to obtain the data quality adjustment coefficient K, and calculates the net asset value according to the formula NAV=(V×K)-C; Where V is the total contribution value, K is the quality adjustment coefficient, and C is the total governance cost; finally, the evaluation result NAV is written to the storage cluster and pushed to the management terminal for display.
7. A method for evaluating public data resource assets according to claim 6, characterized in that, The net asset value calculation and output evaluation results can be used for data asset entry, pledge financing, and transaction pricing.