Data asset analysis method and apparatus, and program product

By building a knowledge graph to integrate data assets, the problem of low efficiency in traditional data management is solved. It enables rapid location and in-depth mining of data relationships, provides a reference list to support enterprises in dealing with new data assets, and improves data management efficiency and decision-making speed.

CN121478871APending Publication Date: 2026-02-06BEIJING ELECTRONIC DIGITAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511654016.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Traditional data management and analysis methods are inefficient, making it difficult to accurately and quickly locate the required information in massive and complex data, and unable to deeply explore the inherent relationships between data, thus failing to meet the reference and analysis needs of enterprises when facing new data assets.

Method used

By constructing a knowledge graph, collecting data entity information of data assets through API interfaces, establishing a relationship network of nodes and edges, identifying key data entities, and pushing reference lists based on similarity, we can achieve recursive management and decision support of data value.

Benefits of technology

By integrating data entities through knowledge graphs, the similarity between new key data entities and existing entities can be quickly calculated, providing a reference list to help enterprises quickly develop response plans, reduce decision-making time and costs, and improve the efficiency of data value mining.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121478871A_ABST
    Figure CN121478871A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data asset analysis method and device and a program product. The method comprises the following steps: acquiring data entity information of data assets by adopting an API (Application Program Interface), and constructing a knowledge graph based on the data entity information; wherein nodes in the knowledge graph represent data entities, and edges between the nodes represent association relationships between the data entities; determining a key data entity from all the data entities; and when a new key data entity is acquired, pushing a reference list based on the similarity between the existing key data entity in the knowledge graph and the new key data entity. According to the method, the association relationship between the data entities can be deeply mined by constructing the knowledge graph, the similarity between the existing key data entities and the new key data entities can be conveniently and accurately obtained, the reference list is pushed through the similarity, the potential value of the new key data entities can be conveniently and better mined, and the new data entities can be analyzed and applied.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of statistical technology, and in particular to a data asset analysis method, apparatus, and program product. Background Technology

[0002] In today's digital age, as enterprises continue to deepen their digital transformation, data assets have become a core strategic resource. Enterprises and organizations accumulate massive amounts of data assets over long-term operations. These assets originate from multiple stages, including business transactions, customer interactions, and market research, and encompass a wealth of business information. When faced with new data assets, enterprises often need to acquire related data assets as a reference in order to better understand and utilize this new data.

[0003] Traditional data management and analysis primarily rely on manual searching and filtering, and fixed report queries. However, manual searching is not only extremely inefficient but also prone to errors, making it difficult to accurately and quickly locate the required information within massive and complex datasets. While fixed report queries provide standardized data presentation, they lack flexibility and cannot be dynamically adjusted based on the characteristics of new data assets and the specific needs of the enterprise. More importantly, neither of these methods can deeply explore the inherent relationships between data, nor can they effectively guarantee the stability and accuracy of data quality. Therefore, they fail to meet the actual needs of enterprises for reference and analysis when facing new data assets. Summary of the Invention

[0004] In view of this, the present disclosure provides a data asset analysis method, apparatus and program product, which can deeply mine the relationship between data entities by constructing a knowledge graph, so as to accurately obtain the similarity between existing key data entities and new key data entities, thereby selecting existing key data entities similar to new key data entities as references, thereby facilitating better mining of the potential value of new key data entities and analyzing and applying the new data entities.

[0005] In a first aspect, this disclosure provides a data asset analysis method, which adopts the following technical solution: Data entity information of data assets is collected using API interfaces, and a knowledge graph is constructed based on the data entity information; In the knowledge graph, nodes represent data entities, and edges between nodes represent the relationships between data entities. The data entities include business scenarios, business objectives, data products, and data resources; The relationships mentioned include the dependency relationship between business scenarios and data products, the value distribution relationship between business objectives and data products, the derivative relationship between data products, and the usage relationship between data products and data resources; Identify key data entities from all data entities; When a new key data entity is acquired, a reference list is pushed based on the similarity between the existing key data entities in the knowledge graph and the new key data entity.

[0006] Optionally, the data asset analysis method further includes: Identify the core data entities from all data entities; Based on the knowledge graph, a management strategy for the core data entity is obtained, and the core data entity is managed based on the management strategy.

[0007] Optionally, the construction of the knowledge graph includes: Based on the list of associated products, identify the data products that the business scenario depends on, and establish dependency edges between the data products and the dependent data products. Based on the upstream element list, identify the parent products that the data products depend on, and establish derivative edges between the data products and the parent products; Based on the derived edge, determine the data product that is the source of data value transfer, and establish a value distribution edge between the data product that is the source of data value transfer and the business objective; Based on the upstream element list, identify the data resources that the data product depends on, and establish a usage edge between the data product and the dependent data resources. The knowledge graph is constructed based on the edges between the business scenario, the business objective, the data product, and the data resource.

[0008] Optionally, the management strategy includes delisting and removal decisions; the step of obtaining a management strategy for the core data entity based on the knowledge graph, and managing the core data entity based on the management strategy, includes: By leveraging the relationships between data entities presented in the knowledge graph, the value of the data is recursively analyzed from top to bottom to obtain the value contribution parameters of each core data entity. Based on the value contribution parameters, a decision is made to remove or delist the core data, and the core data entity is managed based on the removal or delisting decision.

[0009] Optionally, the core data entity includes data products; the decision-making management of the core data entity based on the value contribution parameter includes: When the data product is a parent product, determine whether the data product meets multiple preset conditions for removing the parent product from the platform; If any parent product's delisting condition is met, a recommended delisting decision will be made for the data product. If the conditions for removing all parent products are not met, a decision will be made to continue listing the data product. When the data product is a sub-product, determine whether the data product meets the preset conditions for removing multiple sub-products from the platform; If any sub-product is deemed eligible for removal from the platform, a decision to remove the data product from the platform will be made. If the conditions for removing all sub-products are not met, a decision will be made to continue listing the data product.

[0010] Optionally, when the data product is a parent product, the value contribution parameters of the data product include direct scenario matching degree, technology output, technology output to technology cost ratio, health index, resource and technology cost transmission rate of direct sub-products, and technology output contribution of all sub-products. The preset conditions for removing parent products include: The direct scenario matching degree of the data product is less than the preset direct scenario matching degree threshold; The technology output-to-technology cost ratio of the data product is less than a preset technology output-to-technology cost ratio threshold; The health index of the data product is less than the preset health index threshold; The sum of the resource and technology cost transmission rates of all direct sub-products under the data product is greater than the product of the technology output of the data product and a preset first ratio; The technological output contribution of all sub-products under the data product is less than the preset first technological output contribution threshold.

[0011] Optionally, when the data product is a sub-product, the value contribution parameters of the data product include indirect scenario matching degree, technical output, resource and technical cost transmission rate, compliance risk transmission value, parent product status, and technical output contribution to the parent product. The preset conditions for removing several sub-products include: The indirect scenario matching degree of the data product is less than the preset indirect scenario matching degree threshold; The resource and technology cost transmission rate of the data product is greater than the product of the technology output of the data product and the preset second ratio; The compliance risk transmission value of the data product is greater than the preset compliance risk transmission threshold; The parent product of the data product is in an unavailable state, and the data product's contribution to the parent product's technological output is not greater than a preset second technological output contribution threshold.

[0012] Optionally, when the direct scenario matching degree of the data product is less than the preset direct scenario matching degree threshold, and / or the health index of the data product is less than the parent product delisting condition of the preset health index threshold, the data product is used as the starting node to locate whether there are high-tech cost sub-products and high-risk data resources downward along the knowledge graph, and whether there are inefficient scenarios upward along the knowledge graph; When the data product meets the sub-product delisting condition that the indirect scenario matching degree of the data product is less than the preset indirect scenario matching degree threshold, and / or the resource technology cost transmission rate of the data product is greater than the product of the technology output of the data product and the preset second ratio, the data product is used as the starting node to locate whether there are high technology cost data resources downward along the knowledge graph, and whether there are inefficient parent products upward along the knowledge graph. Based on the location results, the suggested reasons for removing the data product from the platform are obtained.

[0013] Optionally, the core data entity includes data resources, and the management strategy includes execution actions; the step of obtaining the management strategy for the core data entity based on the knowledge graph, and managing the core data entity based on the management strategy, includes: The value contribution parameters of the data resources include usage intensity, technology output-to-technology cost ratio, and compliance risk level. The data resource usage intensity, technology output-to-technology cost ratio, and compliance risk level are matched with each decision rule in the preset decision rule base to obtain the execution action for the data resource. Decision management is performed on the data resources according to the actions taken to access them.

[0014] Optionally, when new data resources are added to the knowledge graph, reference data resources similar to the new data resources are selected from the existing data resources within a preset observation period; The value contribution parameter of the reference data resource is used as the value contribution parameter of the new data resource; The threshold values ​​of the indicators in the decision rule base are reduced by a preset percentage to obtain a temporary decision rule base. The value contribution parameter of the new data resource is matched with each decision rule in the temporary decision rule base to obtain the execution action for the new data resource; Decision management is performed on the new data resources according to the actions taken to access them; When the preset observation period ends, the new data resources are adjusted to existing data resources, and the value contribution parameters of the existing data resources are obtained based on the knowledge graph to support decision management.

[0015] Optionally, the key data entities include business scenarios; when a new business scenario is acquired, a reference list is pushed based on the similarity between existing business scenarios in the knowledge graph and the new business scenario, including: When a new business scenario is acquired, the first entity attribute of the new business scenario is acquired; Based on the knowledge graph, obtain the second entity attribute for each existing business scenario; Based on the first entity attribute and the second entity attribute, obtain the first comprehensive similarity between the new business scenario and each existing business scenario; Based on the first comprehensive similarity, at least one reference business scenario similar to the new business scenario is selected from all existing business scenarios; Based on at least one of the aforementioned reference business scenarios, a reference list is constructed.

[0016] Optionally, the data asset analysis method further includes: Based on the historical value and historical call frequency of each reference business scenario in the reference list; Based on the historical value and the historical call frequency, the technical value index of the new business scenario is obtained.

[0017] Optionally, the first entity attributes include the scenario type, user entity, timeliness requirement parameters, expected data volume, and target compliance risk level of the new business scenario; The second entity attribute includes the existing business scenario type, user entity, timeliness requirement parameters, resource call volume, and the highest compliance risk level of associated data resources; The formula for calculating the first comprehensive similarity is: In the formula, This indicates the first overall similarity between the new business scenario and the existing business scenario; , , , , Indicates weight; This indicates the similarity between the scenario types of the new business scenario and the scenario types of existing business scenarios. This indicates the similarity between the users of the new business scenario and the users of the existing business scenario. This indicates the similarity between the timeliness requirement parameters of the new business scenario and the timeliness requirement parameters of the existing business scenario. This indicates the similarity between the expected data volume of the new business scenario and the resource call volume of the existing business scenario; This indicates the similarity between the target compliance risk level of the new business scenario and the highest compliance risk level of the associated data resources in the existing business scenario.

[0018] Optionally, the formula for calculating the technology value index of the new business scenario is: In the formula, This represents the technological value index of new business scenarios; This represents the historical value benchmark for new business scenarios; This represents the coefficient for the change in technology costs for new business scenarios; This represents the scenario difference coefficient for new business scenarios; This indicates the target utilization rate of the new business scenario.

[0019] Optionally, the key data entity includes data products; when a new data product is acquired, a reference list is pushed based on the similarity between the existing data products in the knowledge graph and the new data product, including: When a new data product is acquired, the third entity attribute of the new data product is acquired; Based on the knowledge graph, obtain the fourth entity attribute of each existing data product; Based on the third entity attribute and the fourth entity attribute, a second comprehensive similarity is obtained between the new data product and each existing data product. Based on the second comprehensive similarity, at least one reference data product similar to the new data product is selected from all existing data products; Based on at least one reference data product, construct a reference list.

[0020] Optionally, the data asset analysis method further includes: Based on the unit technology cost of each reference data product in the aforementioned reference list; The technological value index of the new data product is obtained based on the unit technological cost of each reference data product.

[0021] Optionally, the third entity attribute includes the product type of the new data product, the data resources it depends on, and the business scenario of the services it provides; The fourth entity attribute includes the product type of the existing data product, the data resources it depends on, and the business scenarios of the services it provides. The formula for calculating the second comprehensive similarity is: In the formula, This indicates the second comprehensive similarity between the new data product and existing data products. , , Indicates weight; This indicates the similarity between the product types of the new data product and the product types of existing data products. This indicates the degree of overlap between the data resources relied upon by new data products and those relied upon by existing data products. This indicates the similarity between the business scenarios of new data products and services and those of existing data products and services.

[0022] Optionally, the formula for calculating the technology value index of the new data product is: In the formula, Indicating the technological value index of new data products; Indicates the quantity of products referenced in the data; Indicates the first Unit technology cost of a reference data product; Indicates the new data product and the first The second comprehensive similarity between reference data products; This indicates the demand premium percentage for new data products; This indicates the target profit margin for the new data product.

[0023] Secondly, this disclosure also provides a data asset analysis system, which adopts the following technical solution: The knowledge graph construction module is used to collect data entity information of data assets using API interfaces, and to construct a knowledge graph based on the data entity information. In the knowledge graph, nodes represent data entities, and edges between nodes represent the relationships between data entities. The data entities include business scenarios, business objectives, data products, and data resources; The relationships mentioned include the dependency relationship between business scenarios and data products, the value distribution relationship between business objectives and data products, the derivative relationship between data products, and the usage relationship between data products and data resources; The key identification module is used to identify key data entities from all data entities; The list push module is used to push a reference list based on the similarity between existing key data entities in the knowledge graph and the new key data entity when a new key data entity is acquired.

[0024] Thirdly, this disclosure also provides a computer device, which adopts the following technical solution: The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which enables the at least one processor to perform any of the data asset analysis methods described above.

[0025] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing computer instructions for causing a computer to perform any of the data asset analysis methods described above.

[0026] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of any of the methods described above.

[0027] The data asset analysis method provided in this disclosure embodiment By constructing a knowledge graph, data entity information of data assets can be effectively integrated. In traditional enterprise data management models, data is often scattered across different systems and departments, lacking unified organization and connections, making it difficult to clearly present the relationships between data and to extract data value. The construction of a knowledge graph, however, represents each data entity as a node, with the relationships between entities represented by edges, forming an intuitive and comprehensive data network. Furthermore, this method, through the knowledge graph, allows enterprises to recursively analyze data value from top to bottom, clearly understanding the value contribution of each core data entity in the entire business system. The advantages of the knowledge graph become even more apparent when new key data entities emerge. Because the knowledge graph stores a large amount of existing key data entities and their associated information, the system can quickly calculate the similarity between new key data entities and existing entities based on this data. The reference list pushed based on this similarity provides valuable experience for new data entities, including analysis methods, processing strategies, and business application cases related to similar data entities. With the help of these reference lists, companies can quickly develop solutions when faced with new data assets. Analysts can directly refer to the mature methods and strategies in the lists without having to figure it out from scratch, thereby reducing decision-making time and costs.

[0028] The above description is merely an overview of the technical solution disclosed herein. In order to better understand the technical means of this disclosure and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0029] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 A flowchart illustrating the data asset analysis method provided in this embodiment of the disclosure; Figure 2 A flowchart illustrating the novel data resource management method provided in this embodiment of the disclosure; Figure 3 A flowchart illustrating the method for obtaining a reference list of new business scenarios provided in this disclosure embodiment; Figure 4 A flowchart illustrating the method for obtaining a reference list of new data products provided in this embodiment of the disclosure; Figure 5 A schematic diagram of the data asset analysis system provided in this embodiment of the disclosure; Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure. Detailed Implementation

[0031] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0032] It should be understood that the following specific examples illustrate the implementation of this disclosure, and those skilled in the art can easily understand other advantages and effects of this disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. This disclosure can also be implemented or applied through other different specific implementation methods, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0033] Reference Figure 1 This disclosure provides a data asset analysis method, including the following steps: S1: Collect data entity information of data assets using API interfaces, and construct a knowledge graph based on the data entity information; In the knowledge graph, nodes represent data entities, and edges between nodes represent the relationships between data entities. Among them, data entities include business scenarios, business objectives, data products, and data resources; The relationships include the dependency relationship between business scenarios and data products, the value distribution relationship between business objectives and data products, the derivative relationship between data products, and the usage relationship between data products and data resources; S2: Identify key data entities from all data entities; S3: When a new key data entity is acquired, a reference list is pushed based on the similarity between the existing key data entities in the knowledge graph and the new key data entity.

[0034] This publicly disclosed data asset analysis method, through the construction of a knowledge graph, effectively integrates data entity information of data assets. In traditional enterprise data management models, data is often scattered across various systems and departments, lacking unified organization and connections, making it difficult to clearly present the relationships between data and to extract data value. The construction of a knowledge graph, however, represents each data entity as a node, with the relationships between entities represented by edges, forming an intuitive and comprehensive data network. Furthermore, this method, through the knowledge graph, allows enterprises to recursively analyze data value from top to bottom, clearly understanding the value contribution of each core data entity within the entire business system.

[0035] The advantages of knowledge graphs become even more apparent when new key data entities emerge. Because knowledge graphs store a large amount of existing key data entities and their associated information, the system can quickly calculate the similarity between new key data entities and existing entities based on this data. The reference list pushed based on this similarity provides valuable experience for new data entities, including analytical methods, processing strategies, and business application cases. With the help of these reference lists, enterprises can quickly develop solutions when facing new data assets. Analysts can directly refer to the mature methods and strategies in the lists, eliminating the need for trial and error and reducing decision-making time and costs.

[0036] In S1, data entities and their core attributes are extracted from the data entity information of data assets. Data entities include business scenarios, business objectives, data products, and data resources. Their relevant information is shown in Table 1 below.

[0037] Table 1 The relationships between data entities are analyzed, with each data entity treated as a node. Edges are constructed between nodes based on these relationships, and the attributes of the edges are set based on the core attributes of each node. After these settings are complete, a complete knowledge graph is built. These relationships include the dependency relationship between business scenarios and data products, the value distribution relationship between business goals and data products, the derivative relationship between data products, and the usage relationship between data products and data resources. Relevant information regarding these relationships is shown in Table 2 below.

[0038] Table 2 The relationships can be further subdivided into derivative relationships between data products and usage relationships between data products and data resources. The direction of derivative relationships is data product → data product, and the direction of usage relationships is data product → data resource. Based on the list of associated products, the data products that the business scenario depends on are identified, and dependency edges are established between the data products and the dependent data products. Based on the list of upstream elements, the parent products that the data products depend on are identified, and derivative edges are established between the data products and the parent products. Based on the derivative edges, the data products that are the source of data value transfer are determined, and value allocation edges are established between the data products that are the source of data value transfer and the business objectives. Based on the list of upstream elements, the data resources that the data products depend on are identified, and usage edges are established between the data products and the dependent data resources. Based on the edges between business scenarios, business objectives, data products, and data resources, a knowledge graph is constructed.

[0039] Furthermore, core data entities are identified from all data entities; and these core data entities are managed based on a knowledge graph.

[0040] By leveraging the recursive capabilities of knowledge graphs to assess data value, enterprises can make precise decisions. For high-value core data entities, enterprises can choose to retain them and increase investment to further explore their potential; while for low-value core data entities, enterprises can consider delisting or optimizing them to avoid wasting resources. This helps enterprises optimize resource allocation, concentrating limited resources on more valuable business processes, thereby improving operational efficiency and effectiveness.

[0041] The management strategy includes delisting and listing decisions. Core data entities include data products and data resources. Through the relationships between data entities presented in a knowledge graph, data value is recursively transferred from top to bottom to obtain the value contribution parameters of each core data entity. Based on these value contribution parameters, delisting and listing decisions are made for core data entities, and management decisions are implemented accordingly. The primary manifestation of data value is technological output, which originates from the core attributes of business objectives. Therefore, the technological output of business objectives needs to be recursively transferred from top to bottom along the relationships in the knowledge graph to obtain the technological output of each data product and data resource. Furthermore, both technological output and technological costs are statistically analyzed on a monthly basis. For data products with a value allocation relationship to business objectives, the technological output equals the technological output of the business objective multiplied by the allocation weight from the business objective to the data product. The formula for calculating this weight is as follows: In the formula, Indicates business objectives up to the The allocation weight of each data product; , This indicates configurable weights, and ; Indicates the first The contribution rate of each data product to the improvement of business objectives refers to the quantitative value of the degree to which the data product contributes to the improvement of business-related indicators in the process of helping to achieve business objectives. It reflects the proportion of the contribution of the data product to the improvement of various indicators under the business objectives. Indicates the first in the business scenario The percentage of calls to a specific data product is equal to the number of times that data product is called within the business scenario divided by the total number of calls to all data products within the business scenario.

[0042] When there is a derivative relationship between data products, the parent product and the child product can be determined based on the derivative relationship. The technical output of the child product is equal to the technical output of the parent product multiplied by the allocation weight from the parent product to the child product. This weight is also called the derivative contribution, which can be extracted from the data entity information.

[0043] When there is a usage relationship between data products and data resources, the technical output of all data products that have a usage relationship with the data resource is multiplied by the allocation weight of that data product to the data resource, and then summed to obtain the technical output of the data resource. The formula for calculating this weight is as follows: In the formula, Indicates the first The first data product to the first The allocation weight of each data resource; Indicates the first The resource base weight of each data resource; Indicates the first The first data product to the first The collaborative weight of each data resource; Indicates the relationship with the first The total number of data products that have usage relationships between data resources.

[0044] In the formula, Indicates the first Dependency type coefficient of each data resource; Indicates the first The quality coefficient of each data resource; Indicates the first The scarcity coefficient of each data resource. Information regarding the dependency type coefficient, quality coefficient, and scarcity coefficient is provided in Table 3 below.

[0045] Table 3 For core data entities such as data products, decision management includes the following steps: When a data product is a parent product, determine whether the data product meets multiple preset conditions for removing parent products; if it meets any one of the parent product removal conditions, make a recommendation to remove the data product; if it does not meet all the parent product removal conditions, make a decision to continue listing the data product. When a data product is a child product, determine whether the data product meets multiple preset conditions for removing child products; if it meets any one of the child product removal conditions, make a recommendation to remove the data product; if it does not meet all the child product removal conditions, make a decision to continue listing the data product.

[0046] When a data product is the parent product, its value contribution parameters include direct scenario matching degree, technological output, technological output-to-technology cost ratio, health index, resource-technology cost transmission rate of direct sub-products, and technological output contribution of all sub-products. Several preset conditions for removing a parent product include: the direct scenario matching degree of the data product is less than a preset direct scenario matching degree threshold (e.g., 0.5) (this is a direct cause); the technological output-to-technology cost ratio of the data product is less than a preset technological output-to-technology cost ratio threshold (e.g., 1.0) (this is a direct cause); the health index of the data product is less than a preset health index threshold (e.g., 0.4) (this is a direct cause); the sum of the resource-technology cost transmission rates of all direct sub-products under the data product is greater than the product of the data product's technological output and a preset first ratio (e.g., 0.6) (this is an indirect cause, indicating that the technological cost ratio of the sub-products is too high); and the technological output contribution of all sub-products under the data product is less than a preset first technological output contribution threshold (e.g., 15%).

[0047] The direct scenario matching degree of data products refers to the semantic similarity between the data product and the business scenario requirements (e.g., the cosine value of the BERT vector, ranging from 0 to 1); the technology output-to-technology cost ratio of data products refers to the ratio of the technology output to the technology cost of the data product. The formula for calculating the health index of data products is as follows: In the formula, Indicates the first Health index of individual data products; , , This represents the weights, with values ​​of 0.4, 0.3, and 0.3 respectively. Indicates the first The first data product Each sub-product for the first The technological output contribution of a data product is equal to its derivative contribution. Indicates the relationship with the first The ratio between the number of business scenarios with which a data product has dependencies and the number of preset target business scenarios is essentially the ratio of the number of business scenarios actually served by the parent product to the total number of business scenario data. Indicates the first The compliance risk transmission value of a data product refers to the first... The product of the highest compliance risk level of the data resources upon which the sub-products of each data product depend and the proportion of data resource usage; Indicates the first The number of sub-products of each data product.

[0048] The formula for calculating the resource and technology cost pass-through rate of data products is as follows: In the formula, Indicates the first Resource and technology cost transmission rate of individual data products; Indicates the first The data product relies on the first The unit technology cost of a data resource is equal to the sum of the data resource’s acquisition technology cost, storage technology cost, and compliance processing technology cost (such as the cost of L4 level resources combined with privacy computing technology). Indicates the first The usage percentage of a data resource refers to the percentage of data resources used. The data product for the first The number of times the data resource is called is the same as the number of times the first data resource is called. The ratio between the total number of data resources accessed by each data product and the total number of data resources it depends on; Indicates the first The parent product of the data product is related to the first... The frequency of calls to each data product. According to this formula, when the data product is a parent product, the resource and technology cost transmission rate of the direct sub-products under the data product can be calculated; when the data product is a sub-product, the resource and technology cost transmission rate of the data product can be calculated directly.

[0049] When a data product is a sub-product, its value contribution parameters include indirect scenario matching degree, technical output, resource and technical cost transmission rate, compliance risk transmission value, parent product status, and contribution to the parent product's technical output. Several preset conditions for sub-product delisting include: the data product's indirect scenario matching degree is less than a preset indirect scenario matching degree threshold (e.g., 0.3) (this condition indicates a severe disconnect between the parent product's scenario requirements and the sub-product's functionality, reflecting a value gap); the data product's resource and technical cost transmission rate is greater than the product of the data product's technical output and a preset second ratio (e.g., 1.2) (this condition indicates that technical costs exceed technical output by 20%, reflecting a technical cost inversion); the data product's compliance risk transmission value is greater than a preset compliance risk transmission threshold (e.g., 0.7) (this condition indicates reliance on high-risk resources that are irreplaceable, reflecting uncontrollable risk); the data product's parent product is in a delisted state and the data product's contribution to the parent product's technical output is not greater than a preset second technical output contribution threshold (this condition indicates the sub-product has lost its value). The formula for calculating the data product's indirect scenario matching degree is as follows: In the formula, Indicates the first Indirect scenario matching degree of individual data products; Indicates the first The direct scenario matching degree of the parent product of each data product; Indicates the first The contribution of each data product to the technological output of its parent product.

[0050] Every day at midnight, the system aggregates attributes such as call frequency, technical output, and compliance risk level of data products to sub-products and then to data resources by performing graph traversal on the knowledge graph. When calculating the indirect scenario matching degree of sub-products, the system automatically inherits the previous day's calculation results of the parent product.

[0051] Furthermore, for cases where the data product is the parent product, when the direct scenario matching degree of the data product is less than a preset direct scenario matching degree threshold, and / or the health index of the data product is less than a preset health index threshold, the parent product is removed from the platform. Starting with the data product, the platform is used to locate whether there are high-tech cost sub-products and high-risk data resources downwards along the knowledge graph, and whether there are inefficient scenarios upwards along the knowledge graph. Based on the location results, the recommended reasons for removing the parent product are obtained. Specifically, a high-tech cost sub-product refers to a sub-product under the parent product whose technological output contribution is less than the first threshold (e.g., 10%) and whose technological cost ratio is greater than the second threshold (e.g., 30%). A high-risk data resource refers to the unit technological cost of all data resources relied upon by all sub-products under the parent product. If the ratio of the largest unit technological cost to the total technological cost of the parent product is greater than the third threshold, then the data resource is identified as a high-risk data resource. An inefficient scenario refers to a business scenario where the weighted sum of the direct scenario matching degree and the call ratio with the parent product is less than the fourth threshold.

[0052] When a data product meets the sub-product delisting conditions of having an indirect scenario matching degree less than a preset indirect scenario matching degree threshold, and / or having a resource and technology cost transmission rate greater than the product of the data product's technology output and a preset second ratio, the system uses the data product as the starting node to locate whether there are high-tech-cost data resources downwards along the knowledge graph and whether there are inefficient parent products upwards along the knowledge graph. Based on the location results, the recommended reasons for delisting the sub-product are obtained. High-tech-cost data resources refer to sub-product dependent data resources whose technology cost ratio is greater than the fifth threshold (e.g., 60%). Here, the technology cost ratio is the ratio between the unit technology cost of the data resources dependent on the sub-product and the total unit technology cost of all data resources dependent on the sub-product. When the scenario coverage rate of the parent node is less than the seventh threshold or the frequency decay rate of the parent product's calls to the sub-product is less than the eighth threshold, it is considered an inefficient parent product.

[0053] When a data product is a sub-product and a decision is made to continue its deployment, downstream sub-products are simultaneously modified. The status of related sub-products belonging to the same parent product is simultaneously marked, triggering a functional redundancy check on the parent product. By simultaneously marking the status of related sub-products, the specific scope of the parent product's functional redundancy check can be clearly defined. Because there is a clear relationship between sub-products and parent products, marking the status of sub-products allows identification of which sub-products may have redundant functions, avoiding blind checks and improving efficiency.

[0054] The management strategy also includes execution actions. For core data entities such as data resources, decision management includes the following steps: The value contribution parameters of data resources include usage intensity, technology output-to-cost ratio, and compliance risk level; matching the usage intensity, technology output-to-cost ratio, and compliance risk level of data resources with each decision rule in the pre-set decision rule base to obtain the execution actions for the data resources; and performing decision management on the data resources according to the execution actions. Usage intensity refers to the sum of the call intensity of data resources by data products with usage relationships within a statistical period (usually one month). Call intensity is the product of the frequency of data product calls to the data resource (e.g., 100 calls) and the amount of data called per call (e.g., an average of 500 data entries per call). The decision rule base is shown in Table 4 below. Table 4 Furthermore, when new data resources are added to the knowledge graph, effective management is difficult due to the lack of historical data for these new resources. Therefore, this solution designs a method to enable rational decision-making regarding new data resources. (Refer to...) Figure 2 The flowchart illustrating the new data resource management method shows that the decision-making management approach for new data resources includes the following steps: S11: When new data resources are added to the knowledge graph, reference data resources similar to the new data resources are selected from the existing data resources within the preset observation period; S12: Use the value contribution parameters of the reference data resources as the value contribution parameters of the new data resources; S13: Reduce the threshold values ​​of indicators in the decision rule base according to a preset percentage to obtain a temporary decision rule base; S14: Match the value contribution parameters of the new data resources with each decision rule in the temporary decision rule base to obtain the execution action for the new data resources; S15: Make decisions and manage new data resources according to the execution actions of the new data resources; S16: When the preset observation period ends, the new data resources will be adjusted to the existing data resources, and the value contribution parameters of the existing data resources will be obtained based on the knowledge graph to support decision management.

[0055] The preset observation period can be set to 3 months. During these 3 months, existing data resources will be identified from the knowledge graph, and the type similarity and compliance risk level similarity between existing data resources and new data resources will be obtained. The type similarity and compliance risk level similarity will be weighted and summed to obtain the overall similarity between existing data resources and new data resources. The existing data resource with the highest overall similarity will be selected as the reference data resource. The value contribution parameters of the reference data resource are migrated to the value contribution parameters of the new data resource to facilitate cold start observation of the new data resource. During the observation period, the indicator thresholds in the decision rule base are reduced by a preset percentage. For example, all indicator thresholds in the decision rule base are reduced to only 80% of their original values. The decision rule base with updated indicator thresholds is called the temporary decision rule base. This temporary decision rule base is matched with the value contribution parameters of the new data resource to obtain the execution actions for the new data resource, thereby enabling decision management of the new data resource. If the value contribution parameters of the reference data resource change during the preset observation period, the value contribution parameters of the new data resource will also change accordingly. When the preset observation period ends, the new data resource itself has historical data. At this time, the new data resource can be adjusted to be an existing data resource in the knowledge graph, and its own unique value contribution parameters can be obtained based on its historical data.

[0056] In S2 and S3, key data entities include business scenarios and data products. Note that data products can be both core data entities and key data entities. For key data entities such as business scenarios, refer to... Figure 3 The flowchart illustrating the method for obtaining the reference list of new business scenarios demonstrates that "when a new business scenario is obtained, a reference list is pushed based on the similarity between the new business scenario and the existing business scenarios in the knowledge graph." This includes the following steps: S31: When a new business scenario is obtained, the first entity attribute of the new business scenario is obtained; S32: Based on the knowledge graph, obtain the second entity attributes for each existing business scenario; S33: Based on the first entity attribute and the second entity attribute, obtain the first comprehensive similarity between the new business scenario and each existing business scenario; S34: Based on the first comprehensive similarity, select at least one reference business scenario that is similar to the new business scenario from all existing business scenarios; S35: Build a reference list based on at least one reference business scenario.

[0057] The first entity attribute includes the scenario type, user entity, timeliness requirement parameters, expected data volume, and target compliance risk level of the new business scenario; the second entity attribute includes the scenario type, user entity, timeliness requirement parameters, resource call volume, and the highest compliance risk level of the associated data resources of the existing business scenario. The formula for calculating the first comprehensive similarity is: In the formula, This indicates the first overall similarity between the new business scenario and the existing business scenario; , , , , Indicates weight; This indicates the similarity between the scenario types of the new business scenario and the scenario types of existing business scenarios. This indicates the similarity between the users of the new business scenario and the users of the existing business scenario. This indicates the similarity between the timeliness requirement parameters of the new business scenario and the timeliness requirement parameters of the existing business scenario. This indicates the similarity between the expected data volume of the new business scenario and the resource call volume of the existing business scenario; This represents the similarity between the target compliance risk level of the new business scenario and the highest compliance risk level of the associated data resources in the existing business scenario. In calculating the first comprehensive similarity, the similarity levels of each dimension between the new and existing business scenarios can also be observed based on comparisons of various dimensions, as detailed in Table 5 below: Table 5 When enterprises add new business scenarios in their operations, it is often difficult to obtain a reasonable technical value index for users to refer to. Therefore, this method proposes to use a reference list to evaluate the technical value index of new business scenarios: based on the historical value and historical call frequency of each reference business scenario in the reference list; based on the historical value and historical call frequency, the technical value index of the new business scenario is obtained.

[0058] The technology value index for new business scenarios is an evaluation metric that quantifies the technological value of data assets within these scenarios. It assesses the feasibility of technology integration in new business scenarios and drives the automated orchestration strategies and elastic resource scaling of the data asset management system. The formula for calculating the technology value index for new business scenarios is: In the formula, This represents the technological value index of new business scenarios; This represents the historical value benchmark for new business scenarios; This represents the coefficient for the change in technology costs for new business scenarios; This represents the scenario difference coefficient for new business scenarios; This indicates the target utilization rate of the new business scenario.

[0059] Sort the first comprehensive similarity scores from largest to smallest, and select the top X (e.g., 3) existing business scenarios from those with a first comprehensive similarity score greater than the first comprehensive similarity threshold as reference business scenarios.

[0060] The formula for calculating the historical value benchmark of a new business scenario is: In the formula, This represents the historical value benchmark for new business scenarios; Indicates the number of reference business scenarios; Indicates the first The historical value of a reference business scenario; Indicates the first The first comprehensive similarity between the reference business scenario and the new business scenario; Indicates the first The weight of the reference business scenario, the first The higher the first overall similarity between the reference business scenario and the new business scenario, the better. The larger the value, the better. In addition, the weights of all reference business scenarios are added together to get 1.

[0061] Identify the data resources relied upon by the new business scenario, obtain the current total technical cost and historical total technical cost of the data resources relied upon by the new business scenario, and calculate the technical cost variation coefficient of the new business scenario using the following formula: In the formula, This represents the coefficient for the change in technology costs for new business scenarios; This indicates the current total technical cost of the data resources relied upon by the new business scenario; This represents the total historical technical cost of the data resources relied upon by the new business scenario.

[0062] In the formula, This represents the scenario difference coefficient for new business scenarios; Indicates the expected call frequency for the new business scenario; This indicates the historical average call frequency for new business scenarios.

[0063] The formula for calculating the historical average call frequency of new business scenarios is: In the formula, Indicates the first The historical call frequency of each reference business scenario.

[0064] After calculating the technical value index of new key data entities, business personnel can fine-tune the technical value index based on specific needs, and the adjustment records are synchronized to the knowledge graph. The core logic is implemented through Neo4jCypher or application layer scripts, and can be embedded into the original system's "decision support module" without reconstruction. The indicator thresholds, similarity weights, and correlation coefficients are all designed as configurable parameters (such as stored through configuration tables) to adapt to the needs of different industries such as finance and e-commerce.

[0065] For critical data entities such as data products, refer to Figure 4 The flowchart illustrating the method for obtaining a reference list of new data products demonstrates that "when a new data product is obtained, a reference list is pushed based on the similarity between the new data product and existing data products in the knowledge graph," which includes the following steps: S36: When acquiring a new data product, acquire the third entity attribute of the new data product; S37: Based on the knowledge graph, obtain the fourth entity attribute of each existing data product; S38: Based on the third entity attribute and the fourth entity attribute, obtain the second comprehensive similarity between the new data product and each existing data product; S39: Based on the second comprehensive similarity, select at least one reference data product that is similar to the new data product from all existing data products; S310: Construct a reference list based on at least one reference data product.

[0066] The third entity attribute includes the product type of the new data product, the data resources it relies on, and the business scenarios of its services; the fourth entity attribute includes the product type of the existing data product, the data resources it relies on, and the business scenarios of its services. The formula for calculating the second comprehensive similarity is: In the formula, This indicates the second comprehensive similarity between the new data product and existing data products. , , These represent weights, with values ​​of 0.4, 0.3, and 0.3 respectively. This indicates the similarity between the product types of the new data product and the product types of existing data products. This indicates the degree of overlap between the data resources relied upon by new data products and those relied upon by existing data products. This indicates the similarity between the business scenarios of new data products and services and those of existing data products and services.

[0067] When enterprises add new data products during operation, it is often difficult to obtain a reasonable technology value index for users to refer to. Therefore, this method proposes to use a reference list to evaluate the technology value index of new data products: based on the unit technology cost of each reference data product in the reference list, the technology value index of the new data product is obtained based on the unit technology cost of each reference data product. The technology value index of the new data product is a multi-dimensional evaluation parameter that quantifies the technological contribution of data assets. It is used to characterize the comprehensive technology value density of the new data product relative to the reference data products in terms of technical architecture, data governance efficiency, and system compatibility. This index is calculated through an algorithm model, providing automated technology classification and resource scheduling decision support for the data asset management platform. Its value ranges from [0, +∞), and a higher value indicates better technology integration and system adaptability. The calculation formula for the technology value index of the new data product is as follows: In the formula, Indicating the technological value index of new data products; Indicates the quantity of products referenced in the data; Indicates the first Unit technology cost of a reference data product; Indicates the new data product and the first The second comprehensive similarity between reference data products; This indicates the demand premium percentage for new data products; This represents the target profit margin for the new data product. A first value is obtained based on the timeframe specified by the product's timeliness requirements. A second value is obtained based on the highest compliance risk level among the data resources the product relies on. The first and second values ​​are then added together to obtain the demand premium percentage. For example, based on the timeliness requirements, the first value ranges from 10% to 15%, and based on the highest compliance risk level, the second value ranges from 20% to 25%.

[0068] The formula for calculating the unit technology cost of reference data products is as follows: In the formula, This indicates the amount of data resources that the reference data product relies on; The reference data product depends on the first The unit technology cost of a data resource; The reference data product depends on the first The usage percentage of a data resource refers to the percentage of data products used for the first data resource. The ratio between the number of times a data resource is accessed and the total number of times the reference data product accesses all the data resources it depends on.

[0069] This solution enables companies to develop precise budget plans based on a reasonable technology value index, ensuring that the technology inputs and expected technology outputs of various business activities match, thereby improving the efficiency of capital utilization. Simultaneously, during the project decision-making phase, a reasonable technology value index helps companies assess the feasibility and profitability of projects, avoiding blind investment. In terms of technology cost control, the technology value index can serve as a benchmark for technology cost accounting, prompting companies to optimize procurement processes and improve production efficiency to ensure that project technology costs are kept within a reasonable range. Furthermore, the technology value index can also serve as a performance evaluation standard, incentivizing employees to improve work efficiency and quality, ensuring timely project delivery and controllable technology costs. Through these measures, companies can reduce operating technology costs, improve operational efficiency, and enhance profitability, thereby maintaining a competitive edge in the fierce market.

[0070] In summary, with the advancement of data element marketization, enterprises face core technical pain points in data value management, including a broken and insufficiently granular data value attribution chain, making it impossible to accurately break down technical outputs to their source and quantify element contributions; one-sided data usage behavior statistics, with a disconnect between the "actual value contribution" and "usage activity" of resources; and a lack of automated technology support for business decision-making, resulting in opaque and inefficient processes. To address these issues, this method has implemented corresponding improvement measures and achieved significant results: improved attribution granularity and accuracy, supporting product-derived relationships and multi-resource combinations, achieving an attribution accuracy of ≥95% for refined technical outputs of complex data products; enhanced comprehensiveness of usage statistics, achieving full-chain usage frequency linkage statistics through "product-element (product and resource) usage relationships," with an error ≤3%; and automated and intelligent decision-making, providing technical implementation methods based on knowledge graph queries, solidifying decision-making logic within the system, supporting automated decision-making, and improving efficiency by over 50%. When an enterprise, as a data holder, generates the data required for this patented solution during its internal operations or data transactions with other enterprises, it can use knowledge graphs based on this data to complete the decision-making and management of core data entities and the technical value index assessment of new key data entities, and further optimize internal operations and external data transaction circulation.

[0071] Reference Figure 5 This disclosure provides a data asset analysis system, including: The knowledge graph construction module 101 is used to collect data entity information of data assets using API interfaces and to construct a knowledge graph based on the data entity information. In the knowledge graph, nodes represent data entities, and edges between nodes represent the relationships between data entities. Among them, data entities include business scenarios, business objectives, data products, and data resources; The relationships include the dependency relationship between business scenarios and data products, the value distribution relationship between business objectives and data products, the derivative relationship between data products, and the usage relationship between data products and data resources; Key identification module 102 is used to identify key data entities from all data entities; The list push module 103 is used to push a reference list based on the similarity between the existing key data entities in the knowledge graph and the new key data entities when a new key data entity is obtained.

[0072] The various variations and specific examples of the data asset analysis methods provided above are also applicable to the data asset analysis system provided in this disclosure. Through the foregoing detailed description of the data asset analysis methods, those skilled in the art can clearly understand the implementation method of the data asset analysis system. For the sake of brevity, they will not be described in detail here.

[0073] A computer device according to embodiments of the present disclosure includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, etc.

[0074] The processor may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the computer device to perform desired functions. In one embodiment of this disclosure, the processor is used to execute computer-readable instructions stored in the memory, causing the computer device to perform all or part of the steps of the data asset analysis methods of the foregoing embodiments of this disclosure.

[0075] Those skilled in the art will understand that, in order to solve the technical problem of how to achieve a good user experience, this embodiment may also include well-known structures such as communication buses and interfaces, and these well-known structures should also be included within the protection scope of this disclosure.

[0076] like Figure 6This is a schematic diagram of a computer device provided for an embodiment of the present disclosure. It illustrates a structural schematic diagram suitable for implementing the computer device in the embodiments of the present disclosure. Figure 6 The computer device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0077] like Figure 6 As shown, a computer device may include a processor (such as a central processing unit, graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) or programs loaded from storage devices into random access memory (RAM). The RAM also stores various programs and data required for the operation of the computer device. The processor, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0078] Typically, the following devices can be connected to the I / O interface: input devices, such as sensors or visual information acquisition devices; output devices, such as displays; storage devices, such as magnetic tapes or hard drives; and communication devices. Communication devices allow the computer device to communicate wirelessly or wiredly with other devices (such as edge computing devices) to exchange data. Although Figure 6 A computer apparatus with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or included alternatively.

[0079] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processor, all or part of the steps of the data asset analysis method of embodiments of this disclosure are performed.

[0080] For a detailed description of this embodiment, please refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.

[0081] A computer-readable storage medium according to embodiments of the present disclosure stores non-transitory computer-readable instructions. When these non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the data asset analysis methods described in the foregoing embodiments of the present disclosure are performed.

[0082] The aforementioned computer-readable storage media include, but are not limited to: optical storage media (e.g., CD-ROM and DVD), magneto-optical storage media (e.g., MO), magnetic storage media (e.g., magnetic tape or portable hard drive), media with built-in rewritable non-volatile memory (e.g., memory card), and media with built-in ROM (e.g., ROM cartridge).

Claims

1. A data asset analysis method, characterized in that, include: Data entity information of data assets is collected using API interfaces, and a knowledge graph is constructed based on the data entity information; In the knowledge graph, nodes represent data entities, and edges between nodes represent the relationships between data entities. The data entities include business scenarios, business objectives, data products, and data resources; The relationships mentioned include the dependency relationship between business scenarios and data products, the value distribution relationship between business objectives and data products, the derivative relationship between data products, and the usage relationship between data products and data resources; Identify key data entities from all data entities; When a new key data entity is acquired, a reference list is pushed based on the similarity between the existing key data entities in the knowledge graph and the new key data entity.

2. The data asset analysis method according to claim 1, characterized in that, The construction of the knowledge graph includes: Based on the list of associated products, identify the data products that the business scenario depends on, and establish dependency edges between the data products and the dependent data products. Based on the upstream element list, identify the parent products that the data products depend on, and establish derivative edges between the data products and the parent products; Based on the derived edge, determine the data product that is the source of data value transfer, and establish a value distribution edge between the data product that is the source of data value transfer and the business objective; Based on the upstream element list, identify the data resources that the data product depends on, and establish a usage edge between the data product and the dependent data resources. The knowledge graph is constructed based on the edges between the business scenario, the business objective, the data product, and the data resource.

3. The data asset analysis method according to claim 2, characterized in that, The key data entities include business scenarios; when a new business scenario is acquired, a reference list is pushed based on the similarity between existing business scenarios in the knowledge graph and the new business scenario, including: When a new business scenario is acquired, the first entity attribute of the new business scenario is acquired; Based on the knowledge graph, obtain the second entity attribute for each existing business scenario; Based on the first entity attribute and the second entity attribute, obtain the first comprehensive similarity between the new business scenario and each existing business scenario; Based on the first comprehensive similarity, at least one reference business scenario similar to the new business scenario is selected from all existing business scenarios; Based on at least one of the aforementioned reference business scenarios, a reference list is constructed.

4. The data asset analysis method according to claim 3, characterized in that, Also includes: Based on the historical value and historical call frequency of each reference business scenario in the reference list; Based on the historical value and the historical call frequency, the technical value index of the new business scenario is obtained.

5. The data asset analysis method according to claim 3, characterized in that, The first entity attributes include the scenario type of the new business scenario, the user entity, the timeliness requirement parameters, the expected data volume, and the target compliance risk level; The second entity attribute includes the existing business scenario type, user entity, timeliness requirement parameters, resource call volume, and the highest compliance risk level of associated data resources; The formula for calculating the first comprehensive similarity is: In the formula, This indicates the first overall similarity between the new business scenario and the existing business scenario; , , , , Indicates weight; This indicates the similarity between the scenario types of the new business scenario and the scenario types of existing business scenarios. This indicates the similarity between the users of the new business scenario and the users of the existing business scenario. This indicates the similarity between the timeliness requirement parameters of the new business scenario and the timeliness requirement parameters of the existing business scenario. This indicates the similarity between the expected data volume of the new business scenario and the resource call volume of the existing business scenario; This indicates the similarity between the target compliance risk level of the new business scenario and the highest compliance risk level of the associated data resources in the existing business scenario.

6. The data asset analysis method according to claim 4, characterized in that, The formula for calculating the technology value index of the new business scenario is as follows: In the formula, This represents the technological value index of new business scenarios; This represents the historical value benchmark for new business scenarios; This represents the coefficient for the variation of technical costs in new business scenarios; This represents the scenario difference coefficient for new business scenarios; This indicates the target utilization rate of the new business scenario.

7. The data asset analysis method according to claim 6, characterized in that, The formula for calculating the historical value benchmark of a new business scenario is: In the formula, This represents the historical value benchmark for new business scenarios; Indicates the number of reference business scenarios; Indicates the first The historical value of a reference business scenario; Indicates the first The first comprehensive similarity between the reference business scenario and the new business scenario; Indicates the first The weight of each reference business scenario; The formula for calculating the coefficient of variation of technology costs in new business scenarios is as follows: In the formula, This indicates the current total technical cost of the data resources relied upon by the new business scenario; This represents the total historical technical cost of the data resources relied upon by the new business scenario; The formula for calculating the scenario difference coefficient for new business scenarios is: In the formula, This represents the scenario difference coefficient for new business scenarios; Indicates the expected call frequency for the new business scenario; This indicates the historical average call frequency for new business scenarios; The formula for calculating the historical average call frequency of new business scenarios is: In the formula, Indicates the first The historical call frequency of each reference business scenario.

8. The data asset analysis method according to claim 2, characterized in that, The key data entities include data products; when a new data product is acquired, a reference list is pushed based on the similarity between the existing data products in the knowledge graph and the new data product, including: When a new data product is acquired, the third entity attribute of the new data product is acquired; Based on the knowledge graph, obtain the fourth entity attribute of each existing data product; Based on the third entity attribute and the fourth entity attribute, a second comprehensive similarity is obtained between the new data product and each existing data product. Based on the second comprehensive similarity, at least one reference data product similar to the new data product is selected from all existing data products; Based on at least one reference data product, construct a reference list.

9. The data asset analysis method according to claim 8, characterized in that, Also includes: Based on the unit technology cost of each reference data product in the aforementioned reference list; The technological value index of the new data product is obtained based on the unit technological cost of each reference data product.

10. The data asset analysis method according to claim 8, characterized in that, The third entity attribute includes the product type of the new data product, the data resources it depends on, and the business scenarios of the services it provides. The fourth entity attribute includes the product type of the existing data product, the data resources it depends on, and the business scenarios of the services it provides. The formula for calculating the second comprehensive similarity is: In the formula, This indicates the second comprehensive similarity between the new data product and existing data products. , , Indicates weight; This indicates the similarity between the product types of the new data product and the product types of existing data products. This indicates the degree of overlap between the data resources relied upon by new data products and those relied upon by existing data products. This indicates the similarity between the business scenarios of new data products and services and those of existing data products and services.

11. The data asset analysis method according to claim 9, characterized in that, The formula for calculating the unit technical cost of the reference data product is as follows: In the formula, Indicates the first Unit technology cost of a reference data product; This indicates the amount of data resources that the reference data product relies on; The reference data product depends on the first The unit technology cost of a data resource; The reference data product depends on the first The percentage of data resources used.

12. The data asset analysis method according to claim 9, characterized in that, The formula for calculating the technology value index of the new data product is as follows: In the formula, This represents the technological value index of new data products; Indicates the quantity of products referenced in the data; Indicates the first Unit technology cost of a reference data product; Indicates the new data product and the first The second comprehensive similarity between reference data products; This indicates the demand premium percentage for new data products; This indicates the target profit margin for the new data product.

13. A computer device, characterized in that, The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the data asset analysis method according to any one of claims 1-12.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the data asset analysis method according to any one of claims 1-12.

15. A computer program product comprising computer instructions, characterized in that, When executed by a processor, the computer instructions implement the steps of the data asset analysis method according to any one of claims 1-12.

Citation Information

Patent Citations

  • Data asset library access method and device based on knowledge graph

    CN112256884A

  • Data life cycle management method and system based on data consanguinity map

    CN115080758A

  • Method and system for constructing digital asset directory

    CN117076685A

  • System and method for inputting data assets into table based on multi-level knowledge graph

    CN118710084A

  • Data asset classification method and system based on big data

    CN119989129A