A data asset management method and device

By constructing a data asset graph and performing clustering and identification reconstruction, the problem of accurate data asset identity verification was solved, achieving efficient data asset management and stable enterprise security operations.

CN117312597BActive Publication Date: 2026-04-14ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-06
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Traditional data asset management methods suffer from inaccurate configuration information and imperfect lifecycles due to isolated asset information, making it difficult to accurately verify the identity of data assets.

Method used

A data asset map is constructed based on a graph model. Identification markers are generated through clustering, and the data asset map is reconstructed to achieve accurate data asset management.

Benefits of technology

It improved the accuracy and efficiency of data asset positioning, achieved the integration of diverse and heterogeneous data, and enhanced the efficiency of enterprise data asset management and the stability of secure operation activities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117312597B_ABST
    Figure CN117312597B_ABST
Patent Text Reader

Abstract

The one or more embodiments of the specification disclose a data asset management method and device. The method comprises the following steps: firstly, constructing a first data asset graph based on a graph model according to obtained data asset information, wherein the data asset information at least comprises a data asset type and an association relationship between different data assets; then, performing clustering processing on the data asset information based on a connected subgraph contained in the first data asset graph to obtain a plurality of clustering subgraphs; finally, generating an identification for each clustering subgraph, reconstructing a second data asset graph for the identification according to the identification corresponding to each clustering subgraph and each data asset type contained in the data asset information, and performing data asset management on the data asset information based on the second data asset graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to the field of data asset management technology, and in particular to a data asset management method and apparatus. Background Technology

[0002] In enterprise security operations, data assets refer to the internet-related hardware and software products, including applications and interfaces, involved in the enterprise's business operations, as well as the enterprise's software products. Asset identity refers to the owner or actual controller of the data asset. In enterprise security operations scenarios, data asset identity verification is a crucial aspect of data asset management. Traditional data asset identity verification methods typically involve using traffic analysis or proactive scanning to discover the data asset's IP (Internet Protocol), MAC (Media Access Control or Medium Access Control, also known as physical address or hardware address, used to define the location of network devices), device, domain name, and other configuration information at the network or physical layer. This configuration information is then used to determine the actual controller of the data asset, allowing for appropriate actions to be taken.

[0003] However, since data assets are not usually managed centrally through a single system, and issues such as asset information silos and incomplete data asset lifecycles can lead to inaccurate configuration information, making it difficult to confirm the identity of data assets, a more effective data asset management method is needed to more accurately locate the identity of data assets. Summary of the Invention

[0004] On one hand, one or more embodiments of this specification provide a data asset management method, comprising: constructing a first data asset graph based on a graph model according to acquired data asset information, wherein the data asset information includes at least: data asset types and the relationships between different data assets; performing clustering processing on the data asset information based on connected subgraphs contained in the first data asset graph to obtain multiple clustering subgraphs; generating an identification identifier for each clustering subgraph; reconstructing a second data asset graph for the identification identifier according to the identification identifier corresponding to each clustering subgraph and each data asset type contained in the data asset information; and performing data asset management on the data asset information based on the second data asset graph.

[0005] On the other hand, one or more embodiments of this specification provide a data asset management device, comprising: a first data asset graph construction module, which constructs a first data asset graph based on a graph model according to acquired data asset information, wherein the data asset information includes at least: data asset types and the relationships between different data assets; a clustering module, which performs clustering processing on the data asset information based on the connected subgraphs contained in the first data asset graph to obtain multiple clustering subgraphs; and a second data asset graph construction module, which generates an identification identifier for each clustering subgraph, reconstructs a second data asset graph for the identification identifier based on the identification identifier corresponding to each clustering subgraph and each data asset type contained in the data asset information, and performs data asset management on the data asset information based on the second data asset graph.

[0006] In another aspect, one or more embodiments of this specification provide an electronic device, including: a processor; and a memory arranged to store computer-executable instructions, wherein when the executable instructions are executed, the processor is able to: construct a first data asset graph based on a graph model according to acquired data asset information, the data asset information including at least: data asset types and relationships between different data assets; perform clustering processing on the data asset information based on connected subgraphs contained in the first data asset graph to obtain multiple clustering subgraphs; generate an identification identifier for each clustering subgraph; reconstruct a second data asset graph for the identification identifier according to the identification identifier corresponding to each clustering subgraph and each data asset type contained in the data asset information; and perform data asset management on the data asset information based on the second data asset graph, including:

[0007] Furthermore, one or more embodiments of this specification provide a storage medium for storing computer-executable instructions. When executed by a processor, the executable instructions implement the following process: constructing a first data asset graph based on a graph model according to acquired data asset information, the data asset information including at least: data asset types and relationships between different data assets; performing clustering processing on the data asset information based on connected subgraphs contained in the first data asset graph to obtain multiple clustering subgraphs; generating an identification identifier for each clustering subgraph; reconstructing a second data asset graph for the identification identifier based on the identification identifier corresponding to each clustering subgraph and each data asset type contained in the data asset information; and performing data asset management on the data asset information based on the second data asset graph. Attached Figure Description

[0008] To more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in one or more embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 This is a schematic flowchart of a data asset management method according to an embodiment of this specification;

[0010] Figure 2 This is a schematic diagram illustrating the implementation principle of a data asset management method according to an embodiment of this specification;

[0011] Figure 3 This is a schematic block diagram of a data asset management device according to an embodiment of this specification;

[0012] Figure 4 This is a schematic block diagram of an electronic device according to an embodiment of this description. Detailed Implementation

[0013] This specification provides a data asset management method and apparatus through one or more embodiments.

[0014] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this document.

[0015] like Figure 1 As shown in the embodiments of this specification, a data asset management method is provided. The execution subject of this method can be a terminal device or a server. The terminal device can be a computer device such as a laptop or desktop computer. The server can be a single server or a server cluster composed of multiple servers. The server can be a backend server for financial services or online shopping services, or a backend server for a certain application. This embodiment uses a server as an example for detailed description. For the execution process of the terminal device, please refer to the relevant content below, which will not be repeated here. The method specifically includes the following steps:

[0016] In step S102, a first data asset map is constructed based on the acquired data asset information and a graph model. The data asset information includes at least the data asset type and the relationship between different data assets.

[0017] This specification's embodiments employ a knowledge graph approach to construct a first data asset graph and a second data asset graph for data asset management. The first data asset graph is an initial data asset graph constructed using a knowledge graph approach based on the acquired data asset information. In implementation, the first data asset graph can be constructed based on the acquired data asset information, either according to the function of the data assets, the upstream and downstream relationships of the data assets, or the data asset type and the relationships between different data assets. The first data asset graph includes nodes and edges, which can be constructed in various ways. For example, nodes can be constructed using data asset identifiers, and edges can be constructed based on the relationships between different data assets.

[0018] Data asset information is typically scattered across different data asset management platforms, and methods for obtaining this information can be based on these different platforms. In the embodiments of this specification, the data asset information can be internet technology-related hardware and software products, including applications and interfaces, or it can be enterprise software products.

[0019] Graphical models, or probabilistic graphical models, are probabilistic models that use a graph structure to describe the conditional independence relationships between multiple random variables. They include directed and undirected graphical models, and correspondingly, the first data asset graph can be either a directed or undirected graph. In practical applications, the choice between directed and undirected graphical models depends on the relationships between different data assets. If the relationships between different data assets involve data transfer or flow, a directed graphical model is used. If there are no data transfer or flow relationships, an undirected graphical model is used. For example, when database A reads a table from database B, or database A imports data into database B, a directed graphical model is used.

[0020] In step S104, the data asset information is clustered based on the connected subgraphs contained in the first data asset map to obtain multiple clustering subgraphs.

[0021] In the first data asset graph, the relationships between different data assets can include direct and indirect relationships. Based on these relationships, one or more connected subgraphs may exist between different nodes in the first data asset graph. Data asset information is then clustered based on these connected subgraphs to obtain multiple clustering subgraphs. This process can also be understood as: in the first data asset graph, data assets with relationships are clustered by searching connected graphs to obtain multiple clustering subgraphs. Typically, in directed graphs, different data assets have strong connections, and the corresponding clustering subgraphs are strongly connected. In undirected graphs, different data assets have weak connections, and the corresponding clustering subgraphs are weakly connected. Each clustering subgraph contains data assets with actual connections, and each clustering subgraph corresponds to a data asset controller, which is the actual controller of each data asset within the clustering subgraph.

[0022] For example, if user a opens a card D at a bank using mobile phone number C, then there is a relationship between C and D. If user a opens a card E at another bank using the same mobile phone number C, then there is a relationship between C and E. By clustering the above data assets and the relationships between them, a cluster subgraph consisting of C, D, and E can be obtained. This cluster subgraph corresponds to a data asset controller, namely user a.

[0023] In practice, the data asset controller can be any natural person, ecological institution, or internal organization of an ecological institution that conducts data asset research and development or manages data assets.

[0024] In step S106, an identification identifier is generated for each cluster subgraph. A second data asset map is reconstructed based on the identification identifier corresponding to each cluster subgraph and each data asset type contained in the data asset information. Data asset management is then performed on the data asset information based on the second data asset map.

[0025] The identification identifier of the cluster subgraph is also known as the identification identifier of the data asset controller, or identity identifier. The same data asset may be registered or operated under different identities in different systems. For example, the same data asset may be registered using an IP account in one system and operated using a domain name in another system. The data assets corresponding to different identities are the same data asset. Step S106 generates a unique identification identifier for the cluster subgraph corresponding to the data asset, and the identification identifiers are different for different data asset types. Based on the identification identifiers corresponding to each cluster subgraph and each data asset type, a second data asset graph is reconstructed for each identification identifier. In the second data asset graph, based on the identification identifiers, different types of data assets form multiple different data asset clusters. After reconstructing the first data asset graph to obtain the second data asset graph, data asset management can be performed based on the reconstructed second data asset graph. Similar to the principle of the first data asset graph, the second data asset graph can be a directed graph or an undirected graph.

[0026] In implementation, the identification identifier of a cluster subgraph can be the user identifier corresponding to the cluster subgraph, the identifier of the application corresponding to the cluster subgraph, or the identifier related to the terminal device used by the user corresponding to the cluster subgraph.

[0027] This specification provides a data asset management method. First, based on the acquired data asset information, a first data asset graph is constructed using a graph model, thereby modeling the data asset information in a knowledge graph manner and applying a graph-based identity localization method to data asset management. Then, based on the connected subgraphs contained in the constructed first data asset graph, the data asset information is clustered to obtain multiple clustering subgraphs. Finally, an identification identifier is generated for each clustering subgraph. Based on the identification identifier corresponding to each clustering subgraph and each data asset type contained in the data asset information, a second data asset graph is reconstructed for the identification identifier. Data asset management is then performed based on the second data asset graph, thereby reorganizing the data asset information through graph analysis. By reconstructing a second data asset graph, data asset information of different types can be linked together. Successful matching of any data asset information completes the identification of the data asset, thus avoiding data asset information silos (e.g., missing registration of any data asset information or missing registration of relationships between data assets) or poor data asset information quality (e.g., inconsistent descriptions of data asset information). This avoids the failure of data asset relationship association and identification due to relying solely on single data asset configuration information for data asset management. This significantly improves the accuracy and efficiency of data asset identification, thereby enhancing the efficiency of data asset management. The construction of the second data asset graph also enables data asset identity authentication within the data asset management system, making various attribute information of data assets more accurate, further improving data asset management efficiency. Furthermore, the data asset management method in the embodiments of this specification, by applying a graph-based identification method to the management of data assets with complex data types, can integrate diverse and heterogeneous data, thereby improving the efficiency of enterprise data asset management and ultimately enhancing the stability of enterprise security operations.

[0028] Furthermore, the processing of step S102 above can be varied. Two optional processing methods are provided below. For the first implementation method, please refer to the processing of steps S10202-S10206.

[0029] In step S10202, nodes corresponding to different data asset types are generated based on the data asset type.

[0030] Data asset types can include: App (Application), server, interface, domain name, IP, etc.

[0031] In step S10204, corresponding edges are constructed based on the relationships between different data assets.

[0032] In step S10206, a first data asset graph is constructed based on the generated nodes and constructed edges.

[0033] For the principles of graph construction in practical applications, please refer to [link to relevant documentation]. Figure 2 The diagram illustrates the implementation principle of the data asset management method. The implementation principles of steps S10202-S10206 can be found in [link to documentation]. Figure 2 The first block diagram (1) in the diagram constructs the first data asset map part.

[0034] In the second embodiment, the data asset information acquired in step S102 includes high-timeliness data asset information and low-timeliness data asset information. Specifically, high-timeliness data asset information and low-timeliness data asset information can be distinguished based on the practicality of the data and the timeliness of information acquisition. High-timeliness data asset information is data asset information acquired within a period less than or equal to a preset first time. For example, the first time can be set to 1 minute, and the acquisition period for high-timeliness data asset information can be 1 minute, 30 seconds, 1 second, etc. Low-timeliness data asset information is data asset information acquired within a period greater than or equal to a preset second time. For example, the second time can be set to 12 hours, and the acquisition period for low-timeliness data asset information can be 1 day or 12 hours, etc.

[0035] Accordingly, in the second embodiment, step S102 can be implemented by the following steps S10212-S10216.

[0036] If the acquired data asset information is high-time-sensitivity data asset information, then based on the acquired data asset information, a first data asset map is constructed based on a graph model, including step S10212: based on the acquired data asset information, the first data asset map is constructed online based on a graph model.

[0037] If the acquired data asset information is low-time-sensitivity data asset information, then based on the acquired data asset information, a first data asset map is constructed based on a graph model, including step S10214: based on the acquired data asset information, the first data asset map is constructed offline based on a graph model.

[0038] For low-time-sensitive data asset information, assuming the acquisition period of low-time-sensitive data asset information is T, the time for offline construction of the first data asset map is usually T+1. That is, the construction of the first data asset map begins after acquiring a complete period of low-time-sensitive data asset information. However, there is no such restriction in the application scenarios of high-time-sensitive data asset information.

[0039] If the acquired data asset information includes high-timeliness data asset information and low-timeliness data asset information, then based on the acquired data asset information, a first data asset map is constructed based on a graph model, including step S10216: Based on the acquired data asset information, a first data asset map is constructed using a flow-batch integrated data processing rule based on a graph model.

[0040] The unified stream and batch data processing rule uses the same set of APIs (Application Program Interfaces) and the same development paradigm to implement stream computing and batch computing of big data, thereby ensuring the consistency between the processing process and the results. For example, in an application scenario that includes both transaction and registration operations, the transaction operation typically updates data in real time, while the registration operation involves data that is entered once and remains unchanged. Therefore, the unified stream and batch data processing rule can be used to handle this situation.

[0041] Specifically, in the embodiments of this specification, the integrated stream-batch data processing rule refers to combining streaming processing for high-time-sensitivity data asset information with batch processing for low-time-sensitivity data asset information. It employs both streaming and batch processing methods, maintaining integrated computation (the same set of computational logic can be applied to both streaming and batch processing modes) and integrated storage (data is stored on the same medium throughout both streaming and batch processing; regardless of the processing mode, data flow and storage are completed on the same medium). Since a data asset management scenario typically involves hundreds of different systems, some requiring real-time monitoring and others requiring data acquisition over a period of time, using the integrated stream-batch data processing rule to construct the first data asset map can achieve better data asset utilization efficiency, improve data asset utilization, reduce the complexity of data asset management, and increase data asset management efficiency.

[0042] Furthermore, in step S106 above, generating an identification identifier for each cluster subgraph and reconstructing the second data asset map for the identification identifier based on the identification identifier corresponding to each cluster subgraph and each data asset type contained in the data asset information can be done in various ways. The following provides an optional processing method, which can be found in the following steps S1062-S1066.

[0043] In step S1062, a corresponding identification identifier is generated for each cluster subgraph. This identification identifier is used to determine the data asset controller corresponding to the cluster subgraph.

[0044] In step S1064, based on the knowledge graph reasoning strategy, the direct association between each data asset type and the identification identifier corresponding to each cluster subgraph in the data asset information is reconstructed.

[0045] In implementation, knowledge graph reasoning strategies can include rule-based reasoning methods and algorithm-based reasoning methods. Specifically, rule-based reasoning methods can determine whether the flow direction of data assets conforms to logic by examining the relationships between different data assets in the first data asset graph; or they can determine whether a current association conforms to preset relationship rules by examining the association relationships between different data assets in the first data asset graph. Algorithm-based reasoning methods can use various algorithms based on graph models for reasoning. The specific algorithm used depends on the specific scenario of data asset management, and this specification does not limit this aspect in the embodiments.

[0046] In step S1066, one or more data asset clusters corresponding to different data asset controllers are generated based on direct association relationships, and a second data asset map is constructed based on one or more data asset clusters.

[0047] The second data asset map may include one or more data asset clusters, each data asset cluster corresponding to a data asset controller, and the data asset cluster contains direct associations between related data assets. Corresponding to the different types of identification identifiers in step S102, in the same second data asset map, one or more data asset clusters use the same type of identification identifier, for example: all are user identifiers corresponding to cluster subgraphs, all are application identifiers corresponding to cluster subgraphs, or all are identifiers related to information about the terminal devices used by the users corresponding to cluster subgraphs.

[0048] Figure 2 (2) Connectivity graph clustering corresponds to the implementation principle of step S104, and (3) Knowledge reconstruction corresponds to the implementation principle of step S106. Figure 2 The identification markers in the second data asset map are the identifiers of the applications corresponding to the cluster subgraphs.

[0049] Furthermore, the data asset information in step S102 also includes: data asset configuration information, data asset logs, and data asset codes. Here, data asset configuration information refers to configuration information in a narrow sense, including: data asset type, the proportion of data assets of different data asset types, IP address, domain name, MAC address, etc. Data asset logs include static logs and dynamic logs.

[0050] Furthermore, the data asset management method in the embodiments of this specification also includes step S108: storing the data asset cluster in the form of a graph in the graph database corresponding to the second data asset graph.

[0051] By storing one or more data asset clusters in the second data asset map, in practical applications, data asset information of any data asset type can be quickly associated with its corresponding identification mark, thereby confirming its unique data asset identity. Moreover, other related data asset information can be quickly queried through its identification mark.

[0052] This specification provides a data asset management method. First, based on the acquired data asset information, a first data asset graph is constructed using a graph model, thereby modeling the data asset information in a knowledge graph manner and applying a graph-based identity localization method to data asset management. Then, based on the connected subgraphs contained in the constructed first data asset graph, the data asset information is clustered to obtain multiple clustering subgraphs. Finally, an identification identifier is generated for each clustering subgraph. Based on the identification identifier corresponding to each clustering subgraph and each data asset type contained in the data asset information, a second data asset graph is reconstructed for the identification identifier. Data asset management is then performed based on the second data asset graph, thereby reorganizing the data asset information through graph analysis. By reconstructing a second data asset graph, data asset information of different types can be linked together. Successful matching of any data asset information completes the identification of the data asset, thus avoiding the failure of data asset relationship association and identification due to data silos (e.g., missing registration of any data asset information or missing registration of relationships between data assets) or poor data asset information quality (e.g., inconsistent descriptions of data asset information). This significantly improves the accuracy and efficiency of data asset identification, thereby enhancing the efficiency of data asset management. The construction of the second data asset graph also enables data asset identity authentication within the data asset management system, making various attribute information of data assets more accurate and further improving data asset management efficiency. Furthermore, the data asset management method in the embodiments of this specification, by applying a graph-based identification method to the management of data assets with complex data types, can integrate diverse and heterogeneous data, thereby improving the efficiency of enterprise data asset management and ultimately enhancing the stability of enterprise security operations.

[0053] In summary, specific embodiments of this subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing can be advantageous.

[0054] The above describes a data asset management method provided by one or more embodiments of this specification. Based on the same idea, one or more embodiments of this specification also provide a data asset management device, such as... Figure 3 As shown.

[0055] The data asset management device includes: a first data asset map construction module 210, a clustering module 220, and a second data asset map construction module 230, wherein:

[0056] The first data asset graph construction module 210 constructs a first data asset graph based on a graph model according to the acquired data asset information. The data asset information includes at least: data asset types and the relationships between different data assets.

[0057] Clustering module 220 performs clustering processing on data asset information based on the connected subgraphs contained in the first data asset graph, and obtains multiple clustering subgraphs;

[0058] The second data asset graph construction module 230 generates identification labels for each cluster subgraph, reconstructs the second data asset graph for the identification labels based on the identification labels corresponding to each cluster subgraph and each data asset type contained in the data asset information, and performs data asset management based on the second data asset graph.

[0059] Furthermore, in one embodiment, the first data asset map construction module 210 includes:

[0060] The node generation unit generates nodes corresponding to different data asset types based on the data asset type.

[0061] Edge construction unit: Construct corresponding edges based on the relationships between different data assets;

[0062] The first graph construction unit constructs the first data asset graph based on the generated nodes and constructed edges.

[0063] In another implementation, the acquired data asset information includes high-timeliness data asset information and low-timeliness data asset information. Accordingly, the first data asset map construction module 210 includes:

[0064] The high-timeliness data asset graph construction unit, if the acquired data asset information is high-timeliness data asset information, constructs the first data asset graph online based on the acquired data asset information and a graph model;

[0065] The low-timeliness data asset graph construction unit, if the acquired data asset information is low-timeliness data asset information, constructs the first data asset graph offline based on the acquired data asset information and the graph model;

[0066] If the acquired data asset information includes both high-timeliness and low-timeliness data asset information, the first data asset graph is constructed based on the acquired data asset information and the data processing rules of the integrated batch processing model.

[0067] Furthermore, the second data asset graph construction module 230 includes:

[0068] The identification identifier generation unit generates a corresponding identification identifier for each cluster subgraph. The identification identifier is used to determine the data asset controller corresponding to the cluster subgraph.

[0069] The reconstruction unit, based on the knowledge graph reasoning strategy, reconstructs the direct relationship between each data asset type and the identification identifier corresponding to each cluster subgraph contained in the data asset information.

[0070] The second data asset graph construction unit generates one or more data asset clusters corresponding to different data asset controllers based on direct relationships, and constructs a second data asset graph based on one or more data asset clusters.

[0071] Furthermore, the data asset management device also includes a storage module that stores the data asset clusters in the form of a graph in the graph database corresponding to the second data asset graph.

[0072] This specification provides a data asset management device. First, a first data asset graph construction module constructs a first data asset graph based on the acquired data asset information using a graph model. This models the data asset information in a knowledge graph manner, applying a graph-based identity localization method to data asset management. Then, a clustering module clusters the data asset information based on the connected subgraphs contained in the constructed first data asset graph, obtaining multiple clustered subgraphs. Finally, a second data asset graph construction module generates an identification identifier for each clustered subgraph. Based on the identification identifier corresponding to each clustered subgraph and each data asset type contained in the data asset information, a second data asset graph is reconstructed for the identification identifier. Data asset management is then performed based on the second data asset graph, thereby reorganizing the data asset information through graph analysis. By reconstructing a second data asset graph, data asset information of different types can be linked together. Successful matching of any data asset information completes the identification of the data asset, thus avoiding data asset information silos (e.g., missing registration of any data asset information or missing registration of relationships between data assets) or poor data asset information quality (e.g., inconsistent descriptions of data asset information). This avoids the failure of data asset relationship association and identification due to relying solely on single data asset configuration information for data asset management. This significantly improves the accuracy and efficiency of data asset identification, thereby enhancing the efficiency of data asset management. The construction of the second data asset graph also enables data asset identity authentication within the data asset management system, making various attribute information of data assets more accurate, further improving data asset management efficiency. Furthermore, the data asset management method in the embodiments of this specification, by applying a graph-based identification method to the management of data assets with complex data types, can integrate diverse and heterogeneous data, thereby improving the efficiency of enterprise data asset management and ultimately enhancing the stability of enterprise security operations.

[0073] Those skilled in the art will understand that the above-described data asset management device can be used to implement the data asset management method described above. The detailed description therein should be similar to the method description above, and will not be repeated here to avoid being cumbersome.

[0074] Based on the same idea, one or more embodiments of this specification also provide an electronic device, such as... Figure 4As shown. Electronic devices can vary considerably due to differences in configuration or performance, and may include one or more processors 301 and memory 302. Memory 302 may store one or more application programs or data. Memory 302 may be temporary or persistent storage. The application programs stored in memory 302 may include one or more modules (not shown), each module may include a series of computer-executable instructions for the electronic device. Furthermore, processor 301 may be configured to communicate with memory 302 and execute the series of computer-executable instructions in memory 302 on the electronic device. The electronic device may also include one or more power supplies 303, one or more wired or wireless network interfaces 304, one or more input / output interfaces 305, and one or more keyboards 306.

[0075] Specifically, in this embodiment, the electronic device includes a memory and one or more programs, wherein one or more programs are stored in the memory, and one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for use in the electronic device, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following:

[0076] Based on the acquired data asset information, a first data asset map is constructed using a graph model. The data asset information includes at least: data asset types and the relationships between different data assets.

[0077] Based on the connected subgraphs contained in the first data asset graph, the data asset information is clustered to obtain multiple clustering subgraphs;

[0078] A recognition identifier is generated for each cluster subgraph. A second data asset map is reconstructed based on the recognition identifier corresponding to each cluster subgraph and each data asset type contained in the data asset information. Data asset management is then performed on the data asset information based on the second data asset map.

[0079] This specification provides one or more embodiments of a storage medium for storing computer-executable instructions, which, when executed by a processor, implement the following process:

[0080] Based on the acquired data asset information, a first data asset map is constructed using a graph model. The data asset information includes at least: data asset types and the relationships between different data assets.

[0081] Based on the connected subgraphs contained in the first data asset graph, the data asset information is clustered to obtain multiple clustering subgraphs;

[0082] A recognition identifier is generated for each cluster subgraph. A second data asset map is reconstructed based on the recognition identifier corresponding to each cluster subgraph and each data asset type contained in the data asset information. Data asset management is then performed on the data asset information based on the second data asset map.

[0083] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0084] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in one or more software and / or hardware.

[0085] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0086] This specification describes one or more embodiments of methods, apparatus (systems), and computer program products according to embodiments of this specification with reference to flowchart illustrations and / or block diagrams. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, produce a machine for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0087] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0088] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0089] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0090] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0091] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0092] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0093] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. This specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0094] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0095] The above description is merely one or more embodiments of this specification and is not intended to limit this application. Various modifications and variations can be made to the one or more embodiments of this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of one or more embodiments of this specification should be included within the scope of the claims of one or more embodiments of this specification.

Claims

1. A data asset management method, comprising: Based on the acquired data asset information, a first data asset map is constructed using a graph model. The data asset information includes at least: data asset types and the relationships between different data assets. Based on the connected subgraphs contained in the first data asset map, the data asset information is clustered to obtain multiple clustering subgraphs; A corresponding identification identifier is generated for each of the clustering subgraphs, and the identification identifier is used to determine the data asset controller corresponding to the clustering subgraph; Based on the knowledge graph reasoning strategy, the direct association between each data asset type contained in the data asset information and the identification identifier corresponding to each cluster subgraph is reconstructed; Based on the direct association, one or more data asset clusters corresponding to different data asset controllers are generated, and a second data asset map is constructed based on the one or more data asset clusters. Data asset management is performed on the data asset information based on the second data asset map.

2. The method according to claim 1, wherein constructing a first data asset map based on a graph model according to the acquired data asset information comprises: Nodes corresponding to different data asset types are generated based on the data asset types described. Construct corresponding edges based on the relationships between the different data assets; Based on the generated nodes and constructed edges, a first data asset graph is built.

3. The method according to claim 1, wherein the data asset information further includes: Data asset configuration information, data asset logs, and data asset codes.

4. The method according to claim 1, wherein the identification identifier comprises: The user identifier corresponding to the cluster subgraph, the application identifier corresponding to the cluster subgraph, or the identifier of the terminal device information used by the user corresponding to the cluster subgraph are all one of these.

5. The method according to claim 1, wherein the first data asset map and the second data asset map are directed graph maps or undirected graph maps.

6. The method according to claim 1, further comprising: The data asset clusters are stored in the graph database corresponding to the second data asset graph in the form of a graph.

7. The method according to claim 1, wherein the acquired data asset information includes high-timeliness data asset information and low-timeliness data asset information; If the acquired data asset information is high-time-sensitivity data asset information, then the step of constructing a first data asset map based on a graph model according to the acquired data asset information includes: Based on the acquired data asset information, a first data asset map is constructed online using a graph model; If the acquired data asset information is low-time-sensitivity data asset information, then the step of constructing a first data asset map based on a graph model according to the acquired data asset information includes: Based on the acquired data asset information, the first data asset map is constructed offline using a graph model; If the acquired data asset information includes both high-timeliness data asset information and low-timeliness data asset information, then the step of constructing a first data asset map based on a graph model according to the acquired data asset information includes: Based on the acquired data asset information, a first data asset graph is constructed using a graph model and integrated batch processing data rules.

8. A data asset management device, comprising: The first data asset graph construction module constructs a first data asset graph based on a graph model according to the acquired data asset information. The data asset information includes at least: data asset types and the relationships between different data assets. The clustering module performs clustering processing on the data asset information based on the connected subgraphs contained in the first data asset map, and obtains multiple clustering subgraphs; The second data asset graph construction module generates a corresponding identification identifier for each cluster subgraph, the identification identifier being used to determine the data asset controller corresponding to the cluster subgraph; based on a knowledge graph reasoning strategy, it reconstructs the direct association between each data asset type contained in the data asset information and the identification identifier corresponding to each cluster subgraph; based on the direct association, it generates one or more data asset clusters corresponding to different data asset controllers, constructs a second data asset graph based on the one or more data asset clusters, and performs data asset management on the data asset information based on the second data asset graph.

9. An electronic device, comprising: processor; as well as A memory configured to store computer-executable instructions, which, when executed, enable the processor to: Based on the acquired data asset information, a first data asset map is constructed using a graph model. The data asset information includes at least: data asset types and the relationships between different data assets. Based on the connected subgraphs contained in the first data asset map, the data asset information is clustered to obtain multiple clustering subgraphs; A corresponding identification identifier is generated for each of the clustering subgraphs, and the identification identifier is used to determine the data asset controller corresponding to the clustering subgraph; Based on the knowledge graph reasoning strategy, the direct association between each data asset type contained in the data asset information and the identification identifier corresponding to each cluster subgraph is reconstructed; Based on the direct association, one or more data asset clusters corresponding to different data asset controllers are generated, and a second data asset map is constructed based on the one or more data asset clusters. Data asset management is performed on the data asset information based on the second data asset map.

Citation Information

Patent Citations

  • Asset data processing method, related device and storage medium

    CN116383223A