A system for enterprise data asset inventory and data relationship mining

Through a system architecture consisting of a data ontology module, a relation module, and a triple module, combined with natural language processing and security domain partitioning, the high cost, low efficiency, and security issues of data asset inventory and mining for SMEs have been resolved, achieving low-cost, efficient data management and secure data relation mining.

CN117009515BActive Publication Date: 2025-10-31CHINA INFOMRAITON CONSULTING & DESIGNING INST CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310762816.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-27
Publication Date
2025-10-31
Estimated Expiration
2043-06-27

AI Technical Summary

Technical Problem

Small and medium-sized enterprises (SMEs) face high costs and low efficiency when conducting data asset inventory and data mining, especially in the separate processing of structured and unstructured data, difficulties in data ontology alignment, insufficient data relationship mining, and difficulty in ensuring data security.

Method used

The system architecture adopts a data ontology module, a data relationship module, a data triple module, and a data display module. Through ontology alignment, relationship alignment, and triple construction, combined with natural language processing, it realizes the graph-based display of data and the division of security domains. It uses a network gateway for data transmission and automates data inventory and relationship mining.

Benefits of technology

It enables low-cost and efficient data asset inventory and relationship mining, improves data management efficiency, ensures data security, transforms cold data into hot data, and enhances data availability and value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117009515B_ABST
    Figure CN117009515B_ABST
Patent Text Reader

Abstract

This invention provides a system for enterprise data asset inventory and data relationship mining, including a data ontology module, a data relationship module, a data triple module, and a data visualization module. The data ontology module is used to establish an ontology database. The data relationship module establishes a database of data relationship modules, generated by an attribute relationship matrix between ontologies, with different ontology attribute relationship definitions established between every two different ontology domains. The data triple module automatically constructs association relationships between the ontology module and the data relationship module according to subject-verb-object rules. The data visualization module implements automatic image rendering, graph-based data display, and a data graph-to-natural language converter. This invention can protect data security and accessibility, and improve the efficiency of enterprise data management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a system for enterprise data asset inventory and data relationship mining. Background Technology

[0002] Statistics show that the winning bids for knowledge graph projects often reach several million yuan, with clients concentrated in a few industries such as power, finance, and public security. The construction cost of knowledge graphs is very high, and the implementation workload is also very large. During the implementation process, a large number of experts are needed to annotate the text, and the workload of later maintenance is also very large. This investment approach is not suitable for small and medium-sized enterprises (SMEs). SMEs typically store their structured data using software module forms or Excel data storage. The correspondence between data ontologies is relatively certain, but because the data exists in different application systems and personal office computers, the phenomenon of data silos is very common, and the data becomes cold data. How to conduct low-cost inventory of these data assets and mine the relationships between data is a common concern for SMEs.

[0003] A search using "data asset inventory" as a keyword yields results primarily focused on machine learning methods. For example, the Chinese patent, "A Method and System for Automated Data Asset Inventory" (patent number CN 113792081B), defines attribute requirements corresponding to metadata standards to obtain a set of data asset metadata. It then defines an automated method and model for extracting metadata from business systems to obtain a set of metadata within those systems. Based on deep learning algorithms, it defines and trains a model for automatic identification and similarity matching of metadata standards and metadata to obtain an automated matching algorithm. The shortcomings of this patent in applying data asset inventory and data mining are: 1. It doesn't consider low-cost implementation methods for structured and unstructured data separately; 2. Machine learning relies on extensive expert annotation of documents, but in practice, only unstructured data needs annotation, while structured data can be inventoried and relationships between data points can be mined through ontology alignment.

[0004] Alternatively, one could conduct an inventory of the source data, such as the Chinese patent: "A Data Asset Inventory Method" (patent number CN115269589A). This method inventories data at the system, table, and field levels. By organizing the descriptive information of database tables and fields, it supplements the missing information with accurate descriptions, forming a data resource catalog. However, this patent's application of data inventory has the following shortcomings: 1. It cannot accurately organize the statistical data of the ontology, resulting in alignment issues between ontologites; 2. It cannot find relationships between data, preventing data mining and the generation of new information or knowledge, thus leaving the data as "cold data."

[0005] A search using "data asset inventory" and "data mining" as keywords yielded no relevant technology patents.

[0006] Currently, data has become a new factor of production. How to make good use of data to generate new value is a common challenge faced by small and medium-sized enterprises. Current data collection methods, including databases and data lakes, cannot identify the ontology and relationships of the data.

[0007] Data, as a factor of production, differs from other assets such as raw materials and production equipment. It is replicable, can be transmitted online, and involves information such as a company's trade secrets. Therefore, data security must be a primary concern during data inventory and mining. Consequently, different types of data need to be categorized and grouped into different security domains based on the company's needs. Furthermore, the more relationships between data entities, the greater the data's value. Therefore, it is necessary to move data from security domains with lower security requirements to security domains with higher security requirements, necessitating a one-way data transmission mechanism. Summary of the Invention

[0008] Purpose of the invention: The technical problem to be solved by the present invention is to address the existing problems of data asset inventory and data mining, and to provide a system for enterprise data asset inventory and data relationship mining.

[0009] The system of this invention includes a data ontology module, a data relationship module, a data triplet module, and a data display module;

[0010] The data ontology module is used to establish an ontology database, which includes customer names, employee names, regions, technical terms, project names, departments, contract names, patents, papers, soft science research projects, and awards.

[0011] The data relationship module is used to establish a data relationship module database. This database is generated by establishing an attribute relationship matrix between ontologies (such as customer name, employee name, project name, etc.) and the ontologies. Different ontology attribute relationship definitions are established between every two different ontology domains. These ontology attribute relationship definitions originate from the header fields of different structured data, automatically extracted by the system from various information systems. Taking four ontologies (X1, X2, X3, X4) as an example: the ontologies have a specific order; different orders result in different ontology relationship descriptions. Types such as X1X1, X2X2, X3X3, and X4X4 indicate the frequency of occurrence of the ontology data asset, enabling data inventory.

[0012]

[0013] The data triple module is used to construct triples according to the rules of subject, predicate, and object, as well as the association relationships automatically constructed by the data ontology module and the data relation module: (Entity, relation, Entity), where Entity is the ontology and relation is the relation; for example, employee "Zhang San", attribute relation "graduated from", and customer name "Nanjing University" constitute (Zhang San, graduating institution, Nanjing University).

[0014] The data display module is used for automatic image rendering and data graph display, and provides a data graph to natural language converter, which converts triples into natural language to form text.

[0015] The system also includes a network gateway;

[0016] The data ontology module includes an ontology server; the data relationship module includes a relationship server and an operation and maintenance server; the data triplet module includes a triplet server; and the data display module includes a large screen subsystem, a medium screen subsystem, a small screen subsystem, and a dedicated screen subsystem.

[0017] The large screen subsystem includes a seamless LED display screen and a handheld computer;

[0018] The aforementioned mid-screen subsystem refers to the computer-side display of three main parts: the ontology module, the relationship module, and the triplet module; and the operation and maintenance end includes the ontology management module, the data relationship management module, the data triplet management module, and the data inventory module.

[0019] The data asset inventory includes modules for inventorying the quantity of ontologies, inventorying data relationships, and data retrieval.

[0020] The data ontology module performs data mapping on ontology features, the data relationship management module automatically associates the data mapped by the data ontology module, the data triplet module realizes the automatic presentation of the automatically associated modules, and the data inventory module realizes the automatic statistics of data volume for the data ontology module, the data relationship management module, and the data triplet module.

[0021] All servers and terminals have a modular structure, and server devices at the same level can be horizontally expanded through data buses and data interfaces.

[0022] The server equipment is divided into security domain 1 and security domain 2 according to the type of data. Security domain 1 and security domain 2 are connected through a network gateway, and data can only flow from security domain 1 to security domain 2 in a one-way manner.

[0023] The LED display screen's default display interface is an automatic dynamic display of the triplet layer relationship. The handheld computer can project information onto the LED display screen and can automatically retrieve the subject, perform data penetration according to the technical field subject presented by the triplet layer, and perform data penetration based on the relationship.

[0024] There is a correspondence between the data interface and the system hierarchy. The data ontology module, data relationship module, data triplet module, and data display module correspond to the input data interfaces A, B, C, and D respectively from the input end.

[0025] The input data interfaces A, B, C, and D are the data logic interfaces between different levels of the system.

[0026] The D interface includes four types: D1 interface, D2 interface, D3 interface, and D4 interface. The D1 interface, D2 interface, D3 interface, and D4 interface represent different data levels. Different interface converters are used to realize the conversion of data between different levels and the same level. Each data level can be horizontally expanded and compatible.

[0027] The algorithms between data conversions are direct matching relationships of one-to-many or many-to-one, thus avoiding information bias in fuzzy data matching;

[0028] The data in the ontology database is transformed via the A interface. The same ontology may have different names; in this case, ontology alignment and transformation are performed.

[0029] TRUTH=(BFOn1,BFOn2,BFOn3,…,BFOn N (1)

[0030] Where TRUTH represents the real ontology dataset; BFOn N This represents the specific data of the Nth entity; the value of N ranges from 1 to 10000.

[0031] The specific data in the aligned dataset of the ontology database represents a unique, real ontology. Data sets collected from different data sources via different A interfaces (typically 1-4 A interfaces) are repeatedly aligned with the ontology and then stored in the ontology server.

[0032] ALL BFO = (TRUTH1, TRUTH1, TRUTH3...TRUTH M (2)

[0033] Among them, ALL BFO refers to the ontology data collected from different A interfaces; TRUTH M This represents the Mth real ontology dataset;

[0034] The data relationship module database collects data from the ontology server via the B interface. The transformation of the data relationship module database is automatically performed based on field names. During data collection, there is a relationship alignment issue. ER (Entity-Relationship Data Set) data set alignment:

[0035] FN=(FNn1, FNn2, FNn3,..., FNn N (3)

[0036] Where FN represents the ontology relation dataset; FNn N This represents the Nth ontology relation data, where N ranges from 1 to 10000;

[0037] The data relationship module database ensures that the specific data in the aligned dataset represents a truly unique relationship. Data sets from different data sources collected by different data interfaces are repeatedly aligned and then stored in the relationship server.

[0038] ALL FN = (FN1, FN2, FN3,...FN M (4)

[0039] Where ALL FN represents all relations, and FN M The Mth data point collected from the same A interface and processed by formula (3) is the Mth data point, where M ranges from 1 to 10000.

[0040] The ontology and relationships are displayed graphically by forward data reading between data interfaces A and B, and by reverse data clustering and statistical analysis, and then the clustered data is displayed graphically.

[0041] In addition to graphs, it can also describe text paragraphs using natural language processing (NLP), avoiding the poor user experience and insufficient knowledge extraction problems caused by dense graphs.

[0042] Using natural language processing algorithms, the relationships between ontologies are converted into text paragraphs. The formula for converting from graph form to natural language is as follows:

[0043] PAR = PNL(“ontology”, “relation”, “ontology”) (5)

[0044] PAR stands for Natural Language Paragraph; PNL stands for Natural Language Processing Algorithm.

[0045] The system performs the following steps:

[0046] Step 1: In order to establish relationships between structured data from different sources, the data is first divided into two categories: one is non-classified business and internal enterprise operation data; the other is classified business data. The method of division is to automatically classify the information into security domain 1 and security domain 2 based on the machine source of the information. For example, data generated by non-classified computers, servers and other devices is automatically classified into security domain 1, while data generated by classified machines and servers is automatically classified into security domain 2. In particular, it automatically assigns the two types of data to two different security domains: security domain 1 and security domain 2.

[0047] Step 2: The first entity server in security domain 1 collects data from the data entity source device through interface A, and synchronously transmits the data to the second entity server in security domain 2 through the network gateway; the second entity server in security domain 2 collects data from the data entity source device through interface A.

[0048] Step 3: The data of the first ontology server in security domain 1 is aligned by ontology alignment rules (ontology alignment rules refer to matching or mapping concepts between two or more different ontology to achieve interoperability between ontology), and after being manually verified, it is automatically transmitted to the first relation server through the B interface, and synchronously transmitted to the second relation server in security domain 2 through the network gateway. The second relation server in security domain 2 collects the data of the second ontology server through the B interface.

[0049] Step 4: After the data of the first relation server in security domain 1 is aligned according to the relation alignment rules and manually verified, it is automatically transmitted to the first triple server through the C interface, and synchronously transmitted to the second triple server in security domain 2 through the network gateway. The second triple server in security domain 2 collects the data of the relation server in security domain 2 through the C interface.

[0050] Step 5: The first triplet server in security domain 1 transmits the data processing results (which come from servers at each level) to the caches of the large screen subsystem, medium screen subsystem, and small screen subsystem through interfaces D1, D2, and D3, facilitating quick retrieval for the three types of terminal users; the second triplet server in security domain 2 transmits the triplet data to the cache of the dedicated screen subsystem through interface D4.

[0051] Step 6: The large screen subsystem, medium screen subsystem, small screen subsystem, and dedicated screen subsystem will synchronously present the corresponding data mining results in the form of knowledge graphs and natural language according to the user's search needs.

[0052] This invention belongs to the field of information technology, focusing on enterprise data asset inventory and data relationship mining. The enterprises mentioned in this invention include various large, medium, and small enterprises, with particular suitability for small and medium-sized enterprises.

[0053] The present invention has the following beneficial effects:

[0054] (1) Inventory of enterprise data assets by collecting data from different subsystems: Currently, enterprises generally lack effective ways to inventory data assets. This method can automatically collect the quantity and relational data of the data ontology, turning cold data into hot data, and realizing the secondary flow and processing of data.

[0055] (2) A simple way to achieve data relationship mining: Enterprises mainly conduct data mining by mining data from various subsystems through big data algorithms and machine learning, which has the problem of high cost and is unsustainable. This invention, by drawing on the knowledge graph approach, uses a non-machine learning method to achieve low-cost data collection and alignment in the form of data ontology, data relationships, and triples, and automatically finds the data relationships within them.

[0056] (3) Protect the security and accessibility of data: Enterprise data is divided into different security types. For data with high security requirements, this invention divides different security domains to protect the special business information of enterprises when the network gateway collects data.

[0057] (4) Improve the efficiency of enterprise data management: Through the layered approach of data asset inventory and data mining, the system has horizontal scalability and can include all structured data resources, thereby improving the efficiency of enterprise data resource management. Attached Figure Description

[0058] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, and the advantages of the present invention in the above and / or other aspects will become clearer.

[0059] Figure 1 This is the system architecture diagram of the present invention. Detailed Implementation

[0060] like Figure 1 As shown, the present invention provides a system for enterprise data asset inventory and data relationship mining, including a data ontology module, a data relationship module, a data triple module, and a data display module;

[0061] The data ontology module is used to establish an ontology database, which includes customer names, employee names, regions, technical terms, project names, departments, contract names, patents, papers, soft science research projects, and awards.

[0062] The data relationship module is used to establish a data relationship module database. This database is generated by establishing an attribute relationship matrix between ontologies (such as customer name, employee name, project name, etc.) and the ontologies. Different ontology attribute relationship definitions are established between every two different ontology domains. These ontology attribute relationship definitions originate from the header fields of different structured data, automatically extracted by the system from various information systems. Taking four ontologies (X1, X2, X3, X4) as an example: the ontologies have a specific order; different orders result in different ontology relationship descriptions. Types such as X1X1, X2X2, X3X3, and X4X4 indicate the frequency of occurrence of the ontology data asset, enabling data inventory.

[0063]

[0064] The data triple module is used to construct triples according to the rules of subject, predicate, and object, as well as the association relationships automatically constructed by the data ontology module and the data relation module: (Entity, relation, Entity), where Entity is the ontology and relation is the relation; for example, employee "Zhang San", attribute relation "graduated from", and customer name "Nanjing University" constitute (Zhang San, graduating institution, Nanjing University).

[0065] The data display module is used for automatic image rendering and data graph display, and provides a data graph to natural language converter, which converts triples into natural language to form text.

[0066] The system also includes a network gateway;

[0067] The data ontology module includes an ontology server; the data relationship module includes a relationship server and an operation and maintenance server; the data triplet module includes a triplet server; and the data display module includes a large screen subsystem, a medium screen subsystem, a small screen subsystem, and a dedicated screen subsystem.

[0068] The large screen subsystem includes a seamless LED display screen and a handheld computer;

[0069] The aforementioned mid-screen subsystem refers to the computer-side display of three main parts: the ontology module, the relationship module, and the triplet module; and the operation and maintenance end includes the ontology management module, the data relationship management module, the data triplet management module, and the data inventory module.

[0070] The data asset inventory includes modules for inventorying the quantity of ontologies, inventorying data relationships, and data retrieval.

[0071] The data ontology module performs data mapping on ontology features, the data relationship management module automatically associates the data mapped by the data ontology module, the data triplet module realizes the automatic presentation of the automatically associated modules, and the data inventory module realizes the automatic statistics of data volume for the data ontology module, the data relationship management module, and the data triplet module.

[0072] All servers and terminals have a modular structure, and server devices at the same level can be horizontally expanded through data buses and data interfaces.

[0073] The server equipment is divided into security domain 1 and security domain 2 according to the type of data. Security domain 1 and security domain 2 are connected through a network gateway, and data can only flow from security domain 1 to security domain 2 in a one-way manner.

[0074] The LED display screen's default display interface is an automatic dynamic display of the triplet layer relationship. The handheld computer can project information onto the LED display screen and can automatically retrieve the subject, perform data penetration according to the technical field subject presented by the triplet layer, and perform data penetration based on the relationship.

[0075] There is a correspondence between the data interface and the system hierarchy. The data ontology module, data relationship module, data triplet module, and data display module correspond to the input data interfaces A, B, C, and D respectively from the input end.

[0076] The input data interfaces A, B, C, and D are the data logic interfaces between different levels of the system.

[0077] The D interface includes four types: D1 interface, D2 interface, D3 interface, and D4 interface. The D1 interface, D2 interface, D3 interface, and D4 interface represent different data levels. Different interface converters are used to realize the conversion of data between different levels and the same level. Each data level can be horizontally expanded and compatible.

[0078] The algorithms between data conversions are direct matching relationships of one-to-many or many-to-one, thus avoiding information bias in fuzzy data matching;

[0079] The data in the ontology database is transformed via the A interface. The same ontology may have different names; in this case, ontology alignment and transformation are performed.

[0080] TRUTH=(BFOn1,BFOn2,BFOn3,…,BFOn N (1)

[0081] Where TRUTH represents the real ontology dataset; BFOn N This represents the specific data of the Nth entity; the value of N ranges from 1 to 10000.

[0082] The specific data in the aligned dataset of the ontology database represents a unique, real ontology. Data sets collected from different data sources via different A interfaces (typically 1-4 A interfaces) are repeatedly aligned with the ontology and then stored in the ontology server.

[0083] ALL BFO = (TRUTH1, TRUTH1, TRUTH3...TRUTH M (2)

[0084] Among them, ALL BFO refers to the ontology data collected from different A interfaces; TRUTH M This represents the Mth real ontology dataset;

[0085] The data relationship module database collects data from the ontology server via the B interface. The transformation of the data relationship module database is automatically performed based on field names. During data collection, there is a relationship alignment issue. ER (Entity-Relationship Data Set) data set alignment:

[0086] FN=(FNn1, FNn2, FNn3,..., FNn N (3)

[0087] Where FN represents the ontology relation dataset; FNn N This represents the Nth ontology relation data, where N ranges from 1 to 10000;

[0088] The data relationship module database ensures that the specific data in the aligned dataset represents a truly unique relationship. Data sets from different data sources collected by different data interfaces are repeatedly aligned and then stored in the relationship server.

[0089] ALL FN = (FN1, FN2, FN3,...FN M (4)

[0090] Where ALL FN represents all relations, and FN M The Mth data point collected from the same A interface and processed by formula (3) is the Mth data point, where M ranges from 1 to 10000.

[0091] The ontology and relationships are displayed graphically by forward data reading between data interfaces A and B, and by reverse data clustering and statistical analysis, and then the clustered data is displayed graphically.

[0092] In addition to graphs, it can also describe text paragraphs using natural language processing (NLP), avoiding the poor user experience and insufficient knowledge extraction problems caused by dense graphs.

[0093] Using natural language processing algorithms, the relationships between ontologies are converted into text paragraphs. The formula for converting from graph form to natural language is as follows:

[0094] PAR = PNL(“ontology”, “relation”, “ontology”) (5)

[0095] PAR stands for Natural Language Paragraph; PNL stands for Natural Language Processing Algorithm.

[0096] By categorizing enterprise data into different security types, this invention divides data with high security requirements into different security domains, enabling network gateways to protect specific business information during data collection. It utilizes non-machine learning methods to achieve low-cost data collection and alignment in the form of data ontology, data relationships, and triples, automatically identifying data relationships. By dividing the system into a data ontology module, a data relationship module, a data triples module, a data display module, and two major security domains, corresponding servers and interfaces are used to perform layer-by-layer alignment, collection, statistics, and mining of data, enabling data asset inventory and data relationship discovery.

[0097] The system performs the following steps:

[0098] Step 1: In order to establish relationships between structured data from different sources, the data is first divided into two categories: one is non-classified business and internal enterprise operation data; the other is classified business data. The method of division is to automatically classify the information into security domain 1 and security domain 2 based on the machine source of the information. For example, data generated by non-classified computers, servers and other devices is automatically classified into security domain 1, while data generated by classified machines and servers is automatically classified into security domain 2. In particular, it automatically assigns the two types of data to two different security domains: security domain 1 and security domain 2.

[0099] Step 2: The first entity server in security domain 1 collects data from the data entity source device through interface A, and synchronously transmits the data to the second entity server in security domain 2 through the network gateway; the second entity server in security domain 2 collects data from the data entity source device through interface A.

[0100] Step 3: The data of the first ontology server in security domain 1 is aligned by ontology alignment rules (ontology alignment rules refer to matching or mapping concepts between two or more different ontology to achieve interoperability between ontology), and after being manually verified, it is automatically transmitted to the first relation server through the B interface, and synchronously transmitted to the second relation server in security domain 2 through the network gateway. The second relation server in security domain 2 collects the data of the second ontology server through the B interface.

[0101] Step 4: After the data of the first relation server in security domain 1 is aligned according to the relation alignment rules and manually verified, it is automatically transmitted to the first triple server through the C interface, and synchronously transmitted to the second triple server in security domain 2 through the network gateway. The second triple server in security domain 2 collects the data of the relation server in security domain 2 through the C interface.

[0102] Step 5: The first triplet server in security domain 1 transmits the data processing results (which come from servers at each level) to the caches of the large screen subsystem, medium screen subsystem, and small screen subsystem through interfaces D1, D2, and D3, facilitating quick retrieval for the three types of terminal users; the second triplet server in security domain 2 transmits the triplet data to the cache of the dedicated screen subsystem through interface D4.

[0103] Step 6: The large screen subsystem, medium screen subsystem, small screen subsystem, and dedicated screen subsystem will synchronously present the corresponding data mining results in the form of knowledge graphs and natural language according to the user's search needs.

[0104] This invention provides a system for enterprise data asset inventory and data relationship mining. Many methods and approaches exist for implementing this technical solution; the above description is merely a preferred embodiment of the invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention. All components not explicitly stated in this embodiment can be implemented using existing technologies.

Claims

1. A system for enterprise data asset inventory and data relationship mining, characterized in that, It includes a data ontology module, a data relationship module, a data triple module, and a data visualization module; The data ontology module is used to establish an ontology database, which includes customer names, employee names, regions, technical terms, project names, departments, contract names, patents, papers, soft science research projects, and awards. The data relationship module is used to establish a data relationship module database. The data relationship module database is generated by establishing an attribute relationship matrix between ontology and ontology. Different ontology attribute relationship definitions are established between every two different ontology domains. The ontology attribute relationship definitions originate from the header fields of different structured data and are automatically extracted by the system from various information systems. The data triple module is used to construct triples (Entity, relation, Entity) according to the rules of subject, predicate, and object, as well as the association relationships automatically constructed by the data ontology module and the data relation module, where Entity is the ontology and relation is the relation. The data display module is used for automatic image rendering and data graph display, and provides a data graph to natural language converter, which converts triples into natural language to form text. The system performs the following steps: Step 1: In order to establish relationships between structured data from different sources, the data is first divided into two categories: one is non-confidential business and internal enterprise operation data; the other is confidential business data. The method of division is to automatically classify the information into security domain 1 and security domain 2 based on the machine source of the information. Step 2: The first entity server in security domain 1 collects data from the data entity source device through interface A, and synchronously transmits the data to the second entity server in security domain 2 through the network gateway; the second entity server in security domain 2 collects data from the data entity source device through interface A. Step 3: After the data of the first ontology server in security domain 1 is aligned with the ontology alignment rules and verified, it is automatically transmitted to the first relation server through the B interface and synchronously transmitted to the second relation server in security domain 2 through the network gateway. The second relation server in security domain 2 collects the data of the second ontology server through the B interface. Step 4: After the data of the first relation server in security domain 1 is aligned according to the relation alignment rules and verified, it is automatically transmitted to the first triplet server through the C interface, and synchronously transmitted to the second triplet server in security domain 2 through the network gateway. The second triplet server in security domain 2 collects the data of the relation server in security domain 2 through the C interface. Step 5: The first triplet server in security domain 1 transmits the data processing results to the caches of the large screen subsystem, medium screen subsystem, and small screen subsystem through interfaces D1, D2, and D3; The second triplet server in security domain 2 transmits triplet data to the cache of the dedicated screen subsystem via the D4 interface; Step 6: The large screen subsystem, medium screen subsystem, small screen subsystem, and dedicated screen subsystem will synchronously present the corresponding data mining results in the form of knowledge graphs and natural language according to the user's search needs.

2. The system according to claim 1, characterized in that, The system also includes a network gateway; The data ontology module includes an ontology server; the data relationship module includes a relationship server and an operation and maintenance server; the data triplet module includes a triplet server; and the data display module includes a large screen subsystem, a medium screen subsystem, a small screen subsystem, and a dedicated screen subsystem. The large screen subsystem includes a seamless LED display screen and a handheld computer; The aforementioned mid-screen subsystem refers to the computer-side display of three main parts: the entity module, the relationship module, and the triplet module.

3. The system according to claim 2, characterized in that, All servers and terminals have a modular structure, and server devices at the same level can be horizontally expanded through data buses and data interfaces. The server equipment is divided into security domain 1 and security domain 2 according to the type of data. Security domain 1 and security domain 2 are connected through a network gateway, and data can only flow from security domain 1 to security domain 2 in a one-way manner. The LED display screen's default display interface is an automatic dynamic display of the triplet layer relationship. The handheld computer can project information onto the LED display screen and can automatically retrieve the subject, perform data penetration according to the technical field subject presented by the triplet layer, and perform data penetration based on the relationship.

4. The system according to claim 3, characterized in that, There is a correspondence between the data interface and the system hierarchy. The data ontology module, data relationship module, data triplet module, and data display module correspond to the input data interfaces A, B, C, and D respectively from the input end. The input data interfaces A, B, C, and D are the data logic interfaces between different levels of the system. The D interface includes four types: D1 interface, D2 interface, D3 interface, and D4 interface. The D1 interface, D2 interface, D3 interface, and D4 interface represent different data levels. Different interface converters are used to realize the conversion of data between different levels and the same level. Each data level can be horizontally expanded and compatible.

5. The system according to claim 4, characterized in that, The algorithms for data conversion are direct matching relationships, either one-to-many or many-to-one. The data in the ontology database is transformed via interface A. The same ontology may have different names; in this case, ontology alignment and transformation are performed. TRUTH=(BFOn1,BFOn2,BFOn3, ……,BFOn N ) (1) Where TRUTH represents the real ontology dataset; BFOn N This represents the specific data of the Nth entity; the value of N ranges from 1 to 10000. The specific data in the aligned dataset of the ontology database represents a unique, real ontology. Data sets collected from different A interfaces from different sources are repeatedly aligned with the ontology and then stored in the ontology server. ALL BFO =(TRUTH1,TRUTH1,TRUTH3……TRUTH M ) (2) Among them, ALL BFO refers to the ontology data collected from different A interfaces; TRUTH M This represents the Mth real ontology dataset; The data relationship module database collects data from the ontology server via the B interface. The transformation of the data relationship module database is automatically performed based on field names. During data collection, relationship alignment occurs; ER dataset alignment is as follows: FN=(FNn1,FNn2,FNn3, ……,FNn N ) (3) Where FN represents the ontology relation dataset; FNn N This represents the Nth ontology relation data; The data relationship module database ensures that the specific data in the aligned dataset represents a truly unique relationship. Data sets from different data sources collected by different data interfaces are repeatedly aligned and then stored in the relationship server. ALL FN =(FN1,FN2,FN3, ……FN M ) (4) Where ALL FN represents all relations, and FN M The Mth data point collected from the same A interface and processed by formula (3).

6. The system according to claim 5, characterized in that, The ontology and relationships are displayed graphically by forward data reading between data interfaces A and B, and by reverse data clustering and statistical analysis, and then the clustered data is displayed graphically. In addition to graphs, it can also describe text segments using Natural Language Processing (NLP).

7. The system according to claim 6, characterized in that, Using natural language processing algorithms, the relationships between ontologies are converted into text paragraphs. The formula for converting from graph form to natural language is as follows: PAR = PNL("ontology", "relationship", "ontology") (5) PAR stands for Natural Language Paragraph; PNL stands for Natural Language Processing Algorithm.

Citation Information

Patent Citations

  • A method and system for automating data asset inventory

    CN113792081B

  • Data asset checking method

    CN115269589A

  • Industry process field knowledge graph construction method and device

    CN111444351A

  • Data mining system based on semantic network

    CN114357175A