A method and system for collecting metadata of a relational database
By collecting metadata from relational databases in the entire library, the problems of inconcentrated distribution of metadata and complex acquisition dependencies in the existing technology are solved, and centralized collection and efficient maintenance of metadata are realized.
Patent Information
- Application Number
- CN202111635901.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-29
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-12-29
AI Technical Summary
The metadata collection method of existing relational databases is mainly single library collection, which makes the metadata distributed in various nodes, which is difficult to maintain, and it is difficult to collect dependencies, so it is necessary to establish a connection relationship across nodes.
The method of collecting metadata in the entire library is adopted, and the schema list is obtained by pre-processing the data source parameters, the collector object is initialized, and the metadata is combined according to the preset metamodel, and the metadata is collected and cached, the metadata is established, and the metadata is finally saved and counted.
Centralized acquisition of metadata is realized, maintenance costs are reduced, acquisition efficiency is improved, and the complexity of cross-node connection relationships are avoided.
Smart Images

Figure CN114218229B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data collection, and in particular discloses a method and system for collecting metadata of a whole relational database. Background Art
[0002] Data has gradually become the core asset of enterprises. Data-driven business operations have played a huge role in different industries and enterprises. So, as the core asset data of enterprises, how to manage it is an important thing that different enterprises need to consider when carrying out comprehensive digital transformation.
[0003] Metadata is the most original dictionary of enterprise data assets. Metadata management is an important part of the construction of big data platforms and an important foundation for enterprises to realize data assets and asset services. In the data management environment, it is inextricably linked to data security, data quality, data architecture, data models, etc. It is also a bridge for business and technology to communicate. Therefore, the quality of metadata construction will have an important impact on the overall data and management of the enterprise.
[0004] Metadata collection is the data source for metadata management. Most of the existing relational database metadata collection methods are single-database collection. Multiple data sources need to be created for multiple databases. The metadata collected in this way is distributed in various nodes, which is not easy to maintain. It is also difficult to collect dependencies and it is necessary to establish connections across nodes.
[0005] Therefore, the existing metadata collection methods are difficult to maintain, have difficulty in collecting dependencies, and need to establish connections across nodes, which is a technical problem that needs to be solved urgently. Summary of the invention
[0006] The present invention provides a method and system for collecting metadata of a whole relational database, aiming to solve the technical problems existing in the existing metadata collection methods, such as difficulty in maintenance, difficulty in collecting dependency relationships, and the need to establish connection relationships across nodes.
[0007] One aspect of the present invention relates to a method for collecting metadata of a relational database, comprising the following steps:
[0008] Preprocess the data source parameters and obtain the schema list;
[0009] Initialize the collector object according to the obtained schema list and obtain the available collector objects;
[0010] According to the preset meta-model combination relationship, metadata information is collected in sequence, all required meta-models are obtained, metadata information is cached, and metadata cache is merged;
[0011] According to the preset metamodel dependencies, metadata dependencies are established in the cache;
[0012] All metadata are saved according to the established metadata dependencies, and statistical collection information is performed.
[0013] Furthermore, the steps of preprocessing the data source parameters and obtaining the schema list include:
[0014] Read schemas parameters;
[0015] Determine whether the schemas parameter read is the set symbol;
[0016] If the schema parameter read is a set symbol, the entire schema list is obtained from the database; if the schema parameter read is not a set symbol but a comma-separated schema parameter, the schema parameter is split by commas to obtain the schema list.
[0017] Furthermore, the collector object is initialized according to the obtained schema list, and the steps of obtaining the available collector object include:
[0018] According to the database type, obtain the SQL for switching schemas from the dialect object and use this SQL as the initialization SQL for the connection schema list;
[0019] Create a connection pool object according to the schema name, test whether the SQL and schema list are connected successfully, and return the collector object if the connection is successful.
[0020] Furthermore, according to the preset meta-model combination relationship, the steps of sequentially collecting metadata information, acquiring all required meta-models, caching metadata information, and merging metadata caches include:
[0021] According to the preset metamodel combination relationship, the collector objects are traversed in sequence until all required metamodels are obtained;
[0022] Traverse the metamodel and call all previous metadata collectors to collect metadata, and cache all collected metadata information in the Map object of the metadata collector;
[0023] When caching metadata information, the combination path is calculated based on the combination relationship of the metamodel;
[0024] After collecting all meta-models, merge the Map objects of each metadata collector to integrate the metadata information.
[0025] Furthermore, according to the preset metamodel dependency, the step of establishing metadata dependency in the cache includes:
[0026] When collecting metadata's metamodel dependency, determine whether the collected metadata and the metamodel are metadata dependencies based on the calculated combined path, and exclude erroneous relationships between metadata with the same name;
[0027] A caching mechanism is used to cache fields, storing the combined path of the field as the key value and the metadata information of the field as the value.
[0028] Another aspect of the present invention relates to a system for collecting metadata of a whole relational database, comprising:
[0029] The first acquisition module is used to preprocess the data source parameters and obtain the schema list;
[0030] The second acquisition module is used to initialize the collector object according to the acquired schema list and obtain the available collector object;
[0031] A processing module, used to collect metadata information in sequence according to a preset meta-model combination relationship, obtain all required meta-models, cache metadata information and merge metadata caches;
[0032] A building module for building metadata dependencies in a cache according to preset metamodel dependencies;
[0033] The statistics module is used to save all metadata and collect statistics information based on the established metadata dependencies.
[0034] Furthermore, the first acquisition module includes:
[0035] Reading unit, used to read schema parameters;
[0036] A first judging unit, used to judge whether the read schemas parameter is a set symbol;
[0037] The execution unit is used to obtain the entire schema list from the database if the schema parameter read is a set symbol; if the schema parameter read is not a set symbol but a comma-separated schema parameter, the schema parameter is split according to the comma to obtain the schema list.
[0038] Furthermore, the second acquisition module includes:
[0039] The acquisition unit is used to obtain the SQL for switching schemas from the dialect object according to the database type, and use this SQL as the initialization SQL for the connection schema list;
[0040] The test unit is used to create a connection pool object according to the schema name, test whether the SQL and schema list are connected successfully, and return the collector object if the connection is successful.
[0041] Furthermore, the processing module includes:
[0042] A traversal unit, used to traverse the collector objects in sequence according to the preset metamodel combination relationship until all required metamodels are obtained;
[0043] The first cache unit is used to traverse the meta-model and call all previous metadata collectors to collect metadata, and cache all collected metadata information into a Map object of the metadata collector;
[0044] A calculation unit, used for calculating a combination path according to a combination relationship of a meta-model when caching metadata information;
[0045] The merging unit is used to merge the Map objects of each metadata collector and integrate the metadata information after collecting all meta-models.
[0046] Further, the building module includes:
[0047] The second judgment unit is used to judge whether the collected metadata and the metamodel are metadata dependencies according to the calculated combination path when collecting the metadata's metamodel dependencies, and exclude erroneous relationships of metadata with the same name;
[0048] The second cache unit is used to cache the field by adopting a cache mechanism, storing the combined path of the field as a key value, and storing the metadata information of the field as a value.
[0049] The beneficial effects achieved by the present invention are:
[0050] The present invention provides a method and system for collecting metadata for the entire database of a relational database, which obtains a schema list by preprocessing data source parameters; initializes a collector object according to the obtained schema list to obtain an available collector object; collects metadata information in sequence according to a preset metamodel combination relationship, obtains all required metamodels, caches metadata information and merges metadata caches; establishes metadata dependencies in the cache according to a preset metamodel dependency relationship; saves all metadata according to the established metadata dependency relationship, and collects statistics of the collected information. The method and system for collecting metadata for the entire database of a relational database provided by the present invention have relatively concentrated metadata after collection, low maintenance cost and high efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 The present invention provides a flow chart of an embodiment of a method for collecting metadata from the entire relational database;
[0052] Figure 2 for Figure 1 A detailed flow chart of an embodiment of step 1 of preprocessing data source parameters and obtaining a schema list as shown in FIG.
[0053] Figure 3 for Figure 1 A detailed flow chart of an embodiment of the step 1 of initializing a collector object according to an acquired schema list and acquiring an available collector object;
[0054] Figure 4 for Figure 1 A detailed flow chart of an embodiment of step 1 of collecting metadata information, acquiring all required meta-models, caching metadata information and merging metadata caches in sequence according to a preset meta-model combination relationship as shown in FIG;
[0055] Figure 5 for Figure 1 A detailed flow chart of an embodiment of step 1 of establishing metadata dependency in a cache according to a preset metamodel dependency as shown in FIG.
[0056] Figure 6 A functional block diagram of an embodiment of a system for collecting metadata from the entire relational database provided by the present invention;
[0057] Figure 7 for Figure 6 A functional module diagram of an embodiment of a first acquisition module 1 shown in FIG.
[0058] Figure 8 for Figure 6 A functional module diagram of an embodiment of a second acquisition module 1 shown in FIG.
[0059] Fig. 9 for Figure 6 A functional module diagram of an embodiment of a processing module 1 shown in FIG.
[0060] Fig.10 for Figure 6 A functional module diagram of an embodiment of a building module 1 is shown in FIG.
[0061] Description of Figure Numbers:
[0062] 10. First acquisition module; 20. Second acquisition module; 30. Processing module; 40. Establishment module; 50. Statistics module; 11. Reading unit; 12. First judgment unit; 13. Execution unit; 21. Acquisition unit; 22. Testing unit; 31. Traversal unit; 32. First cache unit; 33. Calculation unit; 34. Merging unit; 41. Second judgment unit; 42. Second cache unit. DETAILED DESCRIPTION
[0063] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0064] like Figure 1 and Figure 2 As shown, the first embodiment of the present invention proposes a method for collecting metadata of a relational database, including the following steps:
[0065] Step S100: pre-process the data source parameters to obtain a schema list.
[0066] The JDBC (Java Data Base Connectivity) link is modified. Usually, JDBC does not have the ability to build a database, so a default schema must be specified when filling in the link. In this embodiment, you can enter an asterisk in the interface, or specify the database name and use commas to separate them, to indicate that multiple databases need to be collected. When performing the collection initialization work, the connection pool will be initialized according to the schemas parameter. If it is an asterisk, the database dialect object is obtained, and the SQL (Structured Query Language) of all schemas is obtained from the dialect object. The corresponding SQL is executed to obtain the schema list. If the schema parameters are separated by commas, the schema parameters are cut by commas to obtain the schema list.
[0067] Step S200: Initialize a collector object according to the acquired schema list to obtain an available collector object.
[0068] After getting the schema list, start to initialize the collector object. First, get the SQL for switching schemas from the dialect object according to the database type, and use this SQL as the initialization SQL for the connection. Create a connection pool object according to the schema name, and test to get the connection after creation. If the connection is successfully obtained, return the collector object, which contains the connection pool and metadata cache objects.
[0069] Step S300: According to the preset meta-model combination relationship, metadata information is collected in sequence, all required meta-models are acquired, metadata information is cached, and metadata cache is merged.
[0070] After obtaining the available collection objects, the collector objects are traversed in sequence according to the combination relationship of the metamodel until all the required metamodels are obtained. The metamodels are traversed and all the previous metadata collectors are called to collect, and all metadata information is cached in the collector's Map object. When caching metadata information, the combination path is calculated according to the combination relationship of the metamodel. After all metamodels are collected, the Map objects of each collector are merged to integrate the metadata information.
[0071] Step S400: Establish metadata dependencies in the cache according to preset metamodel dependencies.
[0072] When collecting relationships, it will determine whether it is a dependency relationship based on the combined path, and exclude erroneous relationships with metadata of the same name. In order to integrate multiple libraries, considering that fields are usually the most numerous in metadata information, the fields are cached. The caching mechanism is to store the combined path of the field as the key value and the metadata information of the field as the value, avoiding the situation where loop nesting is used when collecting field relationships, which causes a significant drop in multi-library collection.
[0073] Step S500: save all metadata according to the established metadata dependency and collect statistics.
[0074] All collected metadata are saved according to the established metadata dependency, and the collected metadata information is counted.
[0075] Compared with the prior art, the method for collecting metadata for the entire relational database provided in this embodiment obtains a schema list by preprocessing the data source parameters; initializes the collector object according to the obtained schema list to obtain an available collector object; collects metadata information in sequence according to the preset metamodel combination relationship, obtains all required metamodels, caches metadata information and merges metadata caches; establishes metadata dependencies in the cache according to the preset metamodel dependencies; saves all metadata according to the established metadata dependencies, and collects statistics on the collected information. The method for collecting metadata for the entire relational database provided in this embodiment realizes the generation of corresponding metadata collectors according to schemas; the collected metadata is relatively concentrated, with low maintenance costs and high efficiency.
[0076] Further, see Figure 2 , Figure 2 for Figure 1 The detailed flow chart of step S100 of an embodiment shown in FIG. 1 is a schematic diagram of a detailed flow chart of an embodiment. In this embodiment, step S100 includes:
[0077] Step S110: read the schemas parameter.
[0078] When performing collection initialization work, the connection pool will be initialized according to the schema parameters and the schema parameters will be read.
[0079] Step S120: Determine whether the read schemas parameter is a set symbol.
[0080] Determine whether the schemas parameter read is a set symbol, for example, whether the set symbol is *.
[0081] Step S130: if the read schemas parameter is a set symbol, then the entire schema list is obtained from the database; if the read schemas parameter is not a set symbol but a comma-separated schemas parameter, then the schemas parameter is split by commas to obtain the schema list.
[0082] If the schema parameter read is a set symbol, the database dialect object is obtained from the database, the SQL (Structured Query Language, Structured QueryLanguage) of all schemas is obtained from the dialect object, the corresponding SQL is executed, and the list of all schemas is obtained; if the schema parameter read is not a set symbol but a comma-separated schema parameter, the schema parameter is split by commas to obtain the schema list.
[0083] Compared with the prior art, the method for collecting metadata for the entire relational database provided in this embodiment reads the schemas parameter; determines whether the read schemas parameter is a set symbol; if the read schemas parameter is a set symbol, obtains the entire schema list from the database; if the read schemas parameter is not a set symbol but a comma-separated schemas parameter, the schemas parameter is split by comma to obtain the schema list. The method for collecting metadata for the entire relational database provided in this embodiment describes the need for collecting multiple databases from a relational database by filling in schemas in the data source, and solves the problem of significant performance degradation during multi-database collection, thereby making maintenance easier and avoiding the need for operation and maintenance personnel to establish multiple data sources to achieve the desired effect.
[0084] Preferably, see Figure 3 , Figure 3 for Figure 1 The detailed flow chart of step S200 of an embodiment is shown in FIG. 1 . In this embodiment, step S200 includes:
[0085] Step S210: According to the database type, the SQL for switching the schema is obtained from the dialect object, and this SQL is used as the initialization SQL for connecting the schema list.
[0086] After obtaining the schema list, start initializing the collector object. First, according to the database type, obtain the SQL for switching schemas from the dialect object and use this SQL as the initialization SQL for the connection.
[0087] Step S220: Create a connection pool object according to the schema name, test whether the SQL and the schema list are connected successfully, and return the collector object if the connection is successful.
[0088] Create a connection pool object according to the schema name. After creation, test to obtain the connection. If the acquisition is successful, return the collector object, which contains the connection pool and metadata cache objects.
[0089] Compared with the prior art, the method for collecting metadata for the entire relational database provided in this embodiment obtains the SQL for switching schemas from the dialect object according to the database type, and uses this SQL as the initialization SQL for connecting to the schema list; creates a connection pool object according to the schema name, tests whether the SQL and the schema list are successfully connected, and returns the collector object if the connection is successful. The method for collecting metadata for the entire relational database provided in this embodiment integrates the connection pool into the collector, obtains the SQL for switching schemas using the dialect object, and avoids the situation where a single JDBC connection can only connect to the same schemas; and the collected metadata is relatively centralized, with low maintenance costs and high efficiency.
[0090] Further, see Figure 4 , Figure 4 for Figure 1 The detailed flow chart of step S300 of an embodiment is shown in FIG. 1 . In this embodiment, step S300 includes:
[0091] Step S310: traverse the collector objects in sequence according to the preset meta-model combination relationship until all required meta-models are obtained.
[0092] After obtaining the available collection objects, the collector objects are traversed in sequence according to the combination relationship of the metamodels until all the required metamodels are obtained.
[0093] Step S320: traverse the meta-model and call all previous metadata collectors to collect metadata, and cache all collected metadata information in the Map object of the metadata collector.
[0094] Traverse the metamodel and call all previous metadata collectors to collect, and cache all metadata information in the collector's Map object.
[0095] Step S330: When caching metadata information, a combination path is calculated according to the combination relationship of the meta-model.
[0096] When caching metadata information, the combination path will be calculated based on the combination relationship of the metamodel.
[0097] Step S340: After all meta-models are collected, the Map objects of each metadata collector are merged to integrate the metadata information.
[0098] After all meta-models are collected, the Map objects of each collector will be merged to integrate the metadata information.
[0099] Compared with the prior art, the method for collecting metadata for the entire relational database provided by this embodiment is to traverse the collector objects in sequence according to the preset metamodel combination relationship until all required metamodels are obtained; traverse the metamodels and call all previous metadata collectors to collect metadata, and cache all collected metadata information in the Map object of the metadata collector; when caching metadata information, calculate the combination path according to the combination relationship of the metamodel; after collecting all metamodels, merge the Map objects of each metadata collector to integrate the metadata information. The method for collecting metadata for the entire relational database provided by this embodiment has relatively concentrated metadata after collection, low maintenance cost and high efficiency.
[0100] Preferably, see Figure 5 , Figure 5 for Figure 1 As shown in FIG. 1 , a detailed flow chart of an embodiment of step 1 of establishing metadata dependency in a cache according to a preset metamodel dependency is shown. In this embodiment, step S400 includes:
[0101] Step S410: When collecting the metamodel dependency of metadata, determine whether the collected metadata and the metamodel are in a metadata dependency relationship based on the calculated combined path, and exclude erroneous relationships of metadata with the same name.
[0102] When collecting relationships, it will determine whether it is a dependency relationship based on the combined path, thereby eliminating erroneous relationships with metadata of the same name.
[0103] Step S420: Use a cache mechanism to cache the field, store the combined path of the field as the key value, and use the metadata information of the field as the value.
[0104] In order to integrate multiple databases, considering that fields are usually the largest in metadata information, fields are cached. The cache mechanism is to store the combined path of the field as the key value and the metadata information of the field as the value, to avoid the situation where loop nesting is used when collecting field relationships, which causes a significant drop in multi-database collection.
[0105] Compared with the prior art, the method for collecting metadata for the entire relational database provided in this embodiment determines whether the collected metadata and the metamodel are metadata dependencies based on the calculated combined path when collecting metadata's metamodel dependencies, thereby eliminating erroneous relationships of metadata with the same name; a cache mechanism is used to cache fields, the combined path of the field is stored as the key value, and the metadata information of the field is used as the value. The method for collecting metadata for the entire relational database provided in this embodiment integrates the connection pool into the collector, uses the dialect object to obtain the SQL for switching the schema, and avoids the situation where a single JDBC connection can only connect to the same schemas; the collected metadata is relatively centralized, with low maintenance costs and high efficiency.
[0106] like Figure 6 As shown, Figure 6 The present invention provides a functional block diagram of an embodiment of a system for collecting metadata for a relational database of the present invention. In this embodiment, the system for collecting metadata for a relational database of the present invention comprises a first acquisition module 10, a second acquisition module 20, a processing module 30, an establishment module 40 and a statistical module 50, wherein the first acquisition module 10 is used to pre-process data source parameters and obtain a schema list; the second acquisition module 20 is used to initialize a collector object according to the obtained schema list and obtain an available collector object; the processing module 30 is used to sequentially collect metadata information, obtain all required metamodels, cache metadata information and merge metadata caches according to a preset metamodel combination relationship; the establishment module 40 is used to establish metadata dependencies in the cache according to a preset metamodel dependency; the statistical module 50 is used to save all metadata according to the established metadata dependency and to count the collected information.
[0107] The first acquisition module 10 is modified for the JDBC (Java Data Base Connectivity) link. Usually, JDBC does not have the ability to build a database, so a default schema must be specified when filling in the link. In this embodiment, you can enter an asterisk in the interface, or specify the database name and use commas to separate them, to indicate that you need to collect multiple databases. When performing the collection initialization work, the connection pool will be initialized according to the schemas parameter. If it is an asterisk, the database dialect object will be obtained, and the SQL (Structured Query Language) of all schemas will be obtained from the dialect object. The corresponding SQL will be executed to obtain the schema list. If the schema parameters are separated by commas, the schema parameters will be cut by commas to obtain the schema list.
[0108] After the second acquisition module 20 obtains the schema list, it starts to initialize the collector object. First, according to the database type, it obtains the SQL for switching the schema from the dialect object, and uses this SQL as the initialization SQL for the connection. It creates a connection pool object according to the schema name, and tests the connection after creation. If the connection is successfully obtained, it returns the collector object, which contains the connection pool and metadata cache objects.
[0109] After the processing module 30 obtains the available collection objects, it traverses the collector objects in sequence according to the combination relationship of the metamodel until all the required metamodels are obtained, traverses the metamodels and calls all the previous metadata collectors to collect, and caches all metadata information in the collector's Map object. When caching metadata information, the combination path will be calculated according to the combination relationship of the metamodel. After all metamodels are collected, the Map objects of each collector will be merged to integrate the metadata information.
[0110] The establishment module 40 will then determine whether it is a dependency relationship based on the combined path when collecting relationships, and exclude erroneous relationships with metadata of the same name. In order to integrate multiple libraries, considering that fields are usually the largest in metadata information, the fields are cached. The cache mechanism is to store the combined path of the field as the key value and the metadata information of the field as the value, avoiding the situation where loop nesting is used when collecting field relationships, which causes a significant drop in multi-library collection.
[0111] The statistics module 50 saves all the collected metadata according to the established metadata dependency and counts the collected metadata information.
[0112] Compared with the prior art, the system for collecting metadata for the entire relational database provided in this embodiment obtains a schema list by preprocessing the data source parameters; initializes the collector object according to the obtained schema list to obtain an available collector object; collects metadata information in sequence according to the preset metamodel combination relationship, obtains all required metamodels, caches metadata information and merges metadata caches; establishes metadata dependencies in the cache according to the preset metamodel dependencies; saves all metadata according to the established metadata dependencies, and counts the collected information. The system for collecting metadata for the entire relational database provided in this embodiment realizes the generation of corresponding metadata collectors according to schemas; the collected metadata is relatively concentrated, with low maintenance costs and high efficiency.
[0113] Further, see Figure 7 , Figure 7 for Figure 6 As shown in the functional module diagram of the first acquisition module embodiment, in this embodiment, the first acquisition module 10 includes a reading unit 11, a first judgment unit 12 and an execution unit 13, wherein the reading unit 11 is used to read schemas parameters; the first judgment unit 12 is used to judge whether the read schemas parameters are set symbols; the execution unit 13 is used to obtain the entire schema list from the database if the read schemas parameters are set symbols; if the read schemas parameters are not set symbols but are comma-separated schemas parameters, the schemas parameters are split by commas to obtain the schema list.
[0114] When performing the collection initialization work, the reading unit 11 will initialize the connection pool according to the schemas parameters and read the schemas parameters.
[0115] The first determination unit 12 determines whether the read schemas parameter is a set symbol, for example, whether the set symbol is *.
[0116] If the read schemas parameter is a set symbol, the execution unit 13 obtains a database dialect object from the database, obtains SQL (Structured Query Language, StructuredQuery Language) of all schemas from the dialect object, executes the corresponding SQL, and obtains a list of all schemas; if the read schemas parameter is not a set symbol but a comma-separated schemas parameter, the schemas parameter is split by commas to obtain a schema list.
[0117] The system for collecting metadata for the entire relational database provided in this embodiment, compared with the prior art, reads the schemas parameter; determines whether the read schemas parameter is a set symbol; if the read schemas parameter is a set symbol, obtains the entire schema list from the database; if the read schemas parameter is not a set symbol but a comma-separated schemas parameter, the schemas parameter is cut by comma to obtain the schema list. The system for collecting metadata for the entire relational database provided in this embodiment describes the need for collecting multiple databases from a relational database by filling in schemas in the data source, and solves the problem of significant performance degradation during multi-database collection, thereby making maintenance easier and avoiding the need for operation and maintenance personnel to establish multiple data sources to achieve the desired effect.
[0118] Preferably, see Figure 8 , Figure 8 for Figure 6 The functional module diagram of the second acquisition module 1 embodiment shown in the figure, in this embodiment, the second acquisition module 20 includes an acquisition unit 21 and a test unit 22, wherein the acquisition unit 21 is used to obtain the SQL for switching the schema from the dialect object according to the database type, and use this SQL as the initialization SQL for connecting the schema list; the test unit 22 is used to create a connection pool object according to the schema name, test whether the SQL and the schema list are connected successfully, and return the collector object if the connection is successful.
[0119] After obtaining the schema list, the acquisition unit 21 starts to initialize the collector object. First, according to the database type, the SQL for switching the schema is obtained from the dialect object, and this SQL is used as the initialization SQL for the connection.
[0120] The test unit 22 creates a connection pool object according to the schema name, and tests the connection acquisition after the creation. If the acquisition is successful, the collector object is returned. The collector object includes the connection pool and the metadata cache object.
[0121] The system for collecting metadata for the entire relational database provided in this embodiment, compared with the prior art, obtains the SQL for switching schemas from the dialect object according to the database type, and uses this SQL as the initialization SQL for connecting to the schema list; creates a connection pool object according to the schema name, tests whether the SQL and the schema list are successfully connected, and returns the collector object if the connection is successful. The system for collecting metadata for the entire relational database provided in this embodiment integrates the connection pool into the collector, obtains the SQL for switching schemas using the dialect object, and avoids the situation where a single JDBC connection can only connect to the same schemas; and the metadata after collection is relatively centralized, with low maintenance costs and high efficiency.
[0122] Further, see Fig. 9 , Fig. 9 for Figure 6 The functional module diagram of the processing module 1 embodiment shown in the figure, in this embodiment, the processing module 30 includes a traversal unit 31, a first cache unit 32, a calculation unit 33 and a merging unit 34, wherein the traversal unit 31 is used to traverse the collector objects in sequence according to the preset metamodel combination relationship until all required metamodels are obtained; the first cache unit 32 is used to traverse the metamodel and call all the previous metadata collectors to collect metadata, and cache all the collected metadata information into the Map object of the metadata collector; the calculation unit 33 is used to calculate the combination path according to the combination relationship of the metamodel when caching the metadata information; the merging unit 34 is used to merge the Map objects of each metadata collector after all metamodels are collected to integrate the metadata information.
[0123] After the traversal unit 31 obtains the available collection objects, it traverses the collector objects in sequence according to the combination relationship of the meta-models until all the required meta-models are obtained.
[0124] The first cache unit 32 traverses the meta-model and calls all previous metadata collectors to collect data, and caches all metadata information into the Map object of the collector.
[0125] When caching metadata information, the calculation unit 33 calculates a combination path according to the combination relationship of the meta-model.
[0126] After collecting all meta-models, the merging unit 34 will merge the Map objects of each collector and integrate the metadata information.
[0127] Compared with the prior art, the system for collecting metadata for the entire relational database provided by this embodiment traverses the collector objects in sequence according to the preset metamodel combination relationship until all required metamodels are obtained; traverses the metamodels and calls all previous metadata collectors to collect metadata, and caches all collected metadata information in the Map object of the metadata collector; when caching metadata information, calculates the combination path according to the combination relationship of the metamodel; after collecting all metamodels, merges the Map objects of each metadata collector to integrate the metadata information. The system for collecting metadata for the entire relational database provided by this embodiment has relatively concentrated metadata after collection, low maintenance cost and high efficiency.
[0128] Preferably, see Fig.10 , Fig.10 for Figure 6 The functional module diagram of an embodiment of the establishment module 1 shown in the figure, in this embodiment, the establishment module 40 includes a second judgment unit 41 and a second cache unit 42, wherein the second judgment unit 41 is used to judge whether the collected metadata and the metamodel are metadata dependencies according to the calculated combined path when collecting the metamodel dependency of the metadata, and exclude the erroneous relationship of metadata with the same name; the second cache unit 42 is used to cache the fields using a cache mechanism, store the combined path of the fields as the key value, and the metadata information of the fields as the value.
[0129] When collecting relationships, the second judgment unit 41 will judge whether it is a dependency relationship according to the combination path, thereby eliminating erroneous relationships with metadata of the same name.
[0130] In order to integrate multiple libraries, the second cache unit 42 caches the fields, taking into account that fields are usually the largest in metadata information. The cache mechanism is to store the combined path of the fields as the key value and the metadata information of the fields as the value, so as to avoid the situation where the multi-library collection is greatly reduced due to the use of loop nesting when collecting field relationships.
[0131] Compared with the prior art, the system for collecting metadata for the entire relational database provided in this embodiment can exclude erroneous relationships of metadata with the same name by judging whether the collected metadata and the metamodel are metadata dependencies based on the calculated combined path when collecting metadata's metamodel dependencies; a caching mechanism is used to cache fields, the combined path of the field is stored as the key value, and the metadata information of the field is used as the value. The system for collecting metadata for the entire relational database provided in this embodiment can integrate the connection pool into the collector, use the dialect object to obtain the SQL for switching the schema, and avoid the situation where a single JDBC connection can only connect to the same schemas; the collected metadata is relatively centralized, with low maintenance costs and high efficiency.
[0132] Although preferred embodiments of the present invention have been described, additional changes and modifications may be made to these embodiments by those skilled in the art once the basic inventive concepts are known. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention. Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A method for collecting metadata of a relational database, characterized in that: The following steps are involved: Preprocess the data source parameters and obtain the schema list; Initialize the collector object according to the obtained schema list and obtain the available collector objects; According to the preset meta-model combination relationship, metadata information is collected in sequence, all required meta-models are obtained, metadata information is cached, and metadata cache is merged; According to the preset metamodel dependencies, metadata dependencies are established in the cache; Save all metadata and collect statistics based on the established metadata dependencies; The step of preprocessing the data source parameters and obtaining the schema list includes: Read schemas parameters; Determine whether the schemas parameter read is the set symbol; If the schema parameter read is a set symbol, the entire schema list is obtained from the database; if the schema parameter read is not a set symbol but a comma-separated schema parameter, the schema parameter is split by commas to obtain the schema list; The step of initializing the collector object according to the acquired schema list and acquiring the available collector object comprises: According to the database type, obtain the SQL for switching schemas from the dialect object and use this SQL as the initialization SQL for the connection schema list; Create a connection pool object according to the schema name, test whether the SQL and schema list are connected successfully, and return the collector object if the connection is successful; The steps of sequentially collecting metadata information, acquiring all required meta-models, caching metadata information and merging metadata caches according to the preset meta-model combination relationship include: According to the preset metamodel combination relationship, the collector objects are traversed in sequence until all required metamodels are obtained; Traverse the metamodel and call all previous metadata collectors to collect metadata, and cache all collected metadata information in the Map object of the metadata collector; When caching metadata information, the combination path is calculated based on the combination relationship of the metamodel; After collecting all meta-models, merge the Map objects of each metadata collector to integrate the metadata information; The step of establishing metadata dependencies in the cache according to the preset metamodel dependencies includes: When collecting metadata's metamodel dependency, determine whether the collected metadata and the metamodel are metadata dependencies based on the calculated combined path, and exclude erroneous relationships between metadata with the same name; A caching mechanism is used to cache fields, storing the combined path of the field as the key value and the metadata information of the field as the value.
2. A system for collecting metadata of the entire relational database, characterized in that: include: A first acquisition module (10) is used to pre-process the data source parameters and obtain a schema list; A second acquisition module (20) is used to initialize a collector object according to the acquired schema list and acquire an available collector object; A processing module (30) is used to collect metadata information in sequence according to a preset meta-model combination relationship, obtain all required meta-models, cache metadata information and merge metadata caches; An establishment module (40) is used to establish metadata dependencies in the cache according to a preset metamodel dependency; A statistics module (50), used to save all metadata according to the established metadata dependency and collect statistics information; The first acquisition module (10) comprises: A reading unit (11), used for reading schema parameters; A first judgment unit (12), used for judging whether the read schemas parameter is a set symbol; The execution unit (13) is used for obtaining the entire schema list from the database if the schema parameter read is a set symbol; if the schema parameter read is not a set symbol but a comma-separated schema parameter, splitting the schema parameter according to the comma to obtain the schema list; The second acquisition module (20) comprises: An acquisition unit (21) is used to acquire the SQL for switching the schema from the dialect object according to the database type, and use the SQL as the initialization SQL for the connection schema list; The test unit (22) is used to create a connection pool object according to the schema name, test whether the SQL and the schema list are connected successfully, and return the collector object if the connection is successful; The processing module (30) comprises: A traversal unit (31), used for traversing the collector objects in sequence according to the preset metamodel combination relationship until all required metamodels are obtained; A first cache unit (32) is used to traverse the meta-model and call all previous metadata collectors to collect metadata, and cache all collected metadata information in a Map object of the metadata collector; A calculation unit (33), used to calculate a combination path according to the combination relationship of the meta-model when caching metadata information; A merging unit (34), used to merge the Map objects of each metadata collector after collecting all meta-models, and integrate the metadata information; The establishment module (40) comprises: A second judgment unit (41) is used to judge whether the collected metadata and the metamodel are metadata dependencies according to the calculated combination path when collecting metadata dependencies, and to exclude erroneous relationships of metadata with the same name; The second cache unit (42) is used to cache the field using a cache mechanism, storing the combined path of the field as a key value and the metadata information of the field as a value.
Citation Information
Patent Citations
Metadata collection method and device
CN110377568A
Method and apparatus for historical analysis analytics
US9336259B1