Database management method, data lineage analysis method, and related devices
By analyzing the tasks to be executed generated by the structured query language, extracting and synchronizing the table structure information, the problem that the table details management module cannot be updated in time is solved, and the database performance is improved.
Patent Information
- Application Number
- CN202011200865.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-02
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-11-02
AI Technical Summary
The existing database table details management module cannot obtain the table structure information created through the structured query language in a timely manner, resulting in frequent query of the database and affecting performance.
By analyzing the tasks to be executed generated by the structured query language, the table structure information in the data structure information is extracted, and the table details management module is synchronized when the table structure changes.
The table details management module is implemented to update the table structure information in a timely manner, reducing the number of frequent query of the database and improving the performance of the database.
Smart Images

Figure CN112015722B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular to a database management method, a data lineage analysis method, and related devices. Background Art
[0002] In the existing table details management module in a database, it is impossible to obtain the table structure information created by Structured Query Language (SQL). Therefore, when the table structure information created by SQL is needed, the database has to be queried to obtain the table structure information created by SQL, and frequent database queries will affect the performance of aspects such as database connection and query. Summary of the Invention
[0003] The main technical problem to be solved by this application is to provide a database management method, a data lineage analysis method, and related devices, which can synchronize the changed table structure information to the table details management module to avoid the problem of frequent database queries caused by the table details management module not updating the table structure information in time, and improve the performance of the database.
[0004] To solve the above technical problem, in a first aspect of this application, a database management method is provided. The method includes: generating a to-be-executed task of the database, where the to-be-executed task is generated by Structured Query Language SQL; parsing the Structured Query Language corresponding to the to-be-executed task to obtain the data structure information corresponding to the Structured Query Language, where the data structure information includes table structure information; in response to the table structure of the database being changed by the to-be-executed task, obtaining the table structure information changed by the to-be-executed task, and using the changed table structure information to synchronously update the table details management module, where the table details management module stores the table structure information corresponding to the database before the update.
[0005] Among them, the to-be-executed task includes at least one of a create table task, a delete table task, a modify table task, and a lineage generation task. Before the step of generating the to-be-executed task of the database, it further includes: creating a corresponding statement type for each type of task, and the statement type corresponding to each type of task can be parsed into the same structural form.
[0006] Before the step of parsing the Structured Query Language corresponding to the to-be-executed task, it includes: performing a syntax check on the Structured Query Language corresponding to the to-be-executed task; in response to the Structured Query Language corresponding to the to-be-executed task not conforming to the preset syntax rules, identifying and prompting the position of the Structured Query Language that does not conform to the preset syntax rules.
[0007] Among them, the step of parsing the structured query language corresponding to the to-be-executed task to obtain the data structure information corresponding to the structured query language includes: obtaining the structured query language corresponding to the to-be-executed task; parsing the structured query language, and extracting the data structure information corresponding to the structured query language in the structural form.
[0008] Among them, the step of synchronously updating the table details management module by using the changed table structure information includes: perceiving the statement type; in response to perceiving the statement type corresponding to the create table task / the statement type corresponding to the delete table task / the statement type corresponding to the modify table task, obtaining the table structure information in the data structure information and synchronizing it to the table details management module.
[0009] To solve the above technical problems, a second aspect of the present application provides a data lineage analysis method, which includes: generating a to-be-executed task for a database, where the to-be-executed task is generated by a structured query language SQL; parsing the structured query language corresponding to the to-be-executed task to obtain the data structure information corresponding to the structured query language, where the data structure information includes table structure information; in response to the table structure of the database being changed by the to-be-executed task, obtaining the changed table structure information, and synchronously updating the table details management module by using the changed table structure information, where the table details management module stores the table structure information corresponding to the database before the update; in response to the lineage information of the database being changed by the to-be-executed task, obtaining the changed lineage information, and updating the entire lineage information of the database by using the changed lineage information and the updated table details management module.
[0010] Among them, the to-be-executed task includes at least one of a create table task, a delete table task, a modify table task, and a lineage generation task. Before the step of generating the to-be-executed task for the database, it further includes: creating a corresponding statement type for each type of task, and the statement type corresponding to each type of task can be parsed into the same structural form.
[0011] Among them, updating the entire lineage information of the database by using the changed lineage information and the updated table details management module includes: perceiving the statement type; in response to perceiving the statement type corresponding to the lineage generation task, obtaining the data structure information corresponding to the lineage generation task; using the table structure information in the data structure information corresponding to the lineage generation task to generate the lineage relationship between the source table and the destination table corresponding to the lineage generation task; obtaining the updated table details management module and combining the lineage relationship generated by the lineage generation task to generate the complete lineage relationship graph corresponding to the current database; storing the complete lineage relationship graph corresponding to the current database in the graph database.
[0012] To solve the above technical problems, a third aspect of the present application provides an electronic device, which includes a memory and a processor coupled to each other. Among them, the memory stores program data, and the processor calls the program data to execute the database management method of the first aspect or the data lineage analysis method of the second aspect described above.
[0013] To solve the above technical problems, a fourth aspect of the present application provides a computer storage medium, on which program data is stored, and when the program data is executed by a processor, it implements the database management method of the first aspect or the data lineage analysis method of the second aspect described above.
[0014] The beneficial effect of the present application is that the present application parses the structured query language corresponding to the task to be executed, extracts the data structure information corresponding to the structured query language, and the data structure information includes table structure information. After obtaining the changed table structure information caused by the task to be executed, the changed table structure information is used to synchronously update the table details management module, so that after the table structure of the database is changed by the task to be executed generated by the structured query language, the table details management module can update the table structure information in time, avoiding the problem of frequent database queries caused by the table details management module not updating the table structure information in time, and improving the performance of the database. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. Among them:
[0016] Figure 1 is a schematic flowchart of an embodiment of the database management method provided by the present application;
[0017] Figure 2It is a schematic flowchart of another embodiment of the database management method provided by this application;
[0018] Figure 3 It is a schematic structural diagram of an embodiment corresponding to the table structure information obtained after parsing the table creation task;
[0019] Figure 4 It is a schematic flowchart of an embodiment of the data lineage analysis method provided by this application;
[0020] Figure 5 It is a schematic structural diagram of an embodiment corresponding to the table structure information obtained after parsing the lineage generation task;
[0021] Figure 6 It is Figure 4 a schematic flowchart of an embodiment corresponding to step S404 in
[0022] Figure 7 It is a schematic structural diagram of an embodiment of the electronic device provided by this application;
[0023] Figure 8 It is a schematic structural diagram of an embodiment of the computer storage medium provided by this application. Detailed Embodiments
[0024] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts shall fall within the protection scope of this application.
[0025] The terms "system" and "network" are often used interchangeably in this article. The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after. In addition, "multiple" in this article means two or more than two.
[0026] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the database management method provided by this application, and the method includes:
[0027] Step S101: Generate a to-be-executed task for the database, where the to-be-executed task is generated by Structured Query Language.
[0028] Specifically, the Structured Query Language (SQL) is a database language with various functions such as data manipulation and data definition. Among them, the task to be executed is generated by the SQL and submitted from the client side to the server side.
[0029] Furthermore, the SQL can specifically be Hive-SQL. Among them, Hive is an application tool based on a data warehouse and is used to process structured data in Hadoop. It is architected on top of Hadoop and operates on data through SQL. To achieve the technical objectives of this application, corresponding improvements have been made to the SQL in this application.
[0030] Specifically, this application reconstructs the grammar file based on the native Antlr3 grammar file of Hive and modifies it into a new version of the Antlr4 grammar file. Antlr4 is used for lexical and syntactic analysis, and its full English name is Another Tool for Language Recognition.
[0031] Furthermore, this application defines a data structure corresponding to the Hive grammar rules, is compatible with the native Hive SQL grammar, and generates the Antlr4 grammar file. The reconstructed Antlr4 grammar file is fully compatible with the native Hive grammar rules, and compared with the native parsing module of Hive, it introduces the visitor and listener patterns, separating the parsing from the application code.
[0032] Step S102: Parse the structured query language corresponding to the task to be executed to obtain the data structure information corresponding to the structured query language, where the data structure information includes table structure information.
[0033] Specifically, the parsing module is used to parse the structured query language corresponding to the task to be executed, and then obtain the data structure information corresponding to the structured query language.
[0034] In an application mode, the client uses the Antlr4 grammar file to perform lexical and syntactic analysis on the structured query language, obtains the source code of the structured query language, sequentially reads the character stream of the source code, identifies the lexemes in the character stream, and after obtaining the lexemes, maps the lexemes to tokens. One token corresponds to a type in the grammar of the structured query language. Further, perform syntactic analysis on all tokens, combine all tokens to generate an abstract syntax tree, and traverse the nodes of the abstract syntax tree to obtain the data structure information corresponding to the structured query language. Among them, the data structure information includes table structure information, such as table names, field names, and field contents, etc.
[0035] Step S103: In response to the table structure of the database being changed by the to-be-executed task, obtain the table structure information changed by the to-be-executed task, and use the changed table structure information to synchronously update the table details management module, where the table details management module stores the table structure information corresponding to the database before the update.
[0036] Specifically, the table details management module stores the table structure information of the database before the to-be-executed task is executed. When the table structure of the database is changed by the to-be-executed task, obtain the table structure information changed by the to-be-executed task parsed in step S102 above, and synchronize the changed table structure information to the table details management module, so that the table details management module can obtain the table structure information of the entire database and update it in real time. Furthermore, when the table structure information created by the structured query language needs to be used, the table structure information of the entire database can be obtained from the table details management module, without frequently querying the entire database, improving the performance in aspects such as database connection and query.
[0037] In this embodiment, the structured query language corresponding to the to-be-executed task is parsed to extract the data structure information corresponding to the structured query language, and the data structure information includes the table structure information. After obtaining the table structure information changed by the to-be-executed task, use the changed table structure information to synchronously update the table details management module, so that after the to-be-executed task generated by the structured query language changes the table structure of the database, the table details management module can update the table structure information in time, avoiding the problem of frequently querying the database due to the table details management module not updating the table structure information in time, and improving the performance of the database.
[0038] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of another implementation manner of the database management method provided by this application. The method includes:
[0039] Step S201: Create a corresponding statement type for each type of task.
[0040] Specifically, the to-be-executed tasks include at least one of a create table task, a delete table task, and a modify table task. Each type of task corresponds to a task purpose. Create respective statement types for tasks with different purposes, and distinguish tasks with different purposes through different statement types. Finally, the statement types corresponding to each type of task can be parsed into the same structural form and temporarily stored in the storage space provided in the statement type.
[0041] Specifically, among the above different types of to-be-executed tasks, the create table task, the delete table task, and the modify table task are tasks that will affect the table structure. Moreover, since corresponding statement types are set for different types of tasks, the task type can be known by perceiving the statement type, and thus what impact the task has on the database can be known according to the task type.
[0042] Step S202: Generate the tasks to be executed for the database, where the tasks to be executed are generated by Structured Query Language (SQL).
[0043] Specifically, the tasks to be executed for the database are generated by inputting code that conforms to the syntax rules of Structured Query Language. The tasks to be executed can include one or more types of tasks, and multiple tasks of the same type can be set. After the tasks to be executed are generated, they are uploaded by the client to the server side.
[0044] In a specific application scenario, the code of Structured Query Language corresponding to a create table task is: createtable table1 (column1 int, column2 string); where table1 is the table name, and column1 and column2 are the fields in table1.
[0045] Furthermore, before parsing the Structured Query Language corresponding to the tasks to be executed, it also includes: performing syntax verification on the Structured Query Language corresponding to the tasks to be executed; in response to the Structured Query Language corresponding to the tasks to be executed not conforming to the preset syntax rules, identifying and prompting the position of the Structured Query Language that does not conform to the preset syntax rules.
[0046] Specifically, this application defines a data structure corresponding to the Hive syntax rules, which fully compatible with the native Hive syntax rules. Therefore, when the input code does not conform to the native Hive syntax rules, an error will be reported when parsing the Structured Query Language. When the parsing fails, the position where the syntax error occurs can be identified according to the prompts in the syntax file, assisting developers in locating and modifying syntax problems, thereby reducing the probability of database crashes caused by executing code that does not conform to the syntax rules and improving the security of the database.
[0047] Step S203: Parse the Structured Query Language corresponding to the tasks to be executed to obtain the data structure information corresponding to the Structured Query Language, where the data structure information includes table structure information.
[0048] Specifically, the client obtains the Structured Query Language corresponding to the tasks to be executed through a parsing module, and then parses the Structured Query Language to extract the data structure information corresponding to the Structured Query Language in the above structural form.
[0049] Specifically, different types of tasks correspond to their respective statement types. The client perceives the statement type through the keywords of the statement type, and analyzes the Structured Query Statements corresponding to different types of tasks respectively to extract the corresponding data structure information.
[0050] In a specific application scenario, the code of the Structured Query Language corresponding to the create table task is: createtable table1 (column1 int, column2 string). Parsing the code of the Structured Query Language corresponding to the create table task, the data structure information arranged in the structural form as shown in Figure 3 can be finally obtained. Among them, CreateTableStatement is the statement type corresponding to the parsing result of the create table task. tableInfo in CreateTableStatement is the table information, columnInfos is the field set, and selectStatement is the source table query statement. SelectStatement is the source table query statement. columnInfos in SelectStatement is the source table field set, fromTableInfo is the source table set, and the source table may be a multi-table association. TableInfo is the table information, and TableInfo contains the database name databaseName and the table name tableName. ColumnInfo is the field information, and ColumnInfo contains the table name tableName and the field name columnName. In addition, for a specific statement that specifies the source table query statement, such as createtable... as..., the value of the source table query statement is not empty.
[0051] It can be understood that after parsing the delete table task and the modify table task, the parsing results with the statement types of DeleteTableStatement and AlterTableStatement can be obtained respectively. Among them, the parsing results will also include information related to the table structure such as the table name and the field name. Therefore, after being parsed by the parsing module, the Structured Query Statements corresponding to each type of task can be parsed into data entities in a fixed structural form, and the data entities contain the complete information of the statements.
[0052] Further, before step S204, when the server successfully executes the task to be executed, it feeds back the successful execution result to the client and sends the number of data records affected by the task to be executed. The client receives the result fed back by the server and the number of data records affected by the task to be executed. Only when the execution result is successful does it enter step S204. When the execution result is failed, the process ends to avoid information mismatch between the client and the server.
[0053] Step S204: In response to the table structure of the database being changed by the task to be executed, obtain the table structure information changed by the task to be executed, and use the changed table structure information to synchronously update the table details management module, where the table details management module stores the table structure information corresponding to the database before the update.
[0054] Specifically, the client searches for and perceives the statement type corresponding to the parsing result to know whether the table structure of the database has been changed by the structured query language. In response to perceiving the statement type corresponding to the create table task / the statement type corresponding to the delete table task / the statement type corresponding to the modify table task, the table structure information in the data structure information is obtained and synchronized to the table details management module.
[0055] Specifically, if the statement type corresponding to the create table task is perceived, the database name, table name, and field name information in the data structure information corresponding to the create table task are added to the table details management module; and / or, if the statement type corresponding to the delete table task is perceived, the table name in the data structure information corresponding to the delete table task is deleted from the table details management module; and / or, if the statement type corresponding to the modify table task is perceived, the renamed table name, renamed fields, and field details in the data structure information corresponding to the modify table task are updated to the table details management module.
[0056] In a specific application scenario, if it is the statement type corresponding to the create table task, the table structure information in the CreateTableStatement is obtained, and the database name, table name, and field name information are synchronized and modified to the table details management module to implement the synchronization of the create table information; if it is the statement type corresponding to the delete table task, the table structure information in the DeleteTableStatement is obtained, and the table name is synchronized to the table details management module for deletion operation to implement the synchronization of the delete table information; if it is the statement type corresponding to the modify table task, the table structure information in the AlterTableStatement is obtained, the table name is renamed, the fields are renamed, and the details of the added and deleted fields are synchronized and modified to the table details management module to implement the synchronization of the modify table structure information.
[0057] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of an implementation manner of the data lineage analysis method provided by this application. The method includes:
[0058] Step S401: Generate a to-be-executed task for the database, where the to-be-executed task is generated by the structured query language.
[0059] Step S402: Parse the structured query language corresponding to the to-be-executed task to obtain the data structure information corresponding to the structured query language, where the data structure information includes table structure information.
[0060] Step S403: In response to the table structure of the database being changed by the to-be-executed task, obtain the table structure information changed by the to-be-executed task, and use the changed table structure information to synchronously update the table details management module, where the table details management module stores the table structure information corresponding to the database before the update.
[0061] Specifically, the above steps S401 - S403 are similar to the above embodiments, and for details, reference can be made to any of the above embodiments, which will not be elaborated here.
[0062] Further, the to-be-executed task includes at least one of a table creation task, a table deletion task, a table modification task, and a lineage generation task. Before the step of generating the to-be-executed task of the database, it further includes: creating a corresponding statement type for each type of task, and the statement type corresponding to each type of task can be parsed into the same structural form.
[0063] Specifically, each task corresponds to a task purpose. Create respective statement types for tasks with different purposes, and distinguish tasks with different purposes through different statement types. Eventually, the statement type corresponding to each type of task can be parsed into the same structural form and temporarily stored in the storage space provided in the statement type. Among them, the lineage generation task will have an impact on the lineage relationship. Moreover, since corresponding statement types are set for different types of tasks, the task type can be known by perceiving the statement type, and thus what impact the task has on the database can be known according to the task type.
[0064] In a specific application scenario, the code of the structured query language corresponding to the lineage generation task is: insert into dest1(column3, column4) select column1, column2 from src1. Where dest1 is the destination table name and src1 is the source table name. Through the insert into statement, data can be directly appended to the table or a partition of the table. Parse the code of the structured query language corresponding to the lineage generation task, and finally, data structure information arranged in the structure form as shown can be obtained. Figure 5 Among them, InsertStatement is the statement type corresponding to the parsing result of the lineage generation task. The overwrite in InsertStatement indicates whether it contains the overwrite keyword. The insert overwrite statement is different from the insert into statement. The insert overwrite statement will first clear the original data in the table and then insert data into the table or a partition of the table. Therefore, it is necessary to judge whether it contains the overwrite keyword during parsing.
[0065] Further, the tableInfo in the InsertStatement is the destination table information, the columnInfos is the set of destination fields, and the selectStatement is the source table query statement. The SelectStatement is the source table query statement, the columnInfos in the SelectStatemen is the set of source fields, the fromTableInfo is the set of source tables, and the source tables may be multiple related tables. The TableInfo is the table information, which includes the databaseName database name and the tableName table name. The ColumnInfo is the field information, which includes the tableName table name and the columnName field name. In addition, for a specific statement that specifies the source table query statement, such as create table…as…, the value of the source table query statement is not empty.
[0066] Step S404: In response to the lineage information of the database being changed by the task to be executed, obtain the lineage information changed by the task to be executed, and use the changed lineage information and the updated table details management module to update the entire lineage information of the database.
[0067] Specifically, when the task to be executed includes a lineage generation task, obtain the source table name, source field name, destination table name, and destination field name in the data structure information corresponding to the lineage generation task, generate table-level lineage relationships according to the order of the source table name and the destination table name, and generate field-level lineage relationships according to the order of the source field name and the destination field name.
[0068] Further, when the table structure of the database changes due to a create table task / delete table task / modify table task, in addition to the lineage relationships corresponding to the lineage generation task, combine the number of data records affected by the task to be executed feedback by the server side, and the data structure information corresponding to the structured query language obtained in step S402, to update the lineage relationships of the entire database. The lineage relationships can help sort out the mapping relationships of the tables and fields in the database, so as to track the data flow in a large amount of data and make it easier to manage the data.
[0069] In an application mode, please refer to Figure 6 , Figure 6 is Figure 4 a schematic flowchart of an implementation corresponding to step S404 in
[0070] Step S601: Sense the statement type.
[0071] Specifically, the client searches for and senses the statement type corresponding to the parsing result to know whether there is a lineage generation task that affects the tables in the database.
[0072] Step S602: In response to sensing the statement type corresponding to the blood relationship generation task, obtain the data structure information corresponding to the blood relationship generation task.
[0073] Specifically, if the statement type corresponding to the blood relationship generation task is sensed, then extract the source table name, source table field name, destination table name, and destination table field name from the data structure information corresponding to the blood relationship generation task.
[0074] In a specific application scenario, when the blood relationship generation task is sensed, the source table name, source table field name, destination table name, and destination table field name in the InsertStatement are extracted.
[0075] Step S603: Use the table structure information in the data structure information corresponding to the blood relationship generation task to generate the blood relationship between the source table and the destination table corresponding to the blood relationship generation task.
[0076] Specifically, generate the blood relationship between tables according to the order of the source table name and the destination table name, and generate the blood relationship between fields according to the order of the source table field name and the destination table field name.
[0077] In a specific application scenario, for the code: insert into dest1(column3, column4) select column1, column2 from src1. Establish a blood relationship between the table name of dest1 and the table name of src1, and establish a blood relationship between the field names of dest1 and src1. Among them, in the blood relationship of fields, column3 corresponds to column1, and column4 corresponds to column2.
[0078] In another specific application scenario, for the code: insert into dest1 select column1,column2 from src1, the destination table fields are not specified. Define the unknown fields unknown1 and unknown2 of the destination table, establish a blood relationship between the table name of dest1 and the table name of src1, and establish a blood relationship between the field names of dest1 and src1. Among them, in the blood relationship of fields, unknown1 corresponds to column1, and unknown2 corresponds to column2.
[0079] Step S604: Obtain the updated table details management module and combine the blood relationship generated by the blood relationship generation task to generate the complete blood relationship diagram corresponding to the current database.
[0080] Specifically, the table structure of the database has been updated in a timely manner in the table details management module. Combining the table structure information in the updated table details management module and the blood relationship generated by the blood relationship generation task, a complete blood relationship diagram corresponding to the entire current database is generated.
[0081] Furthermore, when updating the blood relationship diagram of the entire database, the table names containing unknown fields in the blood relationship are identified. The client only needs to obtain the table structure information from the table details management module, rather than querying the database when it comes to the table structure information created by the structured query language, reducing the number of times the database is accessed and queried.
[0082] Step S605: Store the complete blood relationship diagram corresponding to the current database in the graph database.
[0083] Specifically, the graph database is the Neo4j graph database. The complete blood relationship diagram corresponding to the current database is stored in the graph database so that the blood relationship diagram can be updated in real time and the visualization of the blood relationship diagram can be realized, making the blood relationship diagram more intuitive and easier to use.
[0084] The data blood relationship analysis method provided in this embodiment creates corresponding statement types for different types of tasks, and updates the results after parsing different types of tasks to the table details management module in real time. When it is necessary to generate the blood relationship, the accurate table structure information in the current database can be obtained by querying the table details management module, so as to reduce the number of times of querying the database, and update the complete blood relationship diagram of the current database in real time to improve the efficiency of data management.
[0085] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an embodiment of an electronic device provided by the present application. The electronic device 70 includes a memory 701 and a processor 702 that are coupled to each other. Among them, the memory 701 stores program data (not shown in the figure), and the processor 702 calls the program data to implement the database management method or the data blood relationship analysis method in any of the above embodiments. For the description of related content, please refer to the detailed description of the above method embodiments, and details will not be repeated here.
[0086] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of an embodiment of a computer storage medium provided by the present application. The computer storage medium 80 stores program data 800. When the program data 800 is executed by a processor, it implements the database management method or the data blood relationship analysis method in any of the above embodiments. For the description of related content, please refer to the detailed description of the above method embodiments, and details will not be repeated here.
[0087] It should be noted that the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0088] In addition, the functional units in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0089] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., which can store program codes.
[0090] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A database management method, characterized in that, the method includes: Refactoring the native grammar file of Hive to obtain a refactored grammar file, the refactored grammar file is compatible with the native grammar rules of Hive, and introduces the visitor and listener patterns to separate parsing from application code; Creating a corresponding statement type for each type of task; wherein, the statement type is used to distinguish tasks with different purposes, and the statement types corresponding to each type of task can be parsed into the same structural form and temporarily stored in the storage space provided in the statement type; Generating a to-be-executed task for the database, wherein the to-be-executed task is generated by Structured Query Language (SQL), and the to-be-executed task includes at least one of a create table task, a delete table task, and an alter table task, and each type of task corresponds to a task purpose; Performing syntax verification on the Structured Query Language corresponding to the to-be-executed task; In response to the Structured Query Language corresponding to the to-be-executed task not conforming to the preset syntax rules, identifying and prompting the position of the Structured Query Language that does not conform to the preset syntax rules; Parsing the Structured Query Language corresponding to the to-be-executed task by means of the refactored grammar file to obtain the data structure information corresponding to the Structured Query Language, wherein the data structure information includes table structure information, and the table structure information is the table name and field name; In response to the table structure of the database being changed by the to-be-executed task, obtaining the parsed table structure information changed by the to-be-executed task, and using the changed table structure information to synchronously update the table details management module, wherein the changed table structure information is created by the Structured Query Language, and the table details management module stores the table structure information corresponding to the database before the update.
2. The method according to claim 1, characterized in that, the step of parsing the Structured Query Language corresponding to the to-be-executed task to obtain the data structure information corresponding to the Structured Query Language includes: Obtaining the Structured Query Language corresponding to the to-be-executed task; Parsing the Structured Query Language and extracting the data structure information corresponding to the Structured Query Language in the structural form.
3. The method according to claim 1, characterized in that, the step of using the changed table structure information to synchronously update the table details management module includes: Perceiving the statement type; In response to perceiving the statement type corresponding to the create table task / the statement type corresponding to the delete table task / the statement type corresponding to the alter table task, obtaining the table structure information in the data structure information and synchronizing it to the table details management module.
4. A data lineage analysis method, characterized in that, the method includes: Refactoring the native grammar file of Hive to obtain a refactored grammar file, the refactored grammar file is compatible with the native grammar rules of Hive, and introduces the visitor and listener patterns to separate parsing from application code; Create corresponding statement types for each type of task; wherein, the statement types are used to distinguish tasks for different purposes, and the statement types corresponding to each type of task can be parsed into the same structural form and temporarily stored in the storage space provided in the statement types; Generate tasks to be executed in the database, wherein the tasks to be executed are generated by Structured Query Language (SQL), and the tasks to be executed include at least one of a table creation task, a table deletion task, and a table modification task, and each type of task corresponds to a task purpose; Perform syntax checking on the Structured Query Language corresponding to the task to be executed; In response to the Structured Query Language corresponding to the task to be executed not conforming to the preset syntax rules, identify and prompt the position of the Structured Query Language that does not conform to the preset syntax rules; Parse the Structured Query Language corresponding to the task to be executed by means of the reconstructed syntax file to obtain the data structure information corresponding to the Structured Query Language, wherein the data structure information includes table structure information, and the table structure information is the table name and field name; In response to the table structure of the database being changed by the task to be executed, obtain the parsed table structure information changed by the task to be executed, and use the changed table structure information to synchronously update the table details management module, wherein the changed table structure information is created by the Structured Query Language, and the table details management module stores the table structure information corresponding to the database before the update; In response to the lineage information of the database being changed by the task to be executed, obtain the lineage information changed by the task to be executed, and use the changed lineage information and the updated table details management module to update the entire lineage information of the database.
5. The method according to claim 4, wherein, the updating the entire lineage information of the database by using the changed lineage information and the updated table details management module includes: Perceive the statement type; In response to perceiving the statement type corresponding to the lineage generation task, obtain the data structure information corresponding to the lineage generation task; Use the table structure information in the data structure information corresponding to the lineage generation task to generate the lineage relationship between the source table and the destination table corresponding to the lineage generation task; Obtain the updated table details management module and combine the lineage relationship generated by the lineage generation task to generate a complete lineage relationship graph corresponding to the current database; Store the complete lineage relationship graph corresponding to the current database in the graph database.
6. An electronic device, wherein, it includes: A memory and a processor that are mutually coupled, wherein the memory stores program data, and the processor calls the program data to execute the method according to any one of claims 1-3 or 4-5.
7. A computer storage medium, on which program data is stored, wherein, when the program data is executed by a processor, the method according to any one of claims 1-3 or 4-5 is implemented.
Citation Information
Patent Citations
Pedigree analysis method and device of data warehouse
CN104899314A
Method and device for determining data consanguinity based on structural data
CN109325078A
Method and device for visually monitoring Hive data warehouse
CN110532261A
Metadata-based data blood relationship analysis method and system
CN110555032A