Graph database mapping file creating method and device and medium
Through an automated method, the acquisition of graph Schema and graph data SQL and structured configuration tables are used to parse and generate graph database mapping files row by row, solving the problems of low manual creation efficiency and low accuracy in the existing technology, and achieving efficient and accurate mapping file creation.
Patent Information
- Application Number
- CN202510238159.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-17
AI Technical Summary
In the prior art, the creation of graph database mapping files depends on manpower, is inefficient and has low accuracy, and has high long-term manual maintenance costs.
Design a graph database mapping file creation method. By obtaining the acquisition SQL and structured configuration tables of the schema extracted by graph Schema, analyzing the configuration table row by row and generating graph database mapping files based on the template, realizing automatic creation.
It improves the efficiency and accuracy of the creation of graph database mapping files, reduces manpower intervention, and avoids the problem of error-prone comparison of graph Schema and graph data.
Smart Images

Figure CN120162328A_ABST
Abstract
Description
Technical Field
[0001] This application is at least related to the field of database technology, and particularly relates to a method, device and medium for creating a graph database mapping file. Background Art
[0002] When building a graph database, there are several very important factors: the graph model, the original data, and the graph mapping file (graph database mapping file). The graph mapping file is a file that associates the original data with the graph model. Currently, the creation of the graph mapping file mainly relies on manual labor. It is necessary to manually check the fields and formats of the graph model and the original data for a long time. The efficiency and accuracy of manual maintenance of the mapping file are low. Summary of the Invention
[0003] In view of the above deficiencies, this application provides a method, device and medium for creating a graph database mapping file to solve the following technical problems: how to automatically create a graph database mapping file and improve the creation efficiency and accuracy of this file.
[0004] In a first aspect, this application provides a method for creating a graph database mapping file, and the method includes:
[0005] Obtain a collection structured query language SQL for extracting quasi-graph data according to the graph Schema. The graph Schema is a graph schema used to build a graph database, and the quasi-graph data is intermediate data that processes the source data into the data structure required by the graph Schema;
[0006] Obtain a structured configuration table. In the structured configuration table, each collection SQL is on a separate line, indicating the output file name of the quasi-graph data for each line of the collection SQL, whether the collection object is an entity, an edge, or an entity + edge, the name of the collection object, and the source entity and target entity connected when the collection object is an edge;
[0007] Parse the structured configuration table line by line and generate a graph database mapping file based on a template. The graph database mapping file is a code file that maps the quasi-graph data into a graph model that conforms to the graph Schema pattern.
[0008] Further, obtaining a collection structured query language SQL for extracting quasi-graph data according to the graph Schema specifically includes:
[0009] Obtain each entity and each edge in the graph Schema;
[0010] Obtain the source database where the source data corresponding to each entity and each edge in the graph Schema is located;
[0011] For each entity and each edge in the design diagram Schema, there is a corresponding collection SQL. Each collection SQL is used to extract the source data corresponding to each entity or each edge in the diagram Schema from the source database and process it into quasi-diagram data that conforms to the data structure required by the diagram Schema.
[0012] Further, obtain the structured configuration table, specifically including:
[0013] Read the latest diagram Schema and the latest collection SQL from the cache;
[0014] If the latest collection SQL is not in the structured configuration table, add a new row to the structured configuration table to add the latest collection SQL, and obtain the quasi-diagram data output file name and, based on the latest diagram Schema, obtain whether the collection object is an entity or an edge or entity + edge, the collection object name, and the source entity and target entity connected when the collection object is an edge.
[0015] If a certain row of collection SQL in the structured configuration table is not in the latest collection SQL, set the enable flag of the certain row of collection SQL in the structured configuration table to closed.
[0016] Further, obtain the quasi-diagram data output file name, specifically including:
[0017] Obtain the quasi-diagram data output file name set manually, and execute each row of collection SQL in the structured configuration table with the enable flag set to enabled, and store the collected data in the file corresponding to the quasi-diagram data output file name of the corresponding row;
[0018] Or, if each row of collection SQL in the structured configuration table with the enable flag set to enabled has been executed and the collected data has been stored in the corresponding quasi-diagram data output file, obtain the file name of the quasi-diagram data output file corresponding to each row.
[0019] Further, parse the structured configuration table row by row and generate a graph database mapping file based on a template, specifically including:
[0020] Parse the collection SQL of the current row to obtain the identification ID and attributes of the entity and / or edge;
[0021] In response to the collection object of the current row including an entity, according to the entity object declaration specification of the graph database mapping file, assemble the identification ID and attributes of the entity corresponding to the current row, the quasi-diagram data output file name, and the collection object name to obtain the first assembled data;
[0022] In response to the fact that the acquisition object of the current row contains an edge, according to the edge object declaration specification of the graph database mapping file, assemble the identification ID and attributes of the edge corresponding to the current row, as well as the name of the quasi-graph data output file, the name of the acquisition object, and the source entity and target entity connected, so as to obtain the second assembled data;
[0023] Combine the first assembled data and the second assembled data according to the template to obtain the graph database mapping file, which is used to map the acquisition object data in the quasi-graph data into entities and / or edges that conform to the graph model of the graph Schema mode.
[0024] Further, parse the acquisition SQL of the current row to obtain the identification ID and attributes of the entity and / or edge, specifically including:
[0025] Parse the keyword fields that need to be obtained from the quasi-graph data contained in the acquisition SQL of the current row;
[0026] Set the first obtained keyword field as the identification ID of the entity and / or edge;
[0027] Set the non-first obtained keyword fields as the attributes of the entity and / or edge.
[0028] Further, combine the first assembled data and the second assembled data according to the template to obtain the graph database mapping file, specifically including:
[0029] Obtain the first assembled data and the second assembled data temporarily stored in the memory;
[0030] Fill the first assembled data and the second assembled data into the preset graph database mapping file template in the order in the structured configuration table to obtain the graph database mapping file.
[0031] Further, the method further includes:
[0032] Execute the code of the graph database mapping file to map the acquisition object data in the quasi-graph data into entities and / or edges that conform to the graph model of the graph Schema mode, including: in the graph model, use the acquisition object name as the classification of the entity and / or edge, mark the content of the quasi-graph data corresponding to the identification ID and attributes of the entity and / or edge, and establish connections between entities according to the source entity and target entity connected by the edge.
[0033] In a second aspect, the present application provides a graph database mapping file creation device, and the device includes:
[0034] A preparation module, configured to obtain an acquisition structured query language SQL for extracting quasi-graph data according to the graph Schema, where the graph Schema is a graph schema for establishing a graph database, and the quasi-graph data is intermediate data obtained by processing source data into a data structure required by the graph Schema;
[0035] A configuration module, connected to the preparation module, is used to obtain a structured configuration table. In the structured configuration table, each collection SQL is on a separate line, indicating the file name of the intended graph data output for each line of the collection SQL, whether the collection object is an entity, an edge, or an entity + edge, the name of the collection object, and the source entity and target entity connected when the collection object is an edge.
[0036] An analysis module, connected to the configuration module, is used to parse the structured configuration table line by line and generate a graph database mapping file based on a template. The graph database mapping file is a code file that maps the intended graph data to a graph model conforming to the graph Schema pattern.
[0037] Thirdly, the present application provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium, and when the computer program is run by a processor, the method for creating a graph database mapping file as described above is implemented.
[0038] The present application provides a method, apparatus, and medium for creating a graph database mapping file. By designing a structured configuration table, according to the characteristics of the graph Schema and the collection of intended graph data, the content of each line is set. The collection SQL for extracting the intended graph data according to the graph Schema is used as the main content of each line, and information such as the file name of the intended graph data output required for generating the graph database mapping file for each line configuration, whether the collection object is an entity, an edge, or an entity + edge, the name of the collection object, and the source entity and target entity connected when the collection object is an edge are set. To achieve the purpose of automatically generating a graph database mapping file by parsing the structured configuration table line by line, the creation efficiency of the graph database mapping file is improved. At the same time, due to the automated process, the problem of error-prone manual comparison between the graph Schema and the intended graph data is avoided, and the accuracy of the graph database mapping file is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is a flowchart of a method for creating a graph database mapping file according to an embodiment of the present application;
[0040] Figure 2 is a schematic structural diagram of an apparatus for creating a graph database mapping file according to an embodiment of the present application;
[0041] Figure 3 is a flowchart of another method for creating a graph database mapping file according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] To enable those skilled in the art to better understand the technical solutions of the present application, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0043] It will be understood that the specific embodiments and drawings described herein are for purposes of explaining the present application only and are not intended to limit the present application.
[0044] It will be understood that, without conflict, the various embodiments and features in the present application may be combined with each other.
[0045] It will be understood that, for ease of description, only the parts related to the present application are shown in the drawings of the present application, and the parts unrelated to the present application are not shown in the drawings.
[0046] It will be understood that each module and unit involved in the embodiments of the present application may correspond to only one entity structure, or may be composed of multiple entity structures, or multiple modules and units may also be integrated into one entity structure.
[0047] It will be understood that, without conflict, the functions and steps marked in the flowcharts and block diagrams of the present application may occur in a different order from that marked in the drawings.
[0048] It will be understood that in the flowcharts and block diagrams of the present application, the possible architectures, functions, and operations of the systems, devices, equipment, and methods according to the various embodiments of the present application are shown. Among them, each block in the flowchart or block diagram may represent a module, unit, program segment, or code, which contains executable instructions for implementing the specified function. Moreover, each block or combination of blocks in the block diagram and flowchart may be implemented by a hardware-based device for implementing the specified function, or may be implemented by a combination of hardware and computer instructions.
[0049] It will be understood that the modules and units involved in the embodiments of the present application may be implemented in software or in hardware. For example, the modules and units may be located in a processor.
[0050] Embodiment 1:
[0051] As Figure 1 shown, the present application provides a method for creating a graph database mapping file, and the method includes:
[0052] S1. Obtain a collection structured query language SQL for extracting pseudo-graph data according to a graph Schema. The graph Schema is a graph schema for establishing a graph database, and the pseudo-graph data is intermediate data obtained by processing source data into a data structure required by the graph Schema;
[0053] S2. Obtain a structured configuration table. In the structured configuration table, each collection SQL is a row, indicating the output file name of the pseudo-graph data for each row of collection SQL, whether the collection object is an entity, an edge, or an entity + edge, the name of the collection object, and the source-end entity and target-end entity connected when the collection object is an edge;
[0054] S3. Parse the structured configuration table line by line and generate a graph database mapping file based on a template. The graph database mapping file is a code file that maps the graph data to be drawn into a graph model conforming to the graph Schema pattern.
[0055] In this embodiment, the method designs a structured configuration table. According to the characteristics of the graph Schema and the collection of the graph data to be drawn, the content of each row is set. The collection SQL for extracting the graph data to be drawn according to the graph Schema is used as the main content of each row, and information such as the output file name of the graph data to be drawn required for generating the graph database mapping file for each row configuration, whether the collection object is an entity, an edge, or an entity + edge, the name of the collection object, and the source entity and target entity connected when the collection object is an edge are configured, so as to achieve the purpose of automatically generating a graph database mapping file by parsing the structured configuration table line by line, improve the creation efficiency of the graph database mapping file, and at the same time avoid the problem of error-prone manual comparison between the graph Schema and the graph data to be drawn due to the automated process, improving the accuracy of the graph database mapping file. As Figure 1 shown in the method is correspondingly applied to the device shown in Figure 2 It should be noted that the template can be a piece of code that controls the generation of a graph database mapping file according to a specified structure, or other methods can be used. Its main function is to automatically generate a graph database mapping file by some programmed methods. The graph database mapping file obtains the parsing result of the structured configuration table according to the specified structure, realizing the accurate connection between the graph Schema and the graph data to be drawn.
[0056] Specifically, this embodiment provides a method for automatically creating a graph database mapping file, and can provide corresponding devices, electronic devices, storage media, etc. The application scenarios of this embodiment include synchronously mapping the original data in the relational database into the graph model in the graph database to improve the query efficiency of relational data.
[0057] More specifically, in the development of modern big data application architectures, to improve query efficiency and enhance the query logic for "relationships", enterprises usually use graph databases for data storage and query optimization. Typically, enterprises adopt the lambda architecture (a big data architecture design pattern for processing large-scale data, aiming to combine the advantages of batch processing and stream processing to provide low-latency, high-fault-tolerance, and real-time data processing capabilities). Under this architecture, the core data of enterprises is usually stored in a structured database or data warehouse and is usually read and written in a structured manner. To improve the query efficiency of relational data, enterprises use a graph database as a slave database and collect and access business data into the graph database through data ETL (Extract-Transform-Load, which extracts, transforms, and loads data from the source end to the destination end).
[0058] The construction of a graph database is different from that of a relational database (where the original data usually comes from a relational database). When a graph database serves as a synchronous data source, several very important factors are required: a graph model, original data, and a graph mapping file. The graph model is the data structure, the original data comes from the core data of the enterprise, and the graph mapping file is the most crucial file that associates the original data with the graph model. This file records information including: original data fields, formats, information, storage addresses, mappings between fields and the graph model, abstractions of entities and edges, etc. The main function of the graph database mapping file is to define and describe how to map external data sources (such as CSV files, databases, etc.) to nodes (vertices) and edges (relationships) in the graph database. The graph database mapping file is the bridge between the external data source and the graph database. It not only simplifies the process of data import and integration but also improves the management efficiency and quality of data. At the same time, it is also crucial for understanding and maintaining the structure and relationships of the graph database.
[0059] The development of the mapping file is a task that requires a great deal of manpower to complete during the data collection and synchronization of the entire graph database. Its workload grows linearly with the complexity of the graph model and the number of fields in the original data. When developing the graph mapping file, it is necessary to continuously check the fields and formats of the graph model and the original data. When the graph model and the original data are modified, the mapping file also needs to be modified accordingly. From the perspective of both graph data construction and graph data iteration, long-term manual maintenance of the mapping file is inefficient, inaccurate, and costly.
[0060] Therefore, a method is needed that can automatically generate a mapping file based on the graph model and the original data under unchanging external conditions to improve the work efficiency of graph data construction and iteration, enhance accuracy, and reduce maintenance costs.
[0061] The general process for generating a graph database mapping file (the standardized process for building a knowledge graph) is as follows:
[0062] a) Data experts analyze the original data (i.e., the main production data. Generally, production data is stored in a general relational database and may come from multiple databases) to determine the data fields and query logic that need to be extracted from the source data (original data, data in the source database) into the graph database. Hereinafter, it will be simply referred to as the quasi-graph structured data;
[0063] b) Data experts need to prepare a graph model design (meeting the requirements of the graph database) for the current business needs, generate a graph model design file, hereinafter simply referred to as the graph Schema, and cache it (in the server or file system). When designing, experts usually need to consider the format of the source data, the storage efficiency and query efficiency of the graph library, and design the most suitable graph Schema;
[0064] c) Data experts will design the graph extraction logic to collect the content of the quasi-graph structured data, form quasi-graph data and cache it (in the server, or file system, or message queue, etc.). The collection statement is usually the collection SQL (Structured Query Language). It is necessary to refer to the design of the Schema, process the source data into the data structure required by the Schema, calculate and extract the data according to the situation of the graph Schema, and generate an intermediate data that conforms to the graph Schema, that is, quasi-graph data;
[0065] d) After the quasi-graph structured data and the graph Schema are both prepared, the most crucial step is to develop a graph mapping file, whose function is to associate the two. The steps are as follows:
[0066] d1) Analyze the entities, edges, and attributes (property graph) existing in the graph Schema. The purpose of analyzing the detailed design information of the graph Schema is to understand the entity, edge, and attribute information of the graph. The generation of the mapping file requires specifying these contents and requires corresponding field attributes;
[0067] d2) The characteristic of the quasi-graph data is that the most suitable data structure has been designed and collected according to the corresponding graph Schema situation. For example, for entity + attribute, a single SQL is used for collection, and for edge + attribute, a single SQL is used for collection. There are many types of entities and edges, which is usually determined by the graph Schema design;
[0068] d3) Generate N (depending on the graph Schema design) entity information (the number of nodes in the graph Schema) for the graph database mapping file, which includes entity name, entity ID, the location of the pseudo-graph data corresponding to the entity (pseudo-graph data is usually a file or a storage, generally at a certain path on the server), pseudo-graph data type, pseudo-graph data field headers, encoding type, etc. For example, if the pseudo-graph data is a table file, then the pseudo-graph data type is: csv, the pseudo-graph data field headers are the table headers, and the encoding type refers to the underlying encoding structure of the computer file. These several pieces of information are for the parsing and interpretation of the pseudo-graph data file, simply put, it is a declaration of the pseudo-graph data file;
[0069] d4) Generate N (depending on the graph Schema design) edge information for the graph database mapping file, which includes edge name, edge source ID, edge target ID, the location of the pseudo-graph data corresponding to the edge, pseudo-graph data type, pseudo-graph data field headers, encoding type, ID field mapping configuration, etc.;
[0070] e) Complete the development of the graph database mapping file according to the specifications of the corresponding graph database (the standard format of the data mapping file).
[0071] During the normal big data business R & D process, including in the design of the graph data model (graph Schema), there are very many types of entities and edges. Maybe a common graph Schema usually has 10 or more entities and 20 or more types of edges. If the above process is developed manually, the workload of developing the graph database mapping file grows linearly with the types of data and the complexity of the graph Schema. The types of data refer to points and edges, and these designs will gradually increase with business iteration and design. For example, originally there were only 5 entities and 6 edges, and due to changes in business requirements, there may be 7 entities and 7 edges. Then, the database mapping file to be developed needs to add 2 entities and 1 edge, and so on. In the future, there will be many, many changes and addition / deletion requirements. Whether it is for a new project construction or continuous graph data iteration, the construction and iteration of the graph data mapping file will require a large amount of time for manual development, review, and proofreading. Moreover, there cannot be any mistakes in the format and fields of this mapping file. Once there are some mistakes and there is a lack of effective error debugging tools in the industry, it will trigger a large amount of troubleshooting work, further reducing work efficiency and productivity.
[0072] To address this pain point, this embodiment adopts a strategic approach, comprehensively considering the tasks required in the overall development cycle. It can automatically generate the mapping file without manual intervention in the figure mapping file, completely eliminating human intervention. At the same time, the automatically generated file is the completely correct result file. Even in the case of later gallery reconstruction or iteration, it can automatically accommodate the logical changes caused by modification and iteration, thus achieving complete automation in the development of the figure mapping file. It is estimated that by using this solution to automatically generate the figure mapping file, the efficiency of figure construction and iteration can be increased by more than 5 times, significantly reducing the workload of data experts in comparing (figure Schema, draft figure data), R & D, and error troubleshooting.
[0073] The method described in this embodiment is specifically as Figure 3 shown. Its core includes: designing a structured configuration table (structured policy configuration table), and parsing the configuration table and generating the figure mapping file through a determined program. In the traditional method, the mapping file is handwritten by humans, while in this method, it is generated by a program algorithm, and the iteration methods are also different. In the traditional method, modifying the mapping file during iteration will lead to a large amount of inspection work and regression testing, and it is not easy to troubleshoot problems. In this method, it is regenerated by modifying the configuration. The cooperation of the structured configuration table + program algorithm can completely eliminate format problems, spelling problems, dirty character problems, etc., realizing the automatic generation and automatic calibration of the mapping file.
[0074] In one embodiment, obtaining the collection structured query language SQL for extracting draft figure data according to the figure Schema in S1 specifically includes:
[0075] Obtaining each entity and each edge in the figure Schema;
[0076] Obtaining the source database where the source data corresponding to each entity and each edge in the figure Schema is located;
[0077] Designing the collection SQL for each entity and each edge in the figure Schema. Each collection SQL is used to extract the source data corresponding to each entity or each edge in the figure Schema from the source database and process it into draft figure data that conforms to the data structure required by the figure Schema.
[0078] In this embodiment, as Figure 3 shown, the method first completes the preparation step, obtains the figure Schema and the draft figure structure data, and then obtains the entity information, edge information, and data collection SQL that need to be filled into the structured policy configuration table. The figure Schema is the benchmark mode of the figure model, and the data collection SQL is designed according to the draft figure structure data and is used to collect data (draft figure data) suitable for figure model construction from the source end.
[0079] The method is applicable to the automatic generation of graph database mapping files in the iterative scenario of large-scale graph databases. The preparatory work specifically includes:
[0080] 1) Determine the graph Schema: This operation is the basis for graph construction and iteration, which can be determined manually and is a design task;
[0081] 2) Prepare the graph data to be drawn: To prepare the graph data to be drawn, first, SQL (Structured Query Language, a database language with various functions such as data manipulation and data definition) needs to be designed according to the business situation. Prepare the SQL for each entity and edge and perform collection. The SQL for entities and edges is a design based on the business scenario and is an important factor in knowledge graph construction. This information needs to be abstracted into a structured configuration table;
[0082] The preparatory work is a prerequisite task and not the core of this embodiment. This embodiment is based on the completion of the prerequisite work. Here, for the sake of logical coherence, the details of the prerequisite work are stated first. However, this embodiment needs to complete the filling of the structured configuration table based on these prerequisite works.
[0083] In one embodiment, obtaining the structured configuration table in S2 specifically includes:
[0084] Read the latest graph Schema and the latest collection SQL from the cache;
[0085] If the latest collection SQL is not in the structured configuration table, add a new row to the structured configuration table to add the latest collection SQL, and obtain the output file name of the graph data to be drawn and the collection object (entity or edge or entity + edge), the name of the collection object, and the source entity and target entity connected when the collection object is an edge according to the latest graph Schema;
[0086] If a certain row of collection SQL in the structured configuration table is not in the latest collection SQL, set the enable flag of the certain row of collection SQL in the structured configuration table to off.
[0087] In this embodiment, as Figure 3 shown, the core work of the method includes:
[0088] 3) Design the structured configuration table (structured policy configuration table): This step is the key step in the automatically generated graph mapping file. In this file (structured policy configuration table), some detailed elements need to be included as follows:
[0089] 3.1) Sql, the SQL for writing the draft graph data. The SQL statement represents the calculation process of data from the source data to the draft graph data, including "select" and "from", indicating retrieving data content from the source database;
[0090] 3.2) comment, the Chinese explanation of this entity / edge for maintenance. For example, the Chinese explanation of the above SQL statement is transmission circuit;
[0091] 3.3) databaseType, the source of the draft graph data, facilitating collection identification. The original data may come from many different databases, and perhaps the SQL also has different grammars. Therefore, when making an automated generation solution, it is necessary to identify the type of the database to facilitate different splitting logics for different databases;
[0092] 3.4) isActive, indicating whether this row is adopted (equivalent to the start switch of this row) for maintenance. In many cases, the graph schema design will change. According to business requirements, perhaps a certain entity or a certain edge is no longer needed. Then, only this field in this configuration table needs to be changed, and the data mapping file is regenerated, and this change operation can be completed very quickly;
[0093] 3.5) sinkName, the output file name. The output file name is an essential parameter of the data mapping file because the data mapping file is used to find the draft graph data, and a file name is needed to indicate the storage location of the draft graph data;
[0094] 3.6) schema, the graph model is an entity or an edge or an entity + edge. This schema is just the field name of the table, representing the configuration of this row. In the graph Schema, there are three types: vertex, edge, and edge + vertex. Vertex represents that this row is an entity, edge represents that this row is an edge, and edge + vertex represents that the data of this row is not only an entity but also an edge;
[0095] 3.7) labelV, the name of the entity. For rows with schema as vertex and edge + vertex, write the corresponding entity name. For example, optical cable, provincial site, router, etc. It may be represented in English for easy coding;
[0096] 3.8) labelE, the name of the edge. For rows with schema as edge and edge + vertex, write the corresponding edge name. For example, optical cable, transmission section, etc. The name of each edge is different, and the name of the edge here corresponds to the name of the edge in the graph schema;
[0097] 3.9) Source. If it is an edge, write the mapping of the source field of the edge, indicating the source entity connected by the edge.
[0098] 3.10) Target. If it is an edge, write the mapping of the target field of the edge, indicating the target entity connected by the edge.
[0099] The above content is configured in a structured manner in the form of a table. The headers of each column are the above 10 items. The specific number of rows to be filled depends on the design of the graph schema. If it is designed to have 5 entities and 5 edges, there will be 10 SQLs used to retrieve the data of entities and edges. SQL is used as a technical means to convert source data into the data structure of the proposed graph. By using SQL, the data is calculated and extracted. Parsing the SQL can help understand the structure of the proposed graph data, thereby obtaining the main information for forming the graph database mapping file.
[0100] In one embodiment, obtaining the output file name of the proposed graph data specifically includes:
[0101] Obtain the output file name of the proposed graph data set by the user, execute each row of the collection SQL with the enabled flag enabled in the structured configuration table, and store the collected data in the file corresponding to the output file name of the proposed graph data for each row.
[0102] Alternatively, if each row of the collection SQL with the enabled flag enabled in the structured configuration table has been executed and the collected data has been stored in the corresponding output file of the proposed graph data, obtain the file name of the output file of the proposed graph data corresponding to each row.
[0103] In this embodiment, as Figure 3 shown, this method is mainly a process of automatically generating a graph mapping file by parsing SQL according to the structured policy configuration table. However, in fact, it also includes the process of collecting proposed graph data according to the data collection SQL. SQL collection and parsing are two separate steps. Use SQL to collect the proposed graph data, and use SQL to parse out the required fields and attributes. There is no order or dependency between the two, that is, the process of collecting proposed graph data by SQL is not limited to being executed before looping through the configuration table. It can also execute SQL collection to generate the proposed graph data file and parse SQL to generate the graph mapping file when reading the configuration table. This is also convenient for using this method during the iteration of the graph database. For graph model iteration, only need to open the configuration table, add a new row of content in the table (when the graph model adds new content, add corresponding description information in the table), and then click to re - output. The program will automatically update the content to the latest and re - output a brand - new graph mapping file. Through a very intuitive and well - organized logical abstraction (logical abstraction refers to the design and implementation of the structured configuration table), the automated generation and iteration of the graph database mapping file can be efficiently completed.
[0104] In one embodiment, parsing the structured configuration table line by line in S3 and generating a graph database mapping file based on a template specifically includes:
[0105] Parsing the collection SQL of the current line to obtain the identification ID and attributes of the entity and / or edge;
[0106] In response to the collection object of the current line containing an entity, according to the entity object declaration specification of the graph database mapping file, assembling the identification ID and attributes of the entity corresponding to the current line, the name of the graph data output file to be generated, and the name of the collection object to obtain the first assembled data;
[0107] In response to the collection object of the current line containing an edge, according to the edge object declaration specification of the graph database mapping file, assembling the identification ID and attributes of the edge corresponding to the current line, the name of the graph data output file to be generated, the name of the collection object, and the source entity and target entity connected to obtain the second assembled data;
[0108] Combining the first assembled data and the second assembled data according to the template to obtain a graph database mapping file, which is used to map the collection object data in the graph data to the entities and / or edges of the graph model that conforms to the graph Schema mode.
[0109] In this embodiment, as Figure 3 shown, the core work of the method further includes:
[0110] 4) Using some program methods to read the entity-edge data in the configuration table, looping until the end of all lines. In the loop, analyze whether the element of the current line is an entity or an edge, and parse information such as the configured SQL, that is, including parsing information such as SQL, points, edges, mappings, storage addresses, etc. For entities and edges respectively, generate object data of a single element according to their respective graph mapping specifications (predefined formats), that is, generate data mapping objects (points, edges) which are the smallest units of the graph mapping file, that is, generate an abstraction of a piece of code required in a graph mapping file, store it in a data set cache, and finally combine all the caches stored in the array and convert them into the code of the graph mapping file, specifically including:
[0111] 4.1) If it is determined that the schema field contains an entity, first parse the SQL. Through some regular expressions, parse the fields in the SQL, find the keywords "select" and "from" in the SQL, parse the text in between, and after parsing, separate each field separated by a comma, output each field, and by convention, set the first field of the SQL as the entity's ID field, and the other fields are the entity's attributes. With the entity's ID and attribute fields, the entity ID is the unique identifier of the entity in the knowledge graph, and the attribute fields are obtained and written according to business needs. For example, for a person, the entity ID is the ID card, and the attribute fields are height, weight, nationality, etc. Combine the sinkname of the current column to obtain the name of the graph data + the name of the entity in labelv. Sinkname, labelv, etc. need to be written into the graph mapping file. Through a fixed data assembly logic (the fixed data splicing logic is based on the format of the graph mapping file. For example, if a vertex entity needs to be declared in the graph mapping file, the labelV field in a row will be retrieved from the configuration table, and the information of this field will be written into the vertex entity. The whole template is very large and there are many fixed assemblies, and an object data that conforms to the graph mapping specification for a single element will be generated (this object data is an abstraction of a graph mapping file, abstracted into a json file). Store the object data of a single element in a list data set. The list data set is an intermediate variable during the conversion process, automatically generated during program operation and cleared after running;
[0112] 4.2) If it is determined that the schema field contains an edge, the SQL parsing logic is the same as above, and the processes of sinkname and labelE are similar. However, for an edge, an additional step is required to write the source and target information. Write it into the list data set, which is clearly defined when designing the configuration table. Because the previous work included source data analysis and graph schema analysis, which source and target fields are determined during design. As long as there is an edge, these two pieces of information need to be supplemented. These two pieces of information are the necessary conditions for an edge, and they are also the edge configuration information specified by the graph mapping file specification. These two mappings will be performed when parsing an edge. Because the logics for writing the graph mapping file for entities and edges are different. An entity only needs to specify its ID and attributes, while an edge needs to tell which two entities it needs to connect. Therefore, these two mappings represent its starting entity and ending entity. The final data regularization and data set writing are similar to that of the entity step.
[0113] 4.3) In the above judgment, the judgment logic is "containment" because some elements may be not only entities but also edges. For example, for a section of optical cable, if we want to query the relationship between routers, the optical cable connects the routers, where the routers are entities and the optical cable is an edge. If we want to query the relationship between the optical cable and the transmission section, the optical cable can be regarded as an entity and can be connected to the transmission section entity through an edge, which is a cross - professional connection query. Therefore, in the knowledge graph, entities and edges are not constant but are dynamically adjusted according to the business scenario. So, the judgment and data output are carried out in the way of containment.
[0114] In one embodiment, parse the acquisition SQL of the current row to obtain the identification ID and attributes of entities and / or edges, specifically including:
[0115] Parse the keyword fields that need to be obtained from the mapping data included in the acquisition SQL of the current row;
[0116] Set the first obtained keyword field as the identification ID of the entity and / or edge;
[0117] Set the non - first obtained keyword fields as the attributes of the entity and / or edge.
[0118] In this embodiment, the parsing of SQL has been described above. A core part of the code example of the method is provided as follows:
[0119]
[0120]
[0121]
[0122] In one embodiment, combine the first assembled data and the second assembled data according to the template to obtain the graph database mapping file, specifically including:
[0123] Obtain the first assembled data and the second assembled data temporarily stored in the memory;
[0124] Fill the first assembled data and the second assembled data into the preset graph database mapping file template in the order of the structured configuration table to obtain the graph database mapping file.
[0125] In this embodiment, as Figure 3As shown, after the loop ends, a complex data set containing edges and entities will be obtained. This data set is already a semi-finished product. Finally, this data set is converted into a final graph mapping data set using a fixed algorithm (in-memory data access algorithm). Finally, the data set is serialized and the graph mapping file is output using some program methods. Serialization means converting the data originally stored in memory into a storable and readable file format and finally outputting it to a file. Using the above method, the entire logic adopts a fixed algorithm or template, and with just one click, a completely correct graph mapping configuration file can be output within seconds. An exemplary code for automatically generating a graph mapping file is as follows:
[0126]
[0127]
[0128] In one embodiment, the method further includes:
[0129] Executing the code of the graph database mapping file to map the acquisition object data in the quasi-graph data into entities and / or edges of a graph model that conforms to the graph Schema pattern, including: in the graph model, using the acquisition object name as the classification of the entity and / or edge, annotating the identification ID of the entity and / or edge and the corresponding quasi-graph data content of the attribute, and establishing connections between entities according to the source-end entity and the target-end entity connected by the edge.
[0130] In this embodiment, the method further includes: constructing a knowledge graph after generating the graph mapping, that is, using the graph database mapping file to map the quasi-graph data into a graph model that conforms to the graph Schema pattern. The graph database mapping file defines a graph data model. For example, it can include multiple types of nodes (such as sites and rooms) and multiple edges (such as the relationship connecting sites and rooms). Each node and edge specifies the data source file and its format, column name, and character set.
[0131] This embodiment includes the following innovations: The configuration of points and edges combines the points and edges designed by the graph Schema and then performs divide-and-conquer SQL parsing and fusion. The multi-dimensional contents such as the fields of SQL, the data mapping fields of edges, and the data mapping fields of points are merged, analyzed, parsed, and fused; for the abstract method of the graph Schema, usually the graph Schema is maintained in a coded manner through Groovy (an agile development language based on the JVM), and when developing the mapping file, it is necessary to constantly read the graph Schema and fully associate the graph Schema with the mapping file. Reading the coded file has poor readability and low work efficiency. Instead, the graph Schema is abstracted into structured data sets, which contain the necessary elements of the graph Schema, and these elements are the core elements of automated analysis; combining graph Schema analysis and automated output of the mapping file from the original data, combining graph Schema configuration + SQL fields + original management data to form a sufficient data set and automatically form the mapping file; graph mapping file iteration and reconstruction. Using the method of automatically outputting the mapping file and the file overwrite method, with the idea of "rebuilding when modified", only need to modify individual fields in the structured file or add individual fields to generate a completely correct mapping file with one click, without the need for review, directly solving the problems of error-proneness, large modification workload, and difficult maintenance caused by iteration and modification.
[0132] Embodiment 2:
[0133] As Figure 2 shown, this application provides a device for creating a graph database mapping file, and the device includes:
[0134] A preparation module 1, configured to obtain a collection structured query language SQL for extracting pseudo-graph data according to the graph Schema. The graph Schema is a graph schema for establishing a graph database, and the pseudo-graph data is intermediate data obtained by processing source data into the data structure required by the graph Schema;
[0135] A configuration module 2, connected to the preparation module 1, configured to obtain a structured configuration table. In the structured configuration table, each collection SQL is a row, indicating the output file name of the pseudo-graph data of each collection SQL, the collection object being an entity or an edge or entity + edge, the name of the collection object, and the source-end entity and target-end entity connected when the collection object is an edge;
[0136] An analysis module 3, connected to the configuration module 2, configured to parse the structured configuration table row by row and generate a graph database mapping file based on a template. The graph database mapping file is a code file that maps pseudo-graph data into a graph model conforming to the graph Schema pattern.
[0137] In an implementation manner, the preparation module 1 specifically includes:
[0138] A graph Schema unit, used to obtain each entity and each edge in the graph Schema;
[0139] A database information unit, connected to the graph Schema unit, used to obtain the source database where the source data corresponding to each entity and each edge in the graph Schema is located;
[0140] An SQL unit, connected to the graph Schema unit and the database information unit, used to design the acquisition SQL for each entity and each edge in the graph Schema respectively. Each acquisition SQL is used to extract the source data corresponding to each entity or each edge in the graph Schema from the source database and process it into quasi-graph data that conforms to the data structure required by the graph Schema.
[0141] In one embodiment, the configuration module 2 specifically includes:
[0142] A cache unit, used to read the latest graph Schema and the latest acquisition SQL from the cache;
[0143] An addition unit, connected to the cache unit, used to add a new row to the structured configuration table to add the latest acquisition SQL if the latest acquisition SQL is not in the structured configuration table, and obtain the output file name of the quasi-graph data and the acquisition object (entity or edge or entity + edge), the acquisition object name, and the source entity and target entity connected when the acquisition object is an edge according to the latest graph Schema;
[0144] An invalidation unit, connected to the cache unit, used to set the enable flag of a certain row of acquisition SQL in the structured configuration table to closed if the acquisition SQL in a certain row of the structured configuration table is not in the latest acquisition SQL.
[0145] In one embodiment, the addition unit includes a quasi-graph file acquisition unit, which is specifically used for:
[0146] Obtain the output file name of the quasi-graph data set by the user, execute each row of acquisition SQL in the structured configuration table with the enable flag set to enabled, and store the acquired data in the file corresponding to the output file name of the quasi-graph data for the corresponding row;
[0147] Or, if each row of acquisition SQL in the structured configuration table with the enable flag set to enabled has been executed and the acquired data has been stored in the corresponding quasi-graph data output file, obtain the file name of the quasi-graph data output file corresponding to each row.
[0148] In one embodiment, the parsing module 3 specifically includes:
[0149] An SQL parsing unit, used to parse the acquisition SQL of the current row to obtain the identification ID and attributes of the entity and / or edge;
[0150] The first assembly unit, connected to the SQL parsing unit, is used to respond to the acquisition object of the current row containing an entity, and assemble the identification ID and attributes of the entity corresponding to the current row, as well as the pseudo-graph data output file name and the acquisition object name according to the entity object declaration specification of the graph database mapping file, so as to obtain the first assembled data;
[0151] The second assembly unit, connected to the SQL parsing unit, is used to respond to the acquisition object of the current row containing an edge, and assemble the identification ID and attributes of the edge corresponding to the current row, as well as the pseudo-graph data output file name, the acquisition object name, and the source entity and target entity connected according to the edge object declaration specification of the graph database mapping file, so as to obtain the second assembled data;
[0152] The combined output unit, connected to the first assembly unit and the second assembly unit, is used to combine the first assembled data and the second assembled data according to a template to obtain a graph database mapping file, and the graph database mapping file is used to map the acquisition object data in the pseudo-graph data into entities and / or edges that conform to the graph Schema model.
[0153] In one embodiment, the SQL parsing unit specifically includes:
[0154] The keyword field unit is used to parse the keyword fields that need to be obtained from the pseudo-graph data included in the acquisition SQL of the current row;
[0155] The identification ID unit, connected to the keyword field unit, is used to set the first obtained keyword field as the identification ID of the entity and / or edge;
[0156] The attribute unit, connected to the keyword field unit, is used to set the non-first obtained keyword field as the attribute of the entity and / or edge.
[0157] In one embodiment, the combined output unit specifically includes:
[0158] The memory unit is used to obtain the first assembled data and the second assembled data temporarily stored in the memory;
[0159] The template file unit, connected to the memory unit, is used to fill the first assembled data and the second assembled data into a preset graph database mapping file template according to the order in the structured configuration table, so as to obtain a graph database mapping file.
[0160] In one embodiment, the device further includes a graph model mapping unit, which is used for:
[0161] Execute the code of the graph database mapping file to map the collected object data in the draft graph data into entities and / or edges of a graph model that conforms to the graph Schema pattern, including: in the graph model, using the collected object name as the classification of the entity and / or edge, marking the identity ID of the entity and / or edge and the content of the draft graph data corresponding to the attributes, and establishing connections between entities according to the source entity and the target entity connected by the edge.
[0162] Embodiment 3:
[0163] Embodiment 3 of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is run by a processor, it implements the method for creating a graph database mapping file as described in Embodiment 1, or implements the device for creating a graph database mapping file as described in Embodiment 2.
[0164] The computer-readable storage medium includes volatile or non-volatile, removable or non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, computer program units, or other data). The computer-readable storage medium includes, but is not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory, or other memory technologies, CD-ROM (Compact Disc Read-Only Memory), digital versatile disc (DVD), or other optical disc storage, magnetic cassette, tape, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer.
[0165] In addition, the present application can also provide a computer device, including a memory and a processor. A computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes the method for creating a graph database mapping file as described in Embodiment 1. The computer device can be the device for creating a graph database mapping file as described in Embodiment 2.
[0166] Among them, the memory is connected to the processor. The memory can adopt flash memory, read-only memory, or other memories, and the processor can adopt a central processing unit or a single-chip microcomputer.
[0167] Embodiments 1-3 of the present application provide a method, apparatus, and medium for creating a graph database mapping file. By designing a structured configuration table, the content of each row is set according to the characteristics of the graph Schema and the pseudo-graph data collection. The collection SQL for extracting the pseudo-graph data according to the graph Schema is used as the main content of each row, and information such as the output file name of the pseudo-graph data required for generating the graph database mapping file, whether the collection object is an entity, an edge, or an entity + edge, the name of the collection object, and the source and target entities connected when the collection object is an edge is configured for each row. To achieve the purpose of automatically generating the graph database mapping file by parsing the structured configuration table row by row, the creation efficiency of the graph database mapping file is improved. At the same time, due to the automated process, the problem of error-prone manual comparison between the graph Schema and the pseudo-graph data is avoided, and the accuracy of the graph database mapping file is improved.
[0168] It can be understood that the above embodiments are merely exemplary embodiments adopted to illustrate the principle of the present application. However, the present application is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present application, and these modifications and improvements are also regarded as the protection scope of the present application.
Claims
1. A method for creating a graph database mapping file, characterized in that: The method comprises: Obtain the collection structured query language SQL to extract pseudo-graph data according to the graph schema. The graph schema is the graph model used to establish the graph database, and the pseudo-graph data is the intermediate data that processes the source data into the data structure required by the graph schema; Obtain a structured configuration table, in which each collected SQL is a row, indicating the output file name of the simulated graph data of each row of collected SQL, whether the collected object is an entity, an edge, or an entity + an edge, the name of the collected object, and the source end entity and the target end entity of the connection when the collected object is an edge; Parse the structured configuration table line by line and generate a graph database mapping file based on the template. The graph database mapping file is a code file that maps the simulated graph data into a graph model that conforms to the graph Schema mode.
2. The method according to claim 1, characterized in that Obtain the collection structured query language SQL for extracting the simulated graph data based on the graph schema, including: Get each entity and each edge in the graph schema; Get the source database where the source data corresponding to each entity and each edge in the graph schema is located; Design the collection SQL for each entity and each edge in the graph schema. Each collection SQL is used to extract the source data corresponding to each entity or each edge in the graph schema from the source database, and process it into simulated graph data with a data structure that meets the requirements of the graph schema.
3. The method according to claim 2, characterized in that Get a structured configuration table, including: Read the latest graph schema and the latest collection SQL from the cache; If the latest collection SQL is not in the structured configuration table, add a new row to the structured configuration table to add the latest collection SQL, obtain the simulated graph data output file name, and obtain the collection object according to the latest graph Schema, whether it is an entity, edge, or entity + edge, the collection object name, and the source end entity and target end entity connected when the collection object is an edge; If a certain row of collection SQL in the structured configuration table is not in the latest collection SQL, the opening flag of the certain row of collection SQL in the structured configuration table is set to closed.
4. The method according to claim 3, characterized in that Get the output file name of the simulated image data, including: Get the artificially set pseudo-graph data output file name, execute each row of collection SQL marked as enabled in the structured configuration table, and store the collected data in the file corresponding to the pseudo-graph data output file name of the corresponding row; Alternatively, if each row of collection SQL marked as enabled in the structured configuration table has been executed and the collected data has been stored in the corresponding simulated image data output file, the file name of the simulated image data output file corresponding to each row is obtained.
5. The method according to any one of claims 1 to 4, characterized in that: Parse the structured configuration table line by line and generate a graph database mapping file based on the template, including: Parse the collection SQL of the current row to obtain the identification ID and attributes of the entity and / or edge; In response to the collection object of the current row containing an entity, according to the entity object declaration specification of the graph database mapping file, assemble the identification ID and attributes of the entity corresponding to the current row and the output file name of the simulated graph data and the collection object name to obtain the first assembled data; In response to the collection object of the current row containing an edge, according to the edge object declaration specification of the graph database mapping file, assemble the identification ID and attributes of the edge corresponding to the current row and the output file name of the simulated graph data, the collection object name, and the source end entity and the target end entity of the connection to obtain second assembled data; The first assembled data and the second assembled data are combined according to the template to obtain a graph database mapping file, and the graph database mapping file is used to map the collected object data in the simulated graph data into entities and / or edges of a graph model that conforms to the graph Schema mode.
6. The method according to claim 5, characterized in that Parse the collection SQL of the current row to obtain the ID and attributes of the entity and / or edge, including: Parse the key fields contained in the collection SQL of the current row that need to be obtained from the simulated data; Set the first key field obtained as the entity and / or edge ID; Set the obtained non-first key fields as attributes of the entity and / or edge.
7. The method according to claim 5, characterized in that Combining the first assembled data and the second assembled data according to the template to obtain a graph database mapping file specifically includes: Acquire the first assembly data and the second assembly data temporarily stored in the memory; Fill the first assembled data and the second assembled data into a preset graph database mapping file template according to the order in the structured configuration table to obtain a graph database mapping file.
8. The method according to claim 5, characterized in that The method further comprises: Execute the code of the graph database mapping file to map the collected object data in the simulated graph data into entities and / or edges of the graph model that conforms to the graph Schema mode, including: in the graph model, using the collected object name as the classification of the entity and / or edge, marking the entity and / or edge identification ID and the simulated graph data content corresponding to the attribute, and establishing connections between entities based on the source end entity and the target end entity connected by the edge.
9. A graph database mapping file creation device, characterized in that: The device comprises: The preparation module is used to obtain the collection structured query language SQL to extract the pseudo-graph data according to the graph schema. The graph schema is the graph model used to establish the graph database, and the pseudo-graph data is the intermediate data that processes the source data into the data structure required by the graph schema; The configuration module is connected to the preparation module and is used to obtain a structured configuration table. In the structured configuration table, each collected SQL is a row, indicating the output file name of the simulated graph data of each row of collected SQL, whether the collected object is an entity, an edge, or an entity + an edge, the name of the collected object, and the source end entity and the target end entity of the connection when the collected object is an edge; The parsing module is connected to the configuration module and is used to parse the structured configuration table line by line and generate a graph database mapping file based on the template. The graph database mapping file is a code file that maps the simulated graph data into a graph model that conforms to the graph Schema mode.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for creating a graph database mapping file as described in any one of claims 1 to 8 is implemented.