A method for constructing an automated batch-stream data labeling framework based on a directed computational graph
By building an automated batch flow data marking framework based on directed computing graphs, the problem of changing the marking rules in the existing technology requires rewriting code, and flexible automatic marking of batch flow data is realized, and production efficiency is improved.
Patent Information
- Application Number
- CN202310300686.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-03-23
AI Technical Summary
The existing technology lacks a flexible open source framework to handle automated marking of batch stream data, resulting in rewriting code when the marking rules change, affecting production efficiency.
By inputting the interface description language, an abstract syntax tree is generated and converted into a relational algebra expression tree, a logical directed calculation graph is defined, a physical directed graph is generated, an automated batch flow data marking framework is built, and an automatic generation logic based on marking rules is realized.
It realizes flexible and automated marking of batch stream data without rewriting code, and improves the productivity of the enterprise.
Smart Images

Figure CN116301755B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a method for constructing an automated batch-stream data tagging framework based on a directed computation graph. Background Art
[0002] The Internet of Things (IoT) refers to the use of various information sensors, radio frequency identification technologies, global positioning systems, infrared sensors, laser scanners and other devices and technologies to collect in real time any objects or processes that need to be monitored, connected and interacted with, collect various required information such as their sound, light, heat, electricity, mechanics, chemistry, biology, location, etc., and through various possible network accesses, realize the ubiquitous connection of things to things and things to people, and realize the intelligent perception, identification and management of items and processes. The Internet of Things is an important part of the smart city project. In the process of promoting the smart city project, various sensors will generate a large amount of streaming structured data in real time. At the same time, each data center will also store static (batch) city-related data (such as street longitude and latitude, sensor geographical location, etc.).
[0003] However, the data collected by sensors alone lacks global semantics. For example, a certain pipeline sensor collects values such as pressure and temperature inside the pipeline and also records the street where the sensor is located, but street-related attributes such as longitude and latitude are stored in the relational database of the data center. Therefore, the data collected by the sensor cannot be associated with the street-related semantic information. In order to add global semantics to the sensor data generated in real time, it is necessary to associate the generated data with the semantic data in the data center each time the sensor data is generated. For the sake of convenience of explanation, this process is called data tagging, that is, adding tags related to global semantics to the streaming data.
[0004] According to research, there is currently no open-source framework specifically for processing batch-stream data tagging scenarios on the market. The current mainstream method for processing batch-stream tagging is hard-coding based on tagging rules. The fatal flaw of this method is that it is not flexible enough. Once the tagging rules change, the code needs to be rewritten and the program needs to be taken offline and online again, which is unacceptable in the production environment of software enterprises. Therefore, it is particularly important to develop a framework that can automatically generate tagging logic based on tagging rules.
[0005] A similar scenario is the labeling of dual-stream data (commonly seen in e-commerce OLAP scenarios). Open-source distributed computing frameworks such as Flink, Spark, and Storm can effectively handle the scenario of automatically labeling dual-stream data. Taking Flink as an example, Flink can automatically generate a DAG computational graph (directed computational graph) based on SQL statements, and use the codegen tool to automatically generate corresponding data processing code on each node of the DAG, thereby realizing the business code for automatically generating dual-stream data labels according to SQL statements.
[0006] In summary, there is still a lack of research on automatic labeling of batch-stream data. Summary of the Invention
[0007] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a method for constructing an automatic batch-stream data labeling framework based on a directed computational graph.
[0008] The purpose of the present invention can be achieved through the following technical solutions:
[0009] A method for constructing an automatic batch-stream data labeling framework based on a directed computational graph includes the following steps:
[0010] Input interface description language, which includes table creation instructions, data source definitions, and data connection instructions;
[0011] Perform lexical analysis on the interface description language to generate an abstract syntax tree;
[0012] Convert the abstract syntax tree into a relational algebra expression tree;
[0013] Define the relational algebra expression tree as a logical directed computational graph;
[0014] Perform code generation operations on the nodes of the logical directed computational graph to obtain a physical directed graph, and complete the construction of the batch-stream data labeling framework.
[0015] Furthermore, the selected interface description language is the SPL language.
[0016] Furthermore, the root node of the relational algebra expression tree is the exit of the logical directed computational graph, and the leaf node of the relational algebra expression tree is the entrance of the logical directed computational graph.
[0017] Furthermore, the logical directed computational graph is composed of multiple logical nodes connected in series, and the logical nodes include a table scanning node, a filtering node, a joining node, and a projection node;
[0018] The table scanning node is the entrance of the logical directed computational graph;
[0019] The projection node is used to specify the mapping relationship between input entries and output entries;
[0020] The association node is used to define the association relationship between input entries and output entries.
[0021] Further, code generation operations are performed on the nodes of the logical directed computation graph to obtain a physical directed graph, which specifically includes the following steps:
[0022] Specific business code is generated on the corresponding nodes according to the logical nodes and their node rules, and code generation operations are sequentially performed on each logical node to obtain corresponding physical nodes;
[0023] All physical nodes are combined and optimized based on rules to obtain a physical directed graph.
[0024] Further, in the table scanning node, a base table class is defined, and the base table class uses the GetRows interface;
[0025] The table scanning node also includes a streaming table class and a batch table class. Both the streaming table class and the batch table class inherit from the base table class and also use the GetRows interface;
[0026] For the involved streaming data sources, corresponding streaming data tables are separately defined, and these corresponding streaming data tables all inherit from the streaming table class; for the involved batch data sources, corresponding batch data tables are separately defined, and these corresponding batch data tables all inherit from the batch table class;
[0027] The specific table scanning code logic is implemented in the streaming data table and the batch data table classes.
[0028] Further, when generating the logical nodes in the logical directed computation graph, it is necessary to save the data source information in the interface description language, and implement the code generation operation of the nodes based on this data source information.
[0029] Further, the projection node performs a mapping operation according to the semantic information attached to the relational algebra expression tree.
[0030] Further, the association node selects a batch data source to construct a hash table for association operations;
[0031] If there is an and logic in the association condition, that is: select * from f1, d1 left join on f1.b = d1.b and f1.c = d1.c, then a hash table is constructed with concat(d1.b, d1.c) as the key when constructing the hash table;
[0032] If there is an "or" logic in the association condition, i.e., "select * from f1, d1 left join on d1.b = f1.b or d1.c = f1.c", then construct a hash table with d1.b as the key and another hash table with d1.c as the key. When associating, take the union of the query results of the two hash tables as the output;
[0033] If there is a nested logic of "and" and "or" in the association condition, i.e., "((f1.b = d1.b or f1.c = d1.c) and f1.d = d1.d)", then use the distributive law of logical operators to transform it into "((f1.b = d1.b and f1.d = d1.d) or (f1.c = d1.c and f1.d = d1.d))". The conversion rule is that there is only an "and" operation within all basic rules, and the basic rules are connected by "or". For constructing the hash table within the basic rules, use splicing. For the connection between basic rules, take the union of the output results.
[0034] Furthermore, perform rule-based optimization to equivalently transform all filtering nodes to the position adjacent to the table-scanning node. After optimization, the direct input of the filtering node is a table-scanning node, so as to remove the filtering-related operations in the interface description language and directly write the filtering logic function at the code level.
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] By performing lexical analysis on the input interface description language, the present invention generates a logical directed computation graph and constructs a framework for automatically generating marking logic based on marking rules, which can flexibly implement automated marking for batch stream data. Even if the marking rules change, there is no need to rewrite the code and redeploy the program, thereby effectively improving the production efficiency of enterprises. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 is the implementation flowchart of the present invention;
[0038] Figure 2 is the framework diagram of automatically generating marking logic in an embodiment of the present invention;
[0039] Figure 3 is the code implementation example of generating a relational algebra expression tree using calcite in an embodiment of the present invention;
[0040] Figure 4 is the relationship schematic diagram of table-scanning node-related classes in an embodiment of the present invention;
[0041] Figure 5 is the result schematic diagram of parsing a projection node using calcite in an embodiment of the present invention;
[0042] Figure 6 Schematic diagram of the result of parsing associated nodes using Calcite in the embodiment of the present invention;
[0043] Figure 7 Schematic diagram of the process of performing RBO optimization in the embodiment of the present invention. Detailed implementation manners
[0044] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.
[0045] Glossary:
[0046] Batch data: refers to the data statically stored in the data center (not changing with time);
[0047] Streaming data: refers to the data generated in real time by devices such as sensors (a new piece of data is generated every time a piece of data is collected);
[0048] Batch-stream data tagging: refers to the process of associating batch data and streaming data according to certain rules (such as the same street id). For example, a piece of streaming data is in the form (id: 001, street_id: 20180010, city: shanghai, pressure: 100kapa, temperature: 452, depth: -1422.3), and a piece of batch data is in the form (id: 001, street_id: 20180010, city: shanghai, crowd: 0.672, polulation: 24732). After tagging, the data is in the form (id: 001, street_id: 20180010, city: shanghai, pressure: 100kapa, temperature: 452, depth: -1422.3, crowd: 0.672, polulation: 24732). By associating through the street_id attribute, semantic information of crowd and population is added to the streaming data (the actual process is much more complex than this example, and a piece of streaming data may be recursively associated with batch data in multiple databases);
[0049] Automated tagging: As in the above example, for association based on street_id, the current mainstream method for processing batch-stream tagging is to hard-code the tagging rules. Once the rules change, the code needs to be rewritten, while automated tagging hopes to automatically generate the code for processing data tagging logic through an intermediate language (such as SQL).
[0050] Aiming at the problem that there is a lack of an open-source framework specifically for processing batch-stream data tagging scenarios in the prior art, the present invention proposes a method for constructing an automated batch-stream data tagging framework based on a directed computational graph, as Figure 1 shown. The method of the present invention includes the following steps:
[0051] An input interface description language (IDL, Interface description language), which includes table creation instructions, data source definitions, and data connection instructions;
[0052] Perform lexical analysis on the interface description language (IDL) to generate an abstract syntax tree;
[0053] Convert the abstract syntax tree into a relational algebra expression tree;
[0054] Define the relational algebra expression tree as a logical directed acyclic graph (DAG, Directed Acyclic Graph);
[0055] Perform code generation operations on the nodes of the logical directed acyclic graph to obtain a physical directed graph, and complete the construction of the batch-stream data tagging framework.
[0056] The architecture design of the framework proposed in this application is as Figure 2 shown. In this embodiment, the framework automatically generates tagging logic based on SQL. That is, the input of the overall framework is a piece of IDL (Interface description language, interface description language), and SQL is specified as the IDL in this embodiment.
[0057] Take Figure 2 as an example. In the interface description language input in this embodiment, three abstract tables F1 (streaming), D1 (batch), and D2 (batch) are defined. Among them, F1 is associated with a kafka message queue (a common streaming data source), and D1 and D2 are associated with mysql (batch data sources). The envisioned scenario is that F1 continuously generates new data over time. The data collected by F1 only has two fields, a and b. Every time F1 generates a piece of data, it needs to be associated with the entries in D1 and D2 (the F1 entry is first associated with the D1 entry according to the b field to generate the F2 entry, and then the F2 entry is associated with the D2 entry according to the c field to generate the F3 entry). After the association is completed, the F3 entry is stored in the persistent layer (such as HDFS) according to relevant business requirements for subsequent OLAP (Online Analytical Processing) operations.
[0058] Figure 2The meaning of the SQL is as follows: First, create a streaming table named F1 with the streaming data source being Kafka and having two fields a and b. Then, create two batch tables D1 and D2 with the data source being MySQL. D1 has two fields b and c, and D2 has two fields c and d. Then execute the SQL. The specific meaning of the SQL is: First, perform a left join with the condition that the b field of F1 is equal to the b field of D1 to obtain a temporary table F2 (F2 has three fields a, b, c). Then, perform a left join with the condition that the c field of F2 is equal to the c field of D1 to obtain the output table, and the output table has a total of 4 fields a, b, c, d (corresponding to f2.a, f2.b, c2.c, d2.d).
[0059] In the framework of this embodiment, the process of automatically generating the tagging logic includes two major stages. The first stage is to convert the SQL statement into a logical DAG. In this stage, use the apache calcite tool (a dynamic data management framework) to perform lexical analysis on the SQL, generate an abstract syntax tree, and finally convert the abstract syntax tree into a relational algebra expression tree (relnodetree). Here, the relational algebra expression tree is the logical DAG (the root node is the exit of the graph, and the leaf node is the entry of the graph). Figure 3 Shows the process of the calcite framework converting the SQL statement in this embodiment into a logical DAG (relational algebra expression tree). Figure 3 The relnode in it completely corresponds to Figure 2 the RelNode Tree part in it.
[0060] It should be noted that the DAG here is only logical, and all the filter, projection, join, and scantable nodes do not have specific code implementations yet. Therefore, in the second stage of the framework, it is necessary to generate specific business code on the corresponding nodes according to the nodes and node rules, perform code generation operations on each logical node to obtain a physical node, combine the physical nodes and perform RBO (Rule-Based Optimization) to obtain a physical directed graph. The physical directed graph is a program that can be directly deployed.
[0061] In the two stages of the framework process, calcite can already handle the requirements of the first stage well. Therefore, the technical details in this application mainly focus on the second stage, that is, how to generate business code on logical nodes.
[0062] In the scenario of this embodiment, the entry (leaf node) of the DAG must be a scantable node. According to the data source type, the scantable node can be divided into a batch scantable node and a streaming scantable node. For exampleFigure 4 As shown, in the framework of this embodiment, a base table class (BaseTable) is defined in this embodiment. The base table class has only one interface, GetRows. The flow table class (FlowTable) and the batch table class (BatchTable) both inherit from BaseTable. For each type of streaming data source involved, a separate class is defined (such as KafkaTable, RBMQTable), and these classes all inherit from the FlowTable class. The same applies to batch data sources (such as mysql, postsql, etc.). The specific table scanning code logic is implemented in the lowest-level class (such as the KafkaTable layer).
[0063] When generating logical nodes in Phase I, this embodiment needs to save the information related to the data source in the IDL (such as Figure 2 create table from kafka and create table from mysql in it. When generating specific business code in Phase II, only the specific implementation class needs to be replaced according to the data source information of the table scanning node (such as from kafka, from mysql, etc.) (because all classes implement the GetRows interface).
[0064] The process of generating business code for the projection node includes the following steps: In Phase I, Calcite has generated logical nodes with rich semantic information. As Figure 5 shown, the semantic information attached to the projection node generated by Calcite indicates that the 0th field ($0) of the input entry should be mapped to the 0th field of the output entry, the 1st field ($1) of the input entry should be mapped to the 1st field of the output entry, and so on. Therefore, when generating business code, this embodiment only needs to perform mapping according to the semantic information attached to the relnode.
[0065] Join node:
[0066] The process of generating physical join nodes from logical join nodes is relatively complex. This embodiment starts with the simplest case. Still taking the Figure 2 case in it as an example, F1 is a flow table and D1 is a batch table. For the IDL statement select a, b, c from F1, D1 left join on F1.b = D1.b. Figure 6 Shows the parsing result of this statement by Calicite in the first phase. The semantic information related to the join operation exists in the condition field of the relnode (taking Figure 6For example, in the condition, it is specified that the association is based on the $1st field (f1.b) and the $2nd field (d1.b) of the concatenated entries (f1.a, f1.b, d1.b, d2.c).
[0067] For the single-field association of dual-static data, the mainstream solution in the industry is to select a data stream to construct a hash table (for example, construct a hash table with d1, where the key of the hash table is d1.b and the value of the hash table is the entire entry). Next, for each entry in f1, use f1.b as the key to look up the corresponding entry in the hash table. If found, concatenate the fields of the entry; if not found, concatenate null values. Generally, for the association of dual-static data sources, it is more efficient to construct a hash table with the data source having a larger data volume. However, in the scenario of batch stream tagging, the present invention must construct a hash table with the batch data source.
[0068] For a slightly more complex situation, such as having an and logic in the association condition, like (select * from f1, d1 left join on f1.b = d1.b and f1.c = d1.c), then construct a hash table with concat(d1.b, d1.c) as the key when constructing the hash table.
[0069] If there is an or logic in the association condition, like (select * from f1, d1 left join on d1.b = f1.b or d1.c = f1.c), then construct one hash table with d1.b as the key and another hash table with d1.c as the key. When associating, take the union of the query results of the two hash tables as the output.
[0070] For the most complex situation: having nested and and or logics in the association logic, like ((f1.b = d1.b or f1.c = d1.c) and f1.d = d1.d), it is necessary to use the distributive law of logical operators to convert it into ((f1.b = d1.b and f1.d = d1.d) or (f1.c = d1.c and f1.d = d1.d)). The conversion rule is that there is only an and operation within all basic rules, and the basic rules are connected by or. The hash table is constructed by concatenation within the basic rules, and the union of the output results is taken for the connection between the basic rules.
[0071] Filter node
[0072] In a normal business scenario, associated nodes only process the "=" logic, while filtering nodes process multiple logics (e.g., f1.a > 100). It is difficult to integrate the code generation process of multiple logical operations into one module. In the framework of this embodiment, this embodiment does not directly solve this problem. Instead, it first borrows RBO (rule-based optimization) to equivalently transform all filtering nodes to adjacent table-scanning nodes (such as Figure 7 shown). After optimization, the direct input of the filtering node must be a table-scanning node. For this, this embodiment can remove the filtering-related operations in the idl and directly write a filtering logic function at the code level (the input and output of the function are both List <object>Type).
[0073] Execution of the directed computation graph:
[0074] After code generation is performed on each logical node, various physical nodes are obtained. By connecting the inputs and outputs of these nodes in series, the final directed acyclic graph (DAG) is obtained. Subsequently, an interval time is set. Every time a unit interval time passes, the DAG is executed once. The execution process is that the root node requests data from its child nodes, and the child nodes then request data from their child nodes until the leaf nodes, and the leaf nodes request data from the data sources they are associated with.
[0075] In summary, a framework for automatically generating tagging logic based on tagging rules can be constructed, which can flexibly achieve automated tagging for batch and streaming data. Even if the tagging rules change, there is no need to rewrite the code and redeploy the program, thus effectively improving the production efficiency of enterprises.
[0076] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention through logical analysis, reasoning, or limited experiments based on the concept of the present invention on the basis of the prior art should fall within the protection scope determined by the claims.< / object>
Claims
1. A method for constructing an automated batch stream data tagging framework based on a directed computation graph, characterized in that, Including the following steps: Input interface description language, which contains table creation instructions, data source definitions, and data connection instructions; Perform lexical analysis on the interface description language to generate an abstract syntax tree; Convert the abstract syntax tree into a relational algebra expression tree; Define the relational algebra expression tree as a logical directed computation graph; Perform code generation operations on the nodes in the logical directed computation graph to obtain a physical directed graph, completing the construction of the batch-stream data tagging framework; The logical directed computation graph is composed of multiple logical nodes connected in series. The logical nodes include table scanning nodes, filtering nodes, association nodes, and projection nodes; The table scanning node is the entry of the logical directed computation graph; The projection node is used to specify the mapping relationship between input entries and output entries; The association node is used to define the association relationship between input entries and output entries; Performing code generation operations on the nodes in the logical directed computation graph to obtain a physical directed graph specifically includes the following steps: Generate specific business code on the corresponding nodes according to the logical nodes and their node rules, and perform code generation operations on each logical node in turn to obtain the corresponding physical nodes; Combine all physical nodes and perform rule-based optimization to obtain a physical directed graph.
2. A method for constructing an automated batch stream data tagging framework based on a directed computational graph according to claim 1, characterized in that, The selected interface description language is the spl language.
3. A method for constructing an automated batch flow data tagging framework based on a directed computational graph according to claim 1, characterized in that The root node of the relational algebra expression tree is the exit of the logical directed computation graph, and the leaf node of the relational algebra expression tree is the entry of the logical directed computation graph.
4. A method for constructing an automated batch stream data tagging framework based on a directed computational graph according to claim 1, characterized in that, In the table scanning node, a base table class is defined, and the base table class uses the GetRows interface; The table scanning node also includes a streaming table class and a batch table class. The streaming table class and the batch table class both inherit from the base table class and also use the GetRows interface; For the involved streaming data sources, corresponding streaming data tables are defined separately, and these corresponding streaming data tables all inherit from the streaming table class; for the involved batch data sources, corresponding batch data tables are defined separately, and these corresponding batch data tables all inherit from the batch table class; The specific table scanning code logic is implemented in the streaming data table and the batch data table class.
5. A method for constructing an automated batch stream data tagging framework based on a directed computational graph according to claim 1, characterized in that When generating the logical nodes in the logical directed computation graph, it is necessary to save the data source information in the interface description language and implement the code generation operation of the nodes based on this data source information.
6. A method for constructing an automated batch stream data tagging framework based on a directed computational graph according to claim 1, characterized in that, The projection node performs a mapping operation based on the semantic information attached to the relational algebra expression tree.
7. A method for constructing an automated batch flow data tagging framework based on a directed computational graph according to claim 1, characterized in that, The association node selects a batch data source to construct a hash table for association operations; If there is an and logic in the association condition, that is: select * from f1,d1 left join on f1.b = d1.b and f1.c = d1.c, then construct a hash table with concat(d1.b, d1.c) as the key when constructing the hash table; If there is an "or" logic in the join condition, i.e., select * from f1,d1 left join on d1.b = f1.b or d1.c = f1.c, then construct a hash table with d1.b as the key and another hash table with d1.c as the key. When joining, take the union of the query results of the two hash tables as the output; If there is a nested logic of "and" and "or" in the join condition, i.e., ((f1.b = d1.b or f1.c = d1.c) and f1.d = d1.d), then use the distributive law of logical operators to transform it into: ((f1.b = d1.b and f1.d = d1.d) or (f1.c = d1.c and f1.d = d1.d)). The transformation rule is that there is only an "and" operation within all basic rules, and the basic rules are connected by "or". The hash table is constructed by splicing within the basic rules, and the union of the output results is taken for the connection between the basic rules.
8. A method for constructing an automated batch stream data tagging framework based on a directed computational graph according to claim 1, characterized in that Perform rule-based optimization to equivalently transform all filter nodes to the position adjacent to the table-scanning node. After optimization, the direct input of the filter node is a table-scanning node, so as to remove the filter-related operations in the interface description language and directly write the filter logic function at the code level.
Citation Information
Patent Citations
System for fully integrated capture, and analysis of business information resulting in predictive decision making and simulation
CN109478296A
Method and device for generating real-time calculation logic data
CN113515285A