Data processing method and device based on real-time change perception
By obtaining SQL statements during database change events and identifying explicit and implicit dependencies and updating the knowledge graph, the problem of dependency update delay in data processing tasks is solved, and accurate update and optimization of real-time dependencies are achieved.
Patent Information
- Application Number
- CN202510639404.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, changes in data dependencies in data processing tasks cannot be captured in real time, resulting in long dependency update time and low recognition accuracy.
By obtaining structured query language SQL statements when detecting database change events, determining explicit and implicit dependencies between multiple fields, and updating dependencies in the original knowledge graph, optimizing data processing tasks.
Real-time update of dependencies between data is achieved, shortening the delay time of dependencies, and improving the accuracy and efficiency of dependency recognition.
Smart Images

Figure CN120470011A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data analysis technology, and in particular to a data processing method and device based on real-time change perception. Background Art
[0002] In related technologies, batch processing technology is used to identify the dependencies between data in data processing tasks, which has delays of minutes and lacks fine-grained mapping capabilities at the field level. Therefore, there are problems such as the inability to capture changes in data dependencies in real time, long dependency update time, and low recognition accuracy.
[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0004] The embodiments of the present application provide a data processing method and device based on real-time change perception, so as to at least solve the technical problem of long dependency update delay time caused by the inability of related technologies to capture changes in data dependencies in data processing tasks in real time.
[0005] According to one aspect of an embodiment of the present application, a data processing method based on real-time change perception is provided, including: when a change event is detected in a target database, obtaining a structured query language SQL statement corresponding to the change event, wherein the change event includes: a data definition language event and a data manipulation language event; determining the dependency relationship between multiple fields in the SQL statement; updating the structure related to the change event in the original knowledge graph according to the dependency relationship, wherein the original knowledge graph stores multiple entities in the target database and the dependency relationship between the multiple entities; optimizing the original task according to the update result of the updated original knowledge graph, wherein the original task is the data processing task to which the entity on which the change event is executed belongs.
[0006] Optionally, determining the dependency relationship between multiple fields in an SQL statement includes: performing grammatical analysis on the SQL statement to generate a syntax tree, and extracting explicit dependency relationships between multiple fields in the SQL statement from the syntax tree, wherein the explicit dependency relationships are recorded in the SQL statement in the form of fields; and using a relationship recognition model to process and analyze the SQL statement to obtain a recognition result output by the relationship recognition model, wherein the recognition result is used to indicate an implicit dependency relationship between multiple fields in the SQL statement, wherein the implicit dependency relationship is not recorded in the SQL statement in the form of fields, and the relationship recognition model is trained using multiple historical SQL statements with known implicit dependencies as training data.
[0007] Optionally, performing grammatical analysis on the SQL statement to generate a grammatical tree includes: performing grammatical analysis on the SQL statement to obtain grammatical elements in the SQL statement, wherein the grammatical elements are used to indicate fields in the SQL statement used to indicate data operations, and the grammatical elements include: keywords, operators, function calls, and conditional expressions; analyzing the logical relationship and data flow between multiple grammatical elements; generating nodes according to the grammatical elements, generating edges according to the logical relationship and data flow, and generating a grammatical tree according to the nodes and edges.
[0008] Optionally, a relational recognition model is used to process and analyze the SQL statement to obtain an identification result output by the relational recognition model, including: identifying dynamic fields in the SQL statement, where a dynamic field is a field in the SQL statement that contains dynamic elements, and a dynamic element is a field in the SQL statement that has a variable value; determining the context of the dynamic field in the SQL statement, and determining the logical relationship between the dynamic field and the context based on the context; adding a completion field representing the logical relationship to the SQL statement to obtain a completion result; processing and analyzing the completion result, and outputting a recognition result.
[0009] Optionally, the completion results are processed and analyzed to output recognition results, including: identifying multiple implicit dependency relationships corresponding to each field in the completion results; for each field, determining a confidence score for each implicit dependency relationship, wherein the confidence score is used to indicate the probability of the existence of the implicit dependency relationship, and the probability is positively correlated with the confidence score; and determining the target implicit dependency relationship corresponding to the confidence score with the largest numerical value as the recognition result.
[0010] Optionally, the structure related to the change event in the original knowledge graph is updated according to the dependency relationship, wherein the structure related to the change event is determined by the following method: determining the change object and change status of the change event according to the SQL statement, wherein the change status is used to indicate the operation performed on the change object, and the change status includes: addition, deletion, and modification; locating the change object in the original bitmap corresponding to the original knowledge graph, and determining the original state of the change object in the original bitmap; generating a change bitmap according to the change status and original state corresponding to the change object, wherein the change bitmap is used to reflect the state change of the change object; performing bit operations on the change bitmap and the original bitmap to obtain the operation results; locating the structure related to the change event according to the operation results.
[0011] Optionally, the original task is optimized according to the update results of the original knowledge graph, including: identifying abnormal nodes in the update results, wherein the abnormal nodes include: isolated nodes, circular dependency nodes; determining upstream nodes and downstream nodes related to the abnormal nodes, and querying multiple target structures containing abnormal nodes, upstream nodes and downstream nodes in the original knowledge graph; determining the performance indicators of each target structure, and updating the original task according to the task corresponding to the optimal performance indicator, wherein the performance indicators include: execution time.
[0012] According to another aspect of an embodiment of the present application, a data processing device based on real-time change perception is also provided, including: an acquisition module, used to obtain a structured query language SQL statement corresponding to a change event when a change event is detected in a target database, wherein the change event includes: a data definition language event, a data operation language event; a determination module, used to determine the dependency relationship between multiple fields in the SQL statement; an update module, used to update the structure related to the change event in the original knowledge graph according to the dependency relationship, wherein the original knowledge graph stores multiple entities in the target database and the dependency relationship between the multiple entities; an optimization module, used to optimize the original task according to the update result of the updated original knowledge graph, wherein the original task is the data processing task to which the entity on which the change event is executed belongs.
[0013] According to another aspect of an embodiment of the present application, a non-volatile storage medium is further provided, in which a computer program is stored. The device where the non-volatile storage medium is located executes the above-mentioned data processing method based on real-time change perception by running the computer program.
[0014] According to another aspect of an embodiment of the present application, an electronic device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the data processing method based on real-time change perception as claimed above through the computer program.
[0015] According to another aspect of an embodiment of the present application, a computer program product is further provided, comprising computer instructions, which, when executed by a processor, implement the steps of the above-mentioned data processing method based on real-time change perception.
[0016] In an embodiment of the present application, when a change event is detected in a target database, a structured query language SQL statement corresponding to the change event is obtained, wherein the change event includes: a data definition language event and a data manipulation language event; the dependency relationship between multiple fields in the SQL statement is determined; the structure related to the change event in the original knowledge graph is updated according to the dependency relationship, wherein the original knowledge graph stores multiple entities in the target database and the dependency relationship between the multiple entities; the original task is optimized according to the update result of the updated original knowledge graph, wherein the original task is a data processing task to which the entity on which the change event is executed belongs. By capturing the change event of the database in real time and updating the dependency knowledge graph corresponding to the database according to the dependency relationship of the data in the change event, the purpose of updating the dependency relationship between data in the database in real time is achieved, thereby realizing the technical effect of shortening the delay in updating the dependency relationship between data, and further solving the technical problem of long dependency update delay time caused by the inability to capture the changes in data dependency relationships in data processing tasks in real time in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0018] Figure 1 is a hardware structure block diagram of a computer terminal for implementing a data processing method based on real-time change perception according to an embodiment of the present application;
[0019] Figure 2 is a flowchart of the steps of a data processing method based on real-time change perception according to an embodiment of the present application;
[0020] Figure 3 This is a schematic diagram of extracting dependency relationships in SQL statements according to an embodiment of the present application;
[0021] Figure 4 is a structural diagram of a data processing device based on real-time change perception according to an embodiment of the present application;
[0022] Figure 5 This is a workflow diagram of a data processing device based on real-time change perception according to an embodiment of the present application. DETAILED DESCRIPTION
[0023] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0025] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are explained as follows:
[0026] Bitmap: A data structure used to efficiently store and retrieve large numbers of Boolean values, often used to represent sets or as part of indexing techniques. In a bitmap, each bit represents the state of an element, such as 1 for presence or true, and 0 for absence or false.
[0027] In related technologies, low-code platforms quickly generate data processing tasks (such as ETL and report generation) through visual configuration, but the data dependencies in the data processing tasks generated by them are difficult to trace due to the following problems: 1) Manual maintenance lag: Data dependencies are identified through manual labeling, which cannot adapt to dynamic task changes; 2) Offline analysis limitations: Analysis of data dependencies based on logs or snapshots has poor analysis timeliness and cannot capture real-time changes; 3) Complex logic blind spots: The analysis accuracy for scenarios such as nested Structured Query Language (SQL) statements and dynamic table associations is insufficient, resulting in a low accuracy rate in identifying dependencies. In order to solve the above problems, relevant solutions are provided in the embodiments of this application, which are described in detail below.
[0028] According to an embodiment of the present application, a method embodiment of a data processing method based on real-time change perception is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0029] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG1 shows a hardware structure block diagram of a computer terminal for implementing a data processing method based on real-time change perception. Figure 1 As shown, the computer terminal 10 may include one or more (illustrated as 102a, 102b, ..., 102n in the figure) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0030] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10. As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0031] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the data processing method based on real-time change perception in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, realizing the above-mentioned data processing method based on real-time change perception. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0032] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0033] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 .
[0034] The embodiment of the present application provides a data processing method based on real-time change perception that can be run in the above operating environment. Figure 2 is a flowchart of the steps of the data processing method based on real-time change perception provided in an embodiment of the present application, such as Figure 2 As shown, the method includes the following steps:
[0035] Step S202 : when a change event is detected in the target database, obtaining a structured query language SQL statement corresponding to the change event, wherein the change event includes: a data definition language event and a data manipulation language event.
[0036] The method provided in the embodiment of the present application monitors the change events in the database in real time, and triggers the update of the dependency relationship in the database according to the occurrence of the change event. In step S202, when a change event occurs in the database being detected / monitored (i.e., the target database), an SQL statement for describing the specific operation of the change event is obtained, wherein the change event of the database is divided into two types: one is a data definition language event (Data Definition Language, DDL) indicating that the table structure in the database has changed, for example, creating a data table, modifying a data table, deleting a data table, etc.; the other is a data manipulation language event (Data Manipulation Language, DML) indicating that the row data in the data has changed, for example, inserting row data, updating row data, deleting row data. The above-mentioned specific operation SQL statement for describing the change event can be obtained from the database log or data management system. The change event is the effect of the execution of the SQL statement, and the specific operation of the change event is achieved by executing the SQL statement description. In this embodiment, detecting / detecting whether there is a change event in the target database can be achieved by deploying a continuous data capture component (Flink-CDC) in the target database, and the Flink-CDC component can detect and capture the change events of the target database in real time; the target database is, for example, a relational database (MySQL) or an Oracle database.
[0037] Step S204: Determine the dependency relationship between multiple fields in the SQL statement.
[0038] In the solution provided in the embodiment of the present application, the SQL statement is the object that the system needs to parse, and the SQL statement contains the logic and rules of data processing. After obtaining the SQL statement corresponding to the change event in step S202, in step S204, the dependency relationship between multiple fields contained in the SQL statement is obtained by parsing the SQL statement, wherein each field in the SQL statement represents a specific operation to be performed: a specific operation performed on an object / entity in the database; the dependency relationship between multiple fields contained in the SQL statement is used to represent the dependency relationship between different specific operations (for example, the execution order of different specific operations).
[0039] Optionally, determining the dependency relationship between multiple fields in an SQL statement includes: performing grammatical analysis on the SQL statement to generate a syntax tree, and extracting explicit dependency relationships between multiple fields in the SQL statement from the syntax tree, wherein the explicit dependency relationships are recorded in the SQL statement in the form of fields; and using a relationship recognition model to process and analyze the SQL statement to obtain a recognition result output by the relationship recognition model, wherein the recognition result is used to indicate an implicit dependency relationship between multiple fields in the SQL statement, wherein the implicit dependency relationship is not recorded in the SQL statement in the form of fields, and the relationship recognition model is trained using multiple historical SQL statements with known implicit dependencies as training data.
[0040] In an embodiment of the present application, when identifying the dependency of multiple fields in an SQL statement, both explicit dependencies and implicit dependencies are identified, wherein the explicit dependencies are recorded in the SQL statement in the form of fields. For example, when performing a join (JOIN) operation, the relationship between the fields of the two tables explicitly specified in the join condition, at this time, the explicit dependency is recorded in the field recording the join condition; for another example, in a data loading and conversion task, the relationship between the fields directly selected from the source table and mapped to the target table, at this time, the explicit dependency is recorded in the field recording the mapping relationship. It can be identified through grammatical analysis, while implicit dependencies are not recorded in the SQL statement in the form of fields and cannot be identified through simple grammatical analysis; therefore, in this embodiment, different methods are used to identify explicit dependencies and implicit dependencies. When identifying explicit dependencies, a grammatical analysis method is used to generate a syntax tree (AST) corresponding to the SQL statement by performing grammatical parsing on the SQL statement. Further, the explicit dependencies in the SQL statement can be identified by traversing the syntax tree. For example, a SQL statement parser (Druid parser) is used to perform syntax analysis on SQL statements in data extraction, transformation, and loading (ETL) tasks to generate an abstract syntax tree (AST). By traversing the abstract syntax tree (AST), the system can directly identify explicitly mentioned dependencies between fields (i.e., explicit dependencies). For example, when a join operation indicates that the "order details table's details ID" (orders.orderId) field and the "order table's order ID" (orderdetails.detailId) field are to be joined, the direct association between the orders.orderId field and the orderdetails.detailId field in the join operation can be directly identified by traversing the syntax tree. When identifying implicit dependencies, a pre-trained relationship recognition model is applied. The relationship recognition model is trained using multiple historical SQL statements with known implicit dependencies as training data. Therefore, when applied, using the relationship recognition model to process SQL statements corresponding to change events can directly extract implicit dependencies in SQL statements.
[0041] According to some optional embodiments of the present application, a grammatical analysis is performed on an SQL statement to generate a grammatical tree, including: performing a grammatical analysis on the SQL statement to obtain grammatical elements in the SQL statement, wherein the grammatical elements are used to indicate fields in the SQL statement used to indicate data operations, and the grammatical elements include: keywords, operators, function calls, and conditional expressions; analyzing the logical relationship and data flow between multiple grammatical elements; generating nodes according to the grammatical elements, generating edges according to the logical relationship and data flow, and generating a grammatical tree according to the nodes and edges.
[0042] Figure 3 It is a schematic diagram for extracting dependencies in SQL statements, such as Figure 3 As shown, extracting dependencies in SQL statements includes extracting explicit dependencies and extracting implicit dependencies, wherein extracting explicit dependencies includes two steps: generating a syntax tree and extracting explicit dependencies. In this embodiment, a syntax tree is generated based on analyzing the SQL statement. Specifically, the SQL statement is parsed to obtain syntax elements in the SQL statement; syntax elements include: keywords (such as select (ELECT), from (FROM), join (JOIN), condition (WHERE)), operators (such as equal sign "=", function calls, conditional expressions (such as A = 'B'), etc., which are used to describe data operations. After identifying all syntax elements, the logical relationship and data flow between the syntax elements are analyzed. For example, in the example used in the above embodiment, the JOIN operation defines the association between the fields orders.orderId and orderdetails.orderId, indicating that data flows between these fields. Furthermore, based on the syntax elements and the logical relationships between them, each syntax element is treated as a node, and edges are generated based on the data flow or other logical relationships between syntax elements to construct a syntax tree for the SQL statement. In the syntax tree, nodes may include table names, field names, function names, etc., and edges represent the data flow or the flow of logical operations between fields. For example, when the SQL statement contains the fields SELECT orders.orderId,orderdetails.detailId FROM orders JOINorderdetails ON orders.orderId=orderdetails.orderId (selecting data from the "orders" table and the "orderdetails" table, with the selected fields including orders.orderId: (the order ID field of the orders table) and orderdetails.detailId: (the detail ID field of the order details table)), the dependency between orders.orderId and orderdetails.detailId is explicitly recorded in the JOIN operation. This dependency is directly reflected in the SQL statement through the field names. In the syntax tree, an edge exists between the node representing orders.orderId and the node representing orderdetails.detailId.
[0043] According to other optional embodiments of the present application, a relationship recognition model is used to process and analyze SQL statements to obtain recognition results output by the relationship recognition model, including: identifying dynamic fields in SQL statements, wherein a dynamic field is a field in an SQL statement that contains dynamic elements, and a dynamic element is a field in an SQL statement that has a variable value; determining the context of the dynamic field in the SQL statement, and determining the logical relationship between the dynamic field and the context based on the context; adding a completion field representing the logical relationship to the SQL statement to obtain a completion result; processing and analyzing the completion result, and outputting a recognition result.
[0044] Still Figure 3 As shown, extracting implicit dependencies from SQL statements is implemented based on a relationship identification model. In this embodiment, when using the relationship identification model to extract implicit dependencies from SQL statements, the relationship identification model specifically processes the SQL statements as follows. The relationship identification model first identifies dynamic fields that may exist in the SQL statement. Dynamic fields refer to fields that contain dynamic elements in the SQL statement, such as variables, functions, parameters, and other fields. The values of these fields may change in different execution environments. For example, the SQL statement is as follows: SELECT orders.orderId,orderdetails.detailId,#Select the order ID, select the order details ID#CASE WHEN@customerType='Premium'THEN customers.firstName#Dynamically select the customer name based on the client type. If the customer type is a premium member, select the customer name#ELSE NULL#Otherwise, return null#. In the above SQL statement, @customerType is identified as a dynamic field. Furthermore, the relationship recognition model determines the context of the dynamic field in the SQL statement, that is, the context in which the dynamic field exists. Based on the context of the dynamic field, it determines its logical relationship with the fields contained in the context and with other static fields. Finally, based on the context and logical relationship of the dynamic field, it generates a completion field and adds it to the original SQL statement to obtain the completed SQL statement (i.e., the completion result). The completion field is used to describe specific operations that exist in the SQL statement but are not recorded as fields. By further analyzing the completion result, implicit dependencies in the SQL statement are identified.
[0045] In this embodiment, the relationship recognition model can be loaded into the memory. For example, the raw data of the relationship recognition model can be loaded from the non-volatile memory into the volatile memory, so that the processor can run the relationship recognition model. The raw data of the relationship recognition model refers to unprocessed data, and generally includes parameters and structural data of the relationship recognition model. The structural data can be a calculation relationship based on the parameters, such as the forward propagation calculation relationship between intermediate layers or between neurons. Specifically, the structural data can include code related to the structure of the relationship recognition model, such as code for performing related calculations between intermediate layers or between neurons.
[0046] In one embodiment, a memory area for loading the relationship recognition model can be divided into a structure data storage area and a parameter storage area. The structure data storage area is used to store structure-related code, and the parameters referenced by the structure data storage area can point to the addresses of specific parameters in the parameter storage area through pointers. During the relationship recognition model training process, parameters may need to be frequently updated, and the parameter values in the parameter storage area can be simply updated.
[0047] Optionally, the completion results are processed and analyzed to output recognition results, including: identifying multiple implicit dependency relationships corresponding to each field in the completion results; for each field, determining a confidence score for each implicit dependency relationship, wherein the confidence score is used to indicate the probability of the existence of the implicit dependency relationship, and the probability is positively correlated with the confidence score; and determining the target implicit dependency relationship corresponding to the confidence score with the largest numerical value as the recognition result.
[0048] After obtaining the completion result in the previous embodiment, the completion result is further processed to extract implicit dependencies. For example, in this embodiment, the completed SQL statement is first analyzed in depth to identify which tables and fields each field is directly connected to, and also to dig out potential implicit dependencies. If a field is identified to be associated with multiple implicit dependencies, the relationship identification model will calculate the confidence score of each dependency. The confidence score reflects the probability of the existence of the implicit dependency. The higher the score, the more likely the dependency is to exist. When performing the confidence score, various factors such as the degree of matching of the context grammar are considered. In this embodiment, for all implicit dependencies of each field, the dependency with the highest confidence score is selected as the target implicit dependency, and this result is output. This means that the system will give priority to those dependencies that are most likely to exist and have the greatest impact on data lineage.
[0049] In the solution provided in the embodiment of the present application, the explicit dependency relationship obtained by grammatical analysis and the implicit dependency relationship output by the relationship recognition model can be directly output as the final result of identifying the dependency relationship, or the implicit dependency relationship output by the relationship recognition model can be directly output as the final result of identifying the dependency relationship. Figure 3The explicit dependency obtained by grammatical parsing and the implicit dependency output by the relationship recognition model are further processed in the manner shown, and the processing results are output as the final result of the dependency recognition. Figure 3 Further processing in the manner shown can improve the accuracy of dependency identification results. Figure 3 As shown, the explicit dependencies obtained by grammatical parsing and the implicit dependencies output by the relationship recognition model are further processed, including: filtering noise relationships based on the rule engine and the confidence score (0-1), wherein the engine rule is a preset rule set for filtering noise relationships, especially for temporary tables that appear in test environments or non-production environments, wherein the engine rule includes: filtering based on environment tags: for example, deleting dependencies with the prefix test (test) or temporary (tmp), filtering based on usage frequency: for example, deleting dependencies with a usage frequency less than a preset frequency. Furthermore, the dependencies that pass the engine verification are filtered again using the confidence score (0-1), and the confidence score of each dependency that passes the engine verification is calculated. The dependencies corresponding to the confidence score greater than the preset score value (such as 0.8) are output as recognition results (i.e., valid dependencies), and the dependencies corresponding to the confidence score less than the preset score value (such as 0.8) are deleted.
[0050] Step S206: Update the structure related to the change event in the original knowledge graph according to the dependency relationship, wherein the original knowledge graph stores multiple entities in the target database and the dependency relationships between the multiple entities.
[0051] When a change event is detected in the target database, in step S206, the original knowledge graph of the target database is updated according to the dependency relationship recorded in the SQL statement that records the specific operation of the change event, wherein, in this embodiment, when updating the original knowledge graph, an incremental update method is adopted, that is, only the structure related to the change event in the original knowledge graph is updated. In this embodiment, the original knowledge graph is used to record all entities and all dependencies in the target database, wherein an entity refers to a data table in the database and a field that exists independently of the data table, and the dependency relationship includes: the flow direction of data between different data represented by different entities, conversion logic, etc.
[0052] According to some optional embodiments of the present application, the structure related to the change event in the original knowledge graph is updated according to the dependency relationship, wherein the structure related to the change event is determined by the following method: determining the change object and change status of the change event according to the SQL statement, wherein the change status is used to indicate the operation performed on the change object, and the change status includes: addition, deletion, and modification; locating the change object in the original bitmap corresponding to the original knowledge graph, and determining the original state of the change object in the original bitmap; generating a change bitmap according to the change status and original state corresponding to the change object, wherein the change bitmap is used to reflect the state change of the change object; performing bit operations on the change bitmap and the original bitmap to obtain the operation results; locating the structure related to the change event according to the operation results.
[0053] As mentioned in the above embodiment, this application adopts an incremental update method for the original knowledge graph, that is, only the structure related to the change event in the original knowledge graph is updated. Therefore, in this embodiment, the structure related to the change event must be determined first. Specifically, first, the SQL statement is parsed to determine the change object and change status of the change event, where the change object can be a data table, a field in an independent data table, or a specific data record; the change status reflects the operation performed on the change object, including addition, deletion, and modification. In this embodiment, the storage form of the original knowledge graph is a compressed bitmap, for example, a bitmap compressed using a state compression algorithm (RoaringBitmap). Therefore, the state of the changed object before the change event is executed (i.e., the original state) can be determined by locating the changed object in the bitmap corresponding to the original knowledge graph (i.e., the original bitmap); wherein, the bitmap is an efficient data structure used to represent the state of a table or field in the graph, such as whether the field has been read or modified by a task; for example, if the original bitmap stores the read and write status of each field of table A, the system will locate the original bitmap representation of the changed field B (i.e., the changed field) and determine its original state (e.g., whether it has been read, modified, or deleted). Further, based on the change state of the change object determined above and the original state of the change object, a new change bitmap (i.e., the change bitmap) is generated. This is an incremental update that only represents the change in field state since the last update. For example, if the original state of field B in the previous example is unmodified, then the generated change bitmap will reflect that the field has changed from unmodified to modified. Next, perform bit operations (such as bitwise AND, bitwise OR, bitwise XOR, etc.) on the original bitmap and the changed bitmap to obtain an operation result bitmap. For example, using the bitwise OR operation, the original bitmap is merged with the changed bitmap to update the modified state of field B in the previous example. Ultimately, the structure related to the change event in the original knowledge graph can be determined based on the results of the above bit operations. For each edge in the knowledge graph (representing a dependency relationship), the bitmap status of its source and target nodes (i.e., fields) is checked. If the state has changed, the structure composed of this edge and the nodes related to this edge (including the source node and the target node) is the structure related to the change event.
[0054] Step S208: Optimize the original task according to the updated result of the updated original knowledge graph, wherein the original task is the data processing task to which the entity on which the change event is executed belongs.
[0055] In step S208, after the original knowledge graph corresponding to the target database is incrementally updated according to the change event, task optimization is performed based on the updated original knowledge graph (i.e., the update result). Specifically, the data processing task to which the execution object of the change event belongs is optimized, wherein the execution object of the change event (i.e., the entity on which the change event is executed) is the field or data table in the database on which the specific operation corresponding to the change event is executed.
[0056] According to some other optional embodiments of the present application, the original task is optimized according to the update results of the original knowledge graph, including: identifying abnormal nodes in the update results, wherein the abnormal nodes include: isolated nodes, circular dependency nodes; determining upstream nodes and downstream nodes related to the abnormal nodes, and querying multiple target structures containing abnormal nodes, upstream nodes and downstream nodes in the original knowledge graph; determining the performance indicators of each target structure, and updating the original task according to the task corresponding to the optimal performance indicator, wherein the performance indicators include: execution time.
[0057] When performing task optimization based on the update results of the original knowledge graph in step S208, the update results of the original knowledge graph are first analyzed to identify abnormal nodes. In this embodiment, abnormal nodes include isolated nodes, circular dependency nodes, etc., where an isolated node is a node without upstream or downstream relationships, and a circular dependency node is a plurality of nodes that form a closed-loop dependency. If there are abnormal nodes, the tasks containing the abnormal nodes are optimized. If there are no abnormal nodes, there is no need to optimize the tasks. In this embodiment, the update results of the original knowledge graph can be deeply analyzed by graph neural network (GNN) technology to identify abnormal nodes. Once the abnormal nodes are identified, the upstream nodes and downstream nodes related to these abnormal nodes are further determined, where the upstream nodes refer to the source of the abnormal nodes, and the downstream nodes are the subsequent data processing tasks or tables that may be affected by the abnormal nodes. In this embodiment, there are abnormal nodes in the update results, indicating that the execution mode of the corresponding task is abnormal after the change event. In order to optimize the tasks corresponding to the abnormal nodes, multiple target structures containing the abnormal nodes, upstream nodes and downstream nodes are queried in the original knowledge graph. This can be achieved through the graph query language of the knowledge graph (such as Cypher), thereby locating the position of the abnormal node in the entire data dependency graph (i.e., the original knowledge graph) and its impact range. For example, for a ring-shaped dependency node, the system may query all paths connected to it to determine which tasks or data flows are affected. Next, the performance indicators of the structure composed of the upstream nodes and downstream nodes of the abnormal node machine (i.e., the target structure) are evaluated. The performance indicators can be analyzed based on historical task execution data. The performance indicators include: task resource consumption, execution time, response time, etc.; by analyzing the performance indicators of the target structure, the links with low execution efficiency are identified; for example, by analyzing the operation log of the ETL task, the average execution time of each task can be calculated, and the impact of the existence of loop dependencies or isolated nodes on task performance can be evaluated. Finally, based on the performance metric evaluation results, optimization recommendations are generated for the tasks associated with the abnormal nodes. These recommendations may include: Handling isolated nodes: For isolated nodes, it is recommended to delete or restructure related tasks, removing unnecessary data read operations and reducing resource waste. Resolving circular dependencies: For nodes with circular dependencies, the system recommends redesigning the data flow to break the closed loop. This may involve adjusting the task sequence or data processing logic to avoid redundant computations.
[0058] Figure 4 is a structural diagram of a data processing device based on real-time change perception provided in an embodiment of the present application, such as Figure 4As shown, the data processing device based on real-time change perception includes: an acquisition module 40, which is used to obtain a structured query language SQL statement corresponding to a change event when a change event is detected in a target database, wherein the change event includes: a data definition language event and a data operation language event; a determination module 42, which is used to determine the dependency relationship between multiple fields in the SQL statement; an update module 44, which is used to update the structure related to the change event in the original knowledge graph according to the dependency relationship, wherein the original knowledge graph stores multiple entities in the target database and the dependency relationship between the multiple entities; an optimization module 46, which is used to optimize the original task according to the update result of the updated original knowledge graph, wherein the original task is the data processing task to which the entity on which the change event is executed belongs.
[0059] Figure 5 It is a workflow diagram of a data processing device based on real-time change perception, such as Figure 5 As shown, a data processing device based on real-time change perception is used to detect a relational database (MySQL). When the acquisition module 40 detects that there is a change event in the database MySQL, it obtains the SQL statement corresponding to the change event, and identifies the dependency relationship in the SQL statement through the determination module 42, wherein the determination module 42 simultaneously identifies the explicit dependency relationship and the implicit dependency relationship in the SQL statement; the update module 44 updates the original knowledge graph stored in the graph database for recording all entities (data tables, fields independent of data tables) in MySQL, that is, the dependency relationship between all entities according to the dependency relationship identified by the determination module 42; the update result of the original knowledge graph is detected by the optimization module 46 to determine whether it is necessary to optimize the data processing task to which the entity to which the change event is executed belongs; and when it is necessary to optimize the data processing task to which the entity to which the change event is executed belongs, optimization processing is performed and the optimization result is output; wherein the optimization result can also be displayed through a visual interface.
[0060] It should be noted that Figure 4 The preferred implementation of the embodiment shown can be found in Figure 2 The relevant description of the illustrated embodiment will not be repeated here.
[0061] An embodiment of the present application further provides a non-volatile storage medium, in which a computer program is stored. The device where the non-volatile storage medium is located executes the above-mentioned data processing method based on real-time change perception by running the computer program.
[0062] The above-mentioned non-volatile storage medium is used to store a program that performs the following functions: when a change event is detected in a target database, obtaining a structured query language SQL statement corresponding to the change event, wherein the change event includes: a data definition language event and a data manipulation language event; determining the dependency relationship between multiple fields in the SQL statement; updating the structure related to the change event in the original knowledge graph according to the dependency relationship, wherein the original knowledge graph stores multiple entities in the target database and the dependency relationship between the multiple entities; optimizing the original task according to the update result of the updated original knowledge graph, wherein the original task is the data processing task to which the entity on which the change event is executed belongs.
[0063] According to another aspect of an embodiment of the present application, an electronic device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the data processing method based on real-time change perception as claimed above through the computer program.
[0064] The processor in the above-mentioned electronic device is used to run a program that performs the following functions: when a change event is detected in a target database, obtaining a structured query language SQL statement corresponding to the change event, wherein the change event includes: a data definition language event and a data manipulation language event; determining the dependency relationship between multiple fields in the SQL statement; updating the structure related to the change event in the original knowledge graph according to the dependency relationship, wherein the original knowledge graph stores multiple entities in the target database and the dependency relationship between the multiple entities; optimizing the original task according to the update result of the updated original knowledge graph, wherein the original task is the data processing task to which the entity on which the change event is executed belongs.
[0065] An embodiment of the present application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the above data processing method based on real-time change perception.
[0066] It should be noted that the various modules in the above-mentioned data processing device based on real-time change perception can be program modules (for example, a set of program instructions that implement a certain specific function) or hardware modules. For the latter, it can be expressed in the following forms, but is not limited to this: the expression form of each of the above-mentioned modules is a processor, or the functions of each of the above-mentioned modules are implemented by a processor.
[0067] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0068] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0069] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0070] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0071] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0072] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the relevant technology or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0073] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A data processing method based on real-time change perception, characterized in that: include: When a change event is detected in the target database, obtaining a structured query language SQL statement corresponding to the change event, wherein the change event includes: a data definition language event and a data manipulation language event; Determine dependencies between multiple fields in the SQL statement; updating a structure related to the change event in an original knowledge graph according to the dependency relationship, wherein the original knowledge graph stores a plurality of entities in a target database and the dependency relationships between the plurality of entities; The original task is optimized according to the updated result of the updated original knowledge graph, wherein the original task is a data processing task to which the entity on which the change event is executed belongs.
2. The method according to claim 1, characterized in that Determining the dependency relationship between multiple fields in the SQL statement includes: Performing grammatical analysis on the SQL statement to generate a syntax tree, extracting explicit dependency relationships between multiple fields in the SQL statement from the syntax tree, wherein the explicit dependency relationships are recorded in the SQL statement in the form of fields; and A relationship recognition model is used to process and analyze the SQL statement to obtain a recognition result output by the relationship recognition model, wherein the recognition result is used to indicate an implicit dependency relationship between multiple fields in the SQL statement, wherein the implicit dependency relationship is not recorded in the SQL statement in the form of a field, and the relationship recognition model is trained using historical SQL statements with multiple known implicit dependencies as training data.
3. The method according to claim 2, characterized in that Performing grammatical analysis on the SQL statement to generate a syntax tree includes: Performing grammatical analysis on the SQL statement to obtain grammatical elements in the SQL statement, wherein the grammatical elements are used to indicate fields in the SQL statement used to indicate data operations, and the grammatical elements include: keywords, operators, function calls, and conditional expressions; Analyzing the logical relationship and data flow between the plurality of grammatical elements; Nodes are generated according to the syntax elements, edges are generated according to the logical relationships and the data flow, and the syntax tree is generated according to the nodes and the edges.
4. The method according to claim 2, characterized in that The SQL statement is processed and analyzed using a relational recognition model to obtain a recognition result output by the relational recognition model, including: Identifying a dynamic field in the SQL statement, wherein the dynamic field is a field in the SQL statement that contains a dynamic element, and the dynamic element is a field in the SQL statement that has a variable value; determining a context of the dynamic field in the SQL statement, and determining a logical relationship between the dynamic field and the context according to the context; Adding the completion field representing the logical relationship to the SQL statement to obtain a completion result; The completion result is processed and analyzed, and the recognition result is output.
5. The method according to claim 4, characterized in that Processing and analyzing the completion result and outputting the recognition result includes: Identifying multiple implicit dependencies corresponding to each field in the completion result; For each field, determine the confidence score of each implicit dependency, wherein the confidence score is used to indicate the probability of the existence of the implicit dependency, and the probability is positively correlated with the confidence score; and determine the target implicit dependency corresponding to the confidence score with the largest value as the recognition result.
6. The method according to claim 1, wherein The structure related to the change event in the original knowledge graph is updated according to the dependency relationship, wherein the structure related to the change event is determined by the following method: Determining a change object and a change state of the change event according to the SQL statement, wherein the change state is used to indicate an operation performed on the change object, and the change state includes: adding, deleting, and modifying; Locating the changed object in the original bitmap corresponding to the original knowledge graph, and determining the original state of the changed object in the original bitmap; generating a change bitmap according to the change state and the original state corresponding to the change object, wherein the change bitmap is used to reflect the state change of the change object; Performing a bit operation on the changed bitmap and the original bitmap to obtain an operation result; A structure related to the change event is located according to the operation result.
7. The method according to claim 1, characterized in that Optimizing the original task according to the updated result of the original knowledge graph includes: Identifying abnormal nodes in the update result, wherein the abnormal nodes include: isolated nodes and ring-dependent nodes; Determine upstream nodes and downstream nodes related to the abnormal node, and query the original knowledge graph for multiple target structures including the abnormal node, the upstream node, and the downstream node; Determine the performance index of each target structure, and update the original task according to the task corresponding to the optimal performance index, wherein the performance index includes: execution time.
8. A data processing device based on real-time change perception, characterized in that: include: an acquisition module, configured to acquire, upon detecting a change event in a target database, a structured query language SQL statement corresponding to the change event, wherein the change event includes: a data definition language event and a data manipulation language event; A determination module, configured to determine dependencies between multiple fields in the SQL statement; An updating module, configured to update a structure related to the change event in an original knowledge graph according to the dependency relationship, wherein the original knowledge graph stores a plurality of entities in a target database and the dependency relationships between the plurality of entities; An optimization module is used to optimize the original task according to the updated result of the original knowledge graph after the update, wherein the original task is the data processing task to which the entity on which the change event is executed belongs.
9. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a computer program, wherein the device where the non-volatile storage medium is located executes the data processing method based on real-time change perception as described in any one of claims 1 to 7 by running the computer program.
10. An electronic device comprising a memory and a processor, characterized in that: The memory stores a computer program, and the processor is configured to execute the data processing method based on real-time change perception according to any one of claims 1 to 7 through the computer program.
11. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the steps of the data processing method based on real-time change perception described in any one of claims 1 to 7 are implemented.