Data task processing method and device, equipment and medium
By constructing a directed dependency network structure through grammatical parsing and semantic pairing of data tasks, the problem of opaque logical transmission paths in traditional data task processing is solved, automatic identification and visualization of data lineage relationships are achieved, and data governance efficiency and system control capabilities are improved.
Patent Information
- Application Number
- CN202510856399.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-10
AI Technical Summary
Traditional data task processing methods have difficulty accurately restoring the logical transfer path between data task fields in complex nested statements and multi-task dependency scenarios, and lack the support of visual lineage maps, which affects the efficiency of data anomaly backtracking and data quality verification.
By performing grammatical parsing and semantic pairing on the query statements in the data task, a directed dependency network structure is generated, and a blood relationship map is constructed to achieve automatic identification and visualization of the transfer relationship between input fields and output fields.
It improves data governance efficiency, enhances the system's overall control over the data flow process, supports data traceability, task optimization and risk control, and improves data management quality and decision-making efficiency in the financial and medical fields.
Smart Images

Figure CN120763192A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, and in particular to a data task processing method and device, equipment and medium. BACKGROUND
[0002] In the current data processing process of the financial and medical industries, with the increasing complexity of the system and the multi-source of the business, the traditional development architecture based on large mainframe platforms still has a large number of residues in some core systems. These legacy systems form a huge logical dependency and program call relationship when performing data extraction, transformation and loading tasks, making the data transmission chain between data tasks highly coupled, the structure opaque, and difficult to maintain.
[0003] However, the processing method of the data task in the related art often relies on manual writing of rules or fixed parsing templates, and cannot cope with complex nested statements, multi-task dependencies or dynamic script execution scenarios. Traditional tools are difficult to accurately restore the logical transmission path between data task fields. In addition, in the data exception backtracking and data quality checking process, there is a lack of visual blood relationship map support, which seriously affects the problem positioning efficiency. Therefore, there is an urgent need for a data task processing method that can adaptively identify structural dependencies and quickly build field transmission paths to meet the traceability requirements of data tasks in high-risk and high-sensitivity industries. SUMMARY
[0004] The present application provides a data task processing method, device, equipment and medium to solve the technical problem of highly coupled data transmission chain between data tasks, opaque structure and difficult to maintain in the related art.
[0005] In a first aspect, a data task processing method is provided, the method comprising:
[0006] obtaining a query statement in the data task, and parsing the query statement to obtain syntax element information corresponding to the query statement; wherein the syntax element information includes table name, field name, operation mode and connection condition;
[0007] performing semantic pairing on the query statement based on the language element information to obtain a set of transmission mapping pairs between input fields and output fields in the query statement;
[0008] According to the transmission mapping pair set, a transmission path between the input field and the output field is constructed to generate a directed dependency network structure;
[0009] mapping the directed dependency network structure to the task nodes of the data task to obtain a blood relationship map corresponding to the data task.
[0010] In a second aspect, a data task processing device is provided, comprising:
[0011] An acquisition module, configured to acquire a query statement in the data task and parse the query statement to obtain syntax element information corresponding to the query statement; wherein the syntax element information includes table name, field name, operation mode and connection condition;
[0012] a pairing module, configured to perform semantic pairing on the query statement based on the language element information to obtain a set of transfer mapping pairs between input fields and output fields in the query statement;
[0013] A construction module, configured to construct a transmission path between the input field and the output field according to the transfer mapping pair set to generate a directed dependency network structure;
[0014] A mapping module is used to map the directed dependency network structure to the task nodes of the data task to obtain a blood relationship map corresponding to the data task.
[0015] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned data task processing method when executing the computer program.
[0016] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned data task processing method are implemented.
[0017] In the scheme implemented by the data task processing method, the data task processing device, the computer device, and the storage medium, the method comprises: obtaining a query statement in the data task, and performing syntax analysis on the query statement to obtain syntax element information corresponding to the query statement; wherein the syntax element information comprises a table name, a field name, an operation mode, and a connection condition. Further, the query statement can be semantically paired based on the language element information to obtain a set of transfer mapping pairs between input fields and output fields in the query statement, and a transmission path between the input fields and the output fields is constructed according to the set of transfer mapping pairs to generate a directed dependency network structure. Finally, the directed dependency network structure is mapped between task nodes of the data task to obtain a blood relationship graph corresponding to the data task. In the present application, by performing syntax analysis and semantic pairing on the query statement in the data task, the transfer relationship between the input fields and the output fields can be efficiently and accurately extracted, and a clear directed dependency network structure can be constructed, thereby realizing automatic identification and visualization of data blood relationship. The method has high automation degree, strong scalability, and the ability to adapt to various complex query statements, and significantly improves the data governance efficiency. At the same time, by constructing the blood relationship graph, subsequent operations such as data tracing, task optimization, and risk control can be effectively supported, and the global control of the system on the data flow process is enhanced. In the financial field, the scheme can be applied to data tracking in complex financial report generation and risk control models to ensure traceability of key indicators, and improve audit compliance ability; in the medical field, it can be used in the processing flow of electronic medical records and medical test data to ensure transparent data transmission path and data use compliance, and help to improve medical data management quality and clinical decision-making efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0019] Figure 1 is a schematic diagram of an application environment of the data task processing method in an embodiment of the present application;
[0020] Figure 2 is a flowchart of the data task processing method in an embodiment of the present application;
[0021] Figure 3 is Figure 1 is a specific embodiment flowchart of step S10 in the method;
[0022] Figure 4 is a structural diagram of the data task processing device in an embodiment of the present application;
[0023] Figure 5 is a structural diagram of a computer device in one embodiment of the present invention;
[0024] Figure 6 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0026] The data task processing method provided by the embodiment of the present invention can be applied in the following situations: Figure 1In an application environment, the client communicates with the server through a network. The server can obtain the query statement in the data task through the client, and parse the query statement to obtain the grammatical element information corresponding to the query statement; wherein the grammatical element information includes table name, field name, operation method and connection condition; semantically pair the query statement based on the language element information to obtain a set of transfer mapping pairs between the input field and the output field in the query statement; construct a transmission path between the input field and the output field according to the transfer mapping pair set to generate a directed dependency network structure; map the directed dependency network structure to the task nodes of the data task to obtain the blood relationship map corresponding to the data task, and feed the blood relationship map back to the client. In the present invention, by grammatically parsing and semantically pairing the query statement in the data task, the transfer relationship between the input field and the output field can be efficiently and accurately extracted, and a clear directed dependency network structure can be constructed, thereby realizing automatic identification and visualization of data blood relationship. This method has a high degree of automation, strong scalability, and the ability to adapt to a variety of complex query statements, which significantly improves data governance efficiency. At the same time, by constructing a blood relationship map, it can effectively support subsequent operations such as data traceability, task optimization and risk control, and enhance the system's overall control over the data flow process. In the financial field, this solution can be applied to data tracking in complex financial report generation and risk control models to ensure that the source of key indicators can be traced and improve audit compliance capabilities; in the medical field, it can be used in the processing flow of electronic medical records and medical test data to ensure that the patient data transmission path is transparent and the data usage is compliant, which helps to improve the quality of medical data management and clinical decision-making efficiency. Among them, the client can be but is not limited to various personal computers, laptops, smart phones, tablets and portable wearable devices. The server can be implemented with an independent server or a server cluster composed of multiple servers. The present invention is described in detail below through specific embodiments.
[0027] See also Figure 2 As shown, Figure 2 A flowchart of a method for processing a data task provided by an embodiment of the present invention is provided. The method includes the following steps:
[0028] S10: Obtain a query statement in the data task, and parse the query statement to obtain grammatical element information corresponding to the query statement.
[0029] The syntax element information includes table name, field name, operation method and connection condition.
[0030] For example, the corresponding query statement can be first extracted from the data task, and then the query statement is parsed at the syntax level to identify the key components therein. Specifically, the parsing operation analyzes the structure and semantics of the query statement to extract syntax element information, including: table name for specifying data source, field name for selecting or filtering data, operation mode representing condition judgment or calculation mode, and connection condition describing the logical relationship between tables or fields.
[0031] The above syntax element information can provide basic support for subsequent data processing, execution optimization or semantic understanding, ensuring that the data task can be accurately identified and correctly executed.
[0032] As shown in FIG. 1, the method comprises the following steps: Figure 3 As shown in FIG. 1, the method comprises the following steps:
[0033] S11: Obtain the query statement in the data task, and analyze the query statement by a pre-trained language model to obtain a corresponding syntax expression unit set.
[0034] S12: Perform structural analysis on the syntax expression unit by using a dependency syntax tree to obtain a syntax structure index in the query statement.
[0035] S13: Perform field mapping on the syntax structure index to obtain syntax element information corresponding to the query statement.
[0036] In step S11, the specific query statement is first extracted from the data task, and the pre-trained language model is used to analyze the query statement at the semantic and syntax levels to generate a syntax expression unit set. For example, in the financial field, when receiving a data task such as "query the loan default customer information of the first quarter of 2024", the language model can identify the semantic components related to time, customer, loan and default, and mark them as respective syntax units; in the medical field, if the query statement is "get the medical records of patients who were admitted in 2023 and were diagnosed with hypertension", the key syntax units such as time condition, diagnosis field and patient entity can also be extracted therefrom, laying a foundation for subsequent structured analysis.
[0037] In steps S12 and S13, the syntax expression units can be further parsed by dependency syntax tree technology to construct a syntax structure index, thereby clarifying the dependency relationship and hierarchical structure between the components in the sentence. For example, in a financial scenario, by analyzing that “the first quarter” depends on “loan default customers”, it can be determined that “loan” and “default” are in a limiting relationship, and “customer” is the subject entity; in a medical application, the system can identify that “high blood pressure” is a limiting condition for “diagnosis”, and “admitted in 2023” is a time filtering condition. Then in step S13, by mapping these syntax structure indexes to the field level, the syntax element information including the table name (such as the customer information table and the medical record data table), the field name (such as the default state and the diagnosis result), the operation method (such as equal to and greater than), and the connection condition (such as the time interval and the diagnosis matching) is finally extracted, and the structured understanding of the query statement is realized.
[0038] S20: performing semantic pairing on the query statement based on the language element information to obtain a set of transfer mapping pairs between input fields and output fields in the query statement.
[0039] In some embodiments, the performing semantic pairing on the query statement based on the language element information to obtain a set of transfer mapping pairs between input fields and output fields in the query statement comprises: analyzing the language element information by a graph neural network to obtain corresponding node semantic embedding features; performing semantic pairing on the query statement based on the node semantic features to obtain a semantic clustering result; and dividing the semantic clustering result into independent path segments to obtain the set of transfer mapping pairs between the input fields and the output fields.
[0040] For example, in step S20, the query statement can be paired based on the language element information extracted in the previous step, and the purpose is to identify the transfer logic relationship between the input fields and the output fields in the sentence. In a specific implementation, by introducing a graph neural network to analyze the language element information, each syntax element (such as a field, a table name, and a connection relationship) is regarded as a node in a graph structure, and semantic embedding features with context meaning are generated for these nodes. For example, in a financial scenario, in the query statement “statistical total loan amount of each region divided by month”, the fields “region”, “month”, and “loan amount” will learn their semantic relationship through the graph neural network, “loan amount” is identified as an output field, and “region” and “month” in its dependency path are input fields; in a medical scenario, when querying “calculate the average hospitalization days of patients over 60 years old with diabetes”, the graph neural network can identify “hospitalization days” as an output field, and “age” and “diagnosis result” as input fields, and establish a semantic connection between them.
[0041] Subsequently, the obtained node semantic embedding features can be used to further perform semantic clustering on the query statement, and nodes with similar semantics or connected paths are classified into the same semantic cluster, and a plurality of independent semantic path segments are divided by the clustering results. Each path segment represents the mapping logic from a certain input field to an output field, forming a final set of transfer mapping pairs. For example, in a financial application, there may be a path between "loan amount" and "customer credit rating" through the "loan approval result" field to form a transfer logic; in a medical application, the calculation path between "discharge date" and "admission date" through the "hospitalization days" field can also be mapped as a transfer pair from input to output.
[0042] In this way, not only semantic understanding is achieved, but also structured field mapping basis is provided for data bloodline analysis, result interpretation and risk control.
[0043] S30: constructing a transmission path between the input field and the output field according to the set of transfer mapping pairs to generate a directed dependency network structure.
[0044] In some embodiments, constructing a transmission path between the input field and the output field according to the set of transfer mapping pairs to generate a directed dependency network structure comprises: constructing a transmission path between the input field and the output field according to the set of transfer mapping pairs; traversing a plurality of nodes in the transmission path to identify a cross structure between the nodes; determining an edge weight in the cross structure, and generating the directed dependency network structure based on the edge weight and the transmission path.
[0045] For example, in step S30, a transmission path between the input field and the output field can be constructed according to the set of transfer mapping pairs generated in the previous step, and a directed dependency network structure is further generated. The transmission path takes fields as nodes and mapping relationships between fields as edges, showing the directionality of data flow. For example, in the financial field, in the query statement "analyze the impact of customer risk level on investment return rate", the system can identify "customer risk level" as the input field and "investment return rate" as the output field, and form a clear data transmission path through intermediate fields such as "asset allocation strategy". In the medical scenario, when querying "evaluate the impact of treatment method on recovery after discharge", the field "treatment method" is taken as the input, and through intermediate variables such as "hospitalization duration" and "complication record", it finally affects "recovery condition". This path can also be presented in the form of directed edges in the dependency network, reflecting the causal or data dependency relationship between fields.
[0046] Further, to optimize the expression effect of the dependency network, multiple nodes in these transmission paths can be traversed to identify the cross structure existing therein, i.e., the convergence or branching of multiple paths at the same node. For example, in the financial application, the "credit score" field can simultaneously affect the "loan approval result" and the "interest rate setting", forming a cross structure; by assigning edge weights to these cross edges (such as based on the correlation strength between fields or the weight in the query statement), the accuracy of the dependency relationship is further enhanced. In the medical field, the "patient age" field can simultaneously affect the "treatment plan selection" and the "surgical risk assessment", also constituting a typical cross path. Finally, by comprehensively integrating all path structures and edge weight information, a directed dependency network is generated, which not only accurately reflects the hierarchical and transmission relationship between fields, but also provides structured support for subsequent data governance, visual analysis, and causal modeling.
[0047] S40: mapping the directed dependency network structure between the task nodes of the data task to obtain a blood relationship graph corresponding to the data task.
[0048] In some embodiments, the mapping of the directed dependency network structure between the task nodes of the data task to obtain a blood relationship graph corresponding to the data task comprises: obtaining the task nodes of the data task and determining the incoming direction and outgoing direction of the task nodes; generating a dependency path relationship of the data task according to the incoming direction and outgoing direction of the task nodes; converting the dependency path relationship and the task nodes into a graph structure edge set; mapping the directed dependency network structure to the graph structure edge set to output the blood relationship graph.
[0049] For example, the constructed directed dependency network structure can be further mapped to a specific data task, and by analyzing the input-output direction relationship between the task nodes, a data dependency path at the task level is generated, and finally a blood relationship graph is formed. Taking the financial field as an example, a data task includes multiple processing nodes such as "customer information extraction", "credit score calculation", "risk assessment modeling", etc., and there is a clear data flow between the task nodes. The system first identifies the incoming fields (such as customer data) and outgoing fields (such as risk level) of each node, and generates the dependency path between the tasks according to the same. Then convert these paths into a graph structure edge set, for example, the "customer income field" output by the "customer information extraction" node is connected to the "credit score calculation" node as an edge, mapping the logical dependency relationship between the two tasks. In the medical field, similar transmission graph structures can also be constructed between "electronic medical record parsing", "diagnosis standard matching", "treatment effect prediction", etc. task nodes according to the field flow, realizing the visual blood management of complex medical data tasks.
[0050] Further, the edge set of the graph structure of the above task path can be mapped with the previously constructed field-level directed dependency network structure to form a unified data task-level blood relationship graph. For example, in the financial scenario, through mapping, it can be obtained that how the "customer asset status" field flows from the "account information extraction" node to the "investment suggestion generation" node in detail, supporting investment model traceability and regulatory audit. In the medical scenario, the graph can reveal how the "diagnosis indicator" is transmitted to the "treatment path recommendation" node via the "disease recognition" node, helping to track the data-driven clinical decision-making process. Such a blood relationship graph not only clearly shows the flow context of fields between tasks, but also provides a structured basis for abnormal data tracing, task impact analysis and process optimization, significantly improving the controllability and transparency of complex data processing processes in the financial and medical fields.
[0051] In some embodiments, the method further comprises: determining a task execution frequency and an update time of the data task based on the blood relationship graph; generating a field activity score model based on the task execution frequency and the update time; analyzing the data task through the field activity score model to obtain a field activity corresponding to the data task; and mapping the field activity to between the task nodes of the data task to output a weighted graph structure corresponding to the data task.
[0052] For example, the dynamic characteristics of the data task can be further analyzed based on the aforementioned blood relationship graph. First, the execution frequency and data update time of each task node are determined. For example, in the financial field, the running period and field refresh frequency of task nodes such as "end-of-day clearing", "risk assessment" and "interest rate adjustment" can be identified, and it is found that fields such as "account balance" or "transaction flow" are updated frequently every day and belong to active fields; while "annual income" or "customer rating" may be updated less frequently and belong to low-active fields. In the medical field, fields related to "patient vital sign monitoring" or "drug reaction recording" are updated frequently in the clinical system, while "admission reason" or "underlying disease information" usually remains stable after initial entry. Accordingly, a field activity score model can be constructed to quantitatively evaluate the update behavior of the field and form a basic activity score system.
[0053] Subsequently, the activity score can be applied to the original bloodline graph structure, mapping the activity value of each field to the data task nodes where it is located, and then constructing a weighted graph structure, so that the originally equally weighted graph edges have a dynamic scoring dimension. For example, in the financial data process, if a path involves the "high-frequency trading volume" field, its high activity will increase the path weight, and the system can optimize the data prefetching or caching strategy accordingly; in medical applications, if a task path depends on the "real-time monitoring of blood oxygen saturation" field, a high activity score indicates that this path is more critical to diagnosis and treatment judgment, and can be used as a key node for process monitoring and abnormal warning. This weighted graph structure not only enhances the expressive power of the graph, but also enables priority judgment and dynamic resource allocation based on data activity in financial risk control or medical decision-making.
[0054] In some embodiments, the method further includes: monitoring real-time change events of the data task to obtain a field change log; performing traceability matching on the field names in the field change log to obtain a set of associated statements; generating a field mapping relationship of the data task based on the set of associated statements; and performing path changes on the blood relationship map through the field mapping relationship to obtain an updated blood relationship map.
[0055] For example, real-time changes to data tasks during execution can be monitored and recorded in a field change log to capture adjustments to field names or structures. Field names in the log are then traced back and matched to their locations in historical data tasks, forming a set of associated statements. For example, in the financial sector, when the "Customer Credit Rating" field is renamed to an equivalent field like "Credit Score," the approval, credit, and risk control calculation statements associated with that field can be automatically identified, and new field mappings can be generated. In the medical field, if the "Blood Glucose Test Value" field is adjusted to "Fasting Blood Glucose" and "Postprandial Blood Glucose," the affected diagnosis and treatment recommendation task paths can similarly be identified through associated statements. Based on these field mappings, the original blood relationship map can be automatically adjusted and reconstructed to generate an updated map, ensuring that the data transmission chain remains accurate and valid after field changes, thereby ensuring the continuity and correctness of financial risk control logic and medical analysis processes.
[0056] It can be seen that in the above scheme, by performing syntax analysis and semantic pairing on the query statement in the data task, the transmission relationship between the input field and the output field can be efficiently and accurately extracted, and a clear directed dependency network structure can be constructed, thereby realizing automatic identification and visualization of the data blood relationship. This method has high automation, strong scalability, and the ability to adapt to various complex query statements, significantly improving data governance efficiency. At the same time, by constructing the blood relationship graph, subsequent operations such as data tracing, task optimization and risk control can be effectively supported, and the global control of the system on the data flow process is enhanced. In the financial field, this scheme can be applied to complex financial report generation and data tracking in risk control models to ensure traceability of key indicators and improve audit compliance. In the medical field, it can be used in the processing flow of electronic medical records and medical test data to ensure transparent patient data transmission path and data use compliance, which helps to improve medical data management quality and clinical decision-making efficiency.
[0057] It should be understood that the size of the serial number of each step in the above embodiment does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the application.
[0058] In an embodiment, a data task processing apparatus is provided, which corresponds to the data task processing method in the above embodiment. As shown in the figure, the data task processing apparatus includes an acquisition module 101, a pairing module 102, a construction module 103 and a mapping module 104. The functions of each functional module are described in detail as follows: Figure 4
[0059] The acquisition module 101 is configured to acquire a query statement in the data task and analyze the query statement to obtain syntax element information corresponding to the query statement; wherein the syntax element information includes table name, field name, operation mode and connection condition;
[0060] The pairing module 102 is configured to perform semantic pairing on the query statement based on the language element information to obtain a set of transmission mapping pairs between input fields and output fields in the query statement;
[0061] The construction module 103 is configured to construct a transmission path between the input field and the output field according to the set of transmission mapping pairs to generate a directed dependency network structure;
[0062] The mapping module 104 is configured to map the directed dependency network structure to the task nodes of the data task to obtain a blood relationship graph corresponding to the data task.
[0063] The acquisition module 101 is configured to acquire a query statement in the data task, analyze the query statement by using a pre-trained language model, and obtain a corresponding syntax expression unit set; perform structural analysis on the syntax expression unit by using a dependency syntax tree, and obtain a syntax structure index in the query statement; perform field mapping on the syntax structure index, and obtain syntax element information corresponding to the query statement.
[0064] The pairing module 102 is configured to analyze the language element information by using a graph neural network, obtain corresponding node semantic embedding features, perform semantic pairing on the query statement based on the node semantic features, and obtain a semantic clustering result; and divide the semantic clustering result into independent path segments, and obtain a transmission mapping pair set between the input field and the output field.
[0065] The construction module 103 is configured to construct a transmission path between the input field and the output field according to the transmission mapping pair set, traverse a plurality of nodes in the transmission path, identify a cross structure between the nodes, determine an edge weight value in the cross structure, and generate the directed dependency network structure based on the edge weight value and the transmission path.
[0066] The mapping module 104 is configured to acquire a task node of the data task, determine an incoming direction and an outgoing direction of the task node, generate a dependency path relationship of the data task according to the incoming direction and the outgoing direction of the task node, convert the dependency path relationship and the task node into a graph structure edge set, and map the directed dependency network structure to the graph structure edge set to output the blood relationship graph.
[0067] In an embodiment, the acquisition module 101 is further configured to determine a task execution frequency and an update time of the data task based on the blood relationship graph, generate a field activity score model based on the task execution frequency and the update time, and analyze the data task by using the field activity score model to obtain a field activity corresponding to the data task.
[0068] The field activity is mapped between the task nodes of the data task to output a weighted graph structure corresponding to the data task.
[0069] In an embodiment, the acquisition module 101 is further configured to listen to a real-time change event of the data task to obtain a field change log, perform provenance matching on a field name in the field change log to obtain an associated statement set, generate a field mapping relationship of the data task based on the associated statement set, and perform path change on the blood relationship graph by using the field mapping relationship to obtain an updated blood relationship graph.
[0070] The present invention provides a data task processing device, which can efficiently and accurately extract the transfer relationship between input fields and output fields by performing grammatical parsing and semantic matching on the query statements in the data tasks, and construct a clear directed dependency network structure, thereby realizing the automatic identification and visualization of data lineage relationships. This method has a high degree of automation, strong scalability, and the ability to adapt to a variety of complex query statements, which significantly improves the efficiency of data governance. At the same time, by constructing a lineage relationship map, it can effectively support subsequent operations such as data traceability, task optimization and risk control, and enhance the system's overall control over the data flow process. In the financial field, this solution can be applied to data tracking in complex financial report generation and risk control models to ensure that the source of key indicators can be traced and improve audit compliance capabilities; in the medical field, it can be used in the processing flow of electronic medical records and medical test data to ensure that the patient data transmission path is transparent and the data use is compliant, which helps to improve the quality of medical data management and clinical decision-making efficiency.
[0071] For the specific definition of the data task processing device, please refer to the definition of the data task processing method above, which will not be repeated here. The various modules in the above-mentioned data task processing device can be implemented in whole or in part by software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0072] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface and database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the server side of a data task processing method.
[0073] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 6As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium, an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the computer device is used to communicate with the external server through the network connection. The computer program is executed by the processor to realize the function or step of the client side of the data task processing method
[0074] In one embodiment, a computer device is provided, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, the processor executing the computer program to implement the following steps:
[0075] Obtain the query statement in the data task, and parse the query statement to obtain the syntax element information corresponding to the query statement; wherein the syntax element information includes table name, field name, operation mode and connection condition;
[0076] Based on the language element information, the query statement is semantically paired to obtain a set of transfer mapping pairs between the input fields and the output fields in the query statement;
[0077] According to the transfer mapping pair set, a transmission path between the input field and the output field is constructed to generate a directed dependency network structure;
[0078] Map the directed dependency network structure between the task nodes of the data task to obtain the corresponding blood relationship graph of the data task.
[0079] In one embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the following steps:
[0080] Obtain the query statement in the data task, and parse the query statement to obtain the syntax element information corresponding to the query statement; wherein the syntax element information includes table name, field name, operation mode and connection condition;
[0081] Based on the language element information, the query statement is semantically paired to obtain a set of transfer mapping pairs between the input fields and the output fields in the query statement;
[0082] According to the transfer mapping pair set, a transmission path between the input field and the output field is constructed to generate a directed dependency network structure;
[0083] The directed dependency network structure is mapped to the task nodes of the data task to obtain a blood relationship map corresponding to the data task.
[0084] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0085] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchl ink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0086] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0087] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A method for processing a data task, characterized in that: The method comprises: Obtaining a query statement in the data task, and parsing the query statement to obtain syntax element information corresponding to the query statement; wherein the syntax element information includes table name, field name, operation method and connection condition; Performing semantic pairing on the query statement based on the language element information to obtain a set of transfer mapping pairs between input fields and output fields in the query statement; constructing a transmission path between the input field and the output field according to the transfer mapping pair set to generate a directed dependency network structure; The directed dependency network structure is mapped to the task nodes of the data task to obtain a blood relationship map corresponding to the data task.
2. The method according to claim 1, characterized in that The obtaining of the query statement in the data task and parsing the query statement to obtain grammatical element information corresponding to the query statement includes: Obtaining a query statement in the data task, and analyzing the query statement using a pre-trained language model to obtain a corresponding grammatical expression unit set; Performing structural analysis on the grammatical expression unit using a dependency syntax tree to obtain a grammatical structure index in the query statement; Field mapping is performed on the grammatical structure index to obtain grammatical element information corresponding to the query statement.
3. The method according to claim 1, characterized in that The semantic pairing of the query statement based on the language element information to obtain a set of transfer mapping pairs between input fields and output fields in the query statement includes: Analyze the language element information through a graph neural network to obtain corresponding node semantic embedding features; Performing semantic pairing on the query statements based on the node semantic features to obtain semantic clustering results; The semantic clustering result is divided into independent path segments to obtain a set of transfer mapping pairs between the input field and the output field.
4. The method according to claim 1, wherein The step of constructing a transmission path between the input field and the output field according to the transfer mapping pair set to generate a directed dependency network structure includes: Constructing a transmission path between the input field and the output field according to the transfer mapping pair set; Traversing a plurality of nodes in the transmission path and identifying a cross structure between the nodes; The edge weights in the cross structure are determined, and the directed dependency network structure is generated based on the edge weights and the transmission paths.
5. The method according to claim 1, wherein Mapping the directed dependency network structure to the task nodes of the data task to obtain a blood relationship graph corresponding to the data task includes: Obtaining a task node of the data task, and determining an incoming direction and an outgoing direction of the task node; Generating a dependency path relationship of the data task according to the incoming direction and the outgoing direction of the task node; Converting the dependency path relationship and the task node into a graph structure edge set; The directed dependency network structure is mapped to the graph structure edge set, and the blood relationship graph is output.
6. The method according to claim 1, characterized in that The method further comprises: Determining the task execution frequency and update time of the data task based on the blood relationship map; Generate a field activity scoring model based on the task execution frequency and the update time; Analyze the data task using the field activity scoring model to obtain the field activity corresponding to the data task; The field activity is mapped to the task nodes of the data task, and a weighted graph structure corresponding to the data task is output.
7. The method according to claim 1, characterized in that The method further comprises: Monitor the real-time change events of the data task and obtain the field change log; Perform source matching on the field names in the field change log to obtain a set of associated statements; Generating a field mapping relationship of the data task based on the associated statement set; The path of the blood relationship map is changed through the field mapping relationship to obtain an updated blood relationship map.
8. A data task processing device, characterized in that: include: An acquisition module, configured to acquire a query statement in the data task and parse the query statement to obtain syntax element information corresponding to the query statement; wherein the syntax element information includes table name, field name, operation mode and connection condition; a pairing module, configured to perform semantic pairing on the query statement based on the language element information to obtain a set of transfer mapping pairs between input fields and output fields in the query statement; A construction module, configured to construct a transmission path between the input field and the output field according to the transfer mapping pair set to generate a directed dependency network structure; A mapping module is used to map the directed dependency network structure to the task nodes of the data task to obtain a blood relationship map corresponding to the data task.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the data task processing method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the data task processing method according to any one of claims 1 to 7 are implemented.