Data flow direction management method and system
The dynamic topology optimization engine analyzes and automatically adjusts the data flow diagram in real time, solving the problems of dynamic asset identification blind spots and compliance verification disconnection in data flow management, achieving efficient and secure data flow management, and improving the overall level of automation.
Patent Information
- Application Number
- CN202510788287.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-26
AI Technical Summary
Existing data flow management technologies have problems such as blind spots in dynamic asset identification, disconnection between compliance verification and processes, and low overall automation. Especially in the financial services sector, undiscovered cross-border API interfaces may pose a risk of data leakage. The rule engine is unable to respond to data flow changes under the microservice architecture in real time. Manually drawn data flow diagrams are inefficient and prone to permission configuration conflicts.
A dynamic topology optimization engine is used to analyze data flow diagrams in real time, generate orchestration plans, monitor the status of key node indicators in real time, periodically obtain data asset changes, generate asset version maps, automatically adjust data flow paths, use preset policy libraries to optimize node distribution and transmission paths, perform real-time compliance verification, generate exception handling strategies, and automatically respond to data source changes.
It improves the automation level of data governance and its adaptability to dynamic business changes, ensures the efficiency and security of data flow, reduces the risk of data leakage, improves system response speed and data processing efficiency, and avoids manual intervention.
Smart Images

Figure CN120705210A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data flow management, and in particular to a data flow management method and system. Background Art
[0002] Current mainstream data governance technologies face multiple challenges, the core of which is centered around traditional data flow diagramming and compliance checking mechanisms. On the one hand, static data flow diagramming tools rely on manual analysis of data relationships. While they can generate an initial data flow layout using diagramming software, they lack effective identification of dynamic assets such as unregistered APIs (Application Programming Interfaces) and data lake resources. This leads to significant blind spots in dynamic asset identification, particularly in the financial services sector, where undiscovered cross-border APIs can pose a serious risk of data leakage. On the other hand, while rule-engine-driven compliance checking can provide compliance reviews along known data flow paths, the rule update process relies on manual effort, making it difficult to respond to frequent data flow changes in a microservices architecture in real time. This creates a disconnect between compliance verification and business processes. Furthermore, limited automation is a major shortcoming of existing technologies. Manually drawn data flow diagrams are not only inefficient but also prone to permission configuration conflicts due to untimely updates to business changes. Existing automation tools are limited to processing structured data and lack support for semi-structured and unstructured data sources such as log files and API call chains.
[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0004] The embodiments of the present application provide a data flow management method and system to at least solve the technical problems of traditional data flow management solutions, such as blind spots in dynamic asset identification, disconnection between compliance verification and processes, and low overall automation.
[0005] According to one aspect of an embodiment of the present application, a data flow management method is provided, including: obtaining a first data flow graph created by a target object in an interactive interface, and performing compliance verification on the first data flow graph; using a dynamic topology optimization engine to analyze the first data flow graph and the compliance verification result to obtain a first data flow orchestration scheme, and executing the first data flow orchestration scheme; in the process of executing the first data flow orchestration scheme, obtaining the indicator status of preset key nodes in the first data flow orchestration scheme in real time, and periodically obtaining the data asset changes of the data source to generate an asset version map; performing an anomaly analysis on the indicator status of the preset key nodes, and when an anomaly exists in the preset key nodes, locating the anomaly and generating an anomaly handling strategy; when a preset key data asset in the data source changes, using a dynamic topology optimization engine to analyze the asset version map and the first data flow orchestration scheme to obtain a reconstructed second data flow orchestration scheme, and executing the second data flow orchestration scheme.
[0006] Optionally, the interactive interface is a visual editor, and obtaining the first data flow graph created by the target object in the interactive interface includes: providing a preconfigured component library in the visual editor, wherein the component library includes at least one of the following components: a data source node, a processing node, and a transmission channel node; generating the first data flow graph in response to the target object's dragging and connection operations on each component in the component library.
[0007] Optionally, performing compliance verification on the first data flow graph includes: utilizing a compliance conflict detection engine to perform compliance verification on the first data flow graph, wherein the type of compliance verification includes at least one of the following: node authority matching status detection, data outbound compliance detection, and privacy data security detection.
[0008] Optionally, a dynamic topology optimization engine is used to analyze the first data flow graph and the compliance verification results to obtain a first data flow orchestration plan, including: obtaining the system resource load status, and using the dynamic topology optimization engine to analyze the first data flow graph, the compliance verification results and the system resource load status, optimizing the node distribution and data transmission path in the first data flow graph based on a preset orchestration policy library to obtain a second data flow graph, displaying the second data flow graph in an interactive interface, and generating a first data flow orchestration plan based on the second data flow graph, wherein the orchestration policy library includes at least one of the following strategies: topology optimization strategy, load balancing strategy, disaster recovery strategy, and data security processing strategy.
[0009] Optionally, obtaining the indicator status of preset key nodes in the first data flow orchestration scheme in real time includes: collecting the indicator status of each preset key node in real time through probes pre-deployed at each preset key node in the first data flow orchestration scheme, wherein the type of the preset key node includes at least one of the following: data source node, processing node, transmission channel node, and the collected indicator type includes at least one of the following: processing rate, delay, resource occupancy, and data integrity.
[0010] Optionally, data asset changes of the data source are periodically obtained to generate an asset version map, including: periodically collecting data asset change events of the data source through a multi-protocol adapter pre-deployed at the data source, wherein the type of data asset change event includes at least one of the following: database change, application interface parameter update, file format iteration; generating an asset version map with a timestamp based on the data asset change event, wherein the asset version map records an unorganized data asset list.
[0011] Optionally, an abnormality analysis is performed on the indicator status of the preset key node. When an abnormality exists in the preset key node, the abnormality is located and an abnormality handling strategy is generated, including: using a pre-trained abnormality analysis model to analyze the indicator status of the preset key node; when an abnormality exists in the preset key node, the abnormality type is determined and the source of the abnormality is located, and abnormal prompt information carrying the abnormality type and the source of the abnormality is generated in the interactive interface; an abnormality handling strategy that matches the abnormality type is determined from a preset abnormality handling strategy library, and the abnormality handling strategy is fed back in the interactive interface.
[0012] Optionally, a dynamic topology optimization engine is used to analyze the asset version map and the first data flow orchestration scheme to obtain a reconstructed second data flow orchestration scheme, including: obtaining the system resource load status, and using the dynamic topology optimization engine to analyze the field-level dependencies in the asset version map, the first data flow orchestration scheme and the system resource load status, reconstructing the first data flow orchestration scheme based on a preset orchestration policy library to obtain a second data flow orchestration scheme, generating a third data flow graph based on the second data flow orchestration scheme, and displaying the third data flow graph in an interactive interface, wherein the orchestration policy library includes at least one of the following strategies: topology optimization strategy, load balancing strategy, disaster recovery strategy, and data security processing strategy.
[0013] According to another aspect of the embodiment of the present application, a data flow management system is also provided, including: a data flow diagram orchestration module, used to obtain a first data flow diagram created by a target object in an interactive interface, and perform compliance verification on the first data flow diagram; an orchestration strategy generation module, used to analyze the first data flow diagram and the compliance verification result by using a dynamic topology optimization engine, obtain a first data flow orchestration scheme, and execute the first data flow orchestration scheme; a monitoring and auditing module, used to obtain the indicator status of preset key nodes in the first data flow orchestration scheme in real time during the execution of the first data flow orchestration scheme; a data resource The production synchronization module is used to periodically obtain data asset changes from the data source and generate an asset version map during the execution of the first data flow orchestration plan; the intelligent risk control and governance module is used to perform anomaly analysis on the indicator status of preset key nodes, locate the anomaly when anomalies exist in the preset key nodes, and generate anomaly handling strategies; the orchestration strategy generation module is also used to use the dynamic topology optimization engine to analyze the asset version map and the first data flow orchestration plan when changes occur to the preset key data assets in the data source, obtain a reconstructed second data flow orchestration plan, and execute the second data flow orchestration plan.
[0014] According to another aspect of an embodiment of the present application, an electronic device is further provided, comprising: a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the above-mentioned data flow management method through the computer program.
[0015] In the embodiment of the present application, through real-time compliance verification, it is ensured that the data flow diagram complies with the preset compliance rules at the beginning of its creation, such as cross-border data transmission restrictions, the principle of minimum authority, etc., to avoid the flow of data in unverified paths. The data flow diagram is analyzed in real time using a dynamic topology optimization engine, and the generated first data flow orchestration scheme can dynamically adjust the data transmission path and node selection according to the current system status, such as node load, network delay, etc., to ensure the high efficiency of data flow and the optimized use of resources. Compared with the traditional static orchestration method, this greatly improves the system response speed and data processing efficiency. When executing the first data flow orchestration scheme During the process, the indicator status of preset key nodes will be monitored in real time. When an abnormal situation is detected, the abnormal node can be quickly located and an exception handling strategy can be generated. When changes occur in key data assets in the data source, these changes can be automatically detected and responded to. The asset version map and orchestration plan are re-analyzed through the dynamic topology optimization engine to obtain and execute the reconstructed second data flow orchestration plan. This process does not require human intervention, which improves the automation level of data governance and its adaptability to dynamic business changes, thereby solving the technical problems of traditional data flow management solutions such as dynamic asset identification blind spots, disconnection between compliance verification and processes, and low overall automation. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0017] Figure 1 This is a flowchart of an optional data flow management method according to an embodiment of the present application;
[0018] Figure 2 This is a schematic diagram of the structure of an optional data flow management system according to an embodiment of the present application;
[0019] Figure 3 This is a schematic diagram of an implementation process of an optional data flow management system according to an embodiment of the present application;
[0020] Figure 4 It is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0021] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0022] It should be noted that the terms "first", "second", etc. in the specification, claims, and drawings of the present application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.
[0023] The information collected in the embodiments of the present application is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with the relevant laws, regulations and standards of the relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0024] Example 1
[0025] According to an embodiment of the present application, a data flow management method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0026] Figure 1 This is a flow chart of a data flow management method provided in accordance with an embodiment of the present application. Figure 1 As shown, the method includes the following steps:
[0027] Step S102: obtaining a first data flow diagram created by the target object in the interactive interface, and performing compliance verification on the first data flow diagram;
[0028] Step S104: Analyze the first data flow graph and the compliance verification result using a dynamic topology optimization engine to obtain a first data flow orchestration scheme, and execute the first data flow orchestration scheme;
[0029] Step S106: During the execution of the first data flow orchestration solution, the indicator status of the preset key nodes in the first data flow orchestration solution is obtained in real time, and the data asset changes of the data source are periodically obtained to generate an asset version map;
[0030] Step S108: Perform abnormality analysis on the indicator status of the preset key nodes. If an abnormality exists in the preset key nodes, locate the abnormality and generate an abnormality handling strategy.
[0031] Step S110: When a preset key data asset in the data source changes, the dynamic topology optimization engine is used to analyze the asset version map and the first data flow orchestration scheme to obtain a reconstructed second data flow orchestration scheme, and the second data flow orchestration scheme is executed.
[0032] Data assets refer to all valuable data resources within an enterprise, including but not limited to data in databases, data warehouses, information in data lakes, data transmitted via various interfaces, and data files in file systems. These are considered valuable assets of the enterprise and require effective management, protection, and utilization to support business operations and decision-making. A data flow orchestration solution can be understood as a set of automated and intelligent strategies for planning, managing, and optimizing the flow and processing logic of data assets within an enterprise IT environment. The orchestration solution defines the data flow path from source to destination, including the various processing nodes and transmission channels that the data passes through, as well as the compliance rules that must be followed during path planning, such as restrictions on cross-border data transmission and privacy protection policies.
[0033] The following describes the various steps of the data flow management method in conjunction with a specific implementation process.
[0034] First, a first data flow diagram created by a target object in an interactive interface is obtained, and compliance verification is performed on the first data flow diagram.
[0035] As an optional implementation, the interactive interface is a visual editor, and the following steps can be taken to obtain the first data flow graph created by the target object in the interactive interface: providing a preconfigured component library in the visual editor, wherein the component library includes at least one of the following components: a data source node, a processing node, and a transmission channel node; and generating a first data flow graph in response to the target object's dragging and connection operations on each component in the component library.
[0036] The data flow diagram is a graphical representation of the data flow path in the system, clearly showing the entire process of data collection, processing and storage, including each node through which the data passes and the data transmission path between nodes.
[0037] For example, in a visual editor, a preconfigured component library is designed, which contains various data management components that may be involved in the system, such as data source nodes (representing databases, data warehouses, etc.), processing nodes (such as data cleaning, conversion, aggregation, etc.), and transmission channel nodes (such as network transmission, message queues, etc.). The target object (usually a data governance administrator) selects components through simple and intuitive drag-and-drop operations, and defines the data flow path and processing logic by connecting different components. Once the target object completes the component selection and connection, the system will automatically parse these operations and generate a graphical representation describing how data flows between different components - a data flow diagram. This diagram not only includes the direction of data flow, but also records in detail the properties of each component and the connection rules between them.
[0038] As an optional implementation, the compliance check of the first data flow graph may be performed by taking the following steps: using a compliance conflict detection engine to perform compliance check on the first data flow graph, wherein the type of compliance check includes at least one of the following: node authority matching status detection, data outbound compliance detection, and privacy data security detection.
[0039] For example, a compliance conflict detection engine is built in, which is responsible for parsing the data flow graph and checking whether each node in it meets the preset compliance rules. These rules can cover aspects such as cross-border data transmission restrictions, user privacy protection policies, and the principle of minimum permissions. Point permission matching status detection refers to checking whether each node in the data flow graph (especially data source nodes and processing nodes) has the correct access rights, for example, ensuring that only authorized processing logic can access specific database tables or fields; data outbound compliance detection means that for nodes involved in cross-border data transmission, the engine will check whether the data flow graph complies with relevant laws and regulations. If any path that may violate these regulations is detected, the system will immediately mark it; privacy data security detection refers to detecting whether security measures are in place during the transmission and processing of data streams containing user personal information or other sensitive data, such as encryption and anonymization.
[0040] After obtaining the first data flow diagram and the compliance verification result, the first data flow diagram and the compliance verification result are analyzed using a dynamic topology optimization engine to obtain a first data flow orchestration plan, and the first data flow orchestration plan is executed.
[0041] As an optional implementation, a dynamic topology optimization engine is used to analyze the first data flow graph and the compliance verification results to obtain a first data flow orchestration plan, and the first data flow orchestration plan is executed. The process can be carried out in the following steps: obtaining the system resource load status, and using the dynamic topology optimization engine to analyze the first data flow graph, the compliance verification results and the system resource load status, optimizing the node distribution and data transmission path in the first data flow graph based on a preset orchestration policy library to obtain a second data flow graph, displaying the second data flow graph in an interactive interface, and generating a first data flow orchestration plan based on the second data flow graph, wherein the orchestration policy library includes at least one of the following strategies: topology optimization strategy, load balancing strategy, disaster recovery strategy, and data security processing strategy.
[0042] For example, before the system starts processing data flow, the dynamic topology optimization engine first calls the resource monitoring module to collect resource load data in the entire system, including key indicators such as CPU (Central Processing Unit) usage, memory usage, network bandwidth, and the current task status of each data processing node. At the same time, the engine receives verification results from the compliance rule engine, which marks all potential compliance conflicts and risk points, such as cross-border data transmission that does not comply with regulations, and permission configuration that violates the principle of minimization. Combined with the resource load status and compliance verification results, the optimization engine uses graph computing theory to conduct an in-depth analysis of the node distribution and data transmission path in the first data flow graph. It uses topology optimization algorithms to identify paths with low efficiency or high risk, as well as nodes with underutilized resources. Based on the analysis results and the preset orchestration policy library, the engine generates a second data flow graph. The orchestration policy library contains a variety of strategies, such as topology optimization strategy (adjusting node layout to reduce latency), load balancing strategy (distributing load to avoid single-point bottlenecks), disaster recovery strategy (building multi-path data backup to enhance system resilience), data security processing strategy (encrypting sensitive data and restricting unauthorized access), etc. These strategies ensure the robustness and security of the data flow graph. The optimized second data flow graph will be displayed to the data governance administrator on the interactive interface, intuitively presenting the optimized path and node distribution of the data flow. Finally, based on the second data flow graph, the dynamic topology optimization engine generates the first data flow orchestration plan. This plan details the specific steps, transmission paths, and processing logic of the data flow. Subsequently, the orchestration plan is distributed to each data processing node to guide the actual flow of data and ensure that the data is smoothly transmitted while meeting efficiency, compliance, and security.
[0043] During the execution of the orchestration plan, the dynamic topology optimization engine will continuously monitor the indicator status of key nodes (such as processing rate, latency, resource usage) and changes in data assets. Once a performance indicator deviation or asset change is detected, it will immediately trigger an exception analysis to locate the problem, select the corresponding plan from the policy library, and automatically adjust the data flow diagram and orchestration strategy to form a closed-loop governance mechanism to ensure continuous compliance of data governance and highly optimized data processing.
[0044] During the execution of the first data flow orchestration plan, the indicator status of the preset key nodes in the first data flow orchestration plan is obtained in real time, and the data asset changes of the data source are periodically obtained to generate an asset version map.
[0045] As an optional implementation, the indicator status of the preset key nodes in the first data flow orchestration scheme is obtained in real time, and the following steps can be taken: the indicator status of each preset key node is collected in real time by probes pre-deployed at each preset key node in the first data flow orchestration, wherein the type of the preset key node includes at least one of the following: data source node, processing node, transmission channel node, and the collected indicator type includes at least one of the following: processing rate, delay, resource occupancy, and data integrity.
[0046] For example, during the execution phase of the first data flow orchestration plan, the system will automatically deploy specially designed lightweight monitoring probes on preset key nodes (such as data source nodes, processing nodes, and transmission channel nodes). These probes are responsible for real-time monitoring and collecting key performance indicators, including processing rate, latency, resource utilization (CPU, memory, etc.), and data integrity. By being tightly integrated into the system, the probes can continuously collect data and provide feedback on indicator status, which helps to promptly discover and respond to potential performance bottlenecks or abnormal situations. The collected indicator status will be imported into a central monitoring platform. Using real-time data analysis and machine learning models, the system can immediately identify situations where the indicator status deviates from the normal range, such as high latency, low processing rate, or excessive resource utilization. Once an anomaly is detected, the system will quickly locate the problem node and may trigger a real-time warning so that management personnel can intervene in time.
[0047] As an optional implementation method, the following steps can be taken to periodically obtain data asset changes of a data source and generate an asset version map: periodically collect data asset change events of the data source through a multi-protocol adapter pre-deployed at the data source, wherein the type of data asset change event includes at least one of the following: database change, application interface parameter update, file format iteration; generate an asset version map with a timestamp based on the data asset change event, wherein the asset version map records an unorganized data asset list.
[0048] In this solution, the asset version map refers to a structured view that records and manages the change history of data assets. It records in detail every change of each data asset (such as database tables, API interfaces, data lake resources, etc.) from creation to the present in the form of a map, including the timestamp of the change, the content of the change (such as field addition, deletion, data format modification, etc.), version comparison before and after the change, and possible dependency changes.
[0049] For example, given that data assets in an enterprise environment may be stored in a variety of systems, from traditional databases to modern data lakes and even external API interfaces, the system can periodically collect data asset change events from these data sources through pre-deployed multi-protocol adapters. The adapters support various standard protocols, such as SQL (Structured Query Language) and FTP (File Transfer Protocol), ensuring that changes in all data sources are captured. The adapters automatically scan the specified data sources at a preset period (such as hourly or daily), searching for and recording any change events that occur, including database table structure modifications, API parameter updates, and file format changes. These change events are marked and transmitted to the core processing module. Based on the collected change events, the system dynamically generates and maintains a timestamped asset version map. This map not only records the historical change history of data assets but also clearly marks which assets have not yet been orchestrated (i.e., the unorganized asset list), which is critical for maintaining the timeliness and accuracy of the data flow diagram.
[0050] An abnormality analysis is performed on the indicator status of the preset key nodes in the first data flow orchestration plan. When an abnormality exists in the preset key nodes, the abnormality is located and an abnormality handling strategy is generated.
[0051] As an optional implementation method, an abnormality analysis is performed on the indicator status of a preset key node. When an abnormality exists in the preset key node, the abnormality is located and an abnormality handling strategy is generated. The following steps can be taken: using a pre-trained abnormality analysis model to analyze the indicator status of the preset key node; when an abnormality exists in the preset key node, the abnormality type is determined and the source of the abnormality is located, and abnormal prompt information carrying the abnormality type and the source of the abnormality is generated in the interactive interface; an abnormality handling strategy that matches the abnormality type is determined from a preset abnormality handling strategy library, and the abnormality handling strategy is fed back in the interactive interface.
[0052] For example, pre-trained anomaly analysis models are used to continuously monitor the indicator status of pre-set key nodes. These models, typically based on machine learning algorithms, can learn a baseline of normal indicator status and identify abnormal indicator trends or mutations based on this. Using data collected in real time by probes deployed at key nodes, the anomaly analysis model instantly analyzes whether each indicator falls within the abnormal range. For example, the model might identify that the CPU utilization of a data processing node suddenly soars to over 90%, or that data transmission latency increases significantly. Once the model detects an abnormal indicator, it further analyzes it to determine the specific type of anomaly. Possible anomaly types include, but are not limited to, performance bottlenecks (such as high latency, low processing rate), resource overloads (such as extremely high CPU or memory usage), data integrity issues (such as missing or corrupted data), and security incidents (such as unauthorized data access attempts). Localizing the anomaly to a specific node or data flow path involves tracing the propagation path of the abnormal indicator from the overall data flow graph until the original source of the anomaly is found. This localization capability is critical for rapid response and problem resolution. A preset exception handling policy library is built in, which contains processing templates for different types of abnormal events. These policies can be pre-formulated based on historical data and expert experience, aiming to provide the most effective exception response plan. After locating the source of the exception, the system will automatically select the processing policy that matches the exception type from the exception handling policy library. For example, for performance bottlenecks, possible strategies are to increase node resources or optimize data processing logic. For security incidents, it may include strengthening access control or activating emergency response processes. All exception analysis results and processing strategies will be intuitively displayed on the interactive interface (usually the visual interface of the data flow diagram orchestration tool). Target objects, such as administrators, can see the exception type, exception source and recommended processing strategy, which facilitates quick understanding and decision-making.
[0053] When the preset key data assets in the data source change, the dynamic topology optimization engine is used to analyze the asset version map and the first data flow orchestration scheme to obtain a reconstructed second data flow orchestration scheme, and then execute the second data flow orchestration scheme.
[0054] As an optional implementation method, a dynamic topology optimization engine is used to analyze the asset version map and the first data flow orchestration scheme to obtain a reconstructed second data flow orchestration scheme. The following steps can be taken: obtaining the system resource load status, and using the dynamic topology optimization engine to analyze the field-level dependencies in the asset version map, the first data flow orchestration scheme and the system resource load status, reconstructing the first data flow orchestration scheme based on a preset orchestration policy library to obtain a second data flow orchestration scheme, generating a third data flow diagram based on the second data flow orchestration scheme, and displaying the third data flow diagram in an interactive interface, wherein the orchestration policy library includes at least one of the following strategies: topology optimization strategy, load balancing strategy, disaster recovery strategy, and data security processing strategy.
[0055] For example, before starting a process to respond to a data asset change, the dynamic topology optimization engine will first collect the current resource load status of the entire system, including but not limited to CPU usage, memory usage, network bandwidth, etc., to provide real-time environmental information for the next step of analysis. The engine will carefully check the field-level dependencies of each changed data asset in the asset version map, and evaluate the possible impact of these changes on the existing data flow diagram and orchestration plan. For example, if a new field is added to a data source table, the engine needs to determine which data processing nodes depend on this table and whether the new field affects the processing logic of these nodes. Based on the collected system resource status and dependencies in the asset map, the dynamic topology optimization engine begins to analyze the current first data flow orchestration plan. It will attempt to identify which parts may no longer be applicable, such as data flow path failures caused by changes in the data source, or data processing bottlenecks caused by newly emerging dependencies. Using the diverse strategies in the orchestration strategy library, such as topology optimization strategies, load balancing strategies, disaster recovery strategies, and data security processing strategies, the engine will propose a series of optimization measures to reconstruct the data flow orchestration plan. The purpose of these strategies is to ensure the consistency of data processing, the balance of resource utilization, the stability of the system, and the security of data are not affected by asset changes. After in-depth analysis and strategy adjustment, the dynamic topology optimization engine will output a reconstructed second data flow orchestration plan. This plan not only reflects the latest data asset status, but also takes into account the real-time situation of system resources, aiming to optimize data processing paths and node layout. Subsequently, the system will automatically execute the second data flow orchestration plan, which may include reconfiguring data processing logic, adjusting data transmission paths, adding new processing nodes or security measures, etc. When the first data flow orchestration plan is reconstructed to obtain the second data flow orchestration plan, the system will generate a third data flow diagram based on the second data flow orchestration plan and display the third data flow diagram on the interactive interface, intuitively showing the optimized data flow path and processing node layout. Administrators can view this new diagram to confirm whether the changed data flow meets expectations and perform necessary verification to ensure compliance.
[0056] In the embodiment of the present application, through real-time compliance verification, it is ensured that the data flow diagram complies with the preset compliance rules at the beginning of its creation, such as cross-border data transmission restrictions, the principle of minimum authority, etc., to avoid the flow of data in unverified paths. The data flow diagram is analyzed in real time using a dynamic topology optimization engine, and the generated first data flow orchestration scheme can dynamically adjust the data transmission path and node selection according to the current system status, such as node load, network delay, etc., to ensure the high efficiency of data flow and the optimized use of resources. Compared with the traditional static orchestration method, this greatly improves the system response speed and data processing efficiency. When executing the first data flow orchestration scheme During the process, the indicator status of preset key nodes will be monitored in real time. When an abnormal situation is detected, the abnormal node can be quickly located and an exception handling strategy can be generated. When changes occur in key data assets in the data source, these changes can be automatically detected and responded to. The asset version map and orchestration plan are re-analyzed through the dynamic topology optimization engine to obtain and execute the reconstructed second data flow orchestration plan. This process does not require human intervention, which improves the automation level of data governance and its adaptability to dynamic business changes, thereby solving the technical problems of traditional data flow management solutions such as dynamic asset identification blind spots, disconnection between compliance verification and processes, and low overall automation.
[0057] Example 2
[0058] According to an embodiment of the present application, a data flow management system for implementing the data flow management method in embodiment 1 is also provided. Figure 2 As shown, the data flow management system includes at least: a data flow diagram orchestration module 21, an orchestration strategy generation module 22, a monitoring and auditing module 23, a data asset synchronization module 24 and an intelligent risk control and governance module 25, wherein:
[0059] The data flow diagram arrangement module 21 is used to obtain the first data flow diagram created by the target object in the interactive interface and perform compliance verification on the first data flow diagram;
[0060] An orchestration strategy generation module 22 is configured to analyze the first data flow graph and the compliance verification result using a dynamic topology optimization engine to obtain a first data flow orchestration solution and execute the first data flow orchestration solution;
[0061] The monitoring and auditing module 23 is used to obtain the indicator status of the preset key nodes in the first data flow arrangement scheme in real time during the execution of the first data flow arrangement scheme;
[0062] The data asset synchronization module 24 is used to periodically obtain data asset changes from the data source and generate an asset version map during the execution of the first data flow orchestration solution;
[0063] The intelligent risk control and management module 25 is used to analyze the indicator status of preset key nodes for abnormalities. If an abnormality exists in a preset key node, it will locate the abnormality and generate an abnormality handling strategy;
[0064] The orchestration strategy generation module 22 is also used to use the dynamic topology optimization engine to analyze the asset version map and the first data flow orchestration plan when the preset key data assets in the data source change, obtain a reconstructed second data flow orchestration plan, and execute the second data flow orchestration plan.
[0065] The following describes the functions of each module of the data flow management device in conjunction with a specific implementation process.
[0066] The data flow graph arrangement module obtains a first data flow graph created by a target object in an interactive interface, and performs compliance verification on the first data flow graph.
[0067] As an optional implementation, the interactive interface is a visual editor, and the following steps can be taken to obtain the first data flow graph created by the target object in the interactive interface: providing a preconfigured component library in the visual editor, wherein the component library includes at least one of the following components: a data source node, a processing node, and a transmission channel node; and generating a first data flow graph in response to the target object's dragging and connection operations on each component in the component library.
[0068] As an optional implementation, the compliance check of the first data flow graph may be performed by taking the following steps: using a compliance conflict detection engine to perform compliance check on the first data flow graph, wherein the type of compliance check includes at least one of the following: node authority matching status detection, data outbound compliance detection, and privacy data security detection.
[0069] After obtaining the first data flow diagram and the compliance verification result, the orchestration strategy generation module uses the dynamic topology optimization engine to analyze the first data flow diagram and the compliance verification result, obtains the first data flow orchestration plan, and executes the first data flow orchestration plan.
[0070] As an optional implementation, a dynamic topology optimization engine is used to analyze the first data flow graph and the compliance verification results to obtain a first data flow orchestration plan, and the first data flow orchestration plan is executed. The process can be carried out in the following steps: obtaining the system resource load status, and using the dynamic topology optimization engine to analyze the first data flow graph, the compliance verification results and the system resource load status, optimizing the node distribution and data transmission path in the data flow graph based on a preset orchestration policy library to obtain a second data flow graph, displaying the second data flow graph in an interactive interface, and generating a first data flow orchestration plan based on the second data flow graph, wherein the orchestration policy library includes at least one of the following strategies: topology optimization strategy, load balancing strategy, disaster recovery strategy, and data security processing strategy.
[0071] During the execution of the first data flow orchestration plan, the monitoring and auditing module obtains the indicator status of the preset key nodes in the first data flow orchestration plan in real time, and the data asset synchronization module periodically obtains the data asset changes of the data source and generates an asset version map.
[0072] As an optional implementation, the indicator status of the preset key nodes in the first data flow orchestration scheme is obtained in real time, and the following steps can be taken: the indicator status of each preset key node is collected in real time by probes pre-deployed at each preset key node in the first data flow orchestration, wherein the type of the preset key node includes at least one of the following: data source node, processing node, transmission channel node, and the collected indicator type includes at least one of the following: processing rate, delay, resource occupancy, and data integrity.
[0073] As an optional implementation method, the following steps can be taken to periodically obtain data asset changes of a data source and generate an asset version map: periodically collect data asset change events of the data source through a multi-protocol adapter pre-deployed at the data source, wherein the type of data asset change event includes at least one of the following: database change, application interface parameter update, file format iteration; generate an asset version map with a timestamp based on the data asset change event, wherein the asset version map records an unorganized data asset list.
[0074] The intelligent risk control and governance module performs anomaly analysis on the indicator status of preset key nodes. If anomalies exist in the preset key nodes, it locates the anomalies and generates anomaly handling strategies.
[0075] As an optional implementation method, an abnormality analysis is performed on the indicator status of a preset key node. When an abnormality exists in the preset key node, the abnormality is located and an abnormality handling strategy is generated. The following steps can be taken: using a pre-trained abnormality analysis model to analyze the indicator status of the preset key node; when an abnormality exists in the preset key node, the abnormality type is determined and the source of the abnormality is located, and abnormal prompt information carrying the abnormality type and the source of the abnormality is generated in the interactive interface; an abnormality handling strategy that matches the abnormality type is determined from a preset abnormality handling strategy library, and the abnormality handling strategy is fed back in the interactive interface.
[0076] When the preset key data assets in the data source change, the orchestration strategy generation module uses the dynamic topology optimization engine to analyze the asset version map and the first data flow orchestration plan, obtains a reconstructed second data flow orchestration plan, and executes the second data flow orchestration plan. This process can be carried out in the following steps:
[0077] Obtain the system resource load status, and use the dynamic topology optimization engine to analyze the field-level dependencies in the asset version map, the first data flow orchestration scheme, and the system resource load status, reconstruct the first data flow orchestration scheme based on the preset orchestration policy library to obtain the second data flow orchestration scheme, generate a third data flow diagram based on the second data flow orchestration scheme, and display the third data flow diagram in the interactive interface, wherein the orchestration policy library includes at least one of the following strategies: topology optimization strategy, load balancing strategy, disaster recovery strategy, and data security processing strategy.
[0078] Figure 3 The implementation process diagram of the above scheme is given as follows: Figure 3As shown, the target object can create a first data flow graph in the data flow graph orchestration module. After obtaining the first data flow graph created by the target object in the interactive interface, the data flow graph orchestration module performs compliance verification on the first data flow graph. In order to fully consider the current resource allocation situation, the data flow graph orchestration module can request a topology suggestion for the first data flow graph from the orchestration strategy generation module. The orchestration strategy generation module uses a dynamic topology optimization engine to analyze the first data flow graph and the compliance verification results, generates a second data flow graph, obtains the first data flow orchestration plan, and executes the first data flow orchestration plan. The target object receives the second data flow graph displayed on the interactive interface. During the execution of the first data flow orchestration plan, the data asset synchronization module periodically obtains data asset changes from the data source and generates an asset version graph. If a preset key data asset in the data source changes, the data flow graph orchestration module requests a reconstruction of the data flow orchestration plan. The orchestration strategy generation module analyzes the asset version graph and the first data flow orchestration plan to obtain a reconstructed second data flow orchestration plan and executes the second data flow orchestration plan. The target object receives a third data flow graph generated based on the second data flow orchestration plan. During the execution of the first data flow orchestration plan, the monitoring and auditing module obtains the indicator status of the preset key nodes in the first data flow orchestration plan in real time. If the indicator status is abnormal, the intelligent risk control and governance module locates the anomaly and generates an anomaly handling strategy. The target object receives an anomaly prompt and an anomaly handling strategy, and can confirm the handling strategy. The sequence numbers in the figure do not represent execution steps. There is no execution order relationship between the real-time acquisition of the indicator status of the preset key nodes in the first data flow orchestration plan and the periodic acquisition of data asset changes from the data source.
[0079] It should be noted that each module in the data flow management system in the embodiment of the present application corresponds one-to-one to each implementation step of the data flow management method in Example 1. Since a detailed description has been given in Example 1, some details not reflected in this embodiment can be referred to Example 1 and will not be elaborated here.
[0080] Example 3
[0081] According to an embodiment of the present application, a computer program product is further provided. The computer program product includes a computer program, wherein when the computer program is executed by a processor, the data flow management method in Example 1 is implemented.
[0082] According to an embodiment of the present application, a non-volatile storage medium is further provided, which includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the data flow management method in Example 1 by running the computer program.
[0083] According to an embodiment of the present application, a processor is further provided, which is used to run a computer program, wherein the data flow management method in Example 1 is executed when the computer program is running.
[0084] According to an embodiment of the present application, an electronic device is also provided, which includes: a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the data flow management method in Example 1 through the computer program.
[0085] Specifically, when the computer program is running, the following steps are executed: obtaining a first data flow diagram created by the target object in the interactive interface, and performing a compliance check on the first data flow diagram; using a dynamic topology optimization engine to analyze the first data flow diagram and the compliance check result to obtain a first data flow orchestration plan, and executing the first data flow orchestration plan; in the process of executing the first data flow orchestration plan, obtaining the indicator status of the preset key nodes in the first data flow orchestration plan in real time, and periodically obtaining the data asset changes of the data source to generate an asset version map; performing an abnormality analysis on the indicator status of the preset key nodes, and when an abnormality exists in the preset key nodes, locating the abnormality and generating an abnormality handling strategy; when the preset key data assets in the data source change, using the dynamic topology optimization engine to analyze the asset version map and the first data flow orchestration plan to obtain a reconstructed second data flow orchestration plan, and executing the second data flow orchestration plan.
[0086] As an optional implementation, the electronic device may be in the form of a mobile terminal, a computer terminal or a similar computing device. Figure 4 The figure shows a hardware structure block diagram of an electronic device for implementing a data flow management method. Figure 4 As shown, the electronic device 40 may include one or more (402a, 402b, ..., 402n are shown in the figure) processors 402 (the processor 402 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 404 for storing data, and a transmission device 406 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 4 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 4 More or fewer components than shown, or with Figure 4 Different configurations shown.
[0087] It should be noted that the one or more processors 402 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the electronic device 40. As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0088] The memory 404 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the data flow management method in the embodiment of the present application. The processor 402 executes various functional applications and data processing by running the software programs and modules stored in the memory 404, that is, implementing the vulnerability detection method of the above-mentioned application. The memory 404 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 404 may further include a memory remotely located relative to the processor 402, and these remote memories may be connected to the electronic device 40 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0089] Transmission device 406 is used to receive or send data via a network. Specific examples of the aforementioned network may include a wireless network provided by the communications provider of electronic device 40. In one embodiment, transmission device 406 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 406 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0090] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the electronic device 40 .
[0091] The serial numbers of the above embodiments are for description only and do not represent the advantages or disadvantages of the embodiments.
[0092] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0093] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0094] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected to achieve the purpose of the present embodiment according to actual needs.
[0095] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0096] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program code.
[0097] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A data flow management method, characterized in that: include: Obtaining a first data flow diagram created by the target object in the interactive interface, and performing compliance verification on the first data flow diagram; Analyzing the first data flow graph and the compliance verification result using a dynamic topology optimization engine to obtain a first data flow orchestration scheme, and executing the first data flow orchestration scheme; During the execution of the first data flow orchestration solution, the indicator status of the preset key nodes in the first data flow orchestration solution is obtained in real time, and the data asset changes of the data source are periodically obtained to generate an asset version map; Performing an abnormality analysis on the indicator status of the preset key nodes, and if an abnormality exists in the preset key nodes, locating the abnormality and generating an abnormality handling strategy; When a preset key data asset in the data source changes, the dynamic topology optimization engine is used to analyze the asset version map and the first data flow orchestration scheme to obtain a reconstructed second data flow orchestration scheme, and the second data flow orchestration scheme is executed.
2. The method according to claim 1, characterized in that The interactive interface is a visual editor, and obtaining a first data flow graph created by a target object in the interactive interface includes: Providing a preconfigured component library in the visual editor, wherein the component library includes at least one of the following components: a data source node, a processing node, and a transmission channel node; In response to the target object's dragging and connecting operations on the components in the component library, the first data flow graph is generated.
3. The method according to claim 1, characterized in that Performing compliance verification on the first data flow graph includes: A compliance conflict detection engine is used to perform compliance verification on the first data flow graph, wherein the type of compliance verification includes at least one of the following: node authority matching status detection, data outbound compliance detection, and privacy data security detection.
4. The method according to claim 1, wherein The first data flow diagram and the compliance verification result are analyzed using a dynamic topology optimization engine to obtain a first data flow arrangement solution, including: Obtain the system resource load status, and use a dynamic topology optimization engine to analyze the first data flow graph, the compliance verification result and the system resource load status, optimize the node distribution and data transmission path in the first data flow graph based on a preset orchestration policy library to obtain a second data flow graph, display the second data flow graph in the interactive interface, and generate the first data flow orchestration plan based on the second data flow graph, wherein the orchestration policy library includes at least one of the following strategies: topology optimization strategy, load balancing strategy, disaster recovery strategy, and data security processing strategy.
5. The method according to claim 1, wherein Obtaining the indicator status of the preset key nodes in the first data flow orchestration scheme in real time includes: Probes pre-deployed at each preset key node in the first data flow orchestration scheme are used to collect indicator status of each preset key node in real time, wherein the type of the preset key node includes at least one of the following: data source node, processing node, transmission channel node, and the collected indicator type includes at least one of the following: processing rate, delay, resource occupancy, and data integrity.
6. The method according to claim 1, wherein Periodically obtain data asset changes from data sources and generate asset version maps, including: Periodically collecting data asset change events of the data source through a multi-protocol adapter pre-deployed on the data source, wherein the type of the data asset change event includes at least one of the following: database change, application program interface parameter update, and file format iteration; An asset version map with a timestamp is generated based on the data asset change event, wherein the asset version map records an unorganized data asset list.
7. The method according to claim 1, characterized in that Performing an abnormality analysis on the indicator status of the preset key node, and locating the abnormality and generating an abnormality handling strategy when an abnormality exists in the preset key node, including: Analyze the indicator status of the preset key nodes using a pre-trained anomaly analysis model; When an abnormality occurs at the preset key node, the abnormality type is determined and the abnormality source is located, and abnormality prompt information carrying the abnormality type and abnormality source is generated in the interactive interface; An exception handling strategy that matches the exception type is determined from a preset exception handling strategy library, and the exception handling strategy is fed back in the interactive interface.
8. The method according to claim 1, characterized in that The dynamic topology optimization engine is used to analyze the asset version graph and the first data flow orchestration scheme to obtain a reconstructed second data flow orchestration scheme, including: Obtain the system resource load status, and use a dynamic topology optimization engine to analyze the field-level dependencies in the asset version map, the first data flow orchestration scheme, and the system resource load status, reconstruct the first data flow orchestration scheme based on a preset orchestration policy library to obtain the second data flow orchestration scheme, generate a third data flow graph based on the second data flow orchestration scheme, and display the third data flow graph in the interactive interface, wherein the orchestration policy library includes at least one of the following strategies: topology optimization strategy, load balancing strategy, disaster recovery strategy, and data security processing strategy.
9. A data flow management system, characterized in that: include: A data flow diagram arrangement module is used to obtain a first data flow diagram created by a target object in an interactive interface and perform compliance verification on the first data flow diagram; an orchestration strategy generation module, configured to analyze the first data flow graph and the compliance verification result using a dynamic topology optimization engine, obtain a first data flow orchestration plan, and execute the first data flow orchestration plan; A monitoring and auditing module, configured to obtain, in real time, the indicator status of preset key nodes in the first data flow orchestration scheme during the execution of the first data flow orchestration scheme; A data asset synchronization module, configured to periodically obtain data asset changes from a data source and generate an asset version map during the execution of the first data flow orchestration solution; An intelligent risk control and management module is used to perform abnormal analysis on the indicator status of the preset key nodes, locate the abnormality and generate an abnormality handling strategy when an abnormality exists in the preset key nodes; The orchestration strategy generation module is further configured to, when a preset key data asset in the data source changes, use the dynamic topology optimization engine to analyze the asset version map and the first data flow orchestration scheme, obtain a reconstructed second data flow orchestration scheme, and execute the second data flow orchestration scheme.
10. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the data flow management method according to any one of claims 1 to 8 through the computer program.
Citation Information
Cited By
Surveying and mapping data quality supervision method and system based on machine learning
CN120975652A
Dynamic monitoring and scheduling method for multi-source heterogeneous computer data flow
CN122160337A