Control method and control system of integrated data intelligent platform

By configuring data dependency analysis mode and node dependency parameters, data flow analysis, integrity detection and quality monitoring are performed, which solves the problem in existing technologies that data dependencies cannot dynamically adapt to changes and achieves accuracy and timeliness in data dependency management.

CN120610948AActive Publication Date: 2025-09-09GUANGZHOU ZHONGZHI SOFTWARE DEV CO LTD

Patent Information

Application Number
CN202510731336.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-09
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

Existing data governance platforms are unable to dynamically adapt to changes in the data flow process and lack a quantitative assessment mechanism for the integrity of data dependencies, making it difficult to accurately assess the scope and extent of the impact of data quality anomalies.

Method used

By configuring the data dependency analysis mode, generating overall data link tracking parameters, obtaining data node information, establishing node dependency parameters, conducting bidirectional data flow analysis, performing integrity detection and quality monitoring, updating node dependency parameters, and identifying the scope and extent of the impact of problem data.

Benefits of technology

It achieves accuracy and timeliness in data dependency management, can promptly detect and locate anomalies, accurately assess the scope of impact, and ensure that dependency analysis results promptly reflect the actual situation of data flow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120610948A_ABST
    Figure CN120610948A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data intelligence, in particular to a control method and a control system of an integrated data intelligent platform. According to the method, the data link tracking parameter is generated by configuring the data dependency analysis mode and the target parameter, the data node information is acquired, and the personalized dependency parameter is established for each node; the system can perform bidirectional data flow analysis, upward traces a data source and a generation process, and downward identifies a data flow and a use condition; calculating and comparing the target parameters through integrity detection to obtain an integrity difference value; when the difference value exceeds a preset threshold value, a quality monitoring mechanism is triggered, and abnormity is found and positioned in time; according to the abnormal information, node dependency relationship parameters are updated, and the influence range and degree of problem data are evaluated; the accuracy and timeliness of data dependency relationship management are improved, accurate positioning and influence range evaluation of anomalies are achieved, and it is ensured that a dependency relationship analysis result reflects the actual situation of data circulation in time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data intelligence, and in particular to a control method and control system for an integrated data intelligence platform. Background Art

[0002] As the demand for data governance continues to grow, enterprises are increasingly demanding the management of their data assets. Analyzing and managing data dependencies has become a core task of data governance, helping enterprises understand data flow, ensure data quality, and improve data utilization efficiency.

[0003] Currently, mainstream data governance platforms primarily build data dependency graphs through metadata collection and rule configuration. These platforms automatically identify and track data dependencies by analyzing data processing task logs, database table structures, and ETL job configurations.

[0004] However, the construction of data dependency relationships in existing technologies is often static and cannot dynamically adapt to changes in the data flow process. There is a lack of a quantitative assessment mechanism for the integrity of data dependency relationships. When data quality anomalies occur, it is difficult to accurately assess the scope and extent of the impact. This situation needs further improvement. Summary of the Invention

[0005] To address the problems in existing technologies where data dependency relationships are often static and cannot dynamically adapt to changes in the data flow process, and lack a quantitative assessment mechanism for the integrity of data dependency relationships, and when data quality anomalies occur, it is difficult to accurately assess the scope and extent of the impact, this application provides a control method and control system for an integrated data intelligence platform, which adopts the following technical solutions: In a first aspect, the present application provides a control method for an integrated data intelligence platform, comprising the following steps: Obtaining a data dependency analysis mode and target parameters, and obtaining overall data link tracking parameters based on the data dependency analysis mode and the target parameters; Acquire data node information, and establish node dependency parameters for each data node based on the data node information and the overall data link tracking parameters; Perform data flow analysis based on the node dependency parameters and the overall data link tracking parameters, including tracing the data source and generation process upwards, and tracing and identifying the data flow and usage downwards; Performing integrity testing on the established data dependency relationship to obtain link integrity testing parameters; comparing the link integrity testing parameters with the target parameters to obtain a link integrity difference value; If the link integrity difference value exceeds a preset threshold, data link quality monitoring is triggered; and data link abnormality information is obtained according to the data link quality monitoring result; Based on the data link anomaly information, node dependency parameters are updated, data impact assessment is performed, and the impact scope and degree of the problem data are identified.

[0006] By adopting the above technical solution, this application first configures the data dependency analysis mode and combines it with the target parameter settings to automatically generate the tracking parameters of the entire data link; then, the system obtains detailed information of each data node and establishes personalized dependency parameters for each node based on the tracking parameters; on this basis, the system can automatically perform two-way data flow analysis, which can not only trace the source and generation process of the data upward, but also identify the flow and usage of the data downward; the system continuously performs integrity detection, calculates the link integrity detection parameters and compares them with the target parameters to obtain the integrity difference value; when the difference value exceeds the preset threshold, the system will automatically trigger the quality monitoring mechanism to promptly discover and locate the abnormal situation; finally, the system will update the node dependency parameters based on the abnormal information, and accurately identify the impact scope and degree of the problem data through data impact assessment; significantly improve the accuracy and timeliness of data dependency management; through the linkage mechanism of integrity detection and quality monitoring, the accurate positioning of anomalies and accurate assessment of the impact scope are achieved; through the dynamic update mechanism of parameters, it is ensured that the dependency analysis results can promptly reflect the actual situation of data flow.

[0007] Optionally, integrity testing is performed on the established data dependency relationship to obtain link integrity testing parameters, specifically including the following steps: Calculating the dependency completeness index of the data node based on the data node information and the node dependency relationship parameters; Calculating the data link coverage of each data node according to the dependency integrity indicator; Analyze the difference between the current dependency relationship and the preset complete dependency relationship according to the data link coverage, and obtain the dependency missing value of the current node; The integrity comprehensive evaluation value of the entire data link is estimated according to the dependency missing value, and the integrity analysis of the data dependency relationship is performed according to the integrity comprehensive evaluation value to obtain the link integrity detection parameter.

[0008] By adopting the above technical solution, the system of this application will first combine the basic information of the data nodes and the established node dependency parameters to construct a calculation model, calculate the dependency completeness index of each data node, and further calculate the data link coverage of each data node based on these dependency completeness indicators. The coverage rate characterizes the degree of establishment of the dependency relationship between the node and other nodes; then, the system compares and analyzes the actual calculated data link coverage rate with the pre-set complete dependency relationship, and obtains the dependency missing value of each node through the difference measurement algorithm. This missing value intuitively reflects the incompleteness of the current dependency relationship; based on the dependency missing value of each node, the system uses a weighted calculation method to estimate the comprehensive evaluation value of the integrity of the entire data link, and generates link integrity detection parameters accordingly; by establishing a hierarchical evaluation system for integrity detection, it is ensured that the integrity detection results can fully reflect the status of the entire data link.

[0009] Optionally, the method further includes: Obtaining a dependency strength difference between adjacent data nodes, and obtaining a data dependency and a data association corresponding to the data link coverage rate according to the dependency strength difference; Constructing a data dependency topology graph corresponding to the dependency strength difference according to the data dependency and the data association; Analyzing the propagation characteristics of the data dependency relationships according to the data dependency topology graph to obtain a dependency integrity threshold for the current data environment; When the actual dependency of the data reaches the dependency integrity threshold, a dependency optimization suggestion is output to the data link quality monitoring module.

[0010] By adopting the above technical solution, the present application first obtains the difference in dependency strength between adjacent data nodes through an analysis algorithm, and calculates the data dependency and data correlation corresponding to the data link coverage based on this difference. The data dependency reflects the strength of the direct dependency relationship between nodes, while the data correlation characterizes the degree of indirect influence between nodes; then, the system uses these dependency and correlation information as a basis to construct a multidimensional data dependency topology map, which not only shows the connection relationship between nodes, but also intuitively expresses the difference in dependency strength through different connection strengths; then, the system conducts an in-depth analysis of this topology map, studies the propagation characteristics of data dependencies, including propagation paths, propagation speeds, and propagation ranges, and calculates the dependency integrity threshold under the current data environment based on this; finally, when the system detects that the actual dependency of the data reaches this integrity threshold, it automatically provides targeted optimization suggestions to the data link quality monitoring module, realizing the visual expression and quantitative analysis of the dependency relationship.

[0011] Optionally, obtaining data link abnormality information based on the data link quality assessment result includes the following steps: Obtaining data quality differences between adjacent data nodes based on the data link quality evaluation result; Segmenting the data link according to the data quality difference to generate data link segments of the same quality level; Marking data link segments whose quality levels are lower than a preset quality threshold as abnormal, and performing problem analysis on the marked abnormal quality link segments to obtain abnormal feature identification parameters of the abnormal quality link segments; According to the abnormal feature identification parameters, the impact range corresponding to the quality abnormal link segment and the processing measures corresponding to each node are analyzed to obtain the quality optimization processing parameters of the quality abnormal link segment.

[0012] By adopting the above technical solution, the present application first obtains the data quality difference between adjacent data nodes based on the data link quality assessment results through a comparative analysis algorithm. This difference can accurately reflect the quality changes of the data during the transmission process; then, the system intelligently segments the entire data link according to the calculated quality difference, and combines nodes with similar quality characteristics into data link segments with consistent quality levels. This segmentation method makes anomaly positioning more accurate; then, the system automatically marks those data link segments whose quality levels are lower than the preset quality threshold as abnormal, and conducts in-depth problem analysis on these marked quality abnormal link segments, and obtains abnormal feature identification parameters of the quality abnormal link segments through multi-dimensional feature extraction. These parameters contain key information such as the type, degree and characteristics of the anomaly; finally, based on these abnormal feature identification parameters, the system comprehensively analyzes the impact range of the quality abnormal link segments, and formulates personalized processing measures for each affected node to generate quality optimization processing parameters; through multi-dimensional analysis of abnormal features, the pertinence and effectiveness of the processing solution are ensured.

[0013] Optionally, updating node dependency parameters based on the data link anomaly information and performing data impact assessment specifically includes the following steps: Obtain the abnormal feature type and abnormal impact degree of the current data node; Analyze the impact area of ​​abnormal nodes on related nodes; Evaluate the extent to which abnormal data affects link quality; Generate optimization suggestions for the dependency relationship based on the abnormal feature type, abnormal impact level, impact area, and diffusion level.

[0014] By adopting the above technical solution, the present application first obtains the characteristic type and abnormal impact degree of the abnormal data node; then analyzes the impact area of ​​the abnormal node on its associated nodes and identifies the range of affected nodes; then evaluates the degree of diffusion of the abnormal data to the entire link quality, including the diffusion speed and range; finally, based on the acquired abnormal characteristic type, abnormal impact degree, impact area and diffusion degree information, the system automatically generates optimization suggestions for the dependency relationship; through the impact area analysis, the comprehensiveness of the impact assessment is ensured; through the diffusion degree analysis, the assessment results are made more objective.

[0015] Optionally, evaluating the extent to which abnormal data affects link quality includes the following steps: Identify the directly related nodes of the abnormal node; Assess the spread of abnormal data; Calculate the quality fluctuation degree of each associated node; A quality impact assessment report is generated based on the directly related nodes, diffusion range and quality fluctuation degree.

[0016] By adopting the above technical solution, the present application first identifies the associated nodes directly connected to the abnormal node through the dependency network to determine the first-level impact range; then evaluates the diffusion range of the abnormal data in the entire network and tracks the transmission path of the abnormal impact; then calculates the quality fluctuation degree of each affected node and quantifies the impact of the anomaly; finally, the system integrates the above analysis results and automatically generates a quality impact assessment report containing detailed impact assessment information; through the diffusion range assessment, the integrity of the impact analysis is ensured; through the quality fluctuation measurement, the assessment results are made more quantitative.

[0017] In a second aspect, the present application provides a control system for an integrated data intelligence platform, comprising: A data acquisition module is configured to obtain a data dependency analysis model and target parameters, and obtain overall data link tracking parameters based on the data dependency analysis model and the target parameters; obtain data node information, and establish node dependency parameters for each data node based on the data node information and the overall data link tracking parameters; A data flow analysis module is used to perform data flow analysis based on the node dependency parameters and the overall data link tracking parameters, including tracing the data source and generation process upwards and tracing downwards to identify the data flow and usage; An integrity detection module is used to perform integrity detection on the established data dependency relationship and obtain link integrity detection parameters; compare the link integrity detection parameters with the target parameters to obtain a link integrity difference value; a quality monitoring module, configured to trigger data link quality monitoring when the link integrity difference value exceeds a preset threshold, and obtain data link abnormality information based on the data link quality monitoring result; The impact assessment module is used to update the node dependency parameters according to the data link abnormality information, perform data impact assessment, and identify the impact scope and impact degree of the problem data.

[0018] Optionally, the integrity detection module includes: An integrity calculation unit, configured to calculate a dependency integrity index of a data node based on the data node information and node dependency parameters; A coverage calculation unit, configured to calculate the data link coverage of each data node according to the dependency integrity indicator; A missing value analysis unit is used to analyze the difference between the current dependency relationship and the preset complete dependency relationship according to the data link coverage, and obtain the dependency missing value of the current node; A comprehensive evaluation unit is used to estimate a comprehensive evaluation value of the integrity of the entire data link based on the dependency missing value, perform integrity analysis on the data dependency relationship based on the comprehensive evaluation value of the integrity, and obtain the link integrity detection parameter.

[0019] Optionally, also include: A dependency strength analysis unit, configured to obtain dependency strength differences between adjacent data nodes, and obtain data dependency and data association corresponding to the data link coverage rate based on the dependency strength differences; A topology map construction unit, configured to construct a data dependency topology map corresponding to the dependency strength difference according to the data dependency and the data association; a propagation characteristic analysis unit, configured to analyze the propagation characteristics of the data dependency relationship according to the data dependency relationship topology graph to obtain a dependency relationship integrity threshold value of the current data environment; The optimization suggestion generating unit is used to output dependency optimization suggestions to the data link quality monitoring module when the actual dependency of the data reaches the dependency integrity threshold.

[0020] Optionally, the quality monitoring module includes: A quality difference analysis unit, configured to obtain a data quality difference between adjacent data nodes based on the data link quality evaluation result; a link segmentation processing unit, configured to segment the data link according to the data quality difference, and generate data link segments of the same quality level; An anomaly identification unit is used to mark data link segments whose quality level is lower than a preset quality threshold as abnormal, and perform problem analysis on the marked quality abnormal link segments to obtain abnormal feature identification parameters of the quality abnormal link segments; The processing parameter generating unit is used to analyze the impact range corresponding to the quality abnormal link segment and the processing measures corresponding to each node according to the abnormal feature identification parameters, and obtain the quality optimization processing parameters of the quality abnormal link segment.

[0021] In summary, this application includes at least one of the following beneficial technical effects: 1. This application configures the data dependency analysis mode and target parameters, generates data link tracking parameters, obtains data node information, and establishes personalized dependency parameters for each node. The system can perform bidirectional data flow analysis, tracing the data source and generation process upward and identifying the data flow and usage downward. Through integrity detection, it calculates and compares the target parameters to obtain the integrity difference value. When the difference value exceeds the preset threshold, the quality monitoring mechanism is triggered to promptly detect and locate anomalies. Based on the anomaly information, the node dependency parameters are updated to assess the scope and extent of the impact of the problem data. This improves the accuracy and timeliness of data dependency management, achieves accurate location of anomalies and impact scope assessment, and ensures that the dependency analysis results promptly reflect the actual data flow situation. 2. This application builds a computational model based on data node information and dependency relationship parameters, calculates the dependency completeness index of each node, and further calculates the data link coverage rate to characterize the degree of establishment of dependency relationships between nodes. The system compares and analyzes the actual coverage rate with the preset complete dependency relationship, and obtains the node dependency missing value through a difference measurement algorithm to intuitively reflect the degree of incompleteness of the dependency relationship. Based on the dependency missing value of each node, a weighted calculation method is used to estimate the comprehensive assessment value of the data link integrity, and the link integrity detection parameters are generated to ensure that the detection results fully reflect the status of the data link. 3. This application obtains the dependency strength differences between adjacent data nodes through an analysis algorithm, calculates the data dependency and correlation corresponding to the data link coverage, and reflects the direct dependency strength and indirect influence degree between nodes; the system constructs a multidimensional data dependency topology map based on this information to intuitively display the node connection relationship and dependency strength differences; by analyzing the propagation path, speed, range and other characteristics of the topology map, the dependency integrity threshold is calculated; when the actual dependency reaches the threshold, optimization suggestions are provided to the quality monitoring module to realize the visual expression and quantitative analysis of the dependency relationship. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is a flow chart of a control method for an integrated data intelligence platform according to an embodiment of the present application; Figure 2This is a flow diagram of step S400 in the control method of the integrated data intelligence platform of the embodiment of the present application. Figure 1 ; Figure 3 This is a flow diagram of step S400 in the control method of the integrated data intelligence platform of the embodiment of the present application. Figure 2 ; Figure 4 This is a flow chart of step S500 in the control method of the integrated data intelligence platform according to an embodiment of the present application; Figure 5 This is a flowchart of step S600 in the control method of the integrated data intelligence platform according to an embodiment of the present application; Figure 6 This is a flowchart of step S630 in the control method of the integrated data intelligence platform according to an embodiment of the present application; Figure 7 It is a module diagram of the control system of the integrated data intelligence platform of the embodiment of the present application. DETAILED DESCRIPTION

[0023] The terms used in the following examples of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and appended claims of this application, the singular expressions "a," "an," "said," "above," "the," and "this" are intended to include plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in this application refers to any or all possible combinations comprising one or more of the listed items.

[0024] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of the technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.

[0025] The embodiments of the present application are described in further detail below with reference to the accompanying drawings.

[0026] In the first aspect, the present application provides a control method for an integrated data intelligence platform, referring to Figure 1 , including the following steps: S100: Obtain a data dependency analysis mode and target parameters, and obtain overall data link tracking parameters based on the data dependency analysis mode and target parameters.

[0027] In this embodiment, the data dependency analysis mode refers to the method for identifying dependencies between data nodes, including data table lineage analysis, field-level mapping analysis, data call relationship analysis, and business process association analysis. Target parameters include target values ​​for node coverage, link connectivity, and data transfer integrity. Overall data link tracking parameters are key configuration information guiding subsequent dependency analysis, including node identification rules, dependency determination criteria, and tracking depth limits.

[0028] Specifically, the system first selects an appropriate dependency analysis mode using a pre-established analysis mode mapping table. This mapping table stores appropriate analysis mode combinations for different data scenarios. For example, in a data warehouse scenario, a combination of table lineage analysis and field-level mapping analysis is selected. The system then sets target parameters based on historical experience, such as 95% for node coverage, 98% for link connectivity, and 90% for data transfer completeness.

[0029] S200: Acquire data node information, and establish node dependency parameters for each data node based on the data node information and overall data link tracking parameters.

[0030] In this embodiment, the data node information includes node identification information, node type information and node content information. The node dependency parameters include the direct association strength between nodes, data transmission frequency and data change impact.

[0031] Specifically, the system obtains node identification information from a configured data dictionary, reads node type information from the metadata management library, and obtains node content information through a data collection interface. Based on this acquired node information, the system uses pre-built feature extraction rules to generate node feature vectors. Combined with overall data link tracking parameters, the system uses a rule-based calculation engine to calculate dependency parameters for each node, including a correlation strength score of 0.8 with adjacent nodes, a data transmission frequency of every 5 minutes, and a moderate impact rating for data changes.

[0032] S300. Perform data flow analysis based on node dependency parameters and overall data link tracking parameters, including tracing the data source and generation process upwards, and tracing downwards to identify data flow and usage.

[0033] In this embodiment, data flow analysis includes upstream tracing analysis and downstream tracking analysis. Upstream tracing analysis is responsible for identifying the source of data generation and processing, while downstream tracking analysis is responsible for tracking the usage scenarios and distribution of data.

[0034] Specifically, the system first constructs a data flow index table to record the data transfer relationships between nodes. In upstream traceability analysis, the system queries this index table and, combined with node dependency parameters, traces the data source layer by layer. In downstream tracing analysis, the system uses the same method to trace the data flow downward and identify data usage. The system saves the complete data flow record in the flow analysis result table.

[0035] S400: Perform integrity detection on the established data dependency relationship to obtain link integrity detection parameters; compare the link integrity detection parameters with target parameters to obtain a link integrity difference value.

[0036] In this embodiment, the link integrity detection parameters reflect the completeness of the data dependency relationship, including node coverage completeness, link connectivity completeness, and data transfer completeness. The link integrity difference value represents the difference between the actual integrity and the target parameter.

[0037] Specifically, the system uses a completeness check rule base for testing. First, it calculates node coverage completeness, matching the ratio of the actual number of nodes to the expected number of nodes using a node mapping table. It then calculates link connectivity completeness, using a connectivity check module to verify the availability of data transmission channels between nodes. Next, it calculates data transmission completeness, using data sampling to measure data consistency during transmission. The system compares these test parameters with the target parameters, resulting in a 7% completeness difference.

[0038] S500: If the link integrity difference value exceeds a preset threshold, data link quality monitoring is triggered; and data link abnormality information is obtained according to the data link quality monitoring result.

[0039] In this embodiment, data link quality monitoring includes node quality monitoring, connectivity quality monitoring, and transmission quality monitoring. Data link abnormality information records specific problems found during the monitoring process and their characteristic descriptions.

[0040] Specifically, when the integrity difference value exceeds the preset 5% threshold, the system initiates the quality monitoring process. Node quality monitoring verifies the existence and standardization of node data through the node status checklist, such as finding missing node data or incorrect format; connectivity quality monitoring checks the connection status between nodes based on the link detection module, such as finding data transmission interruptions and transmission delays; and transmission quality monitoring uses data comparison rules to verify the accuracy of data transmission, such as finding data inconsistencies and data duplications. For example, if the system finds that the payment node data format in a certain order processing link is abnormal, resulting in a decrease in node coverage integrity, the connection between the order node and the payment node is interrupted, resulting in a decrease in link connectivity integrity, and the missing fields in the data transmission process result in a decrease in transmission integrity, these specific abnormal information will be recorded in the monitoring result database.

[0041] S600: Update node dependency parameters based on data link anomaly information, perform data impact assessment, and identify the impact scope and impact degree of the problem data.

[0042] In this embodiment, data impact assessment includes direct impact assessment and cascading impact assessment. Direct impact assessment evaluates the adjacent nodes of the abnormal node, while cascading impact assessment analyzes all nodes affected by the abnormal propagation path.

[0043] In one embodiment, referring to Figure 2 In step S400, integrity detection is performed on the established data dependency relationship to obtain link integrity detection parameters, which specifically includes the following steps: S410: Calculate the dependency completeness index of the data node according to the data node information and the node dependency relationship parameters.

[0044] In this embodiment, the dependency integrity indicators include the data integrity rate, update timeliness, and data quality score of the data node. The data integrity rate reflects the completeness of the node's data fields, the update timeliness indicates the effectiveness of data updates, and the data quality score indicates the accuracy of the data content.

[0045] Specifically, the system calculates dependency completeness metrics using a node evaluation rule base. Data completeness is verified using a pre-set field checklist, with the ratio of the actual number of fields to the standard number of fields serving as the completeness rate. Update timeliness is calculated using a timestamp comparison table, recording the degree to which data update times conform to the standard update cycle. Data quality is assessed using a quality scorecard, which includes scores based on three dimensions: field format standardization, reasonableness of numerical ranges, and uniqueness of key fields. The system uses the weighted sum of these three metrics as the node's dependency completeness metric.

[0046] S420: Calculate the data link coverage of each data node according to the dependency integrity index.

[0047] In this embodiment, the data link coverage refers to the degree of dependency between a data node and its associated nodes, and includes three dimensions: upstream node coverage, downstream node coverage, and associated node coverage.

[0048] Specifically, the system calculates data link coverage based on dependency completeness metrics. First, a standard dependency list is obtained through the node relationship mapping table, which records all the dependencies that nodes should establish. The system then compares the actual dependencies established and calculates coverage in three dimensions: upstream node coverage is calculated using a traceability analysis table to record the integrity of connections with upstream data sources; downstream node coverage is calculated using an application call record table to count the coverage of downstream applications; and associated node coverage is analyzed based on the business mapping table to calculate the degree of connectivity between business-related nodes.

[0049] S430: Analyze the difference between the current dependency relationship and the preset complete dependency relationship according to the data link coverage, and obtain the dependency missing value of the current node.

[0050] In this embodiment, the missing dependency value indicates the degree of incompleteness of the current node's dependency relationship. The preset complete dependency relationship is stored in a dependency standard library, which includes three benchmark indicators: the number of standard connections between nodes, the connection method, and the connection strength. The degree of difference is measured by the deviation of the actual value from the standard value.

[0051] Specifically, the system calculates dependency missing values ​​through a variance analysis module. First, it reads the preset complete dependency baseline value from the dependency standard library. This baseline value is determined based on the business scenario and data importance. The data link coverage is then compared with the baseline value. Differences in the number of connections are calculated using a node connection statistics table, recording the difference between the actual number of connections and the standard number of connections. Differences in connection methods are identified using a connection type comparison table, marking connection types that do not meet the standards. Differences in connection strength are calculated based on a weight matrix, analyzing the deviation between the actual connection strength and the standard strength. The system combines these three variance values ​​to calculate the dependency missing value.

[0052] S440: Estimate the comprehensive integrity evaluation value of the entire data link based on the dependency missing value, perform integrity analysis on the data dependency relationship based on the comprehensive integrity evaluation value, and obtain link integrity detection parameters.

[0053] In this embodiment, the comprehensive integrity evaluation value is an overall measure of the integrity of the entire data link. The link integrity detection parameters include specific detection results in three dimensions: node coverage integrity, link connectivity integrity, and data transmission integrity.

[0054] Specifically, the system first establishes an evaluation weight table, assigning different weights to nodes based on their business importance. Through weighted calculations, the dependent missing values ​​of each node are aggregated to obtain a comprehensive assessment of data link integrity. The system employs a layered evaluation mechanism for integrity analysis: the first layer checks the data status of each node using the node status table; the second layer verifies the connection status between nodes using the connectivity check table; and the third layer confirms the accuracy of data transmission based on the data consistency check table. Ultimately, link integrity test parameters encompassing these three dimensions are generated.

[0055] In one embodiment, referring to Figure 3 , the method further comprises: S450: Obtain dependency strength differences between adjacent data nodes, and obtain data dependency and data association corresponding to the data link coverage rate based on the dependency strength differences.

[0056] In this embodiment, dependency strength differences refer to the disparity in data interaction capabilities between adjacent nodes. Data dependency represents the strength of direct dependencies between nodes and is calculated based on three metrics: data call frequency, data volume, and call stability. Data relevance represents the degree of indirect influence between nodes and is calculated based on three metrics: business relevance, data similarity, and transmission path length.

[0057] Specifically, the system uses the strength calculation module to determine dependency strength differences. First, a node interaction record table is established to record data exchange between nodes: data call frequency is recorded through interface call log statistics, data volume is obtained through the data flow statistics table, and call stability is calculated through the fault record table. Data dependencies are then calculated based on these records. Simultaneously, an association analysis table is established to calculate data associations: business relevance is determined through the business rule configuration table, data similarity is compared through the field mapping table, and the transmission path length is obtained through the path tracking table.

[0058] S460: Construct a data dependency topology graph corresponding to the dependency strength difference based on the data dependency and data association.

[0059] In this embodiment, the data dependency topology diagram is a visual dependency display structure, where nodes represent data processing units, lines represent dependency relationships, line thickness represents dependency strength, and line type represents dependency direction.

[0060] Specifically, the system generates a topology graph through a graph construction module. First, a node layout table is established to specify the position distribution of nodes in the graph, with the core node in the center and the dependent nodes distributed hierarchically. Then, the connection properties are determined through the connection drawing rule table: a thick solid line is used for a dependency greater than 0.8, a thin solid line is used for a dependency between 0.5 and 0.8, and a dotted line is used for a dependency less than 0.5. The correlation is indicated by different colors: red for high correlation, yellow for medium correlation, and green for low correlation. The system applies these rules to the graph generator to construct a complete dependency topology graph.

[0061] S470: Analyze the propagation characteristics of the data dependency relationship based on the data dependency topology diagram to obtain the dependency integrity threshold of the current data environment.

[0062] In this embodiment, propagation characteristics include three dimensions: propagation path length, propagation breadth, and propagation depth. The dependency integrity threshold is a benchmark for measuring the degree of dependency establishment and consists of a base threshold and a dynamic adjustment factor. The base threshold is set based on historical experience, while the dynamic adjustment factor is calculated based on the characteristics of the current data environment.

[0063] Specifically, the system processes the data dependency topology through a feature analysis module. First, a propagation feature analysis table is established: the propagation path length is calculated using a path counter to count the longest dependency chain, the propagation breadth is calculated using a node distribution matrix to calculate the horizontal coverage, and the propagation depth is recorded using a hierarchical statistics table to record the vertical impact level.

[0064] S480: When the actual dependency of the data reaches the dependency integrity threshold, output dependency optimization suggestions to the data link quality monitoring module.

[0065] In this embodiment, the dependency optimization suggestion includes three parts: optimization goal, optimization plan, and implementation suggestion. The optimization goal specifies the specific indicators that need to be improved, the optimization plan provides feasible improvement measures, and the implementation suggestion provides specific execution steps.

[0066] In one embodiment, referring to Figure 4 In step S500, data link abnormality information is obtained according to the data link quality assessment result, which specifically includes the following steps: S510 : Obtain data quality differences between adjacent data nodes according to a data link quality evaluation result.

[0067] In this embodiment, data quality difference refers to the degree of difference in data quality between adjacent nodes. Quality difference includes three dimensions: data standardization difference, data integrity difference, and data consistency difference.

[0068] S520: Segment the data link according to the data quality difference to generate data link segments with the same quality level.

[0069] In this embodiment, a data link segment refers to a continuous sequence of nodes with similar quality characteristics. The quality level is divided into four levels: excellent, good, average, and poor. Each level has clear judgment criteria and classification thresholds.

[0070] Specifically, the system uses a segmentation processing module to perform link segmentation. First, a quality grading table is established, with rules for determining different levels. Then, segment-by-segment evaluation is performed along the data flow, grouping adjacent nodes with similar quality differences into link segments.

[0071] S530: Mark the data link segment whose quality level is lower than the preset quality threshold as abnormal, and perform problem analysis on the marked quality abnormal link segment to obtain abnormal feature identification parameters of the quality abnormal link segment.

[0072] In this embodiment, anomaly marking is a special identification of link segments whose quality level falls below a preset threshold. Anomaly feature identification parameters include three indicators: anomaly type identification, anomaly severity score, and anomaly duration. The preset quality threshold is determined based on the business importance and data sensitivity.

[0073] Specifically, the system uses the anomaly analysis module to flag and analyze problems. First, an anomaly determination table is established, marking link segments with a "poor" or "average" quality rating as abnormal. Analysis is then performed using a problem signature database. Anomaly types, such as data gaps, data skew, and data delays, are identified using a feature matching table. The degree of anomaly is calculated using a scoring rule table, quantifying the impact on a scale of 1 to 10. The duration of anomalies is determined using time window statistics.

[0074] S540: Analyze the impact range of the quality abnormal link segment and the corresponding processing measures of each node according to the abnormal feature identification parameters to obtain quality optimization processing parameters of the quality abnormal link segment.

[0075] In one embodiment, referring to Figure 5 In step S600, based on the data link abnormality information, the node dependency parameters are updated and data impact assessment is performed, which specifically includes the following steps: S610: Obtain the abnormal feature type and abnormal impact degree of the current data node.

[0076] S620: Analyze the impact area of ​​the abnormal node on the associated nodes.

[0077] In this embodiment, the impact area refers to the range of all related nodes that may be affected by the abnormal node. Associated nodes include directly associated nodes and indirectly associated nodes. Directly associated nodes are directly connected through data interaction, while indirectly associated nodes are associated through business logic.

[0078] Specifically, the system uses the association analysis module to determine the scope of impact. First, a node association table is created to record the relationships between nodes. Direct associations are identified through a data flow diagram, identifying nodes that interact with the abnormal node. Indirect associations are determined through a business rule table, identifying nodes that are business-dependent. The likelihood of each associated node being affected is then calculated: for directly associated nodes, the dependency strength matrix is ​​used to assess the probability of impact, while for indirectly associated nodes, the business importance table is used to assess the degree of impact.

[0079] S630: Evaluate the degree of spread of abnormal data on link quality.

[0080] S640: Generate optimization suggestions for the dependency relationship based on the abnormal feature type, abnormal impact level, impact area, and diffusion level.

[0081] In one embodiment, referring to Figure 6 In step S630, the degree of diffusion of abnormal data on link quality is evaluated, which specifically includes the following steps: S631. Identify directly associated nodes of the abnormal node.

[0082] In this embodiment, directly associated nodes refer to nodes that interact with abnormal nodes. Data interactions include three types: data input, data output, and data sharing. Node identification is based on two dimensions: data flow and service calls.

[0083] S632. Evaluate the spread of abnormal data.

[0084] In this embodiment, the diffusion range describes the spatial distribution characteristics of the abnormal impact. It includes two dimensions: horizontal diffusion range and vertical diffusion range. Horizontal diffusion represents the impact coverage of nodes at the same level, while vertical diffusion represents the impact transmission across nodes at different levels.

[0085] S633: Calculate the quality fluctuation degree of each associated node.

[0086] In this embodiment, the degree of quality fluctuation represents the magnitude of change in the quality of the associated node data. This includes three indicators: fluctuation in data accuracy, fluctuation in data integrity, and fluctuation in processing efficiency. Fluctuation is calculated by comparing data before and after the anomaly occurs, and the intensity of the fluctuation is measured using the standard deviation and coefficient of variation.

[0087] Specifically, the system analyzes quality changes through a fluctuation calculation module. First, a quality monitoring table is established to record the quality indicators of associated nodes. Data accuracy fluctuations are calculated using a correctness check table, which records changes in error rates. Data integrity fluctuations are measured using an integrity check table, which tracks changes in data loss rates. Processing efficiency fluctuations are measured using a performance monitoring table, which records changes in processing delays.

[0088] S634. Generate a quality impact assessment report based on directly related nodes, diffusion range and quality fluctuation degree.

[0089] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0090] On the second aspect, the present application provides a control system for an integrated data intelligence platform. The control system of the integrated data intelligence platform of the present application is described below in combination with the control method of the above-mentioned integrated data intelligence platform.

[0091] Reference Figure 7 , a control system for an integrated data intelligence platform, including: The data acquisition module is used to obtain the data dependency analysis mode and target parameters, and obtain the overall data link tracking parameters based on the data dependency analysis mode and target parameters; obtain data node information, and establish node dependency parameters for each data node based on the data node information and the overall data link tracking parameters; The data flow analysis module is used to perform data flow analysis based on node dependency parameters and overall data link tracking parameters, including tracing the data source and generation process upwards and identifying the data flow and usage downwards; The integrity detection module is used to perform integrity detection on the established data dependency relationship and obtain link integrity detection parameters; compare the link integrity detection parameters with the target parameters to obtain a link integrity difference value; A quality monitoring module is used to trigger data link quality monitoring when the link integrity difference value exceeds a preset threshold, and obtain data link abnormality information based on the data link quality monitoring result; The impact assessment module is used to update node dependency parameters based on data link anomaly information, conduct data impact assessment, and identify the scope and extent of impact of problem data.

[0092] In one embodiment, the integrity detection module includes: The integrity calculation unit is used to calculate the dependency integrity index of the data node based on the data node information and the node dependency parameters; A coverage calculation unit, used to calculate the data link coverage of each data node based on the dependency integrity index; A missing value analysis unit is used to analyze the difference between the current dependency relationship and the preset complete dependency relationship according to the data link coverage, and obtain the dependency missing value of the current node; The comprehensive evaluation unit is used to estimate the comprehensive evaluation value of the integrity of the entire data link based on the dependency missing value, perform integrity analysis on the data dependency relationship based on the integrity comprehensive evaluation value, and obtain link integrity detection parameters.

[0093] In one embodiment, it further includes: A dependency strength analysis unit is used to obtain the dependency strength difference between adjacent data nodes, and obtain the data dependency and data association corresponding to the data link coverage based on the dependency strength difference; A topology graph construction unit is used to construct a data dependency topology graph corresponding to dependency strength differences based on data dependency and data association; A propagation feature analysis unit is used to analyze the propagation features of the data dependency relationship based on the data dependency relationship topology diagram to obtain the dependency relationship integrity threshold of the current data environment; The optimization suggestion generating unit is used to output dependency optimization suggestions to the data link quality monitoring module when the actual dependency of the data reaches the dependency integrity threshold.

[0094] In one embodiment, the quality monitoring module includes: A quality difference analysis unit is used to obtain the data quality difference between adjacent data nodes based on the data link quality evaluation result; A link segmentation processing unit, configured to segment the data link according to data quality differences and generate data link segments of the same quality level; An anomaly identification unit is used to mark data link segments whose quality level is lower than a preset quality threshold as abnormal, and perform problem analysis on the marked quality abnormal link segments to obtain abnormal feature identification parameters of the quality abnormal link segments; The processing parameter generation unit is used to identify parameters according to abnormal characteristics, analyze the impact range corresponding to the quality abnormal link segment and the processing measures corresponding to each node, and obtain the quality optimization processing parameters of the quality abnormal link segment.

[0095] In one embodiment, the impact assessment module includes: Anomaly feature analysis unit, used to obtain the abnormal feature type and abnormal impact degree of the current data node; Impact area analysis unit, used to analyze the impact area of ​​abnormal nodes on related nodes; A diffusion degree evaluation unit is used to evaluate the diffusion degree of abnormal data on link quality; The optimization suggestion generation unit is used to generate optimization suggestions for dependency relationships based on the abnormal feature type, abnormal impact level, impact area and diffusion level.

[0096] In one embodiment, the diffusion degree assessment unit includes: A related node identification module is used to identify directly related nodes of abnormal nodes; Diffusion range assessment module, used to assess the diffusion range of abnormal data; The fluctuation degree calculation module is used to calculate the quality fluctuation degree of each associated node; The assessment report generation module is used to generate a quality impact assessment report based on directly related nodes, diffusion range and quality fluctuation degree.

[0097] The above are all preferred embodiments of the present application, and are not intended to limit the scope of protection of the present application. Therefore, any equivalent changes made based on the structure, shape, and principle of the present application should be included in the scope of protection of the present application.

Claims

1. A control method for an integrated data intelligence platform, characterized in that: The steps include: Obtaining a data dependency analysis mode and target parameters, and obtaining overall data link tracking parameters based on the data dependency analysis mode and the target parameters; Acquire data node information, and establish node dependency parameters for each data node based on the data node information and the overall data link tracking parameters; Perform data flow analysis based on the node dependency parameters and the overall data link tracking parameters, including tracing the data source and generation process upwards, and tracing and identifying the data flow and usage downwards; Performing integrity testing on the established data dependency relationship to obtain link integrity testing parameters; comparing the link integrity testing parameters with the target parameters to obtain a link integrity difference value; If the link integrity difference value exceeds a preset threshold, data link quality monitoring is triggered; and data link abnormality information is obtained according to the data link quality monitoring result; Based on the data link anomaly information, node dependency parameters are updated, data impact assessment is performed, and the impact scope and degree of the problem data are identified.

2. The control method of the integrated data intelligence platform according to claim 1, characterized in that: Perform integrity testing on the established data dependency relationship and obtain link integrity testing parameters. The specific steps include the following: Calculating the dependency completeness index of the data node based on the data node information and the node dependency relationship parameters; Calculating the data link coverage of each data node according to the dependency integrity indicator; Analyze the difference between the current dependency relationship and the preset complete dependency relationship according to the data link coverage, and obtain the dependency missing value of the current node; The integrity comprehensive evaluation value of the entire data link is estimated according to the dependency missing value, and the integrity analysis of the data dependency relationship is performed according to the integrity comprehensive evaluation value to obtain the link integrity detection parameter.

3. The control method of the integrated data intelligence platform according to claim 2, characterized in that: The method further comprises: Obtaining a dependency strength difference between adjacent data nodes, and obtaining a data dependency and a data association corresponding to the data link coverage rate according to the dependency strength difference; Constructing a data dependency topology graph corresponding to the dependency strength difference according to the data dependency and the data association; Analyzing the propagation characteristics of the data dependency relationships according to the data dependency topology graph to obtain a dependency integrity threshold for the current data environment; When the actual dependency of the data reaches the dependency integrity threshold, a dependency optimization suggestion is output to the data link quality monitoring module.

4. The control method of the integrated data intelligence platform according to claim 1, characterized in that: Based on the data link quality assessment results, obtain data link abnormality information, which specifically includes the following steps: Obtaining data quality differences between adjacent data nodes based on the data link quality evaluation result; Segmenting the data link according to the data quality difference to generate data link segments of the same quality level; Marking data link segments whose quality levels are lower than a preset quality threshold as abnormal, and performing problem analysis on the marked abnormal quality link segments to obtain abnormal feature identification parameters of the abnormal quality link segments; According to the abnormal feature identification parameters, the impact range corresponding to the quality abnormal link segment and the processing measures corresponding to each node are analyzed to obtain the quality optimization processing parameters of the quality abnormal link segment.

5. The control method of the integrated data intelligence platform according to claim 1, characterized in that: Based on the data link anomaly information, node dependency parameters are updated and data impact assessment is performed, specifically including the following steps: Obtain the abnormal feature type and abnormal impact degree of the current data node; Analyze the impact area of ​​abnormal nodes on related nodes; Evaluate the extent to which abnormal data affects link quality; Generate optimization suggestions for the dependency relationship based on the abnormal feature type, abnormal impact level, impact area, and diffusion level.

6. The control method of the integrated data intelligence platform according to claim 5, characterized in that: Evaluate the extent to which abnormal data affects link quality. This includes the following steps: Identify the directly related nodes of the abnormal node; Assess the spread of abnormal data; Calculate the quality fluctuation degree of each associated node; A quality impact assessment report is generated based on the directly related nodes, diffusion range and quality fluctuation degree.

7. A control system for an integrated data intelligence platform, characterized in that: include: A data acquisition module is used to obtain a data dependency analysis model and target parameters, and obtain overall data link tracking parameters based on the data dependency analysis model and the target parameters; Acquire data node information, and establish node dependency parameters for each data node based on the data node information and the overall data link tracking parameters; A data flow analysis module is used to perform data flow analysis based on the node dependency parameters and the overall data link tracking parameters, including tracing the data source and generation process upwards and tracing downwards to identify the data flow and usage; An integrity detection module is used to perform integrity detection on the established data dependency relationship and obtain link integrity detection parameters; compare the link integrity detection parameters with the target parameters to obtain a link integrity difference value; a quality monitoring module, configured to trigger data link quality monitoring when the link integrity difference value exceeds a preset threshold, and obtain data link abnormality information based on the data link quality monitoring result; The impact assessment module is used to update the node dependency parameters according to the data link abnormality information, perform data impact assessment, and identify the impact scope and impact degree of the problem data.

8. The control system of the integrated data intelligence platform according to claim 7, characterized in that: The integrity detection module includes: An integrity calculation unit, configured to calculate a dependency integrity index of a data node based on the data node information and node dependency parameters; A coverage calculation unit, configured to calculate the data link coverage of each data node according to the dependency integrity indicator; A missing value analysis unit is used to analyze the difference between the current dependency relationship and the preset complete dependency relationship according to the data link coverage, and obtain the dependency missing value of the current node; A comprehensive evaluation unit is used to estimate a comprehensive evaluation value of the integrity of the entire data link based on the dependency missing value, perform integrity analysis on the data dependency relationship based on the comprehensive evaluation value of the integrity, and obtain the link integrity detection parameter.

9. The control system of the integrated data intelligence platform according to claim 7, characterized in that: Also includes: A dependency strength analysis unit, configured to obtain dependency strength differences between adjacent data nodes, and obtain data dependency and data association corresponding to the data link coverage rate based on the dependency strength differences; A topology map construction unit, configured to construct a data dependency topology map corresponding to the dependency strength difference according to the data dependency and the data association; a propagation characteristic analysis unit, configured to analyze the propagation characteristics of the data dependency relationship according to the data dependency relationship topology graph to obtain a dependency relationship integrity threshold value of the current data environment; The optimization suggestion generating unit is used to output dependency optimization suggestions to the data link quality monitoring module when the actual dependency of the data reaches the dependency integrity threshold.

10. The control system of the integrated data intelligence platform according to claim 7, characterized in that: The quality monitoring module includes: A quality difference analysis unit, configured to obtain a data quality difference between adjacent data nodes based on the data link quality evaluation result; a link segmentation processing unit, configured to segment the data link according to the data quality difference, and generate data link segments of the same quality level; An anomaly identification unit is used to mark data link segments whose quality level is lower than a preset quality threshold as abnormal, and perform problem analysis on the marked quality abnormal link segments to obtain abnormal feature identification parameters of the quality abnormal link segments; The processing parameter generating unit is used to analyze the impact range corresponding to the quality abnormal link segment and the processing measures corresponding to each node according to the abnormal feature identification parameters, and obtain the quality optimization processing parameters of the quality abnormal link segment.

Citation Information

Patent Citations

  • Industrial data quality treatment system based on artificial intelligence

    CN118070202A

  • Transaction link tracking method and device and computer equipment

    CN118152227A

  • Full-link monitoring system and method based on business indexes and storage medium

    CN118368212A

  • Multi-source software supply chain intelligent analysis method and system

    CN119720225A

  • Intelligent control with hierarchical stacked neural networks

    US9015093B1

Cited By

  • Real-time monitoring system for electrical parameters of production and operation of wiring harness

    CN120761761A