Black box field-level blood verification method and system for cross-platform data link, and storage medium

CN122838873APending Publication Date: 2026-09-29STATE GRID TIANJIN ELECTRIC POWER COMPANY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611302319.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-26
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

当黑盒组件版本升级、运行参数调整或输入数据模式发生变化后,原有血缘关系是否继续有效难以判断,血缘关系的形成过程也缺少可供复验和审计的依据

Benefits of technology

[0012]本发明中,各探针任务在与生产环境隔离的影子执行环境中执行,并以无干预基准任务的输出波动作为响应判定依据。通过比较编码探针矩阵与输出响应综合征,可以识别受到字段干预影响的目标输出字段,降低仅依据字段取值、数据分布或统计相关性推断血缘关系所造成的误判。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838873A_ABST
    Figure CN122838873A_ABST
Patent Text Reader

Abstract

This invention discloses a black-box field-level lineage verification method, system, and storage medium for cross-platform data links. Based on an initial lineage graph, field-level lineage breakpoints and lineage verification tasks are determined, generating a separable coded probe matrix. In a shadow execution environment, a baseline task and multiple probe tasks are executed. A response syndrome is formed based on the changes in probe output relative to the baseline output, and this syndrome is combined with the coded probe matrix to identify the set of parent fields to be verified. For dependency hypotheses that cannot be directly distinguished, additional discriminative probes are added. The validity of the obtained set of parent fields to be verified is verified using negative and positive controls. After successful verification, input data snapshots, operator versions, running parameters, probe records, and output responses are saved to form lineage evidence. The verified field dependencies are written into the initial lineage graph to obtain the field-level lineage graph. This invention achieves black-box field-level lineage verification for cross-platform data links.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of data governance, data lineage analysis, and data processing link testing, and in particular to a black-box field-level lineage verification method, system, and storage medium for cross-platform data links. Background Technology

[0002] As the scope of power data access continues to expand, cross-professional and cross-system data collaboration is becoming increasingly frequent. The same business data typically undergoes multiple processing stages during its formation and use, including data acquisition interfaces, message transmission channels, extraction, transformation, and loading tasks, stored procedures, user-defined functions, closed computing components, and data service interfaces. These processing stages may be distributed across relational databases, distributed data warehouses, and big data computing engines. Different platforms differ in field naming, data types, task execution cycles, and version management methods, giving the data chain characteristics of being cross-platform, cross-engine, cross-level, and constantly changing.

[0003] Existing methods for constructing data lineages include SQL script parsing, metadata association, task scheduling relationship collection, runtime log analysis, and manual configuration. For nodes whose processing scripts are accessible and whose syntax can be parsed, relationships between fields are identified based on operations such as field references, projections, filtering, joins, and aggregations within the script. However, in actual data chains, there are also processing nodes such as stored procedures, user-defined functions, enclosed extraction and transformation loading components, third-party interfaces, and dynamically generated tasks at runtime. Their internal logic is usually not directly accessible or parsed. In such cases, existing systems can generally only identify task-level relationships between input tables, processing tasks, and output tables, making it difficult to further determine the correspondence between output fields and input fields.

[0004] When the aforementioned black-box processing nodes exist in the data chain, field-level lineage typically breaks at those nodes. For example, the system can confirm that an input table is written to the target table via a black-box component, but it cannot determine which input fields the device identifier, acquisition time, measurement value, or region identifier in the target table originates from. The lineage constructed in this way can only support table-level or task-level data traceability, making it difficult to accurately analyze the downstream fields, business metrics, and data services affected by changes in a particular upstream field.

[0005] Some existing methods infer field relationships based on the consistency of values, data distribution, or statistical correlation between source and target fields. These methods are easily influenced by business characteristics in power data scenarios. Data such as power load, power consumption, current, and equipment status typically exhibit temporal synchronization, spatial correlation, and periodicity. Even if there is no direct data transformation relationship between two fields, they may show a high correlation due to common temporal trends or similar upstream factors. Furthermore, after normalization, truncation, enumeration mapping, aggregation, grouping, or encoding, the input and output values ​​of fields may no longer have a direct correspondence. Therefore, relying solely on historical data value matching or statistical correlation is insufficient to reliably distinguish between actual field dependencies and incidentally formed statistical correlations.

[0006] Another verification method involves modifying each candidate input field sequentially, repeatedly executing the verification process, and determining the field relationships based on whether the output changes. With this method, each candidate field typically requires separate test data construction and one or more processing tasks, increasing the number of verification attempts with the number of candidate fields. Verification costs are high when multiple fields such as device, measurement point, region, time, and status exist upstream of the black-box node. Furthermore, data processing results may be affected by factors such as batch data volume, task concurrency, time windows, and non-deterministic aggregation; a change in output at one point may not be entirely caused by that field modification. Without baseline fluctuation measurement, cross-process identification, and anomaly handling mechanisms, fluctuations in the process itself or task execution failures may be misjudged as field dependencies.

[0007] Directly writing test data into the production chain can contaminate official data, triggering data quality alarms, statistical indicator calculations, or downstream business processing. Therefore, it is not suitable for direct application in scenarios such as power distribution operation and production management. Existing verification methods typically only retain the final lineage relationship, without fully recording the input data snapshots, processing component versions, operating parameters, test fields, and their output responses used during verification. When black-box component versions are upgraded, operating parameters are adjusted, or input data patterns change, it is difficult to determine whether the original lineage relationship remains valid, and the formation process of the lineage relationship lacks evidence for verification and auditing. Summary of the Invention

[0008] The purpose of this invention is to overcome the shortcomings and defects of existing technologies by providing a cross-platform data link black-box field-level lineage verification method, system, and storage medium. This method intervenes in candidate input fields using constrained coding probes, executes the link to be verified in an isolated shadow execution environment, and identifies the dependency relationship between candidate input fields and target output fields based on the response results of the target output field. This is used to supplement missing field-level lineage at black-box nodes, reduce the possibility of misjudgments caused by passive correlation analysis, reduce the repetitive execution required for field-by-field verification, and retain verification evidence for re-verification and auditing.

[0009] According to a first aspect of this application, a black-box field-level lineage verification method for cross-platform data links is provided, comprising: Based on the initial lineage graph, black-box field-level lineage breakpoints corresponding to black-box operators in the cross-platform data link are identified, and a lineage verification task is generated. The lineage verification task includes a snapshot of input data frozen within the verification boundary, the black-box operator version and running parameters, the target output field and the corresponding candidate input field. The verification boundary is determined by the direct input checkpoints and direct output checkpoints of the black-box operators. Based on the candidate input fields and preset field constraints, a separable encoded probe matrix is ​​generated. Based on the baseline values ​​in the input data snapshot, the encoded values ​​in the encoded probe matrix are mapped to probe values ​​for intervention operations on the candidate input fields, forming multiple probe micro-batches. In the shadow execution environment, based on the input data snapshot, black box operator version and running parameters, the response judgment threshold is first calculated based on the difference set formed by repeatedly executing the benchmark task multiple times based on the benchmark value. Then, multiple probe tasks are executed based on multiple probe micro-batches. The normalized difference between the output of multiple probe tasks under the target output field and the benchmark output is calculated and compared with the response judgment threshold to generate the response syndrome of the target output field under multiple probe tasks. Based on the encoded probe matrix, generate the ideal response corresponding to the preset dependency hypothesis under the probe task, calculate the observation inconsistency loss between the ideal response and the response syndrome; based on the observation inconsistency loss, the decoding loss of the optimal dependency hypothesis and the confidence interval, decode the set of parent fields to be verified of the target output field within the verification boundary. For the set of parent fields to be validated, perform validation on negative controls without probe intervention and positive controls with probe intervention, and output the validation results of candidate input fields based on the control pass rate; A lineage evidence summary is generated based on the input data snapshot, black-box operator version and running parameters, encoded probe matrix, response syndrome, optimal dependency hypothesis and control pass rate, and a lineage evidence credential is generated; verification lineage edges pointing to the target output field are established for the verified candidate input fields, and the verification lineage edges are written into the initial lineage graph to obtain a field-level lineage graph with verification status and version range.

[0010] According to a second aspect of this application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the black-box field-level lineage verification method for the cross-platform data link.

[0011] According to a third aspect of this application, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the black-box field-level lineage verification method for cross-platform data links as described above.

[0012] In this invention, each probe task is executed in a shadow execution environment isolated from the production environment, and the output fluctuation of the uninterrupted baseline task is used as the basis for response judgment. By comparing the encoded probe matrix with the output response syndrome, the target output field affected by field intervention can be identified, reducing misjudgments caused by inferring lineage solely based on field values, data distribution, or statistical correlation.

[0013] In this invention, when multiple dependent hypotheses cannot be distinguished based on existing response results, a legitimate probe with high hypothesis separation capability is selected for additional verification. For the set of parent fields to be verified, a negative control is used to check its stability under no-intervention conditions, and a positive control is used to check the effectiveness of probe intervention and response acquisition processes, thereby reducing the impact of environmental anomalies or acquisition anomalies on the verification results.

[0014] In this invention, the input data snapshots, black-box operator versions and operating parameters, encoded probe matrices, output responses, and decoding results used during the verification process are all recorded in the lineage evidence certificate. Verified field-level lineage edges are saved in association with the evidence summary, operator version, and valid range, and can be used for subsequent review and auditing; when the operator version, operating parameters, or input mode changes, the relevant lineage edges are marked as requiring review.

[0015] The method of this invention is applicable to the data processing link between relational databases, distributed data warehouses, big data computing platforms, data governance and scheduling platforms, and data services, and is especially applicable to scenarios where the internal processing logic, such as stored procedures, user-defined functions, closed extraction and transformation loading components, and third-party data interfaces, cannot be directly parsed. Attached Figure Description

[0016] Figure 1 This is a flowchart of the black-box field-level lineage verification method for cross-platform data links according to the present invention. Detailed Implementation

[0017] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0018] like Figure 1 As shown, the black-box field-level lineage verification method for the cross-platform data link includes: S1. Based on the initial lineage graph, identify the black-box field-level lineage breakpoints corresponding to black-box operators in the cross-platform data link and generate a lineage verification task; Specifically, the lineage verification task includes a snapshot of input data frozen within the verification boundary, the black-box operator version and running parameters, the target output field and the corresponding candidate input field; the verification boundary is determined by the direct input checkpoint and direct output checkpoint of the black-box operator. S2. Generate a separable encoded probe matrix based on the candidate input fields and map it into multiple probe microbatches; Specifically, based on the candidate input fields and preset field constraints, a separable encoded probe matrix is ​​generated. Then, based on the baseline values ​​in the input data snapshot, the encoded values ​​in the encoded probe matrix are mapped to probe values ​​for intervention operations on the candidate input fields, forming multiple probe micro-batches for executing probe tasks. S3. Execute the baseline task and probe task in the shadow execution environment, and generate the response syndrome of the target output field under the probe task; Specifically, in a shadow execution environment isolated from the production environment, based on the input data snapshot, black box operator version and running parameters, the response judgment threshold is first calculated based on the difference set formed by repeatedly executing the benchmark task multiple times based on the benchmark value. Then, multiple probe tasks are executed based on multiple probe micro-batches. The normalized difference between the output of multiple probe tasks under the target output field and the benchmark output is calculated and compared with the response judgment threshold to generate the response syndrome of the target output field under multiple probe tasks. S4. Based on the response syndrome, decode the set of parent fields to be verified for the target output field within the verification boundary; Specifically, based on the encoded probe matrix, the ideal response corresponding to the preset dependency hypothesis under the probe task is generated, and the observation inconsistency loss between the ideal response and the response syndrome is calculated; based on the observation inconsistency loss, the decoding loss of the optimal dependency hypothesis, and the confidence interval, the set of parent fields to be verified of the target output field is decoded within the verification boundary; a preset dependency hypothesis consists of at least two candidate input fields; the confidence interval is determined by the difference between the observation inconsistency loss of the suboptimal dependency hypothesis and the optimal dependency hypothesis. S5. Perform positive and negative control verification on the set of parent fields to be verified; Specifically, for the set of parent fields to be validated, perform validation on negative controls without probe intervention and positive controls with probe intervention, and output the validation results of candidate input fields based on the control pass rate; S6. Generate bloodline evidence credentials, establish verification bloodline edges, and write them into the initial bloodline graph to obtain a field-level bloodline graph; Specifically, a lineage evidence summary is formed based on the input data snapshot, black-box operator version, running parameters, coding probe matrix, response syndrome, optimal dependency hypothesis, and control pass rate, and a lineage evidence credential is generated; verification lineage edges pointing to the target output field are established for the verified candidate input fields, and the verification lineage edges are written into the initial lineage graph to obtain a field-level lineage graph with verification status and version range.

[0019] This invention determines field-level lineage breakpoints and lineage verification tasks based on an initial lineage graph, and generates a separable coded probe matrix according to preset field constraints and sparse dependency conditions. Subsequently, in a shadow execution environment isolated from the production environment, a baseline task and multiple probe tasks are executed based on fixed input data snapshots, operator versions, and running parameters. A response syndrome is formed based on the changes in the target output field relative to the baseline output, and the set of parent fields to be verified is identified in conjunction with the coded probe matrix. For dependency hypotheses that cannot be directly distinguished, additional legitimate probes with discriminative capabilities are added. The set of parent fields to be verified is tested for the effectiveness of the verification process using negative and positive controls. After successful verification, information such as input data snapshots, black-box operator versions, running parameters, probe records, and output responses are saved to form lineage evidence. The verified field dependencies are written into the initial lineage graph to obtain a field-level lineage graph, thereby realizing black-box field-level lineage verification of cross-platform data links.

[0020] The technology of this invention enables proactive intervention on candidate input fields that satisfy data constraints and are mutually distinguishable, without parsing the internal processing logic of black-box nodes. It executes the data link to be verified in a shadow execution environment isolated from production operations, and identifies the dependencies between input and output fields based on multiple execution results. For verification results that cannot be directly determined, additional probes are used to further distinguish candidate relationships, and the input data snapshots, black-box operator versions, running parameters, test processes, and output responses used for verification are fully preserved to supplement the missing field-level lineage at the black-box nodes. This provides verifiable field relationships for data traceability, data quality governance, change impact analysis, and link maintenance.

[0021] This invention determines black-box field-level lineage breakpoints and lineage verification tasks based on the initial lineage graph, and freezes verification boundaries, input data snapshots, black-box operator versions, and running parameters. This allows for the verification of field dependencies of black-box nodes within a defined link range, supplementing missing field-level lineage edges in the existing lineage graph.

[0022] The black-box operator is a copy or replica of a closed transformation component that is not accessible to the internal code of a big data computing center. It is used to process the input fields and form the target output fields.

[0023] In this application, the lineage verification task refers to the field relationships in a cross-platform link that cannot be determined by static parsing or runtime logs. By identifying and determining lineage verification tasks with clear boundaries, the range of candidates that subsequent proactive testing needs to process can be reduced. The expression for the lineage verification task is as follows: ; ; In the formula, For the first A bloodline verification task; For the first Target output fields in a bloodline verification task Corresponding candidate input fields; For the first Verification boundaries in a lineage verification task; For the first A snapshot of the input data from a lineage verification task; For the first Black-box operator version in a bloodline verification task; For the first Black-box operator execution parameters in a lineage verification task; For the first In a bloodline verification task, the target output field is... The set of upstream fields obtained by reverse search along the initial lineage graph; For the first The first bloodline verification task One candidate input field; for arrive The shortest reachable path length; The preset maximum verification depth; For field types in metadata; symbol This indicates that the data types of the two fields satisfy the type conversion rules allowed by the link to be verified.

[0024] The initial lineage graph can be constructed based on metadata, scheduling relationships, and runtime logs. It identifies unresolvable black-box field-level breakpoints and forms verification tasks by freezing verification boundaries, input data snapshots, and black-box operator versions. For example, it collects table, field, and data type information from a metadata management system, task dependencies and runtime parameters from a task scheduling system, available scripts from a database or computing engine, and task call and data transmission records from a log system. By unifying node identifiers, field identifiers, task identifiers, and black-box operator version identifiers, and performing syntax tree analysis on the resolvable scripts, it extracts field references, projections, filters, joins, and aggregation relationships. These relationships are then associated with the table structure and field types in the metadata to obtain preset data. The script parsing relationships, scheduling call relationships, and runtime transmission relationships are then written into the initial lineage graph. ,in, It is a collection of field nodes, data object nodes, and processing task nodes. This involves constructing an initial lineage graph by including parsed edges, log-inferred edges, and edges to be verified. Then, based on this initial lineage graph, black-box operators that only have task-level connections but no defined field-level edges are marked as lineage breakpoints. The first checkpoint in the direct output checkpoint of the black-box operator is then used as the lineage breakpoint. Using the target output field as the endpoint, a reverse search is performed along the initial lineage graph to determine the input checkpoint closest to the black-box operator and allowing probe data to be written, thus identifying candidate input fields. Candidate input fields are limited to those directly provided to the black-box operator by the input checkpoint. Other fields upstream of the input checkpoint are only used for link tracing and are not considered as candidate input fields for decoding the direct parent field. Field type is only used to determine the range of candidate input fields and is not used as a basis for determining the validity of field dependencies.

[0025] Specifically, based on the identified individual black-box operators in the link to be verified, the direct input checkpoint and direct output checkpoint of the black-box operator are determined as verification boundaries, and the input data snapshot, the version of the black-box operator and its dependent components (black-box operator version), and the running parameters are frozen. The verification boundaries do not contain other unverified black-box operators; when multiple consecutive black-box operators exist, corresponding verification tasks are established for each.

[0026] By utilizing the upstream reachability and type constraints provided by the initial lineage graph, the full-link search is transformed into a local verification problem with controllable input and observable output. Furthermore, by freezing boundaries and using black-box operator versions, it is ensured that subsequent probes face the same data state and the same processing logic.

[0027] In summary, in this application, the initial lineage graph is used to limit the upstream reachability and type conditions of candidate input fields. After determining the field-level lineage breakpoints, the input checkpoints, output checkpoints, input data snapshots, black-box operator versions, and runtime parameters of the link to be verified are frozen, thereby limiting the verification scope to the link interval where the input is modifiable and the output is observable. Subsequent probe tasks are all executed based on the same input data snapshot and black-box operator version to ensure the comparability of the output results.

[0028] Specifically, the field type, value range, non-null constraint, uniqueness constraint, primary and foreign key relationship, and business validation rules of each candidate input field are obtained from the metadata and data quality rule library to form preset field constraints.

[0029] In this application, the encoded probe matrix is ​​represented as follows: ,satisfy: ; In the formula, Indicates the first Encoding probe matrix for a lineage verification task, number of candidate input fields , For the number of probe missions. This is the index for the number of probe tasks, with a value of [value]. ; and Index the candidate input field column, with a value of ; For the first Does the second probe task affect the first? Candidate input fields A binary indicator of the intervention being applied; Not including the first Columns and their size does not exceed A subset of candidate input fields; This is the maximum number of direct parent fields. The allowed number of abnormal responses. and These are all configuration parameters set based on the historical link structure and the stability of the shadow execution environment, and are not obtained through model training.

[0030] The encoded probe matrix of this application, based on the upper limit of the direct parent field and the allowed number of abnormal responses, makes any field... Relative to any not exceeding The combination of other fields has at least the following characteristics: The isolation response opportunity; the encoded probe matrix of this application can be selected from a preset code library, or it can be obtained by gradually increasing the number of probe tasks. And verify the generation.

[0031] The process of mapping encoded values ​​in the encoded probe matrix to probe values ​​for intervention operations on candidate input fields, forming multiple probe micro-batches, includes: For the Candidate input fields , No. The probe value for the second probe mission is obtained using the following formula: ; In the formula, For the first The baseline values ​​for each candidate input field are frozen in a snapshot of the input data. For the first The second probe mission was for the first The probe values ​​written to each candidate input field; In order to be with the first Constraint mapping functions corresponding to the candidate input field types; For the first The intervention magnitude of each candidate input field belongs to preset parameters or non-training parameters calculated from the statistical range of the field. For the first Candidate input fields The set of valid value ranges and relational constraints are derived from metadata, constraint definitions, and business validation rules; For numerical fields, the intervention range is limited by the following formula. It balances detectability and legality.

[0032] ;

[0033] In the formula, and These are the lower and upper bounds of the field's valid value range, derived from metadata or historical valid samples; The relative intervention ratio is a preset parameter. ; The maximum permissible level of intervention for the business, dimensions and fields Consistent; The probe values ​​are validated for field mode, value range, non-null condition, primary and foreign key relationship and business rules. The probe values ​​that pass the validation are grouped into multiple probe micro-batches according to the row number of the encoded probe matrix.

[0034] Among them, when At that time, the baseline value is frozen in the input data snapshot. Within the limits allowed by field constraints, alternative values ​​that are distinguishable from the baseline value are generated as probe values ​​or probe data. The encoded values ​​can be mapped to probe values ​​or probe data for field intervention operations column by column.

[0035] In this application, based on the determined candidate input fields and verification boundaries, an encoded probe matrix is ​​used to assign a distinguishable intervention mode across multiple executions to different candidate input fields, and the abstract encoding is converted into probe data or probe values ​​that meet the data constraints.

[0036] Specifically, for join keys, pairwise insertion or pairwise modification is used to maintain referential integrity; for group aggregation, probe records with unique group identifiers are generated; and for candidate input fields such as date, enumeration, and identifier, intra-domain substitution values ​​are used. Subsequently, each probe micro-batch undergoes mode validation, primary and foreign key validation, and business rule validation. The probe values ​​that pass the validation are then assembled into a valid probe micro-batch. This forms the corresponding probe execution plan.

[0037] The encoded probe matrix described in this application specifies the intervention state of each candidate input field at different times, enabling multiple candidate input fields to be combined and verified in the same set of probe tasks. When there are many candidate input fields and few actual dependent fields, the number of verification steps required to execute the link field by field can be reduced. After field constraint mapping and validity verification, the probe values ​​can meet the requirements of data type, primary and foreign key relationships, and business rules.

[0038] In summary, in this application, the intervention status of each candidate input field in multiple probe tasks is determined by the encoded probe matrix. Different combinations of candidate input fields correspond to different encoding modes. The encoded values ​​in the encoded probe matrix are converted into valid probe data after constraint mapping to meet the requirements of field type, primary and foreign key relationship and business rules, and are used for subsequent shadow link execution.

[0039] In the embodiments of this application, in step S3, the link to be verified is reproduced in a shadow execution environment isolated from the production environment, based on the link state to be verified frozen in step S1 and the probe micro-batch, without contaminating the production data, and each output change is converted into a response syndrome that can be compared across tasks.

[0040] In this step, the first step is based on the verification boundary and the input data snapshot. Black-box operator version and running parameters The system replicates input checkpoints, black-box operators, and output checkpoints within an isolated namespace to construct a shadow execution environment. A write interceptor can be set after the output checkpoint to attach a batch identifier to the probe batch that is only recognizable within the shadow execution environment, preventing probe data and its output from entering the downstream of the production process. The production environment provides only read-only structure information, and probe data is only written to the input checkpoints of the shadow execution environment.

[0041] Then, the baseline micro-batch (i.e., the baseline task) without intervention is repeatedly executed to estimate the natural fluctuation of the target output field. The probe plan is then executed, and each legal probe micro-batch is executed in sequence. The execution status, output record, and output field value of each probe are recorded. The output under the baseline task and the output under the probe task are aligned using stable business keys, record identifiers, or record fingerprints. If alignment cannot be achieved line by line, multiset statistics, group counting, and verification summaries are used to form a set-level response.

[0042] The generation of the response syndrome of the target output field under the multiple probe tasks includes: Calculate the normalized difference between the target output field and the baseline output under the probe task based on the data type of the target output field; based on The difference set formed by repeated execution of the sub-undisturbed baseline value Calculate the response judgment threshold: ; In the formula, The number of repetitions of the baseline value, taken as an integer not less than 2; The median; This represents the absolute deviation of the median. The preset fluctuation amplification factor is an empirical parameter and ; Output fields for the target The response determination threshold; if the baseline output is stable, then At least the preset minimum detection resolution should be used.

[0043] Based on the response determination threshold and the normalized difference, determine the first... Second response: ; in, For the response under the probe task, This indicates that a detectable change has occurred in the output. This indicates that the output has not changed beyond the natural fluctuation range. This indicates a missing response. To normalize the differences; The response syndrome is obtained by arranging the responses from multiple probe missions in sequence. superscript Indicates transpose; Indicates the first Response under the next probe mission; The calculation of the normalized difference between the probe task output and the baseline output, based on the data type of the target output field, includes: For numerical target output fields, the normalized difference is calculated using the following formula. :

[0044] In the formula, This is the number of output records where the probe task and the baseline task were successfully aligned, and it is a positive integer. To align the record index; and These are the first and second target output fields in the probe task and the benchmark task, respectively. One alignment value; As a robust metric for the baseline output field, it can be the interquartile range or the absolute deviation of the median, with dimensions similar to... Consistent; To prevent stable terms with positive denominators of zero, dimensions and Consistent It is a dimensionless non-negative number; For enumerated, character, or boolean target output fields, the normalized difference is calculated using the following formula. :

[0045] In the formula, This is an indicator function; it takes the value when the condition is true. Otherwise take ; , indicating the proportion of the alignment record that has changed; For target output fields that cause changes to the record set due to joins / aggregations, the normalized difference is calculated using the following formula. : ; In the formula, and These are the number of output records for the probe task and the baseline task, respectively, and are non-negative integers. and The verification digests are calculated from the multiset of sort-independent output records of the probe task and the baseline task, respectively. and To predefine non-negative weights, representing the importance of changes in the number of records and changes in the set content respectively, the sum of the two is: .

[0046] In this application, in step S3, each probe task is executed under the same input data snapshot, black-box operator version, and operating parameters. An interceptor is set at the exit of the link to be verified to prevent the probe output from entering the downstream of the formal business. Before executing the probe task, the output fluctuation of the link to be verified is measured through a non-interventional benchmark task, and the response judgment threshold is determined accordingly. After comparing the output of each probe task with the benchmark output, the difference is calculated according to the data type of the target field, and a response syndrome is generated for subsequent dependency decoding.

[0047] According to the embodiments of this application, with detectable intervention conditions as the premise of decoding, for field combinations that may cancel out numerical values, a unidirectional restricted incremental probe or a pair of positive and negative polarity probes is used, and the case where either of the pair of polarity probes exceeds the response determination threshold is recorded as a response occurring under the encoding probe task; if multiple legitimate interventions to a candidate input field cannot form a detectable response, the relationship of the field is kept in a state pending verification, and is not forcibly confirmed based on subsequent matching results.

[0048] Specifically, in this application, the step of generating the ideal response corresponding to the preset dependency assumptions under the probe task based on the encoded probe matrix, and calculating the observation inconsistency loss between the ideal response and the response syndrome, is mainly achieved by enumerating sparse dependency assumptions whose size does not exceed that of the parent field, including: For any dependency hypothesis and Define field to indicate vector Based on the encoded probe matrix, the ideal response under the probe task is calculated using the following formula:

[0049] In the formula, Indicates the first The candidate input fields belong to the dependency hypothesis. Otherwise ; For the probe task, is it related to the first... A binary indication of the intervention applied to each candidate input field; Indicating dependence on the hypothesis When established, the probe task intervenes in at least one ideal response corresponding to a real input field; Calculate the observation inconsistency loss between the ideal response and the response syndrome under each dependency assumption:

[0050] In the formula, The number of valid responses, which is a positive integer; The observation inconsistency loss under the dependency assumption is denoted by , where represents the dependency assumption. Unexplained percentage of valid responses; This represents the set of valid response counts.

[0051] The optimal dependency hypothesis is determined by ranking the observation inconsistency losses under each dependency hypothesis from smallest to largest. ; In the formula, Output fields for the target The optimal dependency assumption; The confidence interval is determined by the following formula: ; In the formula, The confidence interval is a dimensionless non-negative number. The larger the value, the easier it is to distinguish between the optimal dependency hypothesis and the suboptimal dependency hypothesis. This indicates the suboptimal dependency hypothesis.

[0052] When the observation inconsistency loss of the optimal dependency hypothesis and At that time, the optimal dependency assumption will be... The output is the set of parent fields to be validated; otherwise, the observation inconsistency loss will not exceed [a certain value]. The dependency assumptions constitute the ambiguous hypothesis set. , Preset the range of adjacent losses; The maximum allowable decoding loss threshold, Both the minimum confidence interval threshold and the minimum confidence interval threshold are preset judgment parameters with a value range of [value range missing]. .

[0053] Furthermore, if If the value is empty, the current decoding will not be performed and the process will proceed to the retry branch in step S5.

[0054] Since the coded probe matrix of this application enables different combinations of sparse fields to have different ideal responses, step S4 can use observation inconsistency loss to transform field dependency identification into a combination response matching problem, thereby avoiding inferring lineage based solely on field names, identical field values, or statistical correlation.

[0055] In summary, in step S4, under the condition of detectable intervention, when at least one actual dependency field in the probe task is intervened, the change in the target output field should exceed the baseline fluctuation range; when no actual dependency field is intervened, the target output field should, in principle, not produce a change exceeding the baseline fluctuation range. The encoded probe matrix makes different sparse dependency hypotheses correspond to distinguishable ideal response patterns. By calculating the inconsistency loss between the observed response and each ideal response pattern, the explanatory power of different dependency hypotheses for the actual response results can be compared. Therefore, based on the intervention status of each candidate input field and the response result of the target output field, within the verification boundary, a set of parent fields to be verified that can explain the response results can be determined. Since the candidate input fields are limited to the direct input fields of the black-box operator, the verified candidate input fields can be written into the lineage graph as the direct parent fields of the target output field.

[0056] In this embodiment of the application, in step S4, after the probe micro-batch processing fails to output the set of parent fields to be verified, a set of ambiguous hypotheses is generated. Then, regarding the set of ambiguous hypotheses Add legitimate probes and intervene using the aforementioned method to verify the effectiveness of the shadow execution and response acquisition process until the preset conditions are met: a certain dependency hypothesis meets the verification condition, the number of executed additional probes reaches the preset number, and the remaining dependency hypotheses cannot be distinguished under the existing constraints.

[0057] Specifically, after outputting the set of ambiguous hypotheses, adaptive appending probes are generated for the ambiguous dependency hypotheses in the set. Multiple probe tasks are then executed in the shadow execution environment based on these adaptive appending probes. Among these, appendable probes that satisfy preset field constraints and have not yet been executed are appended by evaluating the hypothesis separation capability of the probe vector, including: First, calculate the separation score for additional probes:

[0058] In the formula, The separation score for additional probes is a dimensionless, non-negative number, representing the number of weighted hypothesis pairs that the probes can distinguish. This is an executable append probe encoding vector. For dependency hypothesis The single predicted response, For dependency hypothesis The single predicted response; This is used to ensure that each pair of different dependency assumptions is calculated only once; and Dependency Hypothesis , Normalized weights, This is a set of ambiguous hypotheses.

[0059] Secondly, based on the separation score of the additional probe, an adaptive additional probe is obtained: ; In the formula, A set of probe vectors that satisfy field constraint validation but have not yet been executed; The actual appended probe encoding vector, i.e. the generated adaptive appended probe; Finally, based on the adaptive append probes, corresponding probe micro-batches are generated through constraint mapping. The probe task is executed in the shadow execution environment to obtain new responses to the target output field. Inconsistent dependency assumptions are deleted based on the new responses, and the decoding loss and confidence interval of the response syndrome and optimal dependency assumption are updated. This process is repeated multiple times until a preset condition is met, resulting in the set of parent fields to be verified. Empty or all executable probes If all values ​​are zero, the ambiguity is retained and the case is transferred to manual review.

[0060] For the set of parent fields to be validated, which consists of candidate input fields to be validated, negative controls without probe intervention and positive controls with probe intervention are performed. The control pass rate is calculated, and the control validation judgment result is output based on the control pass rate, including: In the shadow execution environment, a negative control without probe intervention is executed on the input field, and a positive control with probe intervention is executed on the input field with known field dependencies within the validation boundary; the control pass rate is calculated, and the pass status is output based on the control pass rate; let the total number of control tasks be... , No. The expected response of the control task is The actual response is The control pass rate is:

[0061] In the formula, This indicates the pass rate of the control group. , ; ; Among them, satisfying , and When the verification result is passed, This is a preset control pass rate threshold.

[0062] Based on the updated decoding results and comparison task results, determine the set of parent fields to be verified and their verification status, and record the decoding loss, decoding interval, comparison pass rate, and all appended probes. Verified probe tasks proceed to step S6 to generate lineage evidence credentials; tasks with remaining ambiguity or invalid execution processes enter the review queue, and no verified field edges are written to the initial lineage graph.

[0063] In step S5, for dependency hypotheses that cannot yet be distinguished, additional interventions are selected based on the hypothesis separation capabilities of the additional probes. For the set of parent fields to be verified, a negative control is used to check whether the link to be verified in the shadow execution environment remains stable under no-intervention conditions, and a positive control is used to check whether the probe intervention and output response acquisition process are effective. By confirming the corresponding field dependencies only when both the decoding result and the control task meet the judgment criteria, the accuracy and precision of the verification are improved.

[0064] In this embodiment of the application, based on the results of steps S1 to S5, the verification process and results of steps S1 to S5 are summarized to form a verifiable bloodline evidence certificate, and the verified field dependencies are written back to the bloodline graph with field-level edges containing verification status, version and valid range.

[0065] According to an embodiment of this application, the kinship evidence certificate includes at least the kinship evidence summary; it also includes information related to each probe task, such as execution identifier, output difference metric, response threshold, decoding loss, ambiguous spoofing device selection, execution environment identifier, verification time, and verification status.

[0066] The bloodline evidence summary is derived from the input data snapshot. Black-box operator version and operating parameters Encoding probe matrix Response Syndrome Optimal dependency assumption and control pass rate Formed by serializing and concatenating according to a fixed field order: ; In the formula, A deterministic hash function used for data integrity verification; symbol This indicates that serialization and concatenation are performed according to a fixed field order; For the first A fixed-length summary of kinship evidence for each verification task; For each candidate input field that passes validation Establish a pointer to the target output field. The verification of bloodline edges is determined by associating the bloodline evidence summary and the valid range of the operator version using the following formula: ; In the formula, This is the start time after the black-box operator version verification is passed; The effective end time; To verify the status; To verify blood relations.

[0067] Write the verified lineage edges into the initial lineage graph with valid ranges to obtain the updated field-level lineage graph. Finally, output bloodline evidence and field-level bloodline diagram. Each bloodline edge is associated with and saved along with its corresponding input data snapshot, black-box operator version, probe record, and response result, serving as the basis for subsequent verification and auditing; the applicable scope of the bloodline edge is limited by its version information and effective range.

[0068] In this application, after completing a task verification, when a change is detected in the black-box operator version, running parameters, or input mode, the valid interval of the relevant lineage edge is terminated, marked as pending verification or invalid, and steps S1 to S5 are re-executed for verification according to the changed link state.

[0069] According to a second aspect of this application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the black-box field-level lineage verification method for the cross-platform data link.

[0070] According to a third aspect of this application, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the black-box field-level lineage verification method for cross-platform data links as described above.

[0071] The following describes the implementation process of this invention using a batch processing link in a cross-platform data middleware.

[0072] The data processing chain includes a relational database node A, a distributed data warehouse node B, a big data computing node C, a data governance and scheduling node D, and a data service node E, with an independent test resource pool F. Relational database node A stores device operation records; data governance and scheduling node D submits incremental data to big data computing node C according to a preset period; big data computing node C calls a closed-loop transformation component (black-box operator and operating parameters) whose internal code is not accessible, and writes the processing results to the distributed data warehouse node B; data service node E reads the processing results from node B. The independent test resource pool F deploys the input table (hereinafter referred to as the shadow input table), a copy of the closed-loop transformation component, the output table (hereinafter referred to as the shadow output table), and an exit interceptor of the shadow execution environment to perform lineage verification tasks.

[0073] The input records to be verified include eight fields: device identifier, acquisition time, area identifier, raw quantity A, raw quantity B, status code, model code, and batch identifier. Source data is acquired in 15-minute batches, and the closed-loop transformation component generates normalized indicator fields based on these batches. Existing operation logs confirm the task call relationship between the input table, the closed-loop transformation component, and the output table, but cannot determine which input fields generate the normalized indicators.

[0074] The production task batch trigger cycle and batch size are adopted, and a read-only input data snapshot of the same input batch is used as the baseline data. The shadow output table is located in an independent namespace, and the exit interceptor prevents the corresponding results from being written to the data service node E based on the probe batch identifier to prevent test data from entering the formal business chain.

[0075] Parameter used: Number of candidate input fields 8. Maximum number of direct parent fields 2. Number of abnormal responses =0, number of times the baseline value is repeated 3. Maximum decoding loss 0, maximum number of append times 5. Control pass rate threshold The value is 1.

[0076] 1. Determine validation boundaries and candidate fields Metadata such as table structure, field names, and field types are collected from relational database node A and distributed data warehouse node B. Task dependencies and running parameters are obtained from data governance scheduling node D, and task logs provided by big data computing node C are read.

[0077] Since the enclosed transformation component only provides the input table, output table, component version, and execution status to the outside world, and its internal code or field mapping rules cannot be obtained, only task-level edges from the input table to the component and from the component to the output table can be established in the initial lineage graph. This enclosed transformation component is marked as a field-level lineage breakpoint.

[0078] Using the normalized index field as the target output field, the system searches backward along the initial lineage map to the checkpoint of the input table to determine the path reachability and type compatibility conditions. The device identifier, collection time, area identifier, raw quantity A, raw quantity B, status code, model code, and batch identifier are determined as candidate input fields.

[0079] Once the candidate input fields are determined, input and output checkpoints are recorded, and the current input batch data snapshot, closed-loop transformation component version, and runtime parameters are frozen to form a lineage verification task. The frozen information remains consistent across the baseline, probe, and control tasks in this verification process.

[0080] 2. Generate the encoded probe matrix and probe value data. Among them, the eight candidate fields are divided according to their data types: device identifier, area identifier and batch identifier are identifier fields; collection time is a time field; raw quantity A and raw quantity B are numerical fields; status code and model code are enumeration fields.

[0081] Under the conditions of d=2 and e=0, the encoded probe matrix is ​​generated as shown below.

[0082]

[0083] The nine rows of the encoded probe matrix correspond to the nine probe task counts, and the eight columns correspond to the eight candidate input fields mentioned above. The encoded probe matrix satisfies the aforementioned isolation condition, that is, for any candidate input field and any set consisting of no more than two other candidate input fields, there exists at least one probe task that only intervenes in the candidate input field and does not intervene in the fields in the set.

[0084] The encoded values ​​in the encoded probe matrix are further converted into probe value data that can be executed in a shadow execution environment. For the original variables A and B, the constrained increment is calculated to determine... After intervention, field values ​​must not exceed the legal value range and the maximum variation range specified by business rules. For status codes and model codes, legal values ​​different from the baseline value are selected from the corresponding enumeration set. For device identifiers, region identifiers, and batch identifiers, dedicated identifiers within the shadow namespace are used; when related tables are involved, corresponding records are created synchronously to maintain primary-foreign key relationships. The collection time uses a substitute value within the legal time range, and the constraint relationship between batch time and business time is maintained.

[0085] Before execution, all probe value data generated by the mapping process undergoes validation for field schema, value range, NOT NULL conditions, primary / foreign key relationships, and business rules. Probe value data that fails validation is not written to the shadow input table but is regenerated within the valid value range. The validated probe value data is then grouped into nine probe micro-batches according to the matrix row order.

[0086] 3. Perform baseline and probe tasks. Independent test resource pool F deploys a copy of the closed transformation component according to the frozen link information, and establishes shadow input and shadow output tables. The version and operating parameters of the closed transformation component copy are consistent with the production task to be verified, and the input data is taken from the same batch of input data snapshots.

[0087] Before writing probe data, a baseline micro-batch without field intervention is executed three times consecutively. The results of these three executions are used to measure the output fluctuation of the normalized index field under the same input conditions. After completing the baseline execution, nine probe micro-batches are written sequentially according to the row order of the encoded probe matrix, and the closed-loop transformation component is run. Each execution records the following information: probe batch identifier; execution start and end time; input data snapshot identifier; closed-loop transformation component version and running parameters; the field affected in this execution; probe value and constraint verification results; execution status and output record. The probe output is aligned with the baseline output using a shadow record identifier. After successful alignment, the difference in the normalized index field is calculated, and it is determined whether a detectable response has occurred. Executions that fail are not directly involved in decoding and are re-executed under the same input snapshot and component version.

[0088] Suppose that the normalization metric in this version of the closed-loop transformation component is affected by the original values ​​of the fourth candidate field A and the fifth candidate field B. Nine probe executions yield the following response syndrome. : ; This response syndrome This is only for illustrating the calculation process in this embodiment. The responses obtained by other processing links are determined by their actual processing logic, input snapshots, and probe data.

[0089] 4. Decode the parent field to be inspected Among the eight candidate input fields, dependency hypotheses with a size not exceeding 2 are enumerated. For each dependency hypothesis, its ideal response under nine coding probes is calculated, and the inconsistency loss between the ideal response and the observed response syndrome is calculated. For example, for a dependency hypothesis consisting of original quantities A and B, a logical OR operation is performed row by row on the 4th and 5th columns of the coding probe matrix. The resulting ideal response is consistent with the observed response syndrome, therefore the decoding loss of this dependency hypothesis is 0. All other single-field or double-field hypotheses have at least one valid number of inconsistencies with the observed response. The dependency hypotheses are arranged in ascending order of decoding loss. The dependency hypothesis consisting of original quantities A and B is determined as the optimal dependency hypothesis, and the decoding interval between this optimal dependency hypothesis and the second-best hypothesis is calculated. Since both the decoding loss and decoding interval of the optimal dependency hypothesis satisfy the judgment criteria set in this embodiment, original quantities A and B are included in the set of parent fields to be verified.

[0090] 5. Perform control tasks and handle ambiguous hypotheses. After initial decoding, negative and positive controls are executed separately in the same shadow execution environment. The negative control uses the baseline input value without modifying any candidate input fields, and its expected response is 0. The positive control selects a field relationship within the validation boundary that can be determined by direct projection, writes a valid probe value to the input field, and its expected response is 1. The control pass rate is calculated. In this embodiment, the output change of the negative control did not exceed the baseline threshold, the positive control produced a detectable response, and the control pass rate is 1.

[0091] If multiple dependency hypotheses with similar losses exist in other validation tasks, the hypothesis separation score is calculated from the probe vectors that have not yet been executed and can pass field constraint validation, and additional probes are selected. After constraint mapping, the additional probes are written to the shadow input table, and new responses are incorporated into the original response syndrome. Then, the decoding loss and decoding interval for each hypothesis are recalculated. This process continues until any of the following conditions are met: a dependency hypothesis meets the validation condition; the number of executed additional probes reaches a preset number; or the remaining hypotheses cannot be distinguished under the existing constraints. Tasks that still cannot be distinguished after reaching the maximum number of additional probes retain the ambiguous hypothesis set and enter the manual review queue; validated field edges are not written to the initial lineage graph.

[0092] In this embodiment, the set of parent fields to be verified, composed of the original quantity A and the original quantity B, meets the decoding judgment condition. Both the negative control and the positive control pass, so the status of the verification task is recorded as "verification passed".

[0093] 6. Generate verification credentials and update the lineage chart. After the verification task is completed, the following data is summarized: input data snapshot, component version, running parameters, encoding probe matrix, nine-probe micro-batch identifiers, output response syndrome, parent field set, decoding loss, decoding interval, and control results. An evidence summary is generated and saved as lineage evidence. In the field-level lineage graph, verification lineage edges are established, pointing from the original quantity A to the normalized index field and from the original quantity B to the normalized index field, respectively. Each lineage edge records the verification status, closed-loop transformation component version, effective start time, and evidence summary.

[0094] If a change is subsequently detected in the version, operating parameters, or input mode of the closed-loop conversion component, the valid interval of the relevant verification lineage edge is terminated, and its status is adjusted to pending re-verification. The new link status re-enters step S1, and the verification results obtained under the original version are not directly used.

[0095] For example, when data governance personnel need to analyze the downstream impact of a change in the original quantity A, they can query the verification lineage edge from the original quantity A to the normalized index field in the field-level lineage graph and retrieve the corresponding evidence. The evidence records the input data snapshot, component version, probe matrix, and response syndrome corresponding to the field relationship. If the current component version is still within the valid range of the lineage edge, the normalized index field and its downstream fields are included in the scope of the change's impact; if the current version is inconsistent with the evidence record, a new verification task is initiated first, and then it is determined whether to continue using the original field relationship.

[0096] It should be noted that for links that allow only one direct parent field, different non-zero codewords can be assigned to each candidate field, and the corresponding field can be matched based on the observed response codeword. When multiple direct parent fields are allowed, the described separable coded probe matrix is ​​used.

[0097] For links involving connection operations, matching records with the same dedicated identifier can be written to both sides of the connection. The connection key and field source can be determined based on information such as whether the output record exists and whether the output field value has changed. For group aggregation links, a unique group identifier can be set, and a controlled increment can be applied to the numerical field. The group field and the aggregated field can then be determined based on the group count, changes in the aggregation value, and the set checksum.

[0098] For stream processing pipelines, input micro-batches, water level parameters, and operator versions can be frozen within a fixed event time window, and probe tasks can be executed in an isolated environment. Probe outputs remain isolated from the formal downstream environment through independent namespaces and egress interception rules.

Claims

1. A black-box field-level lineage verification method for cross-platform data links, characterized in that, include: Based on the initial lineage graph, identify the black-box field-level lineage breakpoints corresponding to black-box operators in the cross-platform data link, and generate a lineage verification task. The lineage verification task includes a snapshot of input data frozen within the verification boundary, the black-box operator version and running parameters, the target output field and the corresponding candidate input field; the verification boundary is determined by the direct input checkpoint and direct output checkpoint of the black-box operator. Based on the candidate input fields and preset field constraints, a separable encoded probe matrix is ​​generated. Based on the baseline values ​​in the input data snapshot, the encoded values ​​in the encoded probe matrix are mapped to probe values ​​for intervention operations on the candidate input fields, forming multiple probe micro-batches. In the shadow execution environment, based on the input data snapshot, black box operator version and running parameters, the response judgment threshold is first calculated based on the difference set formed by repeatedly executing the benchmark task multiple times based on the benchmark value. Then, multiple probe tasks are executed based on multiple probe micro-batches. The normalized difference between the output of multiple probe tasks under the target output field and the benchmark output is calculated and compared with the response judgment threshold to generate the response syndrome of the target output field under multiple probe tasks. Based on the encoded probe matrix, generate the ideal response corresponding to the preset dependency hypothesis under the probe task, calculate the observation inconsistency loss between the ideal response and the response syndrome; based on the observation inconsistency loss, the decoding loss of the optimal dependency hypothesis and the confidence interval, decode the set of parent fields to be verified of the target output field within the verification boundary. For the set of parent fields to be validated, perform validation on negative controls without probe intervention and positive controls with probe intervention, and output the validation results of candidate input fields based on the control pass rate; A lineage evidence summary is generated based on the input data snapshot, black-box operator version and running parameters, encoded probe matrix, response syndrome, optimal dependency hypothesis and control pass rate, and a lineage evidence credential is generated; verification lineage edges pointing to the target output field are established for the verified candidate input fields, and the verification lineage edges are written into the initial lineage graph to obtain a field-level lineage graph with verification status and version range.

2. The black-box field-level lineage verification method for cross-platform data links according to claim 1, characterized in that, The expression for the lineage verification task is as follows: ; ; In the formula, For the first A bloodline verification task; For the first Target output fields in a bloodline verification task Corresponding candidate input fields; For the first Verification boundaries in a lineage verification task; For the first A snapshot of the input data from a lineage verification task; For the first Black-box operator version in a bloodline verification task; For the first Black-box operator execution parameters in a lineage verification task; For the first In a bloodline verification task, the target output field is... The set of upstream fields obtained by reverse search along the initial lineage graph; For the first The first bloodline verification task One candidate input field; for arrive The shortest reachable path length; The preset maximum verification depth; For field types in metadata; symbol This indicates that the data types of the two fields satisfy the type conversion rules allowed by the link to be verified.

3. The black-box field-level lineage verification method for cross-platform data links according to claim 2, characterized in that, The encoded probe matrix is ​​represented as follows: ,satisfy: ; In the formula, Indicates the first Encoding probe matrix for a lineage verification task, number of candidate input fields , For the number of probe tasks, This is the index for the number of probe tasks, with a value of [value]. ; and Index the candidate input field column, with a value of ; For the first Does the second probe mission affect the first? Candidate input fields A binary indicator of the intervention being applied; Not including the first Columns and their size does not exceed A subset of candidate input fields; This is the maximum number of direct parent fields. The allowed number of abnormal responses.

4. The black-box field-level lineage verification method for cross-platform data links according to claim 3, characterized in that, The process of mapping encoded values ​​in the encoded probe matrix to probe values ​​for intervention operations on candidate input fields, forming multiple probe micro-batches, includes: For the Candidate input fields , No. The probe value for the second probe mission is obtained using the following formula: ; In the formula, For the first The baseline values ​​for each candidate input field; For the first The second probe mission was for the first The probe values ​​written to each candidate input field; In order to be with the first Constraint mapping functions corresponding to the candidate input field types; For the first The extent of intervention for each candidate input field; For the first Candidate input fields The legal value range and the set of relational constraints; For numerical fields, the intervention range is limited by the following formula. : ; In the formula, and These are the lower and upper bounds of the field's valid value range, respectively. The relative intervention ratio ; The maximum permissible level of intervention; The probe values ​​are validated for field mode, value range, non-null condition, primary and foreign key relationship and business rules. The probe values ​​that pass the validation are grouped into multiple probe micro-batches according to the row number of the encoded probe matrix.

5. The black-box field-level lineage verification method for cross-platform data links according to claim 4, characterized in that, The response syndrome of the target output field under multiple probe tasks is generated through the following steps: Calculate the normalized difference between the target output field and the baseline output under the probe task based on the data type of the target output field; based on The difference set formed by repeated execution of the sub-undisturbed baseline value Calculate the response judgment threshold: ; In the formula, The baseline value is repeated a certain number of times. The median; This represents the absolute deviation of the median. To preset the fluctuation amplification factor, ; Output fields for the target The response determination threshold; Based on the response determination threshold and the normalized difference, determine the first... Second response: ; in, For the response under the probe task, This indicates that a detectable change has occurred in the output. This indicates that the output has not changed beyond the natural fluctuation range. This indicates a missing response. To normalize the differences; The response syndrome is obtained by arranging the responses from multiple probe missions in sequence. superscript Indicates transpose. Indicates the first Response under the next probe mission; The calculation of the normalized difference between the probe task output and the baseline output, based on the data type of the target output field, includes: For numerical target output fields, the normalized difference is calculated using the following formula. : ; In the formula, The number of output records for successful alignment between the probe task and the baseline task; To align the record index; and These are the first and second target output fields in the probe task and the benchmark task, respectively. One alignment value; A robust metric for the baseline output field; It is a positive stable term; For enumerated, character, or boolean target output fields, the normalized difference is calculated using the following formula. : ; In the formula, This is an indicator function; it takes the value when the condition is true. Otherwise take ; , indicating the proportion of the alignment record that has changed; For target output fields that cause changes to the record set due to joins / aggregations, the normalized difference is calculated using the following formula. : ; In the formula, and These represent the number of output records for the probe task and the baseline task, respectively. and The verification digests are calculated from the multiset of sort-independent output records of the probe task and the baseline task, respectively. and The preset non-negative weights represent the importance of changes in the number of records and changes in the content of the set, respectively.

6. The black-box field-level lineage verification method for cross-platform data links according to claim 5, characterized in that, The process of generating the ideal response corresponding to the preset dependency assumption under the probe task based on the encoded probe matrix, and calculating the observation inconsistency loss between the ideal response and the response syndrome, includes: For any dependency hypothesis and Define field to indicate vector Based on the encoded probe matrix, the ideal response under the probe task is calculated using the following formula: ; In the formula, Indicates the first The candidate input fields belong to the dependency hypothesis. Otherwise ; For the probe task, is it related to the first... A binary indication of the intervention applied to each candidate input field; Indicating dependence on the hypothesis When established, the probe task intervenes in at least one ideal response corresponding to a real input field; Calculate the observation inconsistency loss between the ideal response and the response syndrome under each dependency assumption: ; In the formula, The number of valid responses; , representing the observation inconsistency loss under the dependency assumption, indicating the dependency assumption. Unexplained percentage of valid responses; The optimal dependency hypothesis is determined by ranking the observation inconsistency losses under each dependency hypothesis from smallest to largest. ; In the formula, Output fields for the target The optimal dependency assumption; The confidence interval is determined by the following formula: ; In the formula, The confidence interval; The second-best dependency hypothesis is the one with the second-highest loss, after the optimal dependency hypothesis. When the observation inconsistency loss of the optimal dependency hypothesis and At that time, the optimal dependency assumption will be... The output is the set of parent fields to be validated; otherwise, the observation inconsistency loss will not exceed [a certain value]. The dependency assumptions constitute the ambiguous hypothesis set. , Preset the range of adjacent losses; The maximum decoding loss threshold, Both the minimum confidence interval threshold and the minimum confidence interval threshold are preset judgment parameters with a value range of [value range missing]. .

7. The black-box field-level lineage verification method for cross-platform data links according to claim 6, characterized in that, After outputting the set of ambiguous hypotheses, adaptive appending probes are generated for the ambiguous dependency hypotheses in the set. Multiple probe tasks are then executed in the shadow execution environment based on these adaptive appending probes; including: Calculate the separation score for additional probes: ; In the formula, The score represents the separation score for additional probes, and the number of weighted hypothesis pairs that the probe can distinguish. This is an executable append probe encoding vector. For dependency hypothesis The single predicted response, For dependency hypothesis The single predicted response; This is used to ensure that each pair of different dependency assumptions is calculated only once; and Dependency Hypothesis , Normalized weights, A set of ambiguous hypotheses; Based on the separation score of the additional probe, an adaptive additional probe is obtained: ; In the formula, A set of probe vectors that satisfy field constraint validation but have not yet been executed; The actual appended probe encoding vector, i.e. the generated adaptive appended probe; Based on the adaptive append probe, corresponding probe micro-batches are generated through constraint mapping. The probe task is executed in the shadow execution environment to obtain the new response of the target output field. Inconsistent dependency assumptions are deleted according to the new response. The decoding loss and confidence interval of the response syndrome and the optimal dependency assumption are updated. The process is repeated multiple times until the preset conditions are met to obtain the set of parent fields to be verified. The preset conditions include: a certain dependency hypothesis meets the verification condition, the number of additional probes executed reaches a preset number, and the remaining dependency hypotheses cannot be distinguished under the existing constraints. Specifically, for the set of parent fields to be validated, negative controls without probe intervention and positive controls with probe intervention are performed, the control pass rate is calculated, and the control validation judgment result is output based on the control pass rate, including: In the shadow execution environment, negative controls without probe intervention are executed on candidate input fields in the set of parent fields to be validated, and positive controls with probe intervention are executed on input fields with known field dependencies within the validation boundary; the control pass rate is calculated, and the pass status is output based on the control pass rate; let the total number of control tasks be... , No. The expected response of the control task is The actual response is The control pass rate is: ; In the formula, This indicates the pass rate of the control group. , ; ; in, , and When the verification result is passed, This is a preset control pass rate threshold.

8. The black-box field-level lineage verification method for cross-platform data links according to claim 7, characterized in that, The bloodline evidence document shall at least include the bloodline evidence summary; The kinship evidence summary is formed by serializing and concatenating the input data snapshot, black-box operator version and running parameters, encoded probe matrix, response syndrome, optimal dependency hypothesis, and control pass rate in a fixed field order: ; In the formula, For deterministic hash functions; symbol This indicates that serialization and concatenation are performed according to a fixed field order; For the first A fixed-length summary of kinship evidence for each verification task; For each candidate input field that passes validation Establish a pointer to the target output field. The verification of bloodline edges is determined by associating the bloodline evidence summary and the valid range of the operator version using the following formula: ; In the formula, This is the start time after the black-box operator version verification is passed; The effective end time; To verify the status; To verify bloodline boundaries; The validated bloodline edges are written into the initial bloodline graph to obtain the updated field-level bloodline graph. Output bloodline evidence and field-level bloodline diagram. .

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the black-box field-level lineage verification method for cross-platform data links as described in any one of claims 1 to 8.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the black-box field-level lineage verification method for cross-platform data links as described in any one of claims 1 to 8.