Data flow path tracking method and system based on dynamic watermark

By introducing a dynamic watermarking mechanism into the data flow path, and utilizing additive spread spectrum watermarking and third-order central moment tensor analysis, the problem of path indistinguishability in multi-stage data flow using traditional watermarking is solved, enabling accurate tracking and reliable traceability of the data flow path.

CN121980549APending Publication Date: 2026-05-05FUJIAN ZHONGXIN NET SAFETY INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FUJIAN ZHONGXIN NET SAFETY INFORMATION TECHNOLOGY CO LTD
Filing Date
2026-01-09
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Traditional static watermarks cannot maintain their effectiveness in multi-stage, multi-modal data flow processes, which gradually weakens the traceability of data flow paths and creates technical bottlenecks such as unclear sources and difficult-to-identify paths.

Method used

A dynamic watermarking-based method is adopted to inject independent additive spread spectrum watermarks into the multimodal data for each flow edge in the kinship network. Through covariance whitening and third-order central moment tensor analysis, the component magnitudes of the principal feature vectors are extracted as the intensity weights of the flow edges, and the maximum product path is planned to output the data flow path.

Benefits of technology

It achieves the detection and stability of watermark information during multimodal and multi-stage data fusion, accurately tracks the data flow path, avoids path ambiguity and source confusion, and improves the verifiability and credibility of enterprise-level data flow system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121980549A_ABST
    Figure CN121980549A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of dynamic watermarking, and discloses a data flow path tracking method and system based on a dynamic watermarking, and the method comprises the steps: firstly injecting a perturbation dynamic watermark signal in each link of a data flow path, and enabling the watermark to continuously exist in the data transmission, combination and aggregation process; secondly, watermark response is extracted through related detection and statistical modeling, and a multi-dimensional path relation matrix is constructed; identifying path structure information formed by combined processing by using tensorization processing and spectral analysis; and finally, outputting a real flow path of the data according to an edge weight mapping and optimal path solving algorithm. According to the method, completeness and verifiability of tracking information can be kept in a complex data fusion environment, high-precision positioning of data sources and paths is realized, and the intelligent level of enterprise data security management and credit risk monitoring is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of dynamic watermarking technology, and more specifically, to a method and system for tracking data flow paths based on dynamic watermarking. Background Technology

[0002] With the improvement of data governance systems and the popularization of cross-domain data sharing, data is no longer limited to a single source or a single system, but is transmitted in a multi-layered, structured form across different business processes. For example, corporate credit risk analysis typically requires processing information such as financial statements, transaction records, and external news and public opinion simultaneously. After aggregation, this data undergoes multiple cleaning, correlation, aggregation, statistical analysis, or modeling processes, forming a complex multi-path data processing graph. In this process, a single static watermark cannot maintain its effectiveness because when data is recalculated or combined, its embedded identifiers are deformed, superimposed, or even partially disappeared, obscuring the identifiers of the original path. Traditional traceability technologies that rely on hash signatures, unique identifiers, or static fingerprints cannot accurately reflect the actual flow trajectory of data when faced with such dynamic, multi-stage data changes.

[0003] Specifically, the data flow process involves numerous multi-source joint operations. Each joint processing introduces signal components from multiple sources into the result, causing downstream data to simultaneously contain feature patterns from different upstream nodes. Computationally, these joint operations are equivalent to multiple multiplications or cross-mappings of the signals from various sources, resulting in complex nonlinear relationships in the outcome. This processing mechanism undermines the separability of traditional watermarking in a single-signal superposition environment, causing downstream systems to obtain only fuzzy superposition results when detecting watermarks, unable to distinguish specific paths. Simultaneously, joint operations on different modalities (such as the merging and analysis of numerical tables and text representations) further exacerbate this ambiguity, as the statistical patterns of each modality differ significantly during joint processing, leading to unstable responses from the original identifiers after fusion. Ultimately, the traceability of the data flow path is gradually weakened through multiple joint processing steps, creating a technical bottleneck of unclear sources and difficult-to-distinguish paths. Summary of the Invention

[0004] This invention provides a data flow path tracking method and system based on dynamic watermarking, which solves the technical problems mentioned in the background art.

[0005] The first aspect is a data flow path tracing method based on dynamic watermarking, including:

[0006] For each flow edge in the lineage network, inject independent additive spread spectrum watermarks into the multimodal data;

[0007] Watermark detection and covariance whitening are performed on samples formed by the joint multimodal data to obtain whitening detection vectors that eliminate second-order statistics;

[0008] For each sample, perform element-wise pairwise multiplication on the whitening detection vector of different modalities to construct a path edge space vector that reflects the multiplicative participation, and generate a third-order central moment tensor based on this.

[0009] Calculate the principal eigenvalues ​​of the third-order central moment tensor and its corresponding principal eigenvectors to characterize the tensor spikes induced by the join operation;

[0010] The component magnitudes of the main feature vectors are extracted as the strength weights of the flow edges, and the maximum product path is planned in the lineage network to output the data flow path.

[0011] Secondly, a data flow path tracing system based on dynamic watermarking, applied to any of the data flow path tracing methods based on dynamic watermarking, includes:

[0012] The data injection module injects independent additive spread spectrum watermarks into the multimodal data for each flow edge in the bloodline network.

[0013] The data whitening module performs watermark detection and covariance whitening on samples formed by the joint multimodal data to obtain whitening detection vectors that eliminate second-order statistics.

[0014] The central moment tensor construction module performs element-wise pairwise multiplication on the whitening detection vectors of different modalities for each sample, constructs a path edge space vector that reflects the multiplicative participation, and generates a third-order central moment tensor accordingly.

[0015] The main feature extraction module calculates the main eigenvalues ​​of the third-order central moment tensor and its corresponding main eigenvectors to characterize the tensor spikes induced by the join operation.

[0016] The data flow path acquisition module extracts the component magnitude values ​​of the main feature vector as the strength weights of the flow edges, and plans the maximum product path in the lineage network to output the data flow path.

[0017] The beneficial effects of this invention include: by introducing a dynamic watermarking mechanism into the multi-source data flow path, accurate tracking of data is achieved in complex joint processing and multi-node transmission environments. This method maintains the detectability and stability of watermark information during multimodal and multi-stage data fusion, ensuring that the true source and flow path can still be recovered after each data combination, branching, or aggregation. Compared to traditional static identification or single-point tracking schemes, this invention effectively avoids problems such as path ambiguity, source confusion, and traceability gaps, significantly improving the verifiability and credibility of evidence collection in enterprise-level data flow systems, and providing reliable technical support for internal auditing, risk warning, and external compliance supervision. Attached Figure Description

[0018] Figure 1 These are experimental comparison diagrams of the characteristic responses of the KR-λ3 tensor spectrum of this invention;

[0019] Figure 2 This is an experimental diagram of the flow edge strength weight distribution and path recognition of the present invention;

[0020] Figure 3 This is a flowchart of the data flow path tracing method based on dynamic watermarking of the present invention;

[0021] Figure 4 This is a block diagram of the data flow path tracking system based on dynamic watermarking of the present invention. Detailed Implementation

[0022] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, features described in some examples may be combined in other examples.

[0023] like Figure 1 As shown, Figure 1 This demonstrates the responsiveness of the tensor spike (KR-λ3) to different data flow logics. Comparative experiments show that when data undergoes only simple linear superposition or union operations (blue curve), its third-order tensor principal eigenvalues ​​remain flat and low, entirely constrained by the additive noise floor. However, once the data undergoes a multimodal join operation (red curve), significant spike signals immediately emerge at the principal component positions of the tensor spectrum. This proves that the proposed scheme can utilize higher-order statistical properties to penetrate linear noise interference, mathematically capturing multiplicative coupling events corresponding to topological conservation mismatch, thus establishing the foundation for detection.

[0024] like Figure 2 As shown, Figure 2 This paper presents the quantized distribution results of the strength weights of each flow edge in a complex kinship network topology, based on the extracted tensor principal eigenvector component magnitudes. Experimental data shows that the weights of all interfering edges (gray bars) that are not part of the actual flow path are effectively suppressed below the noise floor threshold, while the weights of the actual flow edges (orange bars) exhibit extremely high signal-to-noise ratios, with their strength weights significantly higher than the background noise. This verifies the accuracy of the proposed scheme in mapping from the statistical domain to the physical topological domain, ensuring that the subsequent maximum product path planning algorithm can uniquely and unambiguously lock the actual data flow trajectory, achieving high-precision tracking with strong anti-interference capabilities.

[0025] Example 1: As Figure 3 As shown, the data flow path tracing method based on dynamic watermarking includes:

[0026] For each flow edge in the lineage network, inject independent additive spread spectrum watermarks into the multimodal data;

[0027] Watermark detection and covariance whitening are performed on samples formed by the joint multimodal data to obtain whitening detection vectors that eliminate second-order statistics;

[0028] For each sample, perform element-wise pairwise multiplication on the whitening detection vector of different modalities to construct a path edge space vector that reflects the multiplicative participation, and generate a third-order central moment tensor based on this.

[0029] Calculate the principal eigenvalues ​​of the third-order central moment tensor and its corresponding principal eigenvectors to characterize the tensor spikes induced by the join operation;

[0030] The component magnitudes of the main feature vectors are extracted as the strength weights of the flow edges, and the maximum product path is planned in the lineage network to output the data flow path.

[0031] In one embodiment of the present invention, for each flow edge in the kinship network, independent additive spread spectrum watermarks are injected into the multimodal data, including:

[0032] For any flow edge in the bloodline network, construct a structured modal spread spectrum watermark vector that follows a zero-mean, unit-variance multivariate Gaussian distribution for the structured feature modality, and construct a text embedding modal spread spectrum watermark vector that follows a zero-mean, unit-variance multivariate Gaussian distribution for the text embedding modality. The generated structured modal spread spectrum watermark vector and the text embedding modal spread spectrum watermark vector are statistically independent.

[0033] Obtain the structured modal injection amplitude corresponding to the current flow edge, calculate the product of the structured modal injection amplitude and the structured modal spread spectrum watermark vector, and perform vector addition operation on the product and the original structured feature vector to obtain the watermarked structured feature vector;

[0034] Obtain the text embedding modal injection amplitude corresponding to the current flow edge, calculate the product of the text embedding modal injection amplitude and the text embedding modal spread spectrum watermark vector, and perform vector addition operation on the product and the original text embedding vector to obtain the watermarked text embedding vector.

[0035] Structured feature modalities are data types with fixed formats and quantitative attributes in multimodal data, such as revenue data in corporate financial statements, transaction amounts and times, etc. Their data formats are uniform and can be directly used for numerical calculations, making them stable feature carriers in data flow tracking.

[0036] The zero-mean, unit-variance multivariate Gaussian distribution refers to a vector in which the average value of each dimension is zero, the variance of each dimension is one, and the values ​​of multiple dimensions of the vector jointly exhibit a normal distribution. This distribution makes the watermark signal highly concealed and does not interfere with the normal use of the original data.

[0037] The structured modal spread spectrum watermark vector is a watermark carrier adapted to structured feature modes. The vector dimension is consistent with the dimension of the structured feature vector. It is generated by a zero-mean, unit-variance multivariate Gaussian distribution and is used to carry the identification information required for tracking. It is not easily detected after being embedded in the original data.

[0038] Text embedding modalities are vector data types that are transformed from unstructured text data through embedding algorithms. Examples include semantic vectors obtained by processing news and public opinion texts, and feature vectors transformed from company profiles. They retain the semantic information of the text and have numerical computation capabilities.

[0039] The text embedding modal spread spectrum watermark vector is a watermark carrier adapted to the text embedding modality. The vector dimension is the same as that of the text embedding vector. It is generated based on a zero-mean, unit-variance multivariate Gaussian distribution and will not destroy the semantic features of the text after embedding.

[0040] Statistical independence means that changes in the numerical values ​​of the structured modal spread spectrum watermark vector will not affect the numerical distribution of the text-embedded modal spread spectrum watermark vector. The numerical features of the two vectors are not correlated, which can avoid mutual interference between watermark signals and thus prevent tracking failure.

[0041] The structured modal injection amplitude is a scalar that controls the embedding strength of the structured modal watermark. Its value needs to balance the concealment and detectability of the watermark, and is usually determined based on the data volume of the structured features.

[0042] The original structured feature vector is a vector that carries structured business data without embedded watermarks, such as a vector containing multiple financial indicators of an enterprise. It is the basic carrier for watermark embedding.

[0043] The watermarked structured feature vector is a structured feature carrier after embedding a watermark. It retains the business information of the original structured data and carries the watermark information for path tracing, which can continuously transmit the identifier in the data flow. Specifically, the product of the structured modal injection amplitude and the structured modal spread spectrum watermark vector is first calculated to obtain the structured modal watermark vector. Then, the values ​​at the corresponding positions of this vector and the original structured feature vector are added to the new vector, which is the watermarked structured feature vector.

[0044] The text embedding modal injection amplitude is a coefficient that controls the embedding strength of the text embedding modal watermark. The value must match the numerical range of the text embedding vector to ensure that the semantic expression of the text is not changed after the watermark is embedded.

[0045] The original text embedding vector is a vector derived from the text without embedded watermark. It retains the core semantic information of the text and is the original carrier for watermark embedding in the text modality.

[0046] The watermarked text embedding vector is the product of watermarking after text embedding modality. It has both text semantic information and watermark tracking information, and can maintain the integrity of the identifier during the text data flow. Specifically, the product of the text embedding modality injection amplitude and the text embedding modality spread spectrum watermark vector is first calculated to obtain the text embedding modality watermark vector. Then, this vector is added to the corresponding values ​​of the original text embedding vector to obtain the new vector, which is the watermarked text embedding vector.

[0047] In one embodiment of the present invention, element-wise pairwise multiplication is performed on the whitening detection vectors of different modalities for each sample to construct a path edge space vector reflecting the multiplicative participation, including:

[0048] For any sample, retrieve its structured modal whitening detection vector and text embedded modal whitening detection vector. The structured modal whitening detection vector and the text embedded modal whitening detection vector have the same dimension, which corresponds to the number of flow edges in the lineage network.

[0049] Element-wise multiplication is performed on the structured modal whitening detection vector and the text embedded modal whitening detection vector, that is, the components of the two vectors at the same flow edge position are multiplied together;

[0050] The product vector obtained after performing element-wise multiplication is determined as the path edge space vector of the sample.

[0051] A sample is an independent data unit formed by aligning enterprise ID and time window after joint processing of multimodal data. Each sample contains corresponding data features and embedded bimodal watermark information, and is the basic unit for watermark detection.

[0052] The dot product is the result of an operation that measures the similarity between two vectors. It is used to detect whether a sample contains the watermark signal of the corresponding flow edge. The larger the value, the higher the watermark matching degree. Specifically, the components of the corresponding positions of the two vectors are multiplied respectively, and then all the product results are added together. The sum is the dot product.

[0053] The original structured modality detection vector is a vector that carries the structured modality watermark detection results. Its dimension is consistent with the number of flowing edges in the lineage network. Each component in the vector corresponds to the watermark inner product detection value of a flowing edge, and the arrangement order is consistent with the labeling order of the flowing edges. Specifically, the inner product result of the structured modality watermark vector corresponding to each flowing edge and the watermarked structured feature vector of the sample is arranged in the preset arrangement order of the flowing edges, and the resulting vector is the original structured modality detection vector.

[0054] The original detection vector of the text embedding modality is a vector that carries the detection result of the text embedding modality watermark. Its dimension is equal to the number of flow edges. Each component corresponds to the text modality watermark inner product detection value of a flow edge. The arrangement order is consistent with the original detection vector of the structured modality. Specifically, the inner product result of the text embedding modality watermark vector corresponding to each flow edge and the sample watermarked text embedding vector is arranged in the preset arrangement order of the flow edges. The resulting vector is the original detection vector of the text embedding modality.

[0055] The structured modality mean vector is a vector that reflects the overall average level of the original structured modality detection vectors of all samples. It is used for the subsequent centering processing of the original detection vectors to eliminate the influence of the overall offset. Specifically, the components at the corresponding positions of the original structured modality detection vectors of all samples are accumulated, and then the accumulated result at each position is divided by the total number of samples. The resulting vector is the structured modality mean vector.

[0056] The structured modality detection covariance matrix is ​​a matrix that describes the degree of fluctuation correlation among the components of the original structured modality detection vector. Its dimension is the product of the number of transition edges, which is used to eliminate the correlation interference between the components. Specifically, the structured modality mean vector is subtracted from the original structured modality detection vector of each sample to obtain a centered vector. Then, the centered vector is multiplied by its own transpose to obtain the covariance matrix of each sample. Finally, the covariance matrices of all samples are summed, and the summation result is divided by the total number of samples to obtain the structured modality detection covariance matrix.

[0057] The text embedding modality mean vector is the average level vector of the original detection vectors of all samples' text embedding modalities. It is used to center the original detection vectors of text modalities and remove the overall data offset. Specifically, the components at corresponding positions of the original detection vectors of all samples' text embedding modalities are accumulated separately, and the accumulated result at each position is divided by the total number of samples. The resulting vector is the text embedding modality mean vector.

[0058] The text embedding modality detection covariance matrix is ​​used to describe the fluctuation correlation between the components of the original text modality detection vector. Its dimension is the product of the number of transition edges, providing a basis for text modality data whitening. Specifically, the text embedding modality mean vector is subtracted from the original text embedding modality detection vector of each sample to obtain the text modality centering vector. Then, the centering vector is multiplied by its own transpose vector to obtain the text modality covariance matrix of each sample. Finally, the text modality covariance matrices of all samples are summed, and the summation result is divided by the total number of samples to obtain the text embedding modality detection covariance matrix.

[0059] The inverse square root matrix is ​​a matrix that, when multiplied by itself, equals the inverse of the original matrix. It is used to eliminate component correlation and scale differences caused by the covariance matrix and is a key matrix for data whitening. Specifically, the covariance matrix is ​​first decomposed into eigenvalues ​​to obtain an eigenvalue matrix and an eigenvector matrix. Then, each eigenvalue in the eigenvalue matrix is ​​replaced with the reciprocal of its square root to obtain a new eigenvalue matrix. Finally, the eigenvector matrix is ​​multiplied by the new eigenvalue matrix, and then multiplied by the transpose of the eigenvector matrix to obtain the inverse square root matrix.

[0060] The structured modal whitening matrix is ​​a transformation matrix constructed from the inverse square root of the structured modal detection covariance matrix. It is used to transform the original structured modal detection vector into a whitening vector with no correlation and equal variance.

[0061] The text embedding modality whitening matrix is ​​a transformation matrix adapted to the text modality. It is determined by the inverse square root of the text embedding modality detection covariance matrix and is used to whiten the text modality detection vector.

[0062] The structured modal whitening detection vector is a structured modal detection vector that eliminates the interference of second-order statistics. Its components are uncorrelated and have uniform variance, highlighting the multiplicative interaction characteristics of the watermark. Specifically, the original structured modal detection vector of the sample is first subtracted from the structured modal mean vector to obtain the structured modal centered detection vector. Then, the structured modal whitening matrix is ​​multiplied with the centered detection vector to obtain the structured modal whitening detection vector.

[0063] The text embedding modality whitening detection vector is a text modality detection vector that has eliminated the interference of second-order statistics. Its components are uncorrelated and have consistent variance, which facilitates subsequent pairing operations with the structured modality vector. Specifically, the text embedding modality mean vector is subtracted from the original text embedding modality detection vector of the sample to obtain the text embedding modality centered detection vector. Then, the text embedding modality whitening matrix is ​​multiplied with this centered detection vector to obtain the text embedding modality whitening detection vector.

[0064] In one embodiment of the present invention, generating a third-order central moment tensor includes:

[0065] The path edge space vectors corresponding to all samples are vector-accumulated, and the accumulated vector is divided by the total number of samples to obtain the mean vector of the path edge space.

[0066] For each sample, subtract the mean vector of the path edge space from its corresponding path edge space vector to obtain the centered path edge space vector;

[0067] Calculate the sum of the centralized path edge space vectors themselves and the cubic tensor outer product to obtain the single third-order tensor for this sample;

[0068] The third-order tensors corresponding to all samples are summed, and the summed tensor is divided by the total number of samples to obtain the third-order central moment tensor.

[0069] The path edge space vector is a product vector obtained by pairing and multiplying two whitening detection vectors of the same modality element by element. Its dimension is consistent with the number of flowing edges in the lineage network, and it carries feature information jointly participated in by multimodal data multiplication.

[0070] Vector accumulation is an operation that adds the path edge space vectors of all samples one by one according to their corresponding components. It is used to summarize the multiplicative interaction feature information of all samples and provide a basis for subsequent mean calculation. Specifically, the components at the same position in the path edge space vector of each sample are summed, and the resulting vector is the accumulation result of the path edge space vectors of all samples.

[0071] The total number of samples is the total number of independent data units formed by aligning enterprise IDs with time windows after joint processing of multimodal data. It is a fundamental quantitative indicator for calculating the mean vector and the third-order central moment tensor.

[0072] The mean vector of the path sidespace is the average level vector of all sample path sidespace vectors. It is used to eliminate the interference caused by individual differences in samples and to provide a benchmark for the centering of path sidespace vectors. Specifically, the path sidespace vectors of all samples are first summed to obtain the total vector. Then, the value of each component in the total vector is divided by the total number of samples. The resulting vector is the mean vector of the path sidespace.

[0073] The centered path edge space vector is a path edge space vector that eliminates the overall mean bias and can more accurately reflect the degree to which the multiplicative interaction features of a single sample deviate from the overall average level. Specifically, the centered path edge space vector is obtained by subtracting the component of the path edge space mean vector at the corresponding position from the component at each position in the path edge space vector of a single sample.

[0074] The cubic tensor outer product is an operation that transforms a one-dimensional vector into a three-dimensional tensor by performing a three-fold outer product operation between the centered path edge space vector and itself, thereby preserving and enhancing the higher-order feature information brought about by the multiplicative interaction. Specifically, based on the centered path edge space vector, the first outer product is performed to obtain a two-dimensional matrix. Then, the two-dimensional matrix is ​​performed with the centered path edge space vector in the second outer product. Finally, the result is performed with the centered path edge space vector in the third outer product. The final three-dimensional tensor is the result of the cubic tensor outer product.

[0075] A single-sample third-order tensor is a three-dimensional tensor formed by performing three tensor outer product operations on a single sample. Its dimension is the product of the number of flowing edges and the product of the number of flowing edges, and it carries the high-order multiplicative characteristics of a single sample.

[0076] The third-order central moment tensor is the core tensor that summarizes the higher-order multiplicative features of all samples. It can accurately characterize the tensor spikes induced by the join operation of multimodal data and is the basis for subsequent extraction of principal feature values ​​and principal feature vectors. Specifically, first, the individual third-order tensors corresponding to all samples are summed one by one according to the corresponding position to obtain the total tensor. Then, the value of each element in the total tensor is divided by the total number of samples. The resulting tensor is the third-order central moment tensor.

[0077] In one embodiment of the present invention, calculating the principal eigenvalues ​​of the third-order central moment tensor and its corresponding principal eigenvectors includes:

[0078] Construct an optimization objective for a candidate vector of unit length, which is to calculate the inner product of the third-order central moment tensor and the tensor obtained by performing the outer product of the candidate vector itself three times;

[0079] Under the constraint that the L2 norm of the candidate vectors is 1, find the extreme solution that maximizes the optimization objective.

[0080] The maximum value of the optimization objective is determined as the principal eigenvalue of the tensor, and the candidate vector that obtains the maximum value is determined as the principal eigenvector.

[0081] The unit-length candidate vector is the vector to be optimized for solving the principal eigenvalues ​​and principal eigenvectors of the tensor. Its length (norm 2) is limited to one, and the vector dimension is consistent with the number of flowing edges in the lineage network.

[0082] The optimization objective is a quantitative indicator that measures the degree of matching between a candidate vector of unit length and the third-order central moment tensor. The larger the objective value, the better the candidate vector can represent the core peak features of the tensor. Specifically, the corresponding tensor is obtained by first calculating the cubic outer product of the candidate vector of unit length itself, and then calculating the inner product of the tensor with the third-order central moment tensor. The result is the value of the optimization objective.

[0083] The tensor obtained by performing three outer products on the candidate vector itself is a three-dimensional tensor transformed from the candidate vector of unit length through three outer product operations. Its dimension is consistent with the third-order central moment tensor. It is used to perform inner product operations with the third-order central moment tensor to construct the optimization objective. Specifically, based on the candidate vector of unit length, three outer product operations are performed in sequence. First, a two-dimensional matrix is ​​obtained from the vector, and then a three-dimensional tensor is obtained from the two-dimensional matrix. The final result is the tensor.

[0084] The inner product of tensors is the result of an operation that measures the degree of correlation between corresponding elements of two three-dimensional tensors of the same dimension. It is a key operation for constructing optimization objectives and can accurately reflect the fit between candidate vectors and the core features of tensors. Specifically, the inner product of tensors is obtained by multiplying the corresponding elements of the tensor obtained by the cubic outer product of the third-order central moment tensor and the candidate vectors, and then summing all the multiplication results.

[0085] The L2 norm of a candidate vector is a quantitative indicator of the vector length. Its calculation result reflects the overall magnitude of the vector. It is limited to one purpose to avoid the difference in vector magnitude from interfering with the calculation of the optimization objective. Specifically, the value of each component in the candidate vector is squared, all the squared results are summed, and finally the square root of the sum is taken. The value obtained is the L2 norm of the candidate vector.

[0086] The constraint is a rule that limits the candidate vector of unit length, that is, it forces the candidate vector to have a 2-norm of 1. This constraint can narrow the solution range and ensure the uniqueness and validity of the extreme solution.

[0087] An extreme solution is a candidate vector of unit length that maximizes the optimization objective while satisfying the constraints. This vector is the subsequent principal feature vector. Specifically, the gradient ascent algorithm is used to continuously adjust the component values ​​of the candidate vector with the constraints as the boundary, and iteratively calculate the optimization objective value until the objective value no longer increases. The corresponding candidate vector at this point is the extreme solution.

[0088] The principal eigenvalue of a tensor is the maximum value that the optimization objective can achieve. It is the core indicator that characterizes the peak intensity of the third-order central moment tensor. The larger the value, the more significant the multiplicative interaction of the multimodal data.

[0089] The principal eigenvector is the extreme solution of the principal eigenvalue of the corresponding tensor, and each component in the vector corresponds to the feature strength of a flowing edge in the lineage network.

[0090] In one embodiment of the present invention, extracting the component magnitudes of the main feature vector as the strength weights of the transition edges includes:

[0091] For each transition edge in the bloodline network, extract the vector component corresponding to that transition edge from the main feature vector;

[0092] Perform absolute value operation on the extracted vector components, and determine the resulting value as the flow edge strength weight of the flow edge;

[0093] According to the arrangement order of the flowing edges in the kinship network, the flowing edge strength weights of all flowing edges are combined into a flowing edge strength weight vector.

[0094] A flow edge is a data flow channel connecting two nodes in a lineage network. Each flow edge corresponds to a unique identifier and is the basic building block of the data flow path. Their arrangement order remains consistent throughout the entire tracking process.

[0095] Vector components are numerical units corresponding to a single flow edge in the principal feature vector. Their numerical values ​​reflect the feature contribution of the flow edge in the multiplicative interaction. The positive or negative sign does not affect the strength judgment; only the absolute value needs to be considered. Specifically, according to the preset arrangement order of the flow edges in the lineage network, a single value at the corresponding position in the principal feature vector is extracted, and this value is the vector component corresponding to the flow edge.

[0096] The absolute value operation is an operation that eliminates the influence of the positive or negative sign of vector components. It is used to transform the feature contribution into a non-directional strength index, ensuring that the weight value only represents the participation strength of the flowing edge. Specifically, if the extracted vector component is positive, the absolute value is the component itself; if it is negative, the absolute value is the opposite of the component; if it is zero, the absolute value is still zero.

[0097] The strength weight of a flow edge is a quantitative indicator that measures the importance of a flow edge in the data flow path. The larger the value, the higher the probability that the flow edge belongs to the real path.

[0098] The order of the transition edges is a consistent sorting rule throughout the entire tracking process, and it is consistent with the order of the transition edges of the original detection vector to ensure that the correspondence between vector components and transition edges is not confused.

[0099] The flow edge strength weight vector is a vector that summarizes the strength weights of all flow edges. Its dimension is the same as that of the main feature vector, and it fully carries the strength information of each flow edge in the lineage network. Specifically, according to the preset arrangement order of the flow edges in the lineage network, the strength weights corresponding to each flow edge are arranged in sequence, and the resulting vector is the flow edge strength weight vector.

[0100] In one embodiment of the present invention, planning the maximum product path in the lineage network to output the data flow path includes:

[0101] In the bloodline network, a path search objective is established, which is to select a path that maximizes the product of the strength weights of the flowing edges contained therein.

[0102] Traverse the nodes in the topological order of the bloodline network. For the current node, obtain the cumulative path strength of all its predecessor nodes. Calculate the product of the cumulative path strength of the predecessor node and the flow edge strength weight of the flow edge connecting the predecessor node to the current node. Select the maximum value of the product as the cumulative path strength of the current node.

[0103] Starting from the endpoint node of the kinship network, backtracking is performed based on the cumulative path strength of each node to determine the sequence of flow edges that constitute the maximum chain product. This sequence of flow edges is then output as the data flow path.

[0104] A kinship network is a directed acyclic graph that depicts the relationships between data flows. It consists of nodes and flow edges. Nodes represent data processing steps, and flow edges represent data transmission channels.

[0105] The path search objective is the core criterion of path planning. It is used to select the optimal path from multiple candidate paths in the lineage network, focusing on the maximum value of the product of the strength weights of the flowing edges, to ensure that the most likely real data flow path is selected. Specifically, it traverses all complete paths from the start point to the end point in the lineage network, calculates the product of the strength weights of all flowing edges on each path, and determines the path with the largest product as the optimal solution of the path search objective.

[0106] The chain product is a continuous multiplication operation performed on the strength weights of all flowing edges on a single path. The result is a comprehensive index that measures the overall credibility of the path. The larger the value, the higher the credibility of the path. Specifically, according to the transmission order of the flowing edges in the path, the strength weight of the first flowing edge is multiplied by the strength weights of all subsequent flowing edges in turn. The final result is the chain product of the path.

[0107] Topological order follows the node traversal order determined by the data flow direction in the kinship network, ensuring that all predecessor nodes are traversed before the current node, avoiding data transmission logic reversal, and guaranteeing the correctness of cumulative path strength calculation.

[0108] Nodes are key nodes in the data flow of a kinship network, including data starting nodes, intermediate processing nodes, and ending nodes. Each node corresponds to a data processing stage.

[0109] In the kinship network topology, a predecessor node is the upstream node that directly points to the current node. A node may have multiple predecessor nodes, and each predecessor node is connected to the current node through a flow edge.

[0110] The cumulative path strength is the product of the optimal paths from the starting node to the current node. The cumulative path strength of the current node is inherited from the optimal predecessor node. Specifically, the cumulative path strength of the starting node is initialized to one. For other nodes, the product of the cumulative path strength of each predecessor node and the corresponding connection flow edge strength weight is calculated, and the maximum value among these products is selected as the cumulative path strength of the current node.

[0111] The endpoint node is the final node in the data flow of the kinship network, that is, the target node for data processing. Its cumulative path strength is the result of the largest path product in the entire network, and it is the starting point for reverse backtracking.

[0112] Backtracking is the process of reconstructing the path from the endpoint node back to the starting node. It ensures that the reconstructed path is the optimal path with the largest product by matching the predecessor nodes through the cumulative path strength. Specifically, starting from the endpoint node, find the predecessor node whose cumulative path strength can be obtained by multiplying the flow edge strength weights, and record the flow edge. Then, using the found predecessor node as the current node, repeat the above steps until backtracking to the starting node. The recorded flow edge sequence is the backtracking result.

[0113] The flow edge sequence is a set of flow edges obtained by backtracking and arranged in the forward order of data flow, which fully reflects the transmission path of data from the starting point to the ending point.

[0114] The data flow path is composed of a sequence of flow edges, which accurately reflects the real transmission trajectory of data in the kinship network and can be directly used for data traceability, auditing and risk monitoring.

[0115] Example 2: Figure 4 As shown, the data flow path tracing system based on dynamic watermarking, applied to any of the data flow path tracing methods based on dynamic watermarking, includes:

[0116] The data injection module injects independent additive spread spectrum watermarks into the multimodal data for each flow edge in the bloodline network.

[0117] The data whitening module performs watermark detection and covariance whitening on samples formed by the joint multimodal data to obtain whitening detection vectors that eliminate second-order statistics.

[0118] The central moment tensor construction module performs element-wise pairwise multiplication on the whitening detection vectors of different modalities for each sample, constructs a path edge space vector that reflects the multiplicative participation, and generates a third-order central moment tensor accordingly.

[0119] The main feature extraction module calculates the main eigenvalues ​​of the third-order central moment tensor and its corresponding main eigenvectors to characterize the tensor spikes induced by the join operation.

[0120] The data flow path acquisition module extracts the component magnitude values ​​of the main feature vector as the strength weights of the flow edges, and plans the maximum product path in the lineage network to output the data flow path.

[0121] The embodiments of this example have been described above. However, this example is not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms based on the guidance of this example, and all of them are within the protection scope of this example.

Claims

1. A data flow path tracing method based on dynamic watermarking, characterized in that, include: For each flow edge in the lineage network, inject independent additive spread spectrum watermarks into the multimodal data; Watermark detection and covariance whitening are performed on samples formed by the joint multimodal data to obtain whitening detection vectors that eliminate second-order statistics; For each sample, perform element-wise pairwise multiplication on the whitening detection vector of different modalities to construct a path edge space vector that reflects the multiplicative participation, and generate a third-order central moment tensor based on this. Calculate the principal eigenvalues ​​of the third-order central moment tensor and its corresponding principal eigenvectors to characterize the tensor spikes induced by the join operation; The component magnitudes of the main feature vectors are extracted as the strength weights of the flow edges, and the maximum product path is planned in the lineage network to output the data flow path.

2. The data flow path tracing method based on dynamic watermarking according to claim 1, characterized in that, For each flow edge in the kinship network, inject independent additive spread watermarks into the multimodal data, including: For any flow edge in the bloodline network, construct a structured modal spread spectrum watermark vector that follows a zero-mean, unit-variance multivariate Gaussian distribution for the structured feature modality, and construct a text embedding modal spread spectrum watermark vector that follows a zero-mean, unit-variance multivariate Gaussian distribution for the text embedding modality. The generated structured modal spread spectrum watermark vector and the text embedding modal spread spectrum watermark vector are statistically independent. Obtain the structured modal injection amplitude corresponding to the current flow edge, calculate the product of the structured modal injection amplitude and the structured modal spread spectrum watermark vector, and perform vector addition operation on the product and the original structured feature vector to obtain the watermarked structured feature vector; Obtain the text embedding modal injection amplitude corresponding to the current flow edge, calculate the product of the text embedding modal injection amplitude and the text embedding modal spread spectrum watermark vector, and perform vector addition operation on the product and the original text embedding vector to obtain the watermarked text embedding vector.

3. The data flow path tracing method based on dynamic watermarking according to claim 2, characterized in that, Watermark detection and covariance whitening are performed on samples formed by the joint multimodal data to obtain whitening detection vectors that eliminate second-order statistics, including: For any sample, calculate the dot product of the watermarked structured feature vector of the sample with the structured modal spread spectrum watermark vector of each flow edge in the lineage network, and form the original structured modal detection vector according to the order of the flow edges; calculate the dot product of the watermarked text embedding vector of the sample with the text embedding modal spread spectrum watermark vector of each flow edge in the lineage network, and form the original text embedding modal detection vector according to the order of the flow edges.

4. The data flow path tracing method based on dynamic watermarking according to claim 3, characterized in that, Watermark detection and covariance whitening are performed on samples formed by the joint multimodal data to obtain whitening detection vectors that eliminate second-order statistics. This also includes: The original detection vectors of the structured modalities for all samples are statistically analyzed, and the mean vector and covariance matrix of the structured modalities are calculated. The original detection vectors of the text embedding modalities for all samples are statistically analyzed, and the mean vector and covariance matrix of the text embedding modalities are calculated. The inverse square root matrix of the structured modality detection covariance matrix is ​​obtained by solving for the structured modality whitening matrix; the inverse square root matrix of the text embedding modality detection covariance matrix is ​​obtained by solving for the text embedding modality whitening matrix. The structured modality whitening detection vector is obtained by subtracting the structured modality mean vector from the original structured modality detection vector of the sample and then multiplying it on the left by the structured modality whitening matrix. The text embedding modality whitening detection vector is obtained by subtracting the text embedding modality mean vector from the original text embedding modality detection vector of the sample and then multiplying it on the left by the text embedding modality whitening matrix.

5. The data flow path tracing method based on dynamic watermarking according to claim 4, characterized in that, For each sample, element-wise pairwise multiplication is performed on the whitening detection vectors of different modalities to construct a path edge space vector reflecting the multiplicative participation, including: For any sample, retrieve its structured modal whitening detection vector and text embedded modal whitening detection vector. The structured modal whitening detection vector and the text embedded modal whitening detection vector have the same dimension, which corresponds to the number of flow edges in the lineage network. Element-wise multiplication is performed on the structured modal whitening detection vector and the text embedded modal whitening detection vector, that is, the components of the two vectors at the same flow edge position are multiplied together; The product vector obtained after performing element-wise multiplication is determined as the path edge space vector of the sample.

6. The data flow path tracing method based on dynamic watermarking according to claim 5, characterized in that, Generate the third-order central moment tensor, including: The path edge space vectors corresponding to all samples are vector-accumulated, and the accumulated vector is divided by the total number of samples to obtain the mean vector of the path edge space. For each sample, subtract the mean vector of the path edge space from its corresponding path edge space vector to obtain the centered path edge space vector; Calculate the sum of the centralized path edge space vectors themselves and the cubic tensor outer product to obtain the single third-order tensor for this sample; The third-order tensors corresponding to all samples are summed, and the summed tensor is divided by the total number of samples to obtain the third-order central moment tensor.

7. The data flow path tracing method based on dynamic watermarking according to claim 6, characterized in that, Calculate the principal eigenvalues ​​and corresponding principal eigenvectors of the third-order central moment tensor, including: Construct an optimization objective for a candidate vector of unit length, which is to calculate the inner product of the third-order central moment tensor and the tensor obtained by performing the outer product of the candidate vector itself three times; Under the constraint that the L2 norm of the candidate vectors is 1, find the extreme solution that maximizes the optimization objective. The maximum value of the optimization objective is determined as the principal eigenvalue of the tensor, and the candidate vector that obtains the maximum value is determined as the principal eigenvector.

8. The data flow path tracing method based on dynamic watermarking according to claim 7, characterized in that, Extracting the component magnitudes of the main feature vector as the strength weights of the transition edges includes: For each transition edge in the bloodline network, extract the vector component corresponding to that transition edge from the main feature vector; Perform absolute value operation on the extracted vector components, and determine the resulting value as the flow edge strength weight of the flow edge; According to the arrangement order of the flowing edges in the kinship network, the flowing edge strength weights of all flowing edges are combined into a flowing edge strength weight vector.

9. The data flow path tracing method based on dynamic watermarking according to claim 8, characterized in that, Planning the maximum product path in a kinship network to output the data flow path includes: In the bloodline network, a path search objective is established, which is to select a path that maximizes the product of the strength weights of the flowing edges contained therein. Traverse the nodes in the topological order of the bloodline network. For the current node, obtain the cumulative path strength of all its predecessor nodes. Calculate the product of the cumulative path strength of the predecessor node and the flow edge strength weight of the flow edge connecting the predecessor node to the current node. Select the maximum value of the product as the cumulative path strength of the current node. Starting from the endpoint node of the kinship network, backtracking is performed based on the cumulative path strength of each node to determine the sequence of flow edges that constitute the maximum chain product. This sequence of flow edges is then output as the data flow path.

10. A data flow path tracing system based on dynamic watermarking, applied to the data flow path tracing method based on dynamic watermarking as described in any one of claims 1-9, characterized in that, include: The data injection module injects independent additive spread spectrum watermarks into the multimodal data for each flow edge in the bloodline network. The data whitening module performs watermark detection and covariance whitening on samples formed by the joint multimodal data to obtain whitening detection vectors that eliminate second-order statistics. The central moment tensor construction module performs element-wise pairwise multiplication on the whitening detection vectors of different modalities for each sample, constructs a path edge space vector that reflects the multiplicative participation, and generates a third-order central moment tensor accordingly. The main feature extraction module calculates the main eigenvalues ​​of the third-order central moment tensor and its corresponding main eigenvectors to characterize the tensor spikes induced by the join operation. The data flow path acquisition module extracts the component magnitude values ​​of the main feature vector as the strength weights of the flow edges, and plans the maximum product path in the lineage network to output the data flow path.