Data provenance method, system, and device based on big data analysis
By using big data analysis and locally sensitive hashing algorithm, combined with a binary tree structure to generate traceability tags, the problem of low efficiency of traditional data traceability methods is solved, and efficient and flexible data traceability is achieved.
Patent Information
- Application Number
- CN202510841062.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-06-23
AI Technical Summary
Traditional data traceability methods are inefficient when dealing with large-scale data, are difficult to meet real-time requirements, and traceability labels lack flexibility and uniqueness.
Based on big data analysis, distributed mapping is used to generate hash codes, the local sensitive hashing algorithm is used to cluster the hash codes, and the binary tree structure is combined to generate traceability labels to determine the relationship between data and traceability.
It improves the efficiency of data traceability processing, reduces the amount of calculation and time, generates unique traceability labels, enhances the distinctiveness and flexibility of labels, and meets real-time requirements.
Smart Images

Figure CN120354268B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to a data traceability method, system and device based on big data analysis. BACKGROUND
[0002] In many fields, massive data is continuously generated, stored, transmitted and processed. However, the large-scale flow and complex application of data also bring a series of severe challenges, among which data traceability becomes one of the key problems to be solved.
[0003] Traditional data traceability methods have many limitations when facing big data. The processing efficiency of large-scale data is low, and it is difficult to complete the data traceability analysis within an acceptable time. Traditional methods often rely on one-by-one recording and querying of data, which will lead to excessive consumption of system resources, long response time, low processing efficiency and inability to meet the real-time requirements of business scenarios in the case of large amounts of data. At the same time, the current traceability tags also lack flexibility and uniqueness when generated. SUMMARY
[0004] The main purpose of the present application is to provide a data traceability method, system and device based on big data analysis, which aims to overcome the low processing efficiency of the current data traceability method.
[0005] To achieve the above purpose, the present application provides a data traceability method based on big data analysis, comprising the following steps:
[0006] Based on big data analysis, the original data in the data set is analyzed, and the original data is distributedly mapped into hash codes according to the analysis result to obtain a hash code set;
[0007] The similar hash codes in the hash code set are clustered into a class by using a local sensitive hash algorithm to obtain a plurality of cluster hash code sets; wherein the original data mapped by the hash codes in each cluster hash code set is classified into a class;
[0008] For each of the cluster hash code sets, the cluster features and the data features of the original data mapped by the cluster hash code set are extracted;
[0009] Based on the cluster features and data features, a traceability tag is generated using a binary tree structure;
[0010] The traceability relationship between each original data and the corresponding traceability tag is determined to realize data traceability.
[0011] Further, based on big data analysis, the original data in the data set is analyzed, and the original data is distributedly mapped into hash codes according to the analysis result to obtain a hash code set, comprising:
[0012] The original data in the data set is divided into blocks to obtain multiple data subsets.
[0013] For each data subset, a tensor decomposition algorithm is used to mine correlation information in multiple dimensions to generate a correlation tensor; the multiple dimensions include time, business scenario, and data generation entity;
[0014] Analyzing the correlation strength of multiple dimensions in each of the correlation tensors, and adjusting the parameters of the hash function based on the correlation strength;
[0015] The matched data subsets are mapped based on the hash functions after adjusting the parameters, and the original data is mapped into hash codes to obtain hash code sets.
[0016] Furthermore, similar hash codes in the hash code set are clustered into one category using a locality sensitive hashing algorithm to obtain multiple clustered hash code sets, including:
[0017] For each hash code in the hash code set, mapping is performed using a preset locality sensitive hash function to obtain a hash bucket number corresponding to each hash code after mapping by the locality sensitive hash algorithm;
[0018] If two hash codes are assigned to the same numbered hash bucket after mapping, the two hash codes are marked as potentially similar codes;
[0019] For all hash codes marked as potentially similar codes, calculate their exact similarity;
[0020] The potential similar codes with exact similarity higher than a threshold are clustered into one category to form a clustered hash code set; the above clustering process is repeated until all hash codes are classified into corresponding clustered hash code sets, and finally multiple clustered hash code sets are obtained.
[0021] Furthermore, for each cluster hash code set, cluster features and data features of the original data mapped by the cluster hash code set are extracted, including:
[0022] For each cluster hash code set, extract the cluster features from multiple dimensions; wherein the cluster features include the number of hash codes in each cluster hash code set, the average length of the hash codes, and the character distribution frequency of the hash codes;
[0023] For the original data mapped by the clustered hash code set, identify the data type in the data; count the quantity ratio of each data type as a type quantity ratio feature; count the storage capacity ratio of each data type as a storage capacity ratio feature;
[0024] Combining the type quantity proportion feature and the storage quantity proportion feature, the data feature is obtained.
[0025] Further, the determination of the traceability relationship between each original data and the corresponding traceability label to realize data traceability comprises:
[0026] Analyzing the original data corresponding to each traceability label, and identifying the generation time and source system of the original data;
[0027] According to the generation time and the source system, the original data corresponding to each traceability label is sorted to obtain an original data sequence;
[0028] The original data sequence and the traceability label are established to realize data traceability.
[0029] Further, based on the clustering feature and the data feature, a binary tree structure is generated to generate a traceability label, comprising:
[0030] The clustering feature and the data feature are subjected to principal component analysis to extract principal component features;
[0031] An adaptive binary tree structure is constructed, and the principal component features are taken as root nodes of the binary tree structure;
[0032] According to the correlation between the principal component features and other features, left and right child nodes of the root nodes are generated; wherein, the two features most correlated with the principal component features are divided into features on the left and right child nodes of the root nodes;
[0033] In the iterative growth process of the left and right child nodes of the binary tree, the correlation between the features under each child node and other features is calculated, when the correlation is greater than a preset threshold, the node is further subdivided; when the correlation is less than the threshold, the growth of the child node branch is stopped;
[0034] After the binary tree structure is constructed, a unique hash value is generated as a traceability label according to the path from the last leaf node to the root node.
[0035] The application also provides a data traceability system based on big data analysis, comprising:
[0036] An encoding unit is configured to analyze original data in a data set based on big data analysis, and to distribute and map the original data into hash codes according to the analysis result to obtain a hash code set;
[0037] A clustering unit is configured to cluster similar hash codes in the hash code set into a class by using a local sensitive hash algorithm to obtain a plurality of clustering hash code sets; wherein, the original data mapped by the hash codes in each clustering hash code set is divided into a class.
[0038] extracting, for each of the cluster hash code sets, cluster features and data features of the original data mapped by the cluster hash code sets;
[0039] generating, based on the cluster features and the data features, a traceability label using a binary tree structure;
[0040] determining a traceability relationship between each of the original data and the corresponding traceability label to realize data traceability.
[0041] The application also provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor realizes the steps of the method according to any one of the preceding embodiments when executing the computer program.
[0042] The application also provides a computer readable storage medium storing a computer program, wherein the computer program realizes the steps of the method according to any one of the preceding embodiments when executed by a processor.
[0043] The application provides a data traceability method, system and device based on big data analysis, comprising: analyzing original data in a data set based on big data analysis, and distributing and mapping the original data into hash codes according to an analysis result to obtain a hash code set; using a local sensitive hash algorithm to cluster similar hash codes in the hash code set into a class to obtain a plurality of cluster hash code sets; wherein the original data mapped by the hash codes in each cluster hash code set is classified into a class; extracting, for each of the cluster hash code sets, cluster features and data features of the original data mapped by the cluster hash code sets; generating a traceability label using a binary tree structure based on the cluster features and the data features; and determining a traceability relationship between each of the original data and the corresponding traceability label to realize data traceability. In the application, the original data is distributed and mapped into hash codes according to an analysis result of big data analysis in a distributed manner, which can fully utilize computing resources and process a large amount of data in parallel, thereby avoiding the efficiency bottleneck caused by processing data one by one in a traditional method. The local sensitive hash algorithm is used to cluster similar hash codes in the hash code set into a class, and the local sensitive hash algorithm has the advantage of quickly locating similar data in a high-dimensional data space, which can efficiently perform clustering operation on a large amount of hash codes. Compared with the traditional method, the amount of calculation and processing time are reduced, and the processing efficiency is improved. The traceability label is generated using a binary tree structure based on the cluster features and the data features, which can generate a unique label by combining multi-dimensional features, thereby enhancing the distinguishability. BRIEF DESCRIPTION OF DRAWINGS
[0044] Figure 1 is a step schematic diagram of the data traceability method based on big data analysis in an embodiment of the application;
[0045] Figure 2 This is a structural block diagram of a data tracing system based on big data analysis in one embodiment of the present invention;
[0046] Figure 3 It is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention.
[0047] The implementation, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0049] Reference Figure 1 In one embodiment of the present invention, a data tracing method based on big data analysis is provided, comprising the following steps:
[0050] Step S1, analyzing the original data in the data set based on big data analysis, and distributively mapping the original data into hash codes according to the analysis results to obtain a hash code set;
[0051] Step S2, using a locality sensitive hashing algorithm to cluster similar hash codes in the hash code set into one class, to obtain a plurality of clustered hash code sets; wherein the original data mapped to the hash codes in each clustered hash code set is divided into one class;
[0052] Step S3: for each cluster hash code set, extract cluster features and data features of the original data mapped by the cluster hash code set;
[0053] Step S4, generating a traceability label using a binary tree structure based on the clustering features and data features;
[0054] Step S5: Determine the traceability relationship between each original data and the corresponding traceability tag to achieve data traceability.
[0055] In this embodiment, as described in step S1 above, this step is the starting point of the entire data tracing method, which aims to pre-process the original data and convert it into a form that is more convenient for subsequent processing and analysis, namely hash coding. First, big data analysis technology is applied to the original data. The big data analysis here covers a variety of means, such as data cleaning to remove noise, duplicate data, etc., data exploratory analysis to understand the distribution, correlation and other characteristics of the data, and data standardization or normalization to ensure that data from different sources or types are comparable. Through these analyses, the inherent structure and characteristics of the original data can be deeply explored, providing a basis for subsequent hash code mapping.
[0056] Based on the above analysis results, a distributed computing architecture was adopted for hash code generation. Distributed mapping can fully utilize the resources of multiple computing nodes to process large amounts of raw data in parallel, greatly improving processing efficiency. Hash coding converts raw data into a fixed-length encoding using a specific hash function. This encoding is unidirectional, meaning it is difficult to deducing the original data from the encoding. However, it can retain certain characteristics of the original data to a certain extent, ensuring that similar raw data also has a certain degree of similarity or correlation in the hash code space, laying the foundation for subsequent similarity determination and clustering operations.
[0057] As described in step S2 above, after obtaining the hash code set, a locality-sensitive hashing algorithm is used to cluster it. The core feature of the locality-sensitive hashing algorithm is its ability to map similar data points (here, hash codes) to the same or similar hash buckets with a high probability, thereby enabling rapid approximate clustering of large-scale, high-dimensional data (hash-coded data). This algorithm is based on a carefully designed family of hash functions that consider the data's local structure and similarity characteristics when calculating hash values. After hash codes are calculated using the locality-sensitive hashing function, they are assigned to different hash buckets. Hash codes with similar hash values (falling into the same or similar hash buckets) are considered similar and clustered together to form a clustered hash code set. This clustering ensures that the original data corresponding to the hash codes within each cluster have some degree of similarity or correlation. This provides a preliminary grouping basis for further in-depth analysis of the data's source and flow path. Specifically, the original data is divided into different categories based on the similarity of their hash codes, facilitating subsequent more detailed feature extraction and traceability label generation for each category.
[0058] As described in step S3 above, for each cluster hash code set, two important features need to be extracted. Cluster features reflect the overall characteristics of the entire cluster at the hash code level, such as the number of hash codes in the cluster, the average length of the hash codes, and the distribution pattern of hash code elements. These features help to understand the data size and encoding complexity of the cluster at a macro level, and are important for determining the overall identification and discrimination of the cluster when constructing traceability labels. Simultaneously, data features of the original data mapped by the cluster hash code set need to be extracted. These include the data type distribution of the original data (e.g., the proportion of numeric, text, date, etc. in the original data of the cluster), length features of the original data records (average length, maximum and minimum lengths), numeric range features of the original data (minimum value, maximum value, standard deviation for numeric data), and keyword features of the text data (if text data exists, keywords with high frequency of occurrence are extracted). These original data features can reveal the specific content and characteristics of the original data in the cluster at a micro level, complementing the cluster features and providing a rich source of information for generating accurate and representative traceability labels.
[0059] As described in step S4 above, the extracted cluster features and data features are integrated and a traceability label is generated. The binary tree structure has good hierarchical and logical properties, making it ideal for organizing and expressing complex feature information. First, the most representative or discriminative feature is selected as the root node of the binary tree. For example, a feature combination with a large number of hash codes and a relatively simple original data type in the cluster can be given priority as the root node. Then, based on the relevance and importance of other features to the root node feature, the remaining features are divided into branch features of the left and right subtrees. During the construction of the subtree, this process is repeated continuously, that is, key features in the subtree branches are selected as child nodes, and the next layer of branches is divided again until all cluster features and data features are reasonably integrated into the binary tree structure. Finally, by encoding or labeling the path from the root node to the leaf node of the binary tree, a traceability label is generated for each cluster (and therefore corresponding to each original data). This traceability label not only contains the overall feature information of the cluster, but also covers the key detailed features of the original data. It can serve as a unique identifier and traceability clue for the data in the subsequent data traceability process, accurately reflecting the source and flow history of the data.
[0060] As described in step S5 above, after generating the traceability label, it is necessary to establish a clear traceability relationship between the original data and the traceability label. This process involves associating and mapping each original data with the traceability label of the cluster to which it belongs. Through this mapping relationship, based on the rich information contained in the traceability label (such as the source, processing process, etc. reflected by the cluster characteristics and data characteristics), the reverse tracking along the data flow path can be performed to determine the various links and sources experienced by the data from generation to the current state. For example, when it is necessary to trace the source of a certain original data, the corresponding traceability label can be first found, and then the relevant cluster characteristics and data characteristics information can be parsed from the traceability label, and then the source system, processing time sequence, associated other data, etc. information of the data in the previous processing process can be determined according to these information, so as to build a complete data traceability path, realize accurate data traceability, and meet the data management needs of data quality control, compliance review, error troubleshooting, and security audit, etc.
[0061] In an embodiment, based on big data analysis, the original data in the data set is analyzed, and the original data is distributedly mapped into hash codes according to the analysis result to obtain a hash code set, including:
[0062] The original data in the data set is processed in blocks to obtain a plurality of data subsets.
[0063] For each data subset, a correlation tensor is generated by using a tensor decomposition algorithm to mine correlation information in multiple dimensions, including time, business scenario, and data generation subject.
[0064] The correlation strength of the multiple dimensions in each correlation tensor is analyzed, and the parameters of the hash function are adjusted based on the correlation strength.
[0065] The matched data subsets are respectively mapped based on the hash function with adjusted parameters to map the original data into hash codes to obtain a hash code set.
[0066] In this embodiment, in a big data environment, the amount of raw data is often extremely large, and directly processing the entire raw data set will face many difficulties, such as excessive consumption of computing resources, too long processing time, etc. By dividing the raw data into multiple relatively small data subsets through block processing, the subsequent processing operations can be more operable and efficient. This block method can fully utilize the advantages of distributed computing environment, and different data subsets are allocated to different computing nodes for parallel processing, thereby greatly improving the speed of the entire data processing flow. The specific block method can be determined according to the characteristics and storage method of the data. For example, the raw data generated in a certain time period can be divided into a subset according to the time sequence; or the raw data from the same system can be divided into a subset according to the data source system; or the data can be divided according to the business type, so that each data subset has certain relevance and relative independence in business logic.
[0067] In order to more comprehensively and deeply understand the structure and association relationship of the raw data, it is not enough to analyze the data from a single dimension. By introducing a tensor decomposition algorithm and considering multiple key dimensions (time, business scenario, data generating subject), the correlation information of the raw data under different levels and angles can be mined. These correlation information is of great significance for subsequent accurate hash encoding mapping of raw data and the entire data tracing process, because it can help identify the potential connection and similarity between data in different contexts. The tensor decomposition algorithm is a mathematical tool that can handle multi-dimensional data structures and reveal their internal components and relationships. For each data subset, it is regarded as a multi-dimensional tensor, where each dimension corresponds to a different attribute (such as the time dimension corresponding to the time point or time period of data generation, the business scenario dimension corresponding to the specific business activity to which the data belongs, and the data generating subject dimension corresponding to the individual, department or system that generates the data, etc.). The tensor decomposition algorithm is used to decompose the multi-dimensional tensor, thereby mining the correlation information between different dimensions and between different elements within the same dimension, and representing this information in the form of a correlation tensor. This correlation tensor can clearly show the correlation structure of the data subset under multiple dimensions, providing rich data association information for subsequent steps.
[0068] Hash functions play a key role in mapping raw data into hash codes. Different data subsets, due to their varying multi-dimensional correlation structures, require different hash function parameter settings for more accurate and efficient mapping. By analyzing the multi-dimensional correlation strengths in the correlation tensor, the hash function parameters can be dynamically adjusted based on the specific correlation characteristics of the data subsets. This allows the hash function to better adapt to the characteristics of each data subset, thereby improving the quality and accuracy of the hash codes. The correlation tensor generated for each data subset is analyzed in detail to determine the correlation strengths between dimensions and between elements within the same dimension. For example, correlation strengths can be quantified by calculating statistical metrics such as correlation coefficients. The hash function parameters are then adjusted based on these correlation strength analysis results. For example, if the correlation strength between the time dimension and the business scenario dimension of a data subset is high, the hash function parameters can be appropriately weighted to increase the weight of the relevant features in these two dimensions when adjusting the hash function parameters. This will allow the hash function to prioritize these highly correlated features during the mapping process, thereby more accurately mapping the raw data into hash codes.
[0069] After the previous steps, the next step involves accurately mapping the raw data into hash codes using a hash function with adjusted parameters. Hash coding compresses and transforms raw data, preserving certain characteristics to a certain extent and making subsequent operations such as similarity determination and clustering more convenient and efficient. The resulting hash code set serves as important foundational data for subsequent data traceability. For each previously processed data subset, a mapping operation is performed using a hash function with adjusted parameters that matches it. The raw data within each data subset is input into the hash function one by one. The hash function then calculates and converts the raw data into a fixed-length hash code. The hash codes generated by all the mapped data subsets are then combined to form the final hash code set. In this process, since the hash function parameters for each data subset are adjusted based on the results of its own correlation tensor analysis, the raw data of each data subset is accurately mapped into a hash code, thereby improving the quality and accuracy of the overall hash code set.
[0070] In one embodiment, a locality sensitive hashing algorithm is used to cluster similar hash codes in the hash code set into one category, thereby obtaining a plurality of clustered hash code sets, including:
[0071] For each hash code in the hash code set, mapping is performed using a preset locality sensitive hash function to obtain a hash bucket number corresponding to each hash code after mapping by the locality sensitive hash algorithm;
[0072] If two hash codes are assigned to the same hash bucket after mapping, the two hash codes are marked as potential similar codes;
[0073] For all hash codes marked as potential similar codes, the exact similarity is calculated;
[0074] The potential similar codes with the exact similarity higher than a threshold are clustered into a cluster hash code set; the above clustering process is repeated until all hash codes are classified into corresponding cluster hash code sets, and finally a plurality of cluster hash code sets are obtained.
[0075] In the embodiment, the core of the locality sensitive hashing algorithm is to map the high-dimensional data space (here, the space formed by the hash code set) to a low-dimensional hash bucket space by a specific hash function, so as to quickly make approximate similarity judgment on the data. Each hash code is mapped to a corresponding hash bucket number by a pre-set locality sensitive hash function, which lays a foundation for subsequent discovery of potential similar hash codes. Through the mapping, the hash codes that are difficult to directly compare similarity in the high-dimensional space can be preliminarily judged for similarity according to the hash bucket number to which they belong in the low-dimensional hash bucket space.
[0076] Firstly, a set of pre-set locality sensitive hash functions need to be determined. These functions are designed according to the characteristics of the data and the requirements of the similarity judgment, and they can preserve the similarity features between the data (hash codes) to a certain extent. Then, each hash code in the hash code set is input into these pre-set locality sensitive hash functions in turn for calculation. After the function operation, each hash code will be assigned to a specific hash bucket and obtain a corresponding hash bucket number.
[0077] Based on the characteristics of the locality sensitive hashing algorithm, when two hash codes are mapped to the same hash bucket, although it cannot be conclusively determined that they are completely similar, it can be highly believed that they have a certain similarity. Therefore, such two hash codes are marked as potential similar codes, so as to further accurately judge their similarity in the subsequent process, and thus to screen out the truly similar hash codes for subsequent operations such as clustering.
[0078] After obtaining the hash bucket numbers of all hash codes after mapping by the locality sensitive hash function, the hash codes are compared two by two. If it is found that the hash bucket numbers of two hash codes are the same, for example, hash code A is assigned to hash bucket number "3" and hash code B is also assigned to hash bucket number "3", then hash code A and hash code B are marked as potential similar codes. In this way, all hash codes are traversed to find all pairs of hash codes with similarity, and the corresponding marks are given.
[0079] For each pair of hash codes marked as potentially similar, an appropriate exact similarity calculation method is selected based on the specific hash code format and data characteristics. For example, if the hash code is in binary form, a bitwise similarity calculation method, such as Hamming distance, is used; if the hash code is in other forms, such as string form, an appropriate calculation method, such as edit distance, is used. By performing this exact similarity calculation on all hash code pairs marked as potentially similar, an exact similarity value for each pair is obtained.
[0080] Finally, based on the exact similarity calculated previously, truly similar hash codes are clustered together to form a clustered hash code set. This clustering operation can organize the previously disorganized data in the hash code set into orderly groups based on similarity, ensuring that the hash codes in each clustered hash code set have high similarity, facilitating subsequent operations such as data tracing based on clustering characteristics. This process is repeated until all hash codes are classified, ensuring that the entire hash code set is properly processed and clustered.
[0081] In one embodiment, for each cluster hash code set, extracting cluster features and data features of original data mapped by the cluster hash code set includes:
[0082] For each cluster hash code set, extract the cluster features from multiple dimensions; wherein the cluster features include the number of hash codes in each cluster hash code set, the average length of the hash codes, and the character distribution frequency of the hash codes;
[0083] For the original data mapped by the clustered hash code set, identify the data type in the data; count the quantity ratio of each data type as a type quantity ratio feature; count the storage capacity ratio of each data type as a storage capacity ratio feature;
[0084] The data feature is obtained by combining the type quantity ratio feature and the storage capacity ratio feature.
[0085] In this embodiment, the primary purpose of extracting cluster features is to quantitatively describe the overall characteristics of the cluster hash code set. This allows for a better understanding of the essential characteristics of each cluster based on these features, providing an important basis for generating accurate and representative traceability labels and performing data traceability analysis. Different cluster features reflect the clustering situation from different perspectives, helping to distinguish between different clusters.
[0086] Directly count the number of hash codes contained in each cluster's hash code set. This number can reflect the size of the cluster. For example, if a cluster has a large number of hash codes, the corresponding raw data has a wider distribution or more complex correlations to some extent. This can serve as an important reference factor in considering the overall nature of the cluster in subsequent analysis.
[0087] Calculate the sum of the lengths of all hash codes in the cluster's hash code set and divide it by the number of hash codes to obtain the average length. Hash code lengths vary depending on the characteristics of the data and the mapping method of the hash function. The average length reflects the overall characteristics of the hash code lengths within the cluster. For example, if a cluster has a longer average hash code length, it suggests that the original data corresponding to the cluster contained richer information or a more complex structure before the hash mapping.
[0088] For each hash code (assuming the hash code is in character form; if it's in other forms, the analysis method can be adjusted accordingly), analyze the frequency of each character's occurrence across all hash codes in the cluster's hash code set. By counting the number of times each character appears and dividing it by the total number of times all characters appear, the character distribution frequency is calculated. This character distribution frequency can reveal the distribution pattern of the cluster at the character level. Different clusters have different character distribution characteristics, which can help further distinguish the characteristics of different clusters. For example, some clusters have a higher frequency of occurrence of specific characters, which is related to certain specific attributes of the original data corresponding to the cluster.
[0089] Understanding the data types and the proportion of each type of data in the raw data mapped to the clustered hash code set allows for deeper insight into the characteristics of the raw data itself. The type and storage capacity ratios provide a more comprehensive picture of the raw data's composition, providing more detailed raw data-level information for subsequent operations like generating traceability tags. This helps to more accurately understand the relationship between the raw data and the clustered hash code set, as well as the overall characteristics of the data.
[0090] The raw data mapped to the clustered hash code set is checked one by one, and its data type is determined based on its format, structure, and content. Common data types include numeric types (such as integers and decimals), text types (such as strings), and dates. For example, numeric data is determined by checking whether it conforms to the numeric format specification, and text data is determined by checking whether it contains recognizable text characters.
[0091] After determining the various data types in the original data, the number of each data type is counted respectively, and then the number of each data type is divided by the total number of original data to obtain the proportion of the number of each data type, that is, the type number proportion feature. This feature can intuitively reflect the relative richness of different data types in the original data corresponding to the cluster, for example, if the number of numerical data is high, it means that the original data corresponding to the cluster is related to numerical calculation or quantitative analysis to a large extent.
[0092] In addition to the number proportion, the storage proportion of each data type also needs to be considered. The total storage space occupied by each data type (which can be measured by the number of data storage bytes, etc.) is calculated, and then the storage amount of each data type is divided by the total storage amount of the original data to obtain the storage amount proportion of each data type, that is, the storage amount proportion feature. This feature can reflect the difference in storage resource occupation of different data types, which is of great significance for some application scenarios that are sensitive to storage resources or for analyzing the storage characteristics of data.
[0093] By combining the type number proportion feature and the storage amount proportion feature, a more comprehensive and integrated data feature of the original data is formed. Such combination can more completely present the characteristics of the original data from the number and storage dimensions, providing a more rich and accurate information basis for subsequent operations such as generating a traceability label based on clustering features and data features, so that the generated traceability label can better reflect the true situation of the original data and the relationship with the cluster. For example, they can be arranged into a data structure, such as an array or a structure (depending on the programming language or data processing environment used), where one element stores the type number proportion feature and the other element stores the storage amount proportion feature. Or they can be encoded in a certain format to form a new string or encoded value as a data feature representing the original data. Through such a combination, it is ensured that the data feature can completely contain the key information extracted from the original data, so as to fully utilize these information for related operations in the subsequent process.
[0094] In an embodiment, the determination of the traceability relationship between each original data and the corresponding traceability label to realize data traceability comprises:
[0095] Parsing the original data corresponding to each traceability label and identifying the generation time and source system of the original data;
[0096] Sorting the original data corresponding to each traceability label according to the generation time and source system to obtain an original data sequence;
[0097] Establishing a traceability relationship between the original data sequence and the traceability label to realize data traceability.
[0098] In this embodiment, the traceability label is an identification closely associated with the original data and containing a lot of key traceability information. By analyzing the original data corresponding to each traceability label, information essential to determining the data traceability relationship, such as the generation time and the source system, can be extracted. These information are the basic elements for constructing the data flow path and clarifying the data source, and are helpful for subsequently accurately teasing out the trajectory of the original data in the entire life cycle.
[0099] For the identification of the generation time, the exact generation time point or time period can be determined according to the timestamp field in the data record or by analyzing the specific format representation related to time in the data content. For the identification of the source system, the specific system that produces the original data needs to be determined according to the format characteristics of the data, the storage location identifier, and the specific system identifier field.
[0100] The sorting of the original data according to the generation time and the source system is to establish an orderly arrangement of the data in the time and space dimensions. Such sorting helps to intuitively present the generation sequence of the data and the distribution of the data from different systems, thereby providing a clearer context for subsequently establishing accurate traceability relationship, so that the analysis can be more orderly when tracing the data source and flow path.
[0101] Generally, the original data can be sorted first according to the generation time, with the earlier generated original data arranged in front and the later generated original data arranged in back. If the generation time is the same, the source system can be used for further sorting. For example, the original data with the same generation time can be sorted again according to the alphabetical order of the name of the source system or the pre-set importance level of the system. Through such sorting method, the original data corresponding to each traceability label is arranged into an ordered sequence, i.e. the original data sequence.
[0102] Finally, by associating the original data sequence that has been sorted with the corresponding traceability label, the complete traceability relationship is established. Once this traceability relationship is established, the various links experienced by the data from generation to the current state can be accurately traced according to the rich information contained in the traceability label and the order presented by the original data sequence, realizing effective traceability of the data and meeting the needs of data management, compliance review, error troubleshooting, etc.
[0103] By establishing this traceability relationship, when it is necessary to trace the source of a specific data, we can quickly locate the corresponding traceability tag, and then obtain key information (such as generation time, source system, etc.) and view the position of the data in the original data sequence, so as to understand its previous and subsequent connections in the entire data flow process and restore the complete historical path of the data.
[0104] Specifically, a mapping mechanism can be used to establish the relationship between the original data sequence and the traceability label. This can be done by setting a specific association field in the data storage structure to bind the identifier of the traceability label to each data element in the original data sequence in a one-to-one correspondence; or by establishing an index table, the key information of the traceability label can be associated with the index value of the original data sequence. For example, in a database, a new field can be added to the original data table specifically for storing the number of the corresponding traceability label, and the index range of the original data sequence related to the label can also be recorded in the traceability label table. Through such a two-way association, a strong traceability relationship between the original data sequence and the traceability label is achieved. Once this relationship is established, data traceability operations can be conveniently performed when needed based on these association information.
[0105] In one embodiment, based on the clustering features and data features, a binary tree structure is used to generate traceability labels, including:
[0106] Performing principal component analysis on the clustering features and data features to extract principal component features;
[0107] Constructing an adaptive binary tree structure, and taking the principal component features as the root nodes of the binary tree structure;
[0108] Generate left and right child nodes of the root node based on the correlation between the principal component feature and other features; wherein the two features most correlated with the principal component feature are divided into features on the left and right child nodes of the root node;
[0109] During the iterative growth of the left and right child nodes of the binary tree, the correlation between the features under each child node and other features is calculated. When the correlation is greater than the preset threshold, the node is further subdivided; when the correlation is less than the threshold, the growth of the child node branch is stopped.
[0110] After the binary tree structure is constructed, a unique hash value is generated as a traceability label based on the path from the last leaf node to the root node.
[0111] In the embodiment, when facing a large number of clustering features and data features, there is a complex correlation and information redundancy between the above features. The purpose of principal component analysis (PCA) is to convert the original multiple features into a set of independent principal component features that can maximize the preservation of original data information through mathematical transformation. The data structure can be simplified, the key information of the data can be highlighted, and a more refined and effective feature basis can be provided for subsequent construction of a more representative and discriminative traceability label.
[0112] First, the above features are arranged according to the arrangement, to form a multi-dimensional feature matrix. Then, the standard calculation process of principal component analysis is performed on the feature matrix. This usually involves calculating the covariance matrix of the feature matrix, and then calculating the eigenvalues and eigenvectors of the covariance matrix. According to the size of the eigenvalue, the eigenvectors are sorted, and the features corresponding to the eigenvectors with larger eigenvalues are selected as the principal component features. These principal component features reduce the redundancy between the features while retaining most of the information of the original features, and can more effectively represent the key characteristics of the original data and clustering.
[0113] The binary tree structure has good hierarchy and logic, and is suitable for organizing and displaying complex feature information to generate traceability labels with clear hierarchical relationship and discriminability. The principal component features are selected as the root node because the principal component features are the key information that can best represent the overall characteristics of the original data and the clustering, and placing them in the root node position can lay a core foundation for the construction of the entire binary tree, and the subsequent branch and node growth will revolve around this core feature, so that the generated traceability label can accurately reflect the key traceability elements of the data from the root node.
[0114] An initial structure of a binary tree is created, a node is set as a root node, and the principal component features extracted by principal component analysis are assigned to the root node. For example, if the principal component feature is a feature value that integrates the size of the clustering hash code set and the proportion of the main data type of the original data, the feature value is stored in the data field of the root node as the starting point of the construction of the binary tree.
[0115] To further enrich the structure of the binary tree and refine the information contained in the traceability label, it is necessary to generate the left and right child nodes of the root node based on the correlation between the principal component features and other features. By selecting the two features most correlated with the principal component features as the features of the left and right child nodes, supplementary information closely related to the core principal component features can be introduced into the first-level branches of the binary tree. This allows the traceability label to gradually cover more levels of feature information during the subsequent generation process, enhancing the accuracy and discrimination of the traceability label. Among them, the correlation coefficient between the principal component features and the other remaining features (i.e., clustering features and data features other than the principal component features that have been used as the root node) can be calculated using common correlation calculation methods such as the Pearson correlation coefficient. The calculated correlation coefficients are sorted to find the two features with the highest correlation with the principal component features. These two features are used as the feature values of the left and right child nodes of the root node, and corresponding left and right child nodes are created, with the feature values assigned to the data fields of these two child nodes respectively.
[0116] Then, through an iterative approach, the complete structure of the binary tree is gradually constructed, allowing it to be reasonably subdivided and expanded based on the actual correlation between data features. By setting a correlation threshold to control the growth of nodes, we can ensure that the binary tree does not overgrow, resulting in an overly complex and difficult-to-understand structure, while also ensuring that nodes can be further subdivided to obtain more detailed feature information when there is sufficient correlation. This results in a binary tree structure that accurately reflects the key elements of data traceability while also having a reasonable complexity, providing a good foundation for the final generation of appropriate traceability labels.
[0117] For each child node of the binary tree (including left and right child nodes and child nodes of each level generated by subsequent iterations), repeat the following operation: calculate the correlation coefficient between the features under the child node (i.e., the features represented by the current child node and the features inherited from the parent node) and the other remaining features (i.e., all features except the features involved in the current child node and its parent node). Methods such as the Pearson correlation coefficient can also be used. Compare the calculated correlation coefficient with the preset correlation threshold. If the correlation coefficient is greater than the threshold, it means that the features under the current child node are still strongly correlated with other features, and it is necessary to further subdivide the child node, that is, create new left and right child nodes based on the two features with the highest correlation with the features under the current child node, and assign these two features to the data domains of the new left and right child nodes respectively. If the correlation coefficient is less than the threshold, it means that the features under the current child node are weakly correlated with other features. At this time, the growth of the child node branch is stopped and no new child nodes are created.
[0118] The purpose of finally generating the traceability label is to provide a unique identification for each original data or cluster that can accurately reflect the traceability key elements. By hashing the path information of the binary tree from the last leaf node to the root node, a unique hash value is generated as the traceability label. The uniqueness and irreversibility of the hash value can ensure that each original data or cluster has a unique and tamper-proof identification. This identification contains all the feature information involved in the path from the root node to the leaf node, and can accurately trace the source, characteristics and situation of the data in the clustering process, meeting the data traceability requirements.
[0119] Specifically, after the binary tree structure is constructed, starting from the last leaf node, along the path from the leaf node to the root node, the feature values (or other related information, depending on the storage structure of the binary tree nodes) of each node are extracted in turn. These extracted feature values are combined in a certain order (e.g. from leaf node to root node). It can be simply spliced into a string, or encoded according to a specific encoding rule. Then, the combined result is hashed using a hash function to obtain a unique hash value. This hash value is the traceability label generated for the original data or cluster.
[0120] In an embodiment, based on the cluster features and data features, a binary tree structure is used to generate a traceability label, comprising:
[0121] Encoding the cluster features to obtain cluster feature encoding;
[0122] Obtaining a preset binary tree structure, and sequentially adding characters in the cluster feature encoding to each node of the binary tree structure to generate an encoded binary tree;
[0123] Performing curve simulation on the type quantity proportion feature to generate a first curve;
[0124] Performing curve simulation on the storage quantity proportion feature to generate a second curve;
[0125] According to a preset superposition rule, superimpose the first curve and the encoded binary tree; take the nodes in the encoded binary tree that satisfy the preset positional relationship with the first curve as target nodes, and reinsert the characters on the target nodes from the root node position of the encoded binary tree in sequence, and move the characters on the original nodes backward in sequence to fill the complete binary tree to obtain a changed binary tree;
[0126] According to a preset superposition rule, superimpose the second curve and the changed binary tree; take the nodes in the changed binary tree that satisfy the preset positional relationship with the second curve as label nodes;
[0127] The characters on the tag node are extracted and combined to obtain the provenance label.
[0128] In this embodiment, the cluster features describe the characteristics of the cluster hash code set from multiple dimensions, such as the number of hash codes, the average length, the character distribution frequency, etc. The purpose of encoding these cluster features is to convert their complex and diverse information into a form that is more convenient for processing and storage in the subsequent binary tree structure. Through encoding, each cluster feature can be assigned a specific and unique character representation (or other encoding form), so that these features can be orderly integrated into the binary tree structure according to certain rules, laying the foundation for generating the provenance label.
[0129] The binary tree structure has good hierarchy and logic, and is suitable for organizing and displaying information. The preset binary tree structure is obtained to provide a unified framework for orderly integrating the encoded cluster features. By adding the characters in the cluster feature encoding to each node of the binary tree one by one, a binary tree based on the cluster feature encoding, i.e., the encoding binary tree, can be constructed, so that the cluster features are presented in a structured manner, preparing for further combining other data features to generate the provenance label.
[0130] First, a preset binary tree structure is prepared, which can be an empty binary tree with a certain number of nodes and hierarchical structure. The nodes of the binary tree can be pre-set to store data in a format that can receive and store the characters of the cluster feature encoding. Then, the characters in the cluster feature encoding are taken out in order, and the characters are added to each node of the binary tree one by one, starting from the root node. For example, if the cluster feature encoding is "ABCDE", "A" is added to the root node, "B" is added to the left child node of the root node (assuming a certain preset addition order), "C" is added to the right child node of the root node, and so on, until all the characters are added to the corresponding nodes of the binary tree, thereby generating the encoding binary tree.
[0131] The type quantity ratio feature reflects the proportion of different data types in the original data mapped to the clustered hash code set. The purpose of curve simulation is to convert this discrete ratio data into a continuous curve. This allows for a more intuitive display of the changing trends in the data type ratios and facilitates subsequent overlay with the binary tree structure, thus uniquely integrating this data feature into the traceability label generation process. Based on the data for the type quantity ratio feature, an appropriate curve simulation method is selected. Common methods include polynomial fitting and spline interpolation. Assume that the data for the type quantity ratio feature is a set of discrete points, for example, numeric data accounts for 30%, text data accounts for 50%, and date data accounts for 20%. These data points can be used as input, and a polynomial fitting method can be used to find a polynomial curve that best fits these points. This curve is used as the first curve. This curve can, to a certain extent, reflect the changing trends of the proportions of different data types as a function of some underlying variable (such as an abstract attribute of the data or a sorting criterion).
[0132] The storage capacity ratio feature describes the storage capacity ratio of different data types in the original data mapped by the clustered hash code set. Similar to the purpose of curve simulation for the type quantity ratio feature, the purpose of curve simulation for the storage capacity ratio feature is to convert it into a continuous curve form to more intuitively display the changing trend of the storage capacity ratio and facilitate subsequent overlay operations with the binary tree structure, so that this data feature can effectively participate in the generation of traceability labels.
[0133] Further, the type quantity proportion feature represented by the first curve obtained through the curve simulation is combined with the clustering feature code contained in the constructed encoding binary tree. By performing superposition operation according to the preset superposition rule and adjusting the nodes satisfying the specific position relationship, the type quantity proportion feature can be integrated into the binary tree structure in a unique way, changing the node layout and character distribution of the binary tree, thereby generating a changed binary tree (changed binary tree) to prepare for further combining the storage quantity proportion feature to generate the traceability label. First, the preset superposition rule is defined. This rule can be set according to specific business requirements and data characteristics, for example, it can be set that when the value of the first curve in a certain interval is greater than a certain threshold, the corresponding encoding binary tree node is the node (target node) that satisfies the preset position relationship. Or the matching relationship with the encoding binary tree node is determined according to the geometric characteristics such as the slope and curvature of the curve. Then, the first curve and the encoding binary tree are superimposed according to the set superposition rule. In the superposition process, it is judged one by one whether each node of the encoding binary tree satisfies the preset position relationship with the first curve. Once the target node is determined, the characters on the target node are sequentially reinserted from the root node position of the encoding binary tree. For example, if the target node is the left child node of the encoding binary tree, and the character on it is "B", then "B" is sequentially reinserted from the root node. At the same time, the characters on the original nodes are sequentially moved backward to fill the gaps caused by the reinsertion of characters, thereby ensuring the integrity of the binary tree structure, and finally obtaining the changed binary tree.
[0134] On the basis of generating the changed binary tree in the previous step, the storage quantity proportion feature (represented by the second curve) is further combined with the changed binary tree. By determining the nodes (label nodes) that satisfy the preset position relationship with the second curve, the storage quantity proportion feature can be integrated into the binary tree structure in a specific way to prepare for the characters needed for the final extraction of the traceability label.
[0135] Similarly, the operation is performed according to the preset superposition rule. This rule is similar to the rule when the first curve is superimposed with the encoding binary tree, and can also be adjusted according to specific circumstances, for example, the matching relationship with the changed binary tree node can be determined according to different characteristics of the second curve (such as the highest point, the lowest point, the turning point, etc.). According to the superposition rule, the second curve is superimposed with the changed binary tree, and in the superposition process, it is judged one by one whether each node of the changed binary tree satisfies the preset position relationship with the second curve. Once the label node is determined, these nodes will be used to extract the characters needed for the traceability label in the subsequent steps.
[0136] After the foregoing series of operations, the clustering features, type quantity proportion features, and storage quantity proportion features are integrated into the binary tree structure in different ways, and the label nodes are determined. The purpose of this step is to extract characters from the label nodes and combine them to obtain a traceability label that comprehensively reflects the original data and the clustering hash coding set features, so that in the subsequent data traceability process, the source of the data, the processing process, and other related information can be accurately traced through the traceability label.
[0137] In an embodiment, based on the clustering features and the data features, a binary tree structure is used to generate a traceability label, including:
[0138] For the clustering features, a feature correlation matrix is constructed; the feature correlation matrix is mapped to a two-dimensional plane using a multidimensional scaling algorithm to obtain a plurality of coordinate points; and a feature correlation curve is drawn according to the plurality of coordinate points;
[0139] When constructing the binary tree, the correlation evaluation results of the clustering features and the data features are used as root nodes;
[0140] In the binary tree growth process, a plurality of data are sequentially selected from the data features, and a data feature curve is constructed; the similarity of the feature correlation curve and the data feature curve at a plurality of data points is calculated; and the similarity results are sequentially added to the growing binary tree nodes;
[0141] After the binary tree is constructed, from the leaf nodes to the root nodes, the feature information of each node and the path features are combined to perform an encryption hash to generate a unique traceability label.
[0142] In this embodiment, the clustering features contain a lot of information, such as the distribution characteristics of the data, the aggregation degree of different categories of data, etc. The feature correlation matrix is constructed to clearly present the correlation between these features, facilitating subsequent analysis. A plurality of clustering features are like many different types of items. Each feature is regarded as a position in the matrix, and then the close degree of correlation between different features is determined. If two features are closely related, a value that can reflect the close relationship is assigned to the corresponding position in the matrix; if the correlation is weak, another value that can reflect the weak relationship is assigned. In this way, a matrix that can show the correlation of all features is formed.
[0143] The feature correlation matrix contains a high dimension of information, which is difficult to observe and understand directly. Multidimensional scaling algorithm is a tool that can compress high-dimensional information into a two-dimensional plane, and the relative position relationship between each feature can be seen directly. The above algorithm will find a suitable position for each feature point on the two-dimensional plane according to the correlation of each feature in the feature correlation matrix. The position relationship of these points can reflect the relative distance and correlation degree of the features in high-dimensional space.
[0144] Connecting the coordinate points on the two-dimensional plane into a curve can more clearly show the correlation trend between the clustering features. Just like connecting the scattered points on the map into a line, we can see how these features are related and changed. According to the preset order, the points on the two-dimensional plane are connected in turn to form a curve. In order to make the curve look more natural, some smoothing processing can also be done on the curve, just like polishing a rough line to make it smoother.
[0145] The above root node is the starting point of the binary tree, and choosing a suitable root node is important for the construction of the binary tree and the subsequent classification effect. Taking the correlation evaluation result of the clustering feature and the data feature as the root node can make the binary tree start from the most critical feature correlation and divide it, making the binary tree more effective in classifying data. First, evaluate the correlation between the clustering feature and the data feature, and find the relationship with the closest correlation. Then take this correlation relationship as the root node of the binary tree.
[0146] In order to better analyze the relationship between the data features and the clustering features, the trend of the data features is represented by a curve, which is convenient for comparison with the previously obtained feature correlation curve. Some data points are selected from the data features in a certain order. Then draw these data points on a two-dimensional plane and connect them to form a data feature curve. Similarly, the curve can also be smoothed.
[0147] By comparing the similarity of the two curves, we can know the matching situation between the data features and the clustering features. This helps us make more reasonable decisions in the construction process of the binary tree. Similarity algorithms are used to measure the similarity of the two curves, such as evaluating whether the shapes of the two curves are close, whether the change trends are consistent, etc. By comparing the two curves at multiple data points, a result is obtained that can represent their similarity.
[0148] Each node in the binary tree contains information about the relationship between data features and clustering features. This information can be leveraged when subsequently generating traceability labels, making them more accurate and meaningful. As the binary tree grows, the calculated similarity results are added to each new node as part of the node's information. The traceability label acts like an ID card for the data, uniquely identifying it. By combining the feature information of the binary tree nodes with path characteristics to generate traceability labels, each piece of data is uniquely identified, facilitating data traceability. Starting from a leaf node in the binary tree, the path is followed all the way to the root node, recording the feature information of each node, such as the correlation between features, similarity results, and whether the path was left or right. This information is combined and encrypted to produce a fixed-length traceability label. These traceability labels are unique and difficult to tamper with. The traceability label generation process is closely linked to the clustering and data features described above, and utilizes an innovative binary tree approach to enhance security and uniqueness.
[0149] Reference Figure 2 In one embodiment of the present invention, a data tracing system based on big data analysis is provided, including:
[0150] An encoding unit is used to analyze the original data in the data set based on big data analysis, and distribute the original data into hash codes according to the analysis results to obtain a hash code set;
[0151] A clustering unit, configured to cluster similar hash codes in the hash code set into one class using a locality sensitive hashing algorithm to obtain a plurality of clustered hash code sets; wherein the original data mapped to the hash codes in each clustered hash code set is divided into one class;
[0152] an extraction unit, configured to extract, for each of the cluster hash code sets, cluster features and data features of original data mapped by the cluster hash code sets;
[0153] A generating unit, configured to generate a traceability label using a binary tree structure based on the clustering features and the data features;
[0154] The traceability unit is used to determine the traceability relationship between each original data and the corresponding traceability tag to achieve data traceability.
[0155] In this embodiment, for the specific implementation of each unit in the above system embodiment, please refer to the above method embodiment, which will not be repeated here.
[0156] Reference Figure 3 In an embodiment of the present invention, a computer device is also provided. The computer device may be a server, and its internal structure may be as follows:Figure 3 The computer device includes a processor, a memory, a display screen, an input device, a network interface and a database connected by a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store corresponding data in the embodiment. The network interface of the computer device is configured to communicate with an external terminal through a network connection. The computer program is executed by the processor to implement the above method.
[0157] Those skilled in the art can understand that, Figure 3 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied.
[0158] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program. The computer program is executed by the processor to implement the above method. It can be understood that the computer readable storage medium in the embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0159] In summary, the data traceability method, system and device based on big data analysis provided in the embodiments of the present application include: analyzing original data in a data set based on big data analysis, distributing and mapping the original data into hash codes according to the analysis result to obtain a hash code set; using a local sensitive hash algorithm to cluster similar hash codes in the hash code set into a class to obtain a plurality of clustered hash code sets; wherein the hash codes in each clustered hash code set are mapped to a class of original data; for each of the clustered hash code sets, extracting a clustering feature and a data feature of the original data mapped by the clustered hash code set; generating a traceability label using a binary tree structure based on the clustering feature and the data feature; and determining a traceability relationship between each original data and the corresponding traceability label to achieve data traceability. In the present application, the original data is mapped into hash codes according to the analysis result of big data analysis in a distributed manner, which can fully utilize computing resources and process a large amount of data in parallel, avoiding the efficiency bottleneck caused by processing data one by one in traditional methods. The local sensitive hash algorithm is used to cluster similar hash codes in the hash code set into a class, and the local sensitive hash algorithm has the advantage of quickly locating similar data in a high-dimensional data space, which can efficiently perform clustering operation on a large number of hash codes. Compared with traditional methods, the amount of calculation and processing time is reduced, and the processing efficiency is improved. Based on the clustering feature and the data feature, a traceability label is generated using a binary tree structure, which can generate a unique label by combining multi-dimensional features and enhance its distinctiveness.
[0160] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiments. Any reference to memory, storage, database or other medium provided by the present application and used in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.
[0161] It is to be understood that the terminology "including", "comprising", or any other variation thereof, is intended to cover a non-exclusive inclusion such that process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a... " does not, without more constraints, exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0162] The preferred embodiments of the present application have been described above with the intent to enable those skilled in the art to make and use it. Various modifications to these embodiments will occur to those skilled in the art and are intended to be encompassed by the present application. Therefore, it is to be understood that, within the scope of the present application, the application can be practiced otherwise than as specifically described. For example, the present application can be implemented using software, hardware or both.
Claims
1. A data tracing method based on big data analysis, characterized in that: The following steps are involved: Analyze the original data in the data set based on big data analysis, and distribute and map the original data into hash codes according to the analysis results to obtain a hash code set, including: block processing of the original data in the data set to obtain multiple data subsets; for each data subset, use a tensor decomposition algorithm to mine correlation information under multiple dimensions to generate a correlation tensor; the multiple dimensions include time, business scenario, and data generation subject; analyze the correlation strength of the multiple dimensions in each of the correlation tensors, and adjust the parameters of the hash function based on the correlation strength; map the matching data subsets based on the hash function after adjusting the parameters, map the original data into hash codes, and obtain a hash code set; A locally sensitive hashing algorithm is used to cluster similar hash codes in the hash code set into one category to obtain multiple clustered hash code sets; wherein the original data mapped by the hash codes in each clustered hash code set is divided into one category; the method includes: for each hash code in the hash code set, mapping is performed using a preset locally sensitive hash function to obtain a hash bucket number corresponding to each hash code after mapping by the locally sensitive hashing algorithm; if two hash codes are assigned to hash buckets with the same number after mapping, the two hash codes are marked as potentially similar codes; for all hash codes marked as potentially similar codes, their exact similarity is calculated; the potentially similar codes with an exact similarity higher than a threshold are clustered into one category to form a clustered hash code set; the clustering process is repeated until all hash codes are classified into corresponding clustered hash code sets, and finally multiple clustered hash code sets are obtained; For each of the cluster hash code sets, extract cluster features and data features of original data mapped by the cluster hash code set; Based on the clustering features and data features, a binary tree structure is used to generate traceability labels; Determine the traceability relationship between each original data and the corresponding traceability tag to achieve data traceability.
2. The data tracing method based on big data analysis according to claim 1 is characterized in that: For each of the cluster hash code sets, extracting cluster features and data features of the original data mapped by the cluster hash code set, including: For each cluster hash code set, extract the cluster features from multiple dimensions; wherein the cluster features include the number of hash codes in each cluster hash code set, the average length of the hash codes, and the character distribution frequency of the hash codes; For the original data mapped by the clustered hash code set, identify the data type in the data; count the quantity ratio of each data type as a type quantity ratio feature; count the storage capacity ratio of each data type as a storage capacity ratio feature; The data feature is obtained by combining the type quantity ratio feature and the storage capacity ratio feature.
3. The data tracing method based on big data analysis according to claim 1 is characterized in that: Determining the traceability relationship between each piece of original data and the corresponding traceability tag to achieve data traceability includes: Parse the original data corresponding to each traceability tag and identify the generation time and source system of the original data; Sort the original data corresponding to each traceability tag according to the generation time and source system to obtain an original data sequence; A traceability relationship is established between the original data sequence and the traceability tag to achieve data traceability.
4. The data tracing method based on big data analysis according to claim 1 is characterized in that: Based on the clustering features and data features, a binary tree structure is used to generate traceability labels, including: Performing principal component analysis on the clustering features and data features to extract principal component features; Constructing an adaptive binary tree structure, and taking the principal component features as the root nodes of the binary tree structure; Generate left and right child nodes of the root node based on the correlation between the principal component feature and other features; wherein the two features most correlated with the principal component feature are divided into features on the left and right child nodes of the root node; During the iterative growth of the left and right child nodes of the binary tree, the correlation between the features under each child node and other features is calculated. When the correlation is greater than the preset threshold, the node is further subdivided; when the correlation is less than the threshold, the growth of the child node branch is stopped. After the binary tree structure is constructed, a unique hash value is generated as a traceability label based on the path from the last leaf node to the root node.
5. The data tracing method based on big data analysis according to claim 1 is characterized in that: Based on the clustering features and data features, a binary tree structure is used to generate traceability labels, including: Constructing a feature association matrix for the clustering features; using the feature association matrix as input, mapping it to a two-dimensional plane using a multidimensional scaling algorithm to obtain a plurality of coordinate points; and drawing a feature association curve based on the plurality of coordinate points; When constructing the binary tree, the correlation evaluation result of the clustering feature and the data feature is used as the root node; During the binary tree growth process, a plurality of data are sequentially selected from the data features and a data feature curve is constructed; similarities between the feature association curve and the data feature curve at a plurality of data points are calculated; and the similarity results are sequentially added to the growing binary tree nodes; After the binary tree is constructed, the path from the leaf node to the root node is encrypted and hashed by combining the characteristic information of each node and the path characteristics to generate a unique traceability label.
6. A data tracing system based on big data analysis, characterized in that: include: An encoding unit is configured to analyze the raw data in a data set based on big data analysis, and distribute-map the raw data into hash codes based on the analysis results to obtain a hash code set, including: performing block processing on the raw data in the data set to obtain multiple data subsets; for each data subset, using a tensor decomposition algorithm to mine correlation information in multiple dimensions to generate a correlation tensor; the multiple dimensions include time, business scenario, and data generation entity; analyzing the correlation strength of the multiple dimensions in each of the correlation tensors, and adjusting the parameters of the hash function based on the correlation strength; and mapping the matching data subsets based on the hash function after adjusting the parameters, mapping the raw data into hash codes to obtain a hash code set; A clustering unit is configured to cluster similar hash codes in the hash code set into one category using a local sensitive hashing algorithm to obtain a plurality of clustered hash code sets; wherein the original data mapped by the hash codes in each clustered hash code set is divided into one category; the method comprises: for each hash code in the hash code set, mapping is performed using a preset local sensitive hash function to obtain a hash bucket number corresponding to each hash code after mapping by the local sensitive hashing algorithm; if two hash codes are assigned to hash buckets with the same number after mapping, the two hash codes are marked as potentially similar codes; for all hash codes marked as potentially similar codes, their exact similarity is calculated; the potentially similar codes with an exact similarity higher than a threshold are clustered into one category to form a clustered hash code set; the clustering process is repeated until all hash codes are classified into corresponding clustered hash code sets, and a plurality of clustered hash code sets are finally obtained; an extraction unit, configured to extract, for each of the cluster hash code sets, cluster features and data features of original data mapped by the cluster hash code sets; A generating unit, configured to generate a traceability label using a binary tree structure based on the clustering features and the data features; The traceability unit is used to determine the traceability relationship between each original data and the corresponding traceability tag to achieve data traceability.
7. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Text clustering method and device, terminal and storage medium
CN110413787A
Data tracing method combining big data and multi-dimensional features and big data cloud server
CN112818067A