Multi-source data management method and related device based on machine learning
By processing multi-source data through machine learning, generating standardized data sets with timestamp alignment and merging to analyze blood relationships, the problems of time sequence confusion and dependency location in multi-source data governance are solved, and the accuracy of data governance and the support capabilities for rescue decisions are improved.
Patent Information
- Application Number
- CN202510854764.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Existing multi-source data governance methods fail to effectively and uniformly manage data of different forms, resulting in time sequence confusion, difficulty in locating operation impact chains and cross-entity dependencies, and difficulty in finding the root causes of errors, affecting the accuracy and reliability of rescue decisions.
A machine learning-based method is used to obtain multi-source data sets for preprocessing, generate a standardized data set with time stamp alignment, and call the blood relationship modeling model to extract data metadata and generate blood relationship descriptions, recording the operation process and association relationships of the entire data life cycle.
It realizes unified time base management of multi-source data, clearly displays data paths and dependencies, and improves the reliability of data governance and the support capability for rescue decisions.
Smart Images

Figure CN120429549B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data governance, and in particular to a multi-source data governance method and related devices based on machine learning. Background Art
[0002] With the improvement of the level of informatization of emergency command and rescue, multi-source data governance has become an important data management link. It manages data with mixed data forms from different rescue terminals during the emergency rescue process to ensure the traceability and reliability of the data and provide support for rescue decision-making. At present, the multi-source data governance method simply converts data of different forms into structured tables, or only records the source and generation time of the data, lacking in-depth analysis of the data life cycle operation process and the dependencies between data. Due to the lack of effective data unification, the dynamically changing data during the rescue process may cause temporal confusion, making it difficult to accurately trace the temporal dependencies between data; in addition, the current simple data recording method makes it impossible to effectively locate the data operation impact chain and the correlation relationship across rescue resources, events, and personnel; the lack of data link description also makes it easy to make it difficult to find the root cause of erroneous data and difficult to assess the impact of processing operations on the final data. Summary of the Invention
[0003] In view of this, an embodiment of the present invention provides a multi-source data management method and related device based on machine learning. The technical solution of the embodiment of the present invention is implemented as follows:
[0004] On the one hand, an embodiment of the present invention provides a multi-source data governance method based on machine learning, the method comprising: obtaining a multi-source data set in an emergency command and rescue scenario, the multi-source data set comprising original data units from different rescue terminals, the original data units having a mixed data form of unstructured text, semi-structured logs and structured tables; performing data preprocessing on the multi-source data set to obtain a standardized data set with a unified data form, the standardized data set comprising structured data entries aligned with timestamps; extracting a data meta-information set from the standardized data set, the data meta-information set comprising a data source identifier, a data generation time node, a data operation record and a data association object identifier; calling a pre-trained blood relationship modeling model to perform association analysis processing on the data meta-information set to generate a blood relationship description set between each data entry in the standardized data set, the blood relationship description set comprising a data source path, a processing process node and a dependency chain; generating a data blood relationship tracking result based on the blood relationship description set, the data blood relationship tracking result comprising a historical record of the entire life cycle of the data and dependency graph information across data entries.
[0005] In another aspect, the present invention provides a multi-source data management device, comprising:
[0006] A data acquisition module is used to acquire a multi-source data set in an emergency command and rescue scenario, wherein the multi-source data set includes raw data units from different rescue terminals, and the raw data units have a mixed data form of unstructured text, semi-structured logs, and structured tables;
[0007] a data preprocessing module, configured to perform data preprocessing on the multi-source data set to obtain a standardized data set having a unified data form, wherein the standardized data set includes structured data entries with aligned timestamps;
[0008] A metadata extraction module is used to extract a data metadata set from the standardized data set, wherein the data metadata set includes a data source identifier, a data generation time node, a data operation record, and a data-related object identifier;
[0009] A lineage tracking module is used to call a pre-trained lineage relationship modeling model to perform association analysis on the data metadata information set, and generate a lineage relationship description set between each data item in the standardized data set, wherein the lineage relationship description set includes a data source path, a processing process node, and a dependency chain;
[0010] A result generation module is used to generate data lineage tracing results based on the lineage relationship description set, wherein the data lineage tracing results include historical records of the entire life cycle of the data and cross-data entry dependency graph information.
[0011] The machine learning-based emergency command and rescue multi-source data governance method provided by the present invention obtains a multi-source data set in a mixed data form under an emergency command and rescue scenario, and generates a standardized data set with time stamp alignment through preprocessing, which can solve the problem of dependency breakage caused by temporal confusion under heterogeneous data forms and provide a unified time benchmark for subsequent lineage analysis; by extracting a data metadata set containing operation records and associated object identifiers, it can record the entire life cycle operation process of the data and its association with entities such as rescue resources, events, and personnel, avoiding the shortcomings of traditional data governance in being unable to trace the operation impact chain and locate cross-entity dependencies; by calling the lineage relationship The modeling model analyzes metadata to generate a set of lineage relationship descriptions including source paths, processing nodes and dependency chains, which can clearly display the complete path of data from original collection to final output, specific processing operations and derivative relationships between entries, solving the problems of difficulty in locating the root causes of errors and evaluating the impact of processing operations in rescue data; the final generated lineage tracking results including the historical records of the entire life cycle of data and cross-entry dependencies can provide emergency command and rescue with governance support such as traceable data sources, traceable processing processes and clear dependencies, so that rescue decisions can be made based on reliable data lineage information, effectively improving the support capabilities of data governance for rescue decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 A schematic diagram of the implementation process of a multi-source data governance method based on machine learning provided in an embodiment of the present invention.
[0013] Figure 2 A schematic diagram of the structure of a multi-source data management device provided in an embodiment of the present invention.
[0014] Figure 3 A schematic diagram of a hardware entity of a computer system provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0015] An embodiment of the present invention provides a multi-source data governance method based on machine learning, which can be executed by a processor of a computer system. The computer system can refer to a device with data processing capabilities, such as a server, laptop, tablet computer, or desktop computer.
[0016] Figure 1 A schematic diagram of the implementation process of a multi-source data governance method based on machine learning provided in an embodiment of the present invention is shown as follows: Figure 1 As shown, the method includes:
[0017] Step S100: Acquire a multi-source data set in an emergency command and rescue scenario, where the multi-source data set includes raw data units from different rescue terminals, and the raw data units have a mixed data form of unstructured text, semi-structured logs, and structured tables.
[0018] Emergency command and rescue scenarios encompass command, dispatch, and rescue operations in response to various emergencies, such as earthquakes, fires, and floods. In this scenario, different types of rescue terminals continuously generate data in various forms. Multi-source data sets integrate data from different rescue terminals, each with different formats and characteristics. Raw data units are the basic elements of a multi-source data set and can take on three different data forms: unstructured text, semi-structured logs, and structured tables.
[0019] Unstructured text data lacks a fixed format or standardized structure, such as rescue workers' on-site descriptions of rescue operations and conversations with disaster victims. While this textual information can provide a detailed and rich description of the specific conditions at the rescue site, it lacks a unified organizational structure, making it difficult to directly analyze and process systematically. For example, at an earthquake rescue site, rescue workers might record, "Several houses have collapsed on a certain street, people are trapped, and large-scale rescue equipment is needed." This description is relatively free-form and lacks strict formatting constraints.
[0020] Semi-structured log data lies somewhere between structured and unstructured, containing predefined fields and formats, but not as strict and standardized as structured tables. For example, an operation log for a rescue device might include the timestamp of the operation, the specific operation type (such as start, stop, parameter adjustment), and the result of the operation (success, failure, exception, etc.). For example, a rescue device operation log might record "2024-10-15 14:30:00 Started device, operation successful." While this format is consistent, there may be subtle differences between different devices or systems.
[0021] Structured tabular data has a clear row and column structure, with each row representing a complete record and each column representing a specific attribute. For example, a table listing relief supplies might include columns such as the name of the supply, quantity, storage location, and receipt date, with each row corresponding to a specific type of supply. This data format facilitates storage, querying, and statistical analysis.
[0022] In actual emergency rescue scenarios, this multi-source data can be obtained in a variety of ways. For unstructured text data, rescuers can use their mobile devices, such as dedicated rescue recording apps, to enter relevant information in real time on-site. For semi-structured log data, rescue equipment itself has a built-in logging function that automatically records device operation information and transmits it to a data center via the network. Structured tabular data may be generated by relevant management systems, such as material management systems and personnel management systems, which regularly synchronize data to designated storage locations. For example, during a fire rescue, firefighters can use a mobile app to record unstructured text data such as the fire situation at the scene and estimated casualties. A fire truck's equipment log will record semi-structured log data such as vehicle startup, driving speed, and water spraying operations. A material management system will record structured tabular data such as the inventory quantity and distribution status of firefighting equipment.
[0023] Step S200: performing data preprocessing on a multi-source data set to obtain a standardized data set with a unified data form, wherein the standardized data set includes structured data entries aligned with timestamps.
[0024] Data preprocessing involves a series of cleaning, conversion, and integration operations on multi-source data sets. The goal is to eliminate noise, inconsistencies, and redundant information in the data, thereby aligning different data forms into a unified data format for subsequent analysis and processing. Unifying data forms means converting unstructured text, semi-structured logs, and structured tabular data into a consistent, easy-to-process format. A standardized data set is the result of preprocessing, where the data entries not only have a unified format but also contain structured data with aligned timestamps. Timestamp alignment involves uniformly adjusting the time information of different data entries to make the data comparable and coherent across the time dimension.
[0025] Specifically, step S200 can be implemented as the following steps S210 to S250:
[0026] Step S210: performing text parsing processing on the unstructured text data units in the multi-source data set to extract text segments containing key event descriptions, resource scheduling instructions, and on-site feedback information.
[0027] Text parsing is an analytical operation performed on unstructured text data, aiming to extract valuable content from large amounts of text information. Key event descriptions refer to textual content that accurately describes important events in emergency rescue scenarios, such as the time, location, and scale of the disaster. This information is crucial for understanding the full picture of the incident and formulating rescue strategies. Resource dispatch instructions are specific instructions on the allocation of rescue resources, such as the type of materials to be allocated, the quantity to be allocated, and the destination of the allocation. These instructions are directly related to the resource allocation and coordination of rescue work. On-site feedback information is the real-time feedback from rescue personnel on the scene, such as the needs of the affected people, the progress of the rescue work, and the difficulties encountered. This information can help the command center adjust the rescue strategy in a timely manner.
[0028] When parsing text, a rule-based parsing approach can be employed. First, a series of keywords and rules are defined. For example, by identifying keywords such as "occurrence" and "discovery," combined with subsequent information such as time and location, key event descriptions can be extracted. Resource dispatch instructions can be extracted by searching for keywords such as "deploy" and "dispatch," followed by information such as the name, quantity, and location of the material. For on-site feedback, relevant information can be filtered based on the text description, combined with the actual rescue process and key points. For example, in the text "This morning, a fire broke out in a certain residential complex. The fire is large. Five fire trucks and ten firefighters are needed immediately for rescue. Currently, several people are trapped and urgently need medical assistance," rule matching can be used to extract the key event description "This morning, a fire broke out in a certain residential complex. The fire is large. Five fire trucks and ten firefighters are needed immediately for rescue. Currently, several people are trapped and urgently need medical assistance."
[0029] Step S220: performing log format parsing processing on the semi-structured log data units in the multi-source data set, and extracting log fields including timestamp, operation type and execution result.
[0030] Log format parsing addresses the unique characteristics of semi-structured log data, parsing it into a format with clearly defined fields for subsequent analysis and use. Timestamps record the specific time a log event occurred, which can be used to determine the sequence and time intervals between events. Operation types indicate the specific actions performed during the rescue process, such as starting or stopping equipment, adjusting parameters, etc. Different operation types reflect the operating status of the rescue equipment and the execution of the task. Execution results describe the final outcome of the operation, such as success, failure, or exception.
[0031] When parsing semi-structured log data, pattern matching can be used. Based on common formats and patterns found in log data, corresponding matching rules can be defined. For example, in most log data, timestamps appear in a specific date and time format, such as "YYYY-MM-DDHH:MM:SS." String matching can be used to extract the portion of the log text that matches this format as the timestamp. For operation types and execution results, common operation description terms and result identifiers, such as "start," "stop," "success," and "failure," can be used to extract the corresponding fields based on the log's context and grammatical structure. For example, for the log entry "2024-10-16 10:15:00 Started generator, operation successful," pattern matching can accurately extract the timestamp "2024-10-16 10:15:00," the operation type "Start generator," and the execution result "Operation successful."
[0032] Step S230: performing field consistency verification on the structured table data units in the multi-source data set, and screening out table entries whose field names and data types comply with preset specifications.
[0033] Field consistency verification is performed to ensure the quality and standardization of structured table data. This process screens out qualified table entries by checking whether the field names and data types in the table conform to preset specifications. Preset specifications are pre-established standards based on the needs of data management and analysis, including naming conventions for field names and data type definitions. Field names should have clear meanings and accurately describe the attributes they represent. Data types should match the actual content of the field. For example, fields representing quantities should be numeric, and fields representing names should be string.
[0034] When performing field consistency check, you can follow the steps below. First, define a preset list of field names and data types, and compare each field name in the table with the preset field name to check for consistency. For data types, you can judge based on the data's value range, format, and other characteristics. For example, for a field representing a date, check whether its value meets the date format requirements. If the field name or data type in the table does not meet the preset specifications, the table entry is considered unqualified and needs to be corrected or excluded. For example, in a relief supplies inventory table, the preset specification requires that the data type of the "supply quantity" field be an integer. If there is a non-integer value in this field in the table, the entry does not meet the specifications and needs to be processed.
[0035] Step S240: input the parsed text segments, extracted log fields and filtered table entries into a data format conversion module to generate an intermediate data set with unified field names, data types and storage formats.
[0036] The data format conversion module converts data in different formats into a unified format. Field name standardization eliminates differences in field names across different data sources and prevents data confusion caused by inconsistent names. Data type standardization unifies the data types of different data entries, ensuring consistency during data analysis and processing. Unified storage formats store data in the same format for easier management and use.
[0037] The data format conversion module can be implemented using a mapping and conversion approach. First, define a unified set of field names and data type specifications. Then, map the field names and data types of parsed text snippets, extracted log fields, and filtered table entries. For text snippets, map them to unified field names based on the extracted key information, such as mapping key event descriptions to the "Event Description" field. For log fields, map timestamps, operation types, execution results, and other fields to corresponding unified fields. For table entries, adjust the filtered field names and data types to conform to unified specifications. For storage format, choose a common format, such as CSV, and store the converted data in a file in that format to form an intermediate data set. For example, convert the parsed text, log, and table data into a CSV file containing fields such as "Event Description," "Operation Time," "Operation Type," "Material Name," and "Material Quantity."
[0038] Step S250: performing timestamp alignment processing on the intermediate data set, generating time-aligned data entries with uniform time intervals through a time series interpolation algorithm based on the original timestamp information of each data unit, and obtaining a standardized data set.
[0039] Timestamp alignment is used to ensure temporal consistency and comparability of data entries within intermediate data sets. Original timestamp information represents the actual time of occurrence of each data unit record, but due to varying acquisition frequencies and time accuracy across data sources, timestamp discrepancies may occur. Time series interpolation algorithms are used to fill missing values in time series data or generate uniformly spaced data intervals. These algorithms estimate the data values at missing points based on known timestamp information, thereby generating time-aligned data entries with uniformly spaced intervals.
[0040] The time series interpolation algorithm can adopt linear interpolation and polynomial interpolation. Taking linear interpolation as an example, exemplarily, there are two adjacent data entries in the intermediate data set. One has a timestamp of t1 and a corresponding data value of v1, and the other has a timestamp of t2 and a corresponding data value of v2, and t1 < t2. If it is necessary to estimate the data value v at a certain time point t (t1 < t < t2) between t1 and t2, it can be estimated according to the linear relationship, and the formula is v = v1 + (v2 - v1) × (t - t1) / (t2 - t1). By performing similar interpolation processing on all data entries in the intermediate data set, time-aligned data entries with uniform time intervals can be obtained, and finally a standardized data set is formed.
[0041] Step S300: Extract a data element information set from the standardized data set. The data element information set includes a data source identifier, a data generation time node, a data operation record, and a data associated object identifier. <000
[0046] When extracting source identifiers, a concatenation method can be used to generate a source identifier code. The device number and system identifier are concatenated, separated by a specific symbol, such as "-." For example, if the device number is "001" and the system identifier is "Rescue Command System," the generated source identifier code would be "Rescue Command System-001." This ensures that each data entry has a unique source identifier code, making it easier to trace and manage the data's origin.
[0047] Step S320: Perform generation time node extraction processing on each data entry in the standardized data set, and extract the specific date and time of data generation as the data generation time node by parsing the timestamp field in the data entry.
[0048] Generation time node extraction involves extracting the specific time information of data generation from data entries in a standardized data set. The timestamp field records the time of data generation and can be stored in a specified format, such as "YYYY-MM-DDHH:MM:SS." By parsing this field, the date and time of data generation can be accurately extracted.
[0049] When parsing a timestamp field, you can perform appropriate processing based on its format. If the timestamp field format is a standard date and time format, you can directly extract the year, month, day, hour, minute, and second information. For example, for the timestamp "2024-10-18 15:20:30," you can directly extract the date "2024-10-18" and the time "15:20:30" as the data generation time node. If the timestamp field format does not conform to the standard, you can first convert the format and then extract the corresponding information.
[0050] Step S330: Perform operation record extraction processing on each data entry in the standardized data set, and extract operation record information including the operation type, operation subject and operation time of data collection, transmission, storage and modification by parsing the operation log field in the data entry.
[0051] Operation log extraction involves extracting relevant information about data operations from data entries in a standardized data set. The operation log field records operations throughout the data lifecycle, including data collection, transmission, storage, and modification. The operation type indicates the specific operation, such as collection, transmission, storage, or modification; the operation subject is the object performing the operation, such as a rescuer, device, or system; and the operation time records the specific time when the operation occurred.
[0052] When parsing operation log fields, you can use rule matching. By pre-defining a series of keywords and rules, you can extract the operation type, operator, and operation time based on the description in the operation log. For example, in the operation log "2024-10-19 10:00:00 Rescue worker A collected material inventory data," rule matching can extract the operation type "collection," the operator "rescuer A," and the operation time "2024-10-19 10:00:00."
[0053] Step S340: performing associated object identification extraction processing on each data entry in the standardized data set, and extracting the rescue resource identification, event identification and personnel identification associated with the data entry as the data associated object identification by parsing the associated fields in the data entry.
[0054] The associated object identifier extraction process extracts the identifiers of objects associated with data entries in a standardized data set. An associated field records the relationship between data and other objects. For example, in a rescue supply data entry, the associated field may contain information such as the rescue event associated with the supply and the personnel who used the supply. Rescue resource identifiers uniquely identify resources such as rescue supplies and equipment; event identifiers identify rescue events; and personnel identifiers identify personnel involved in the rescue.
[0055] When parsing associated fields, you can process them based on their content and format. If the associated field is a string separated by specific symbols, such as multiple identifiers separated by commas, it can be split into multiple identifiers. For example, if the associated field content is "Supplies 001, Event 002, Personnel A," it can be split into the rescue resource identifier "Supplies 001," the event identifier "Event 002," and the personnel identifier "Personnel A." If the associated field is a complex structure, the corresponding identifier information can be extracted based on its internal attributes and relationships.
[0056] Step S350: The source identification code, data generation time node, operation record information and data-related object identification are combined and packaged to generate a data metadata information set.
[0057] The combined encapsulation process is to integrate the extracted source identification code, data generation time node, operation record information and data-related object identifier to form a complete data meta-information set. For example, encapsulation can be performed in the form of a data structure, such as using an object or dictionary to store this information. Each data entry corresponds to an object or dictionary, which contains fields such as the source identification code, data generation time node, operation record information and data-related object identifier. For example, a data structure can be created, which contains attributes such as "source identification code", "data generation time node", "operation record information" and "data-related object identifier". The extracted corresponding information is assigned to these attributes, and then the object of each data entry is stored in a list to form a data meta-information set. In this way, the data in the standardized data set can be conveniently described and managed through the data meta-information set.
[0058] Step S400: Call the pre-trained blood relationship modeling model to perform association analysis on the data metadata set to generate a blood relationship description set between each data item in the standardized data set. The blood relationship description set includes the data source path, processing process node and dependency chain.
[0059] The kinship modeling model is a pre-trained model used to analyze relationships between data. Association analysis uses this model to analyze a set of data metadata and determine the kinship relationships between data items. The kinship description set is a detailed description of the kinship relationships between data items. It includes the data source path (the path the data takes from its initial source to its final state); the processing nodes (the various steps the data passes through during processing); and the dependency chain (the dependencies between data items, where one data item may depend on other data items for generation).
[0060] As an implementation manner, step S400 can be specifically implemented as the following steps S410 to S450:
[0061] Step S410: Input the data metadata set into the feature embedding layer of the blood relationship modeling model, perform feature vectorization on the data source identifier, data generation time node, data operation record and data associated object identifier, and generate a metadata feature vector with a unified dimensional representation.
[0062] The feature embedding layer is a component of the kinship modeling model, used to convert different types of information within the data metadata collection into vector representations. Feature vectorization involves converting information such as data source identifiers, data generation time points, data operation records, and data-related object identifiers into vectors with uniform dimensions, enabling the model to effectively process and analyze this information. Uniform dimensionality representation converts different types of information into vectors of the same length, ensuring consistency during model calculations and comparisons.
[0063] For data source identifiers, they can be converted into vectors using encoding. For example, a unique code can be assigned to each different source identifier, and then the code can be converted into a vector representation. For data generation time nodes, they can be digitized, such as converting the date and time into a timestamp value, and then normalizing the value and converting it into a dimension of the vector. For data operation records, information such as the operation type, operation subject, and operation time can be combined and encoded and then converted into a vector. For data-related object identifiers, they can be converted into vectors using an encoding method similar to that of the data source identifier. Finally, these vectors are spliced to form a metadata feature vector with a unified dimension.
[0064] Step S420: Perform time series dependency analysis on the meta-information feature vector through the time series association module of the blood relationship modeling model to identify the generation and derivation relationships between data entries at different time nodes.
[0065] The Time Series Association Module is used within the kinship model to analyze relationships between time series data. Time series dependency analysis uses this module to analyze metadata feature vectors and identify generation and derivation relationships between data items at different time points. Generation and derivation relationships indicate that a data item may be generated from other data items through certain processing or operations, or that one data item may derive from other data items.
[0066] As an implementation manner, step S420 may be specifically implemented as the following steps S421 to S425:
[0067] Step S421: sorting the data generation time nodes in the meta-information feature vector to generate a meta-information feature sequence arranged in chronological order.
[0068] Sorting involves arranging the data generation time nodes in the metadata feature vector in chronological order. This process converts the metadata feature vector into a chronological sequence, facilitating subsequent analysis. Sorting algorithms, such as bubble sort and selection sort, can be used to sort the data generation time nodes. During the sorting process, the order of the metadata feature vectors is adjusted so that they are also arranged in chronological order, forming a metadata feature sequence.
[0069] Step S422: extracting meta-information feature vectors of adjacent time nodes in the meta-information feature sequence, and calculating the cosine similarity value between the feature vectors as a correlation index of temporally adjacent data entries.
[0070] Cosine similarity measures the similarity between two vectors by calculating the cosine of the angle between them. By extracting metadata feature vectors from adjacent time nodes within a metadata feature sequence and calculating the cosine similarity between them, we can determine the correlation index for temporally adjacent data entries. A higher correlation index indicates a higher similarity between two temporally adjacent data entries, potentially indicating a closer relationship.
[0071] Step S423: performing threshold screening processing on the correlation index, and retaining temporally adjacent data item pairs whose correlation index exceeds a preset threshold.
[0072] The purpose of the threshold screening process is to filter out data entry pairs with low correlation and retain only data entry pairs with high correlation. The preset threshold is a value pre-set according to the actual situation and analysis requirements. Only data entry pairs whose correlation index exceeds the threshold will be retained. Through the threshold screening process, unnecessary calculations and analysis can be reduced, and the efficiency and accuracy of the analysis can be improved. By traversing the calculated correlation index list, the indexes of the data entry pairs whose correlation index exceeds the preset threshold can be recorded, and then the corresponding data entry pairs can be extracted based on these indexes. For example, the preset threshold is 0.8. When the calculated correlation index of a data entry pair of two adjacent time nodes is 0.9, the data entry pair will be retained; and when the correlation index is 0.7, the data entry pair will be filtered out.
[0073] Step S424: Identify the modification operation types in the data operation records and extract the time nodes corresponding to the modification operation records including data copying, conversion and merging.
[0074] Modification operation type identification involves identifying records containing modification operations such as data copying, conversion, and merging, and extracting the corresponding time points for these records. Data copying refers to copying a data entry to another location or creating a copy; data conversion refers to performing operations such as format conversion and calculations on data; and data merging refers to combining multiple data entries into a single one. By identifying the time points corresponding to these modification operation records, we can identify the key time points in data generation and derivation relationships.
[0075] For example, pattern matching can be used to identify data operation records. Define a series of keywords, such as "copy," "convert," and "merge," and search for records containing these keywords within the data operation records. For records containing these keywords, extract the operation time as the corresponding time node. For example, in the operation record "2024-10-2014:00:00 data copy operation," pattern matching can identify this as a data copy operation record and extract the time node "2024-10-2014:00:00."
[0076] Step S425: Based on the time nodes corresponding to the modification operation records, the original data entry of the previous time node and the modified data entry of the next time node are associated and matched to identify data generation and derivation relationships.
[0077] Association matching involves associating the original data entry at the previous time node with the modified data entry at the next time node, based on the time nodes corresponding to the modification operation records, to identify the generation and derivation relationships between them. For example, if a data copy operation is performed at a certain time node, the original data entry at the previous time node is the source data for the copy operation, and the modified data entry at the next time node is the copy, thus establishing a generation and derivation relationship between them.
[0078] By traversing the time nodes corresponding to the modification operation records, we can find the original data entry at the previous time node and the modified data entry at the next time node in chronological order. We can then match and associate the data based on its content and characteristics. For example, we can compare the source identifiers and associated object identifiers of the data entries to determine whether they are related. If the source identifiers are the same and the associated object identifiers are also somewhat related, then we can assume that the two data entries have a generation and derivation relationship.
[0079] Step S430: performing cross-analysis processing on the source identification and the associated object identification of the meta-information feature vector through the spatial association module of the blood relationship modeling model, and identifying the interactive relationship between different source data items and the same associated object.
[0080] The spatial association module, within the kinship modeling model, analyzes data associations in the spatial dimension. This module analyzes metadata feature vectors for cross-analysis of source and associated object identifiers, identifying interactions between data entries from different sources and the same associated object. Interactions mean that data entries from different sources may be associated with the same associated object. For example, data collected by different rescue equipment may be related to the same disaster event.
[0081] As an implementation manner, step S430 may be specifically implemented as the following steps S431 to S436:
[0082] Step S431: performing grouping processing on the data source identifiers in the meta-information feature vector to generate a plurality of data entry groups divided by the source identifiers.
[0083] Grouping involves categorizing the meta-information feature vectors according to the data source identifier, grouping data entries with the same source identifier. This allows data to be categorized by source, facilitating subsequent analysis. For example, grouping can be implemented using a hash table or dictionary, with the data source identifier serving as the key and the data entries with the same source identifier stored as the value under the corresponding key. For example, for a list of meta-information feature vectors, data entries with the source identifier "rescue equipment A" are grouped together, while data entries with the source identifier "rescue equipment B" are grouped together.
[0084] Step S432: performing aggregation processing on the data-related object identifiers in the meta-information feature vector to generate a plurality of related object groups divided according to the related object identifiers.
[0085] Aggregation processing involves classifying metadata feature vectors according to the data-related object identifier, grouping data entries associated with the same associated object identifier. This allows data to be classified according to associated objects, facilitating analysis of the relationship between data entries from different sources and the same associated object. Aggregation processing can also be implemented using a hash table or dictionary, using the data-related object identifier as the key and storing data entries with the same associated object identifier as the value under the corresponding key. For example, data entries with the associated object identifier "Event 001" are grouped together, and data entries with the associated object identifier "Event 002" are grouped together.
[0086] Step S433: extracting the meta-information feature vectors of the different source data entry groups contained in each associated object group, and calculating the feature overlap of the different source data entry groups in the same associated object group.
[0087] Feature overlap refers to the degree of overlap between feature vectors of data entry groups from different sources within the same associated object group. By calculating the feature overlap, the degree of association between data entries from different sources and the same associated object can be measured. Similarity calculation methods between vectors, such as cosine similarity and Euclidean distance, can be used to calculate the feature overlap between metadata feature vectors of data entry groups from different sources. For example, for data entry groups with source identifications of "rescue equipment A" and "rescue equipment B" within the same associated object group, the cosine similarity between their metadata feature vectors is calculated and used as the feature overlap.
[0088] Step S434: identifying collaborative record relationships between data items from different sources on the same associated object based on feature overlap.
[0089] A collaborative record relationship refers to a cooperative record relationship between data entries from different sources on the same associated object. If the feature overlap of data entries from different sources on the same associated object is high, it means that there may be a collaborative record relationship between them, that is, these data entries may be jointly recording relevant information of the same associated object. A threshold can be set according to the feature overlap. When the feature overlap exceeds the threshold, it is considered that data entries from different sources have a collaborative record relationship on the same associated object. Exemplarily, the preset threshold is 0.7. When the calculated feature overlap of two groups of data entries from different sources on the same associated object is 0.8, it is considered that there is a collaborative record relationship between them.
[0090] Step S435: Identify the transmission operation type in the data operation record, and extract the source identification pair and the associated object identification corresponding to the operation record containing the cross-source data transmission.
[0091] Transfer operation type identification involves identifying operation records involving cross-source data transfers and extracting the corresponding source ID pairs and associated object IDs from these records. Cross-source data transfer refers to the transfer of data from one source to another, potentially involving different rescue equipment or systems. By identifying the corresponding source ID pairs and associated object IDs in these transfer operation records, key information about data interaction relationships can be identified.
[0092] Pattern matching can be used to identify data operation records. Define a series of keywords, such as "transmit," "send," and "receive," and search for records containing these keywords in the data operation records. For records containing these keywords, extract the source identification pair (sending source identification and receiving source identification) and the associated object identification. For example, in the operation record "2024-10-21 15:00:00 Data was transmitted from rescue device A to rescue device B, and the associated event is event 003," pattern matching can be used to extract the source identification pair "rescue device A, rescue device B" and the associated object identification "event 003."
[0093] Step S436: Based on the source identification pair and the associated object identification, the original data entry of the sending source and the target data entry of the receiving source are associated and matched to identify the data interaction relationship.
[0094] Association matching involves associating the original data entry from the sending source with the target data entry from the receiving source based on the source ID pair and the associated object ID, identifying the data interaction relationship between them. For example, if a transfer operation record shows data transferred from source A to source B, and the associated object is associated object C, then the original data entry from source A can be associated with the target data entry from source B, assuming a data interaction relationship exists between them.
[0095] For example, by traversing the extracted source identification pairs and associated object identifications, the original data entry of the corresponding sending source and the target data entry of the receiving source can be found according to the source identification, and then matching and associating can be performed according to the associated object identification. For example, for the source identification pair "rescue device A, rescue device B" and the associated object identification "event 003", the original data entry with the source identification of "rescue device A" and the associated object identification of "event 003" and the target data entry with the source identification of "rescue device B" and the associated object identification of "event 003" are searched in the data entry, and they are associated to identify the data interaction relationship.
[0096] Step S440: Based on the generated derivative relationship output by the temporal association module and the interactive relationship output by the spatial association module, a dependency graph between data items is constructed. The dependency graph includes edges representing the relationship between data items and attribute labels representing the relationship type.
[0097] A dependency graph is a graphical structure used to visually display the dependencies between data items. This dependency graph is constructed based on the generation and derivation relationships between data items at different time points, identified by the temporal association module, and the interactions between data items from different sources and the same associated object, identified by the spatial association module.
[0098] When constructing a dependency graph, each data entry is first considered a node in the graph. For generation and derivation relationships output by the temporal association module, if one data entry is generated or derived from another, an edge is added between the nodes corresponding to the two data entries. This edge represents the generation or derivation relationship between them, and a corresponding attribute label, such as "generation" or "derived," is added to the edge to clarify the type of relationship.
[0099] For example, in an emergency rescue scenario, if a data entry about the use of rescue supplies is generated by calculating and converting a previous material inventory data entry, then an edge is added between the nodes corresponding to the two data entries, and the attribute label is marked as "generated".
[0100] For the interactive relationships output by the spatial association module, when data entries from different sources have a collaborative recording or data transfer relationship on the same associated object, edges are added between the corresponding data entry nodes, and attribute labels are added based on the specific relationship type. For example, if data entries from two different rescue devices have a collaborative recording relationship because they jointly record a disaster event, an edge is added between the two nodes with the attribute label "Collaborative Recording"; if a data transfer operation occurs from one rescue device to another, an edge is added and the attribute label is "Data Transfer."
[0101] In this way, all identified relationships are converted into edges and attribute labels in the dependency graph, thereby constructing a complete dependency graph between data entries.
[0102] Step S450: Perform path traversal processing on the dependency graph, extract complete path information from the initial data entry to the subsequent derived data entry, and generate a blood relationship description set.
[0103] Path traversal is the process of finding all complete paths from initial data entries to subsequent derived data entries within the constructed dependency graph. Initial data entries are those that have no preceding dependent nodes in the dependency graph and are the source of the data. Subsequent derived data entries are those that are derived from the initial data entries through a series of processing and operations.
[0104] When traversing a path, you can use either a depth-first search (DFS) or breadth-first search (BFS) algorithm. Taking depth-first search as an example, starting from the node corresponding to the initial data entry, the search continues along the edges in the dependency graph until all possible paths are found. During the exploration process, each node and edge traversed, along with the edge attribute labels, is recorded, thus obtaining a complete path.
[0105] For example, starting from an initial data entry node representing the initial inventory of relief supplies, the algorithm explores along the edges to a data entry node for the remaining supplies generated after a supply distribution operation. During this process, information about each intermediate node and edge is recorded, including the specific data entry corresponding to the node and the attribute label of the edge (such as "derived from the distribution operation"). After traversing all possible paths in the dependency graph, the information of each path is organized and summarized to generate a set of lineage descriptions. This set contains the complete path information from the initial source to the final derived state, including the data source path, the processing nodes passed through, and the dependency chain between nodes, providing a detailed description for a comprehensive understanding of the data's lineage relationships.
[0106] Step S500: Generate data lineage tracing results based on the lineage relationship description set. The data lineage tracing results include historical records of the entire life cycle of the data and cross-data entry dependency graph information.
[0107] The data lineage tracking result is the final result obtained after in-depth analysis and organization of the data lineage relationship, which integrates the historical records of the entire life cycle of the data and the dependency graph information across data entries.
[0108] The data lifecycle history records all information from the initial data collection through a series of operations such as processing, conversion, storage, and final use. This information includes the time, location, and personnel of data collection, as well as the time, subject, and content of various operations performed during the processing process, such as cleaning, conversion, and merging. This history records can be used to trace every important step in the data lifecycle, ensuring data quality and reliability.
[0109] The cross-data item dependency graph is a visual diagram constructed using a dependency graph and a set of lineage relationship descriptions. It graphically displays the dependencies between different data items. In this graph, each data item is represented by a node, and the edges between nodes represent the dependencies between them. The attribute labels on the edges clearly indicate the type of relationship. This graph allows you to intuitively see the connections and interactions between data items, quickly locate the source and flow of data, and provide powerful support for data management and analysis.
[0110] As an implementation manner, step S500 can be specifically implemented as the following steps S510 to S550:
[0111] Step S510: Perform hierarchical division processing on the data source path in the blood relationship description set to generate a hierarchical structure description including an original data layer, an intermediate processing layer and a final output layer.
[0112] The hierarchical division process is to more clearly display the entire process of data from the original source to the final output, and divide the data source path in the blood relationship description set into different levels.
[0113] The raw data layer represents the original source of data. This data is collected directly from rescue terminals without extensive processing. For example, photos and videos of the disaster site recorded by rescuers using handheld devices, or operating parameters automatically collected by rescue equipment, all belong to the raw data layer. When performing hierarchical division, identify data entries that do not have preceding dependent nodes in the dependency graph and classify them as raw data.
[0114] The intermediate processing layer encompasses the various intermediate links that data passes through during processing. These links include operations such as cleaning, conversion, and merging raw data to generate new data entries. For example, raw relief supply inventory data is calculated and analyzed to generate usage trend data, which belongs to the intermediate processing layer. Data entries in the intermediate processing layer are often dependent on data entries in the raw data layer and further derive subsequent data entries.
[0115] The final output layer represents the final results of data processing. This data can be directly used for decision-making, analysis, or reporting. For example, rescue effectiveness assessment reports and material demand forecasts generated from various rescue data fall into the final output layer. The data entries in the final output layer depend on the data entries in the intermediate processing layers and are the final product of the entire data processing process.
[0116] After the hierarchical division is completed, the data items within each level are assigned positions and sorted. First, the vertical hierarchical arrangement is carried out, with the original data layer at the top, the intermediate processing layer in the middle, and the final output layer at the bottom. This clearly shows the data processing flow and direction. Then, within the same level, the data items are arranged from left to right according to the data generation time nodes. This makes the hierarchical structure description not only reflect the hierarchical relationship of the data, but also the chronological order of the data, making it easier to observe the data generation and processing process.
[0117] Step S511: extract the starting data entries that do not contain the preceding dependent nodes in the blood relationship description set as the original data layer nodes.
[0118] The starting data item is the starting point in the data source path. They have no dependencies on other data items and are the original source of the data. In the set of lineage descriptions, by examining the connections of each data item in the dependency graph, we identify those data items that have no incoming edges (i.e., no preceding dependent nodes).
[0119] For example, in a data governance scenario related to earthquake rescue, photos of damaged houses taken directly by rescue workers using their mobile phones at the earthquake site, without relying on any other data entries, are considered starting data entries and serve as nodes in the original data layer. By extracting these starting data entries, the scope and content of the original data layer can be determined, providing a foundation for subsequent hierarchical division and analysis.
[0120] Step S512: extracting intermediate data entries in the blood relationship description set that are only dependent on the previous original data layer nodes as intermediate processing layer nodes.
[0121] Intermediate data items are transitional data generated during data processing. They rely on nodes in the original data layer for further processing and transformation. In the lineage description set, each data item's predecessor dependency nodes are checked. If these predecessor dependency nodes all belong to original data layer nodes, then the data item can be used as an intermediate processing layer node.
[0122] For example, in the management of relief supplies, the raw data layer contains records of incoming supplies. The inventory turnover rate data calculated from these incoming supplies relies solely on these raw data layer nodes. Therefore, the inventory turnover rate data entries belong to the intermediate processing layer nodes. By extracting these intermediate data entries, we can clarify the specific content of the intermediate processing layer and understand the data processing and conversion in the intermediate links.
[0123] Step S513: extract the final data entry in the blood relationship description set that is dependent on the post-intermediate processing layer node as the final output layer node.
[0124] Final data entries are the final product of the data processing process. They rely on the data entries in the intermediate processing layer for final processing and integration. In the set of blood relationship descriptions, find those data entries whose preceding dependent nodes all belong to the intermediate processing layer nodes and use them as the final output layer nodes.
[0125] For example, based on the material usage trend data and personnel rescue efficiency data generated by the intermediate processing layer, the optimal rescue resource allocation plan data is generated through comprehensive analysis and calculation. This data entry depends on the corresponding data entries in the intermediate processing layer and belongs to the final output layer node. By extracting the final output layer node, the ultimate goal and results of data processing can be clarified, providing a basis for subsequent decision-making and application.
[0126] Step S514: Perform hierarchical position allocation processing on the original data layer nodes, the intermediate processing layer nodes and the final output layer nodes to generate a vertical hierarchical arrangement structure with the original data layer at the top, the intermediate processing layer in the middle and the final output layer at the bottom.
[0127] The hierarchical position allocation process is to clearly display the data processing flow and hierarchical relationship in a visual hierarchical structure. First, a fixed vertical position range is set for each level. The original data layer is at the top, the intermediate processing layer is in the middle, and the final output layer is at the bottom. For example, the entire vertical space can be divided into three areas, the top area is allocated to the original data layer, the middle area is allocated to the intermediate processing layer, and the bottom area is allocated to the final output layer. Then, the corresponding nodes are placed within the position range of their respective levels, so that the entire hierarchical structure presents a clear vertical arrangement, which is convenient for intuitively observing the flow of data from the original source to the final output.
[0128] Step S515: Horizontally sort the nodes in the same level, and arrange the nodes from left to right based on the data generation time of the nodes to generate a hierarchical structure description.
[0129] In the same level, in order to reflect the time sequence of the data, the nodes need to be sorted horizontally. According to the data generation time node of each node, the nodes are arranged from left to right. For example, in the original data layer, if there are multiple data entries, the nodes corresponding to the earlier generated data entries are placed on the left, and the nodes corresponding to the later generated data entries are placed on the right in the order of data generation time. In this way, in the hierarchical structure description, not only can the hierarchical relationship of the data be seen, but also the time development order of the data in the same level can be understood through horizontal arrangement, which is helpful for analyzing the generation and evolution process of the data.
[0130] Step S520: perform operation type labeling on the processing process nodes in the blood relationship description set to generate processing node labels including collection, cleaning, conversion and storage.
[0131] Processing nodes are the steps that data goes through during processing, representing different types of operations. Operation type annotation is the process of adding a corresponding operation type label to each processing node to clearly demonstrate the specific process of data processing.
[0132] The collection operation refers to the process of obtaining data from rescue terminals or other data sources. For example, rescue workers use sensors to collect ambient temperature and humidity data at the disaster site. The node corresponding to this collection process is labeled "Collect".
[0133] Cleaning is the process of cleaning and filtering raw data to remove noise, duplicate data, and erroneous data. For example, the node corresponding to the cleaning process is labeled "cleaning" when checking the collected inventory data of relief supplies and deleting duplicate records and erroneous data. Conversion is the process of converting data from one form to another by performing format conversion, calculation, or other processing. For example, converting material weight data in kilograms to data in tons, or calculating the total value based on the unit price and quantity of the materials, the nodes corresponding to these conversion processes are labeled "conversion."
[0134] The storage operation saves the processed data to a designated storage medium for subsequent use and query. For example, the node corresponding to the cleansed and converted rescue data is labeled "Store" when storing it in a database.
[0135] Step S530: Perform length statistics on the dependency chain in the blood relationship description set, and extract the longest dependency path and key dependency nodes as the core tracking objects of data blood relationship.
[0136] A dependency chain is a series of edges and nodes connecting different data items in a dependency graph. It reflects the dependencies and processing flow between data. Length statistics are used to identify the longest paths and key nodes in the dependency chain, allowing us to focus on the core parts of the data lineage relationships.
[0137] As an implementation manner, step S530 may be specifically implemented as the following steps S531 to S536:
[0138] Step S531: Perform node number statistics processing on each dependency chain in the blood relationship description set to generate a node counting result representing the length of the dependency chain.
[0139] To count the length of dependency chains, we need to count the number of nodes in each dependency chain. We find all dependency chains in the dependency graph and traverse each chain's nodes in turn, recording the number of nodes.
[0140] For example, in a dependency graph for rescue personnel dispatch data, there's a dependency chain starting with initial personnel registration data, moving through multiple steps like qualification review and task assignment, and ultimately to personnel attendance records. By traversing all nodes in this chain and counting the number of nodes, this number represents the length of the dependency chain. The node count results for all dependency chains are recorded for subsequent screening and analysis.
[0141] Step S532: Based on the node counting result, the dependency chain with the largest number of nodes is selected as the longest dependency path.
[0142] After obtaining the node count results for all dependency chains, compare these results to find the dependency chain with the largest number of nodes. This dependency chain represents the most complex and involved process in the data processing, typically encompassing multiple important steps from the original source to the final output.
[0143] For example, if one of all dependency chains contains 10 nodes, while the other chains all have fewer than 10 nodes, then this dependency chain containing 10 nodes is the longest dependency path. This longest dependency path is extracted and becomes the focus of subsequent analysis.
[0144] Step S533: Calculate the dependency of each node in the longest dependency path. The dependency indicates the number of times the node is referenced by other dependency chains.
[0145] Dependency calculations assess the importance of each node in the longest dependency path within the entire data chain. The higher a node's dependency, the more frequently it is referenced by other dependency chains, and the more critical its role in data processing.
[0146] As an implementation manner, step S533 may be specifically implemented as the following steps S5331 to S5336:
[0147] Step S5331: Count the number of times each node in the longest dependency path is referenced by other dependency chains as a direct reference count.
[0148] The direct reference count refers to the number of times each node in the longest dependency path appears in other dependency chains. Traverse all dependency chains and for each node in the longest dependency path, check whether it appears in other chains. If so, the count is increased by 1.
[0149] For example, if there is a data node representing a relief supplies distribution plan in the longest dependency path, and three other dependency chains reference this node, then the direct reference count of this node is 3. By performing this statistics on each node in the longest dependency path, the direct reference count of each node can be obtained.
[0150] Step S5332: extract the length parameters of all dependency chains containing the node, and calculate the ratio of each length parameter to the longest dependency path length as the path length influencing factor.
[0151] The path length impact factor is designed to consider the impact of the length of the dependency chain on the importance of the node. For each node in the longest dependency path, all dependency chains containing the node are found and the length parameters of these chains are extracted.
[0152] Next, the length parameter of each dependency chain is compared with the length of the longest dependency path, and their ratio is calculated. For example, if the length of the longest dependency path is 10, and the length of a dependency chain containing a node is 5, then the path length impact factor of this chain is 5 / 10 = 0.5. By calculating the path length impact factor of each dependency chain containing a node, we can comprehensively consider the impact of chains of different lengths on the importance of a node.
[0153] Step S5333: normalize the direct reference count to obtain a reference strength index, and perform weighted summation on the path length impact factor to obtain a path impact strength index.
[0154] Normalization is done to convert direct citation counts into relative values, keeping them within a uniform range. A min-max normalization method can be used to convert direct citation counts into values between 0 and 1, yielding a citation strength index.
[0155] To comprehensively consider the impact of different chains, path length influencing factors require a weighted summation. Each path length influencing factor can be assigned a weight, which can be adjusted based on the actual situation, for example, based on the importance or relevance of the dependency chain. Each path length influencing factor is then multiplied by its corresponding weight, and the results are summed to obtain the path influence strength index.
[0156] For example, there are three dependency chains containing a certain node, and the path length influence factors are 0.2, 0.3, and 0.5, respectively. The corresponding weights are 0.2, 0.3, and 0.5, respectively. Then the path influence strength index = 0.2×0.2+0.3×0.3+0.5×0.5=0.38.
[0157] Step S5334: linearly combine the reference strength index and the path influence strength index to generate a dependency value representing the comprehensive importance of the node in multiple paths.
[0158] The linear combination process combines the citation strength index and the path influence strength index to obtain a dependency value that represents the comprehensive importance of the node in multiple paths. The linear combination formula can be used. For example, the dependency value = citation strength index × weight 1 + path influence strength index × weight 2, where weight 1 and weight 2 are weights set according to actual conditions, and their sum is usually 1. For example, if weight 1 is set to 0.6, weight 2 is set to 0.4, the citation strength index is 0.8, and the path influence strength index is 0.38, then the dependency value = 0.8 × 0.6 + 0.38 × 0.4 = 0.632. Through such a linear combination, the influence of the number of direct citations of the node and the length of the dependency chain are comprehensively considered to obtain a more accurate dependency value.
[0159] Step S5335: Divide the longest dependency path into subpaths, extract the continuous subpath segments containing key dependency nodes, and calculate the variance of the dependency values of each node in the subpath segment as the node importance fluctuation parameter.
[0160] Subpath partitioning involves dividing the longest dependency path into multiple consecutive subpath segments, and then finding the subpath segments that contain key dependency nodes. Key dependency nodes are nodes with high dependency values.
[0161] For subpath segments containing key dependent nodes, calculate the variance of the dependency values for each node. Variance is a statistic that measures the degree of data dispersion. Here, variance indicates the fluctuation in node importance within a subpath segment. A larger variance indicates greater disparity in node importance; a smaller variance indicates relatively stable node importance.
[0162] For example, in a sub-path segment containing a key dependency node, there are 5 nodes, and their dependency values are 0.6, 0.7, 0.65, 0.75, and 0.68 respectively. By calculating the variance of these values, the fluctuation parameter of the node importance in the sub-path segment can be obtained.
[0163] Step S5336: Based on the dependency value and the node importance fluctuation parameter, the nodes with the highest dependency value and the lowest fluctuation parameter are selected as the core key dependency nodes. The core key dependency nodes are used to mark the key monitoring objects of data lineage tracking.
[0164] After calculating the dependency value and node importance fluctuation parameters for each node, it is necessary to screen out the core key dependency nodes. Core key dependency nodes are the most important and relatively stable nodes in the entire data lineage relationship. They play a key role in data generation and processing.
[0165] By comparing the dependency values and node importance fluctuation parameters of all nodes, we identify the nodes with the highest dependency values and lowest fluctuation parameters. For example, if one of several nodes has a dependency value of 0.8 and a node importance fluctuation parameter of 0.05, while the other nodes have lower dependency values or higher fluctuation parameters, this node is identified as a core, critical dependency node. These core, critical dependency nodes are marked as key monitoring targets for data lineage tracking, allowing us to focus on their changes and impact in subsequent data management and analysis.
[0166] Step S534: Filter out nodes whose dependency exceeds a preset threshold as key dependency nodes.
[0167] The preset threshold is a value set in advance based on actual conditions and analysis requirements to determine the importance of a node. When the dependency value of a node exceeds this threshold, the node is considered a critical dependency node.
[0168] For example, if the preset threshold is 0.7 and the calculated dependency value of a node is 0.8, exceeding the preset threshold, then the node is screened as a key dependency node. Key dependency nodes play an important role in the data processing process, and their changes may have a significant impact on the entire data lineage relationship, so they require special attention.
[0169] Step S535: performing path decomposition processing on the longest dependency path, extracting the source identification sequence, time node sequence, and operation type sequence contained in the path as core tracking path information.
[0170] Path decomposition involves breaking down the longest dependency path in detail to extract key information. The source identifier sequence records the source identifiers of the data entries corresponding to each node in the longest dependency path. This sequence allows us to understand the different rescue terminals or data sources from which the data originated during processing.
[0171] The time node sequence records the generation time of the data entry corresponding to each node in the longest dependency path, reflecting the chronological order and progress of data processing. The operation type sequence records the operation type corresponding to each node in the longest dependency path, such as collection, cleaning, and conversion. This sequence clearly shows the specific operations that the data undergoes during processing.
[0172] For example, in the longest dependency path, there are data items A, B, and C. Data item A's source identifier is "Rescue Equipment 1," its generation time is "2024-11-10 09:00:00," and its operation type is "Collect." Data item B's source identifier is "Rescue System 2," its generation time is "2024-11-10 10:00:00," and its operation type is "Convert." Data item C's source identifier is "Material Management System," its generation time is "2024-11-10 11:00:00," and its operation type is "Store." Therefore, the source identifier sequence is "Rescue Equipment 1, Rescue System 2, Material Management System," the time node sequence is "2024-11-10 09:00:00, 2024-11-10 10:00:00, 2024-11-10 11:00:00," and the operation type sequence is "Collect, Convert, Store."
[0173] Step S536: Combine and encapsulate the longest dependency path, key dependency nodes, and core tracing path information to generate a core tracing object.
[0174] The combined encapsulation process combines the extracted longest dependency path, key dependency nodes, and core tracing path information to form a complete core tracing object. This object contains key information about data lineage relationships, facilitating subsequent data lineage tracing and analysis.
[0175] This information can be stored in a data structure, such as an object or a dictionary, containing attributes such as "longest dependency path," "critical dependency node," and "core tracking path information," with the corresponding information assigned to these attributes. This way, through the core tracking object, we can fully understand the critical paths, important nodes, and detailed operation and timing information during the data processing process, providing strong support for data management and decision-making.
[0176] Step S540: Perform visual mapping processing on the hierarchical structure description, processing node labels and core tracking objects to generate a kinship visualization map including node icons, edge connection relationships and label annotations.
[0177] Visual mapping processing is to convert abstract data lineage relationship information into intuitive visual maps to more clearly display the data processing process and dependency relationships.
[0178] A hierarchical structure describes the hierarchical relationship of data from its original source to its final output. This is mapped into a visual graph, showing the data flow through different hierarchical positions and arrangements. For example, nodes for the original data layer are placed at the top of the graph, nodes for the intermediate processing layer are placed in the middle, and nodes for the final output layer are placed at the bottom, forming a clear vertical hierarchy.
[0179] Processing node labels clearly identify the type of operation performed by each node during data processing. Adding these labels to the corresponding nodes makes the node's function clear at a glance. For example, for a node labeled "Collect," the "Collect" label will appear next to it in the graph.
[0180] The core tracking object contains the longest dependency path, key dependency nodes, and core tracking path information, which is integrated into the visual map. The longest dependency path is highlighted with a special line or color to make it more eye-catching in the map; key dependency nodes are marked with different icons or colors to focus on them. For the core tracking path information, the source identification sequence, time node sequence, and operation type sequence are displayed in the map in an appropriate manner, such as adding annotations next to the path.
[0181] Additionally, icons are added to each node. These icons can be designed based on the node type or operation type, making the graph more intuitive and visual. Edges are added between nodes to represent dependencies between them. The edge style can be differentiated by the type of relationship, such as solid lines for generative relationships and dashed lines for derivative relationships. Labels are also added to each edge to clarify the specific type of relationship.
[0182] Through such visual mapping processing, a visual map of blood relationship is generated, which includes node icons, edge connection relationships and label annotations, making the data blood relationship clear at a glance and convenient for users to observe, analyze and understand.
[0183] Step S550: Perform historical version management on the blood relationship visualization map, record blood relationship change information at different time points, and generate data blood relationship tracking results.
[0184] The purpose of historical version management is to track the changes in the visual map of lineage relationships over time and record the change information at different time points in order to understand the evolution of data lineage relationships.
[0185] As an implementation manner, step S550 can be specifically implemented as the following steps S551 to S555:
[0186] Step S551: Assign a unique version identification code to the blood relationship visualization map, where the version identification code includes timestamp information.
[0187] In order to distinguish different versions of the kinship visualization map, each map needs to be assigned a unique version identification code. This code contains timestamp information, which records the specific time when the map was generated or updated. For example, the version identification code can adopt the format of "YYYYMMDDHHMMSS-serial number", where "YYYYMMDDHHMMSS" is the timestamp, indicating the year, month, day, hour, minute, and second when the map was generated or updated, and "serial number" is a number set to distinguish multiple versions that may have been generated at the same time point. In this way, the version identification code can accurately identify the version and generation time of each map.
[0188] Step S552: Perform change detection on the nodes and edges in the blood relationship visualization graph to identify change operations such as adding new nodes, deleting nodes, modifying edge connection relationships, and modifying label annotations.
[0189] Change detection identifies changes in the kinship visualization graph at different points in time. By comparing two consecutive graph versions, nodes and edges are carefully examined. Nodes are checked for additions or deletions. For example, if a new node representing a new rescue data processing step is found in a new version, this is a new node. If a node representing old collected data no longer exists in the new version, this is a deleted node.
[0190] For edges, check whether the edge connection has changed. For example, if an edge originally connecting two nodes is removed, or a new edge is added between two nodes, these are considered modifications to the edge connection. Also, check whether the label annotation has been modified, such as changing the node operation type label from "collect" to "transform" or changing the edge relationship type label from "generate" to "derive".
[0191] As an implementation manner, step S552 can be specifically implemented as the following steps S5521 to S5528:
[0192] Step S5521: Generate a unique feature hash value for each node of the blood relationship visualization map. The feature hash value is calculated by combining the source identifier of the node, the generation time node and the associated object identifier.
[0193] The characteristic hash value is a code used to uniquely identify a node. It is calculated by combining the node's source identifier, generation time node, and associated object identifier.
[0194] The source identifier indicates the source of the data entry corresponding to the node, the generation time node records the time when the data entry was generated, and the associated object identifier clarifies the relationship between the node and other objects. This information is combined together and calculated using a hash algorithm to obtain a unique characteristic hash value.
[0195] For example, for a node, its source identifier is "rescue equipment A," its generation time is "2024-11-15 12:00:00," and its associated object identifier is "disaster event 001." This information is combined into a string in a certain order, and then a hash algorithm (such as MD5 or SHA-256) is used to calculate a characteristic hash value. This way, each node has a unique characteristic hash value, making it easier to accurately identify the node during change detection.
[0196] Step S5522: Generate a unique relationship hash value for each edge of the blood relationship visualization map. The relationship hash value is calculated by combining the hash value of the starting node, the hash value of the ending node, and the relationship type label of the edge connection.
[0197] The relationship hash value is used to uniquely identify the edge and is calculated by combining the hash value of the starting node and the end node connected by the edge and the relationship type label.
[0198] The starting node hash value and the ending node hash value are the characteristic hash values of the two nodes connected by the edge, respectively. The relationship type label specifies the type of relationship represented by the edge, such as "generated" or "derived." This information is combined and calculated using a hash algorithm to obtain the relationship hash value.
[0199] For example, consider an edge connection with a hash value of "abc123" for its starting node and "def456" for its ending node, along with a relationship type label of "Generate." This information is combined into a string and then calculated using a hash algorithm to generate the relationship hash value. This relationship hash value accurately identifies each edge and facilitates detection of edge changes.
[0200] Step S5523: Obtain the node hash set and edge hash set of the current version of the graph, and obtain the node hash set and edge hash set of the previous version of the graph.
[0201] The node hash set is the set of feature hash values of all nodes in the current version or the previous version of the graph, and the edge hash set is the set of relationship hash values of all edges.
[0202] During change detection, the node hash sets and edge hash sets of the current and previous versions of the graph are obtained. By comparing these two sets, changes to nodes and edges can be identified. For example, the node hash sets of the current version can be compared with the node hash sets of the previous version to see if new hash values appear (indicating added nodes) or if certain hash values disappear (indicating deleted nodes).
[0203] Step S5524: Identify the newly added node hash value and the deleted node hash value by calculating the symmetric difference between the two version node hash sets. The newly added node hash value corresponds to the newly added node, and the deleted node hash value corresponds to the deleted node.
[0204] A symmetric difference is a set consisting of elements that belong to only one of the two sets. By calculating the symmetric difference between the hash sets of the nodes in the current version and the previous version, you can find newly added and deleted nodes.
[0205] If a hash value exists only in the node hash set of the current version but not in the previous version, then the node corresponding to this hash value is a newly added node. Conversely, if a hash value exists only in the node hash set of the previous version but not in the current version, then the node corresponding to this hash value is a deleted node. In this way, the addition and deletion of nodes can be accurately identified.
[0206] Step S5525: Identify the newly added edge hash values and the deleted edge hash values by calculating the symmetric difference between the two versions of the edge hash sets. The newly added edge hash values correspond to newly added edge connection relationships, and the deleted edge hash values correspond to deleted edge connection relationships.
[0207] Similar to the processing of node hash sets, the symmetric difference between the edge hash sets of the current version and the previous version is calculated to find the newly added edges and the deleted edges.
[0208] If a relationship hash value only exists in the edge hash set of the current version but not in the previous version, then the edge corresponding to this hash value is a newly added edge, representing the addition of an edge connection relationship; if a relationship hash value only exists in the edge hash set of the previous version but not in the current version, then the edge corresponding to this hash value is a deleted edge, representing the deletion of an edge connection relationship.
[0209] Step S5526: Perform feature content comparison processing on the retained node hash value, identify the modification operation of the node label annotation, and generate a node modification record including the node identifier, the content before modification, and the content after modification.
[0210] For nodes that exist in both versions (i.e., nodes with the same node hash value), it is necessary to further check whether their label annotations have been modified.
[0211] Perform a feature content comparison on the nodes corresponding to the retained node hash values, comparing the node's label annotations between the two versions. If a change is found in the label annotation, record the node's ID (which can be represented by the node's feature hash value), the content before the change, and the content after the change, generating a node modification record. For example, if a node's operation type label changes from "collect" to "convert," record the node ID, the content before the change (collection), and the content after the change (convert).
[0212] Step S5527: perform relationship type label comparison processing on the retained edge hash value, identify the modification operation of the edge connection relationship, and generate an edge modification record including the edge start node, edge end node, pre-modification label and post-modification label.
[0213] For edges that exist in both versions (i.e., edges with the same edge hash value), check whether their relationship type labels have been modified.
[0214] Compare the relationship type labels of the edges corresponding to the retained edge hash values, comparing the edge relationship type labels in the two versions. If a label change is found, record the edge's start node, end node, label before the change, and label after the change, generating an edge modification record. For example, if an edge's relationship type label changes from "generated" to "derived," record the edge's start node, end node, label before the change, and label after the change.
[0215] Step S5528: Summarize the newly added nodes, deleted nodes, newly added edge connection relationships, deleted edge connection relationships, node modification records, and edge modification records to generate change operation information.
[0216] All identified change information is summarized, including newly added nodes, deleted nodes, newly added edge connection relationships, deleted edge connection relationships, node modification records, and edge modification records, to form complete change operation information.
[0217] For example, information about all newly added nodes is organized into one list, information about deleted nodes into another, and newly added edge connections, deleted edge connections, node modification records, and edge modification records are organized into separate lists. These lists are then combined to form change operation information. This information provides a comprehensive understanding of the specific changes in the kinship visualization map between two versions.
[0218] Step S553: Record the timestamp, operation type and impact range information of each change operation, and generate a blood relationship change log.
[0219] The lineage change log is used to record detailed information about each change operation, including the timestamp of the change, the type of operation (such as adding a node, deleting a node, modifying a label, etc.), and the scope of the change.
[0220] The timestamp records the specific time when the change occurred, the operation type clarifies the specific content of the change, and the scope of impact describes the extent of the change's impact on the entire graph and the nodes and edges involved. For example, an operation that adds a new node would have a timestamp of "2024-11-16 13:00:00," an operation type of "add node," and a scope of impact of the node's level and related edge connections.
[0221] By generating a blood relationship change log, you can track historical changes in the graph and provide detailed records for subsequent analysis and management.
[0222] Step S554: associate and store the version identification code, the blood relationship visualization map, and the blood relationship change log to generate a set of historical versions arranged in chronological order.
[0223] Associative storage involves associating the three pieces of information—version identification codes, the visual lineage graph, and the lineage change log—and storing them in a suitable storage system to form a chronologically ordered set of historical versions. This facilitates the management, query, and tracing of the visual lineage graph and its changes at different points in time.
[0224] The version identifier is the key to identifying each version. It includes timestamp information, clearly indicating when the graph was generated or updated. When storing, each version identifier is bound to the corresponding kinship visualization graph and kinship change log. For example, when a new version is generated, the version identifier is stored along with the current kinship visualization graph and the change log generated by the change operation.
[0225] The visual lineage graph intuitively displays the lineage relationships of data and is the core content of the entire historical version collection. The graphs of different versions reflect the evolution of the data processing process over time. During storage, ensure that each graph is accurately associated with the corresponding version identifier code, so that the corresponding graph can be quickly found based on the version identifier.
[0226] The lineage change log records the details of each change operation, including the timestamp, operation type, and impact scope. It provides a detailed basis for understanding changes in the graph. By associating the change log with the corresponding version identifier and the graph, when viewing a particular version of the graph, one can simultaneously understand the specific changes compared to the previous version.
[0227] By associating and storing these three pieces of information and sorting them by the timestamp information in the version identifier, an ordered set of historical versions is generated. This way, when you need to view a visual map of lineage relationships and their changes at a specific point in time, you can easily find the corresponding version in the set.
[0228] Step S555: Perform version backtracking on the historical version set, support querying the visual map of the blood relationship and the corresponding change log at a specified time point through timestamp, and generate data blood relationship tracking results.
[0229] Version backtracking processing uses the generated historical version set to realize the function of querying the visual map of lineage relationships and the corresponding change log at a specified time point based on the timestamp, thereby ultimately generating a complete data lineage tracking result.
[0230] In the historical version collection, each version has a unique version identification code that includes timestamp information. When a user needs to query the lineage relationship information at a specific point in time, the system will search the historical version collection based on the input timestamp.
[0231] First, the version identifier closest to the specified timestamp is selected from the collection. If the specified timestamp matches the timestamp of a version, the corresponding lineage visualization map and change log are directly retrieved. If there is no exact timestamp match, the version closest to the specified time that is earlier than the specified time is selected.
[0232] For example, if a user wants to query the blood relationship information at 10:00:00 on November 20, 2024, and the historical version set contains two versions, 09:30:00 on November 20, 2024 and 10:30:00 on November 20, 2024, the version at 09:30:00 on November 20, 2024 will be selected.
[0233] After finding the corresponding version, the system extracts a visual map of the lineage relationships of that version, displaying the lineage relationships of the data at that point in time. It also obtains the lineage change log corresponding to that version, which records all changes from the previous version to the current version, including details of operations such as adding nodes, deleting nodes, modifying edge connections, and modifying label annotations.
[0234] This version backtracking process allows users to clearly understand the state and changes of data lineage relationships at different points in time, comprehensively grasp the processing and dependencies of data throughout its entire lifecycle, from initial acquisition to final output, and generate complete data lineage tracking results. This result, which includes historical records of the data's entire lifecycle and a cross-entry dependency graph, provides strong support for data management, decision-making, and troubleshooting in emergency command and rescue scenarios.
[0235] In actual emergency command and rescue scenarios, data lineage tracking results can help rescue commanders better understand the data's source and processing process. For example, when an anomaly is discovered in a piece of rescue data, version backtracking can be used to query the data's status and changes at different time points, identify the operational links that may have caused the anomaly, and take timely corrective measures. At the same time, for the deployment and use of rescue resources, data lineage tracking can clearly show the flow and usage of resources, providing a basis for subsequent resource management and optimization. Furthermore, when evaluating rescue effectiveness, data lineage tracking results can provide accurate information about the data processing process, making the evaluation results more reliable and convincing.
[0236] To sum up, the multi-source data governance method based on machine learning provided by the embodiment of the present invention processes and analyzes multi-source data through a series of steps, and finally generates comprehensive and accurate data lineage tracking results, which provides effective support for data management and decision-making in emergency command and rescue scenarios, and helps to improve the efficiency and quality of rescue work.
[0237] It is understandable that the various algorithms involved in the above-mentioned introductions of the embodiments of the present invention, such as the Euclidean distance algorithm, the cosine distance algorithm, the hash algorithm, the interpolation algorithm, etc., can all be learned from the relevant content in the prior art. In order to save space, they will not be expanded too much in the embodiments of the present invention. In addition, when implementing the scheme of the present invention, those skilled in the art can supplement the details according to the common knowledge in this field. For example, according to the common knowledge in this field, normalization can be used to eliminate dimensional conflicts before feature fusion, interpolation can be used to eliminate dimensional differences, and thresholds can be reasonably set based on historical data, experience or business scenario requirements. The model can be trained based on a general model training method, and the number of layers in the model structure can be set based on actual needs, the activation function can be selected, etc. The present invention will no longer provide redundant introductions to the overly detailed implementation process.
[0238] Based on the foregoing embodiments, an embodiment of the present invention provides a multi-source data governance device, and the various units included in the device, as well as the various modules included in each unit, can be implemented by a processor in a computer device; of course, they can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.
[0239] Figure 2 A schematic diagram of the structure of a multi-source data management device provided by an embodiment of the present invention is shown in FIG. Figure 2 As shown, the multi-source data management device 200 includes:
[0240] The data acquisition module 210 is used to acquire a multi-source data set in an emergency command and rescue scenario, wherein the multi-source data set includes raw data units from different rescue terminals, and the raw data units have a mixed data form of unstructured text, semi-structured logs, and structured tables;
[0241] A data preprocessing module 220 is configured to perform data preprocessing on the multi-source data set to obtain a standardized data set having a unified data form, wherein the standardized data set includes structured data entries aligned with timestamps;
[0242] Meta-information extraction module 230, used to extract a data meta-information set from the standardized data set, wherein the data meta-information set includes a data source identifier, a data generation time node, a data operation record, and a data-related object identifier;
[0243] The bloodline tracking module 240 is used to call a pre-trained bloodline relationship modeling model to perform association analysis on the data metadata information set to generate a bloodline relationship description set between each data item in the standardized data set, wherein the bloodline relationship description set includes a data source path, a processing process node, and a dependency chain;
[0244] The result generation module 250 is used to generate a data lineage tracing result based on the lineage relationship description set, wherein the data lineage tracing result includes historical records of the entire life cycle of the data and dependency graph information across data entries.
[0245] The description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. In some embodiments, the functions or modules provided by the device provided by the embodiment of the present invention can be used to execute the method described in the above method embodiment. For technical details not disclosed in the device embodiment of the present invention, please refer to the description of the method embodiment of the present invention for understanding. It should be noted that in the embodiment of the present invention, if the above multi-source data governance method based on machine learning is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention or the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present invention is not limited to any specific hardware, software or firmware, or any combination of hardware, software and firmware.
[0246] Figure 3 A hardware entity diagram of a computer system provided by an embodiment of the present invention is as follows Figure 3As shown, the hardware entity of the computer system 1000 includes: a processor 1001 and a memory 1002, wherein the memory 1002 stores a computer program that can be run on the processor 1001, and when the processor 1001 executes the program, the steps in the method of any of the above embodiments are implemented.
[0247] The memory 1002 stores computer programs that can be run on the processor. The memory 1002 is configured to store instructions and applications executable by the processor 1001. It can also cache data to be processed or processed by the processor 1001 and various modules in the computer system 1000 (for example, image data, audio data, voice communication data, and video communication data). It can be implemented through flash memory (FLASH) or random access memory (RAM).
[0248] When the processor 1001 executes the program, the steps of any of the above-mentioned multi-source data governance methods based on machine learning are implemented. The processor 1001 generally controls the overall operation of the computer system 1000.
[0249] The above description is only an embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A multi-source data governance method based on machine learning, characterized in that: The method comprises: Acquire a multi-source data set in an emergency command and rescue scenario, wherein the multi-source data set includes raw data units from different rescue terminals, and the raw data units have a mixed data form of unstructured text, semi-structured logs, and structured tables; Performing data preprocessing on the multi-source data set to obtain a standardized data set with a unified data form, wherein the standardized data set includes structured data entries with aligned timestamps; Extracting a data meta information set from the standardized data set, wherein the data meta information set includes a data source identifier, a data generation time node, a data operation record, and a data-related object identifier; Inputting the data metadata set into the feature embedding layer of the blood relationship modeling model, performing feature vectorization processing on the data source identifier, data generation time node, data operation record and data associated object identifier to generate a metadata feature vector with a unified dimensional representation; Performing time series dependency analysis on the meta-information feature vectors through the time series association module of the blood relationship modeling model to identify generation and derivation relationships between data entries at different time nodes; Performing cross-analysis processing on the source identification and the associated object identification of the meta-information feature vector through the spatial association module of the blood relationship modeling model to identify the interactive relationship between different source data items and the same associated object; Based on the generated derivative relationship output by the temporal association module and the interactive relationship output by the spatial association module, construct a dependency graph between the data items, wherein the dependency graph includes edges representing the relationship between the data items and attribute labels representing the relationship type; Performing path traversal processing on the dependency graph, extracting complete path information from the initial data entry to the subsequent derived data entry, and generating a lineage description set, wherein the lineage description set includes a data source path, a processing node, and a dependency chain; A data lineage tracing result is generated based on the lineage relationship description set, and the data lineage tracing result includes historical records of the entire life cycle of the data and cross-data entry dependency graph information.
2. The multi-source data governance method based on machine learning according to claim 1 is characterized in that: The time series dependency analysis and processing of the meta-information feature vector by the time series association module of the blood relationship modeling model to identify the generation and derivation relationships between data entries at different time nodes includes: Sorting the data generation time nodes in the meta-information feature vector to generate a meta-information feature sequence arranged in chronological order; Extracting meta-information feature vectors of adjacent time nodes in the meta-information feature sequence, and calculating cosine similarity values between the feature vectors as a correlation index of temporally adjacent data entries; Performing threshold screening on the correlation index, and retaining temporally adjacent data item pairs whose correlation index exceeds a preset threshold; Identifying and processing the modification operation types in the data operation records, and extracting time nodes corresponding to modification operation records including data copying, conversion, and merging; Based on the time nodes corresponding to the modification operation records, the original data entries of the previous time node and the modified data entries of the next time node are associated and matched to identify data generation and derivation relationships.
3. The multi-source data governance method based on machine learning according to claim 1 is characterized in that: The spatial association module of the blood relationship modeling model performs cross-analysis processing on the meta-information feature vector with respect to the source identifier and the associated object identifier to identify the interactive relationship between different source data items and the same associated object, including: performing grouping processing on the data source identifiers in the meta-information feature vector to generate a plurality of data entry groups divided by the source identifiers; Aggregating the data-related object identifiers in the meta-information feature vector to generate a plurality of related object groups divided according to the related object identifiers; Extracting meta-information feature vectors of different source data entry groups contained in each associated object group, and calculating the feature overlap of different source data entry groups in the same associated object group; Identifying collaborative record relationships between data items from different sources on the same associated object based on the feature overlap; Identify the transmission operation type in the data operation record, and extract the source identification pair and the associated object identification corresponding to the operation record containing the cross-source data transmission; Based on the source identification pair and the associated object identification, the original data entry of the sending source and the target data entry of the receiving source are associated and matched to identify the data interaction relationship.
4. The multi-source data governance method based on machine learning according to claim 1 is characterized in that: The generating of the data lineage tracing result based on the lineage relationship description set includes: Performing hierarchical division processing on the data source path in the blood relationship description set to generate a hierarchical structure description including an original data layer, an intermediate processing layer, and a final output layer; Performing operation type labeling on the processing process nodes in the blood relationship description set to generate processing node labels including collection, cleaning, conversion and storage; Perform length statistics on the dependency chains in the lineage description set, and extract the longest dependency path and key dependency nodes as core tracking objects of data lineage; Performing visual mapping processing on the hierarchical structure description, processing node labels and core tracking objects to generate a kinship visualization map including node icons, edge connection relationships and label annotations; Performing historical version management on the visual atlas of blood relationship, recording blood relationship change information at different time points, and generating the data blood relationship tracking result.
5. The multi-source data governance method based on machine learning according to claim 4 is characterized in that: The hierarchical division processing of the data source path in the blood relationship description set to generate a hierarchical structure description including an original data layer, an intermediate processing layer and a final output layer includes: Extracting the starting data entry that does not contain the preceding dependent node in the blood relationship description set as the original data layer node; Extracting intermediate data entries in the blood relationship description set that are only dependent on the previous original data layer nodes as intermediate processing layer nodes; Extracting the final data entry in the blood relationship description set that is dependent on the post-intermediate processing layer node as the final output layer node; Performing hierarchical position allocation processing on the original data layer nodes, the intermediate processing layer nodes, and the final output layer nodes to generate a vertical hierarchical arrangement structure with the original data layer at the top, the intermediate processing layer in the middle, and the final output layer at the bottom; The nodes in the same level are sorted horizontally, and the nodes are arranged from left to right based on the data generation time of the nodes to generate the hierarchical structure description.
6. The multi-source data governance method based on machine learning according to claim 4 is characterized in that: The statistical processing of the length of the dependency chain in the blood relationship description set and the extraction of the longest dependency path and key dependency nodes as the core tracking objects of data blood relationship include: Performing node number counting processing on each dependency chain in the blood relationship description set to generate a node counting result representing the length of the dependency chain; Based on the node counting result, the dependency chain with the largest number of nodes is selected as the longest dependency path; Performing dependency calculation processing on each node in the longest dependency path, where the dependency indicates the number of times the node is referenced by other dependency chains; Nodes whose dependency exceeds the preset threshold are selected as key dependency nodes; Performing path decomposition processing on the longest dependency path, extracting the source identification sequence, time node sequence, and operation type sequence contained in the path as core tracking path information; The longest dependency path, key dependency nodes and core tracking path information are combined and packaged to generate the core tracking object.
7. The multi-source data governance method based on machine learning according to claim 4 is characterized in that: The performing of historical version management on the visualized blood relationship map, recording blood relationship change information at different time points, and generating the data blood relationship tracking result includes: Allocating a unique version identification code to the kinship visualization map, wherein the version identification code includes timestamp information; Performing change detection on the nodes and edges in the kinship visualization graph to identify change operations such as adding new nodes, deleting nodes, modifying edge connection relationships, and modifying label annotations; Record the timestamp, operation type, and impact range of each change operation, and generate a lineage change log; The version identification code, the blood relationship visualization map and the blood relationship change log are associated and stored to generate a historical version set arranged in chronological order; The historical version set is subjected to version backtracking processing, and a visual map of lineage relationships and corresponding change logs at a specified time point can be queried through timestamps to generate the data lineage tracking result.
8. The multi-source data governance method based on machine learning according to claim 1 is characterized in that: The step of performing data preprocessing on the multi-source data set to obtain a standardized data set with a unified data form includes: Performing text parsing on the unstructured text data units in the multi-source data set to extract text segments containing key event descriptions, resource scheduling instructions, and on-site feedback information; Performing log format parsing processing on the semi-structured log data units in the multi-source data set, and extracting log fields including timestamps, operation types, and execution results; Performing field consistency verification on the structured table data units in the multi-source data set to screen out table entries whose field names and data types meet preset specifications; Inputting the parsed text fragments, the extracted log fields, and the filtered table entries into a data format conversion module to generate an intermediate data set with unified field names, data types, and storage formats; The intermediate data set is subjected to timestamp alignment processing, and based on the original timestamp information of each data unit, time-aligned data entries with uniform time intervals are generated by a time series interpolation algorithm to obtain the standardized data set.
9. A multi-source data management device, characterized in that: The device comprises: A data acquisition module is used to acquire a multi-source data set in an emergency command and rescue scenario, wherein the multi-source data set includes raw data units from different rescue terminals, and the raw data units have a mixed data form of unstructured text, semi-structured logs, and structured tables; a data preprocessing module, configured to perform data preprocessing on the multi-source data set to obtain a standardized data set having a unified data form, wherein the standardized data set includes structured data entries with aligned timestamps; A metadata extraction module is used to extract a data metadata set from the standardized data set, wherein the data metadata set includes a data source identifier, a data generation time node, a data operation record, and a data-related object identifier; A lineage tracking module is used to input the data meta-information set into the feature embedding layer of the lineage relationship modeling model, perform feature vectorization processing on the data source identifier, data generation time node, data operation record and data associated object identifier, and generate a meta-information feature vector with a unified dimensional representation; perform time series dependency analysis processing on the meta-information feature vector through the temporal association module of the lineage relationship modeling model, and identify the generation and derivation relationship between data entries at different time nodes; perform cross-analysis processing on the meta-information feature vector of the source identifier and the associated object identifier through the spatial association module of the lineage relationship modeling model, and identify the interaction relationship between data entries from different sources and the same associated object; based on the generation and derivation relationship output by the temporal association module and the interaction relationship output by the spatial association module, construct a dependency graph between data entries, the dependency graph including edges representing the relationship between data entries and attribute labels representing the relationship type; perform path traversal processing on the dependency graph, extract complete path information from the initial data entry to the subsequent derived data entry, and generate a lineage description set, the lineage description set including the data source path, processing process node and dependency chain; A result generation module is used to generate data lineage tracing results based on the lineage relationship description set, wherein the data lineage tracing results include historical records of the entire life cycle of the data and cross-data entry dependency graph information.
Citation Information
Patent Citations
Multi-dimensional data blood relationship tracking method and system
CN119807190A
Data Visibility and Quality Management Platform
US20240281419A1