A method for synchronous identification of industrial data identifiers during lineage update
By assigning unique data tags to data sources, establishing a bloodline relationship map, and monitoring and synchronously updating data identifiers in real time, the update confusion problem caused by inconsistent data identifiers is solved, and efficient and accurate data bloodline updates are achieved.
Patent Information
- Application Number
- CN202410460160.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-17
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2044-04-17
AI Technical Summary
During the data lineage update process, inconsistent data identifiers lead to update confusion and low efficiency.
When the data source is first connected to the lineage update system, a unique data tag is assigned to each piece of data, the association between the data identifier and the data is recorded, a lineage relationship map is established, data source changes are monitored in real time, the data identifier synchronization recognition process is triggered, the target data identifier is determined according to the change type and map for synchronization update, and the consistency of the latest data status is verified.
Ensure that the data lineage update process is clear, the update effect is good, and the efficiency is high, so as to achieve the accuracy and consistency of data throughout its life cycle.
Smart Images

Figure CN118349557B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method for synchronously identifying industrial data identifiers during a lineage update process. Background Art
[0002] Data lineage is the process in the big data industry of generating new data sets through operations such as fusion, union, conversion, and transformation of data warehouse data sets. This creates a chain of connections between data, naturally forming upstream and downstream dependencies. When using big data technologies for data processing, data backtracking and impact analysis are often encountered. This requires timely and accurate data lineage that represents the chain relationships in the data production process.
[0003] Currently, when data lineage is updated or converted, inconsistent data identifiers are prone to occur, resulting in a chaotic data lineage update process, poor update results, and low update efficiency. Summary of the Invention
[0004] The present invention provides a method for synchronously identifying industrial data identifiers during a lineage update process, in order to solve the problems raised in the background technology.
[0005] The present invention provides a method for synchronously identifying industrial data identifiers during a lineage update process, comprising:
[0006] S1: When a data source is first connected to the lineage update system, a unique data tag is assigned to each piece of data in the data source, and the association between the data tag and the data is recorded;
[0007] S2: By analyzing the flow of all data, a blood relationship map between all data is established, and based on the association between data identifiers and data, the dependency and transmission relationship between data identifiers is determined;
[0008] S3: Monitors changes in data sources in real time. When changes are detected, the data tag synchronization process is triggered to identify the changed data.
[0009] S4: Based on the data change type and blood relationship map of the data source, combined with the data identifier recognition result, determine the target data identifier that needs to be synchronized, and synchronize and update the target data identifier;
[0010] S5: Based on the synchronization update result, the latest data identifier and the latest data status of the changed data are obtained, and the consistency of the latest data identifier and the latest data status is verified.
[0011] Preferably, in said S1, when a data source is first connected to the lineage update system, a unique data tag is assigned to each piece of data in the data source, and the association relationship between the data tag and the data is recorded, including:
[0012] Obtain the data type and data status of each piece of data in the data source, and obtain the data characteristics of each piece of data;
[0013] Based on the data type and data state, an initial identifier is matched for the corresponding data, and based on the data characteristics, a unique identification mark is performed on the basis of the initial identifier to obtain a data identifier;
[0014] An association relationship between the data identifier and the data is established, and the association relationship is stored.
[0015] Preferably, in S2, by analyzing the flow process of all data, a blood relationship map between all data is established, including:
[0016] Obtaining data identifiers of all data, determining data-related information between the data based on the data identifiers, and establishing a first blood relationship between the data based on the data-related information;
[0017] Obtain the data flow direction of the data from the flow process of all data, and establish a data flow structure diagram based on the data flow direction of all data, and based on the flow direction and flow node characteristics in the data flow structure diagram;
[0018] Based on the flow direction and flow node characteristics, the data flow structure diagram is split, and a data flow identifier is set for each branch from the split result;
[0019] Based on the data flow identifier, determining the data flow pattern of each piece of data, and determining the flow similarity of the data flow patterns between all the data;
[0020] Based on the flow similarity, a second blood relationship between the data is established;
[0021] Based on the first blood relationship, verify the second blood relationship to determine whether the second blood relationship is established on the basis of the first blood relationship;
[0022] If yes, establishing a blood relationship map between the data based on the second blood relationship;
[0023] Otherwise, based on the first blood relationship, the second blood relationship is adjusted, and a blood relationship map between the data is established based on the adjusted second blood relationship.
[0024] Preferably, adjusting the second blood relationship based on the first blood relationship includes:
[0025] Obtaining the blood relationship to be adjusted that conflicts with the first blood relationship in the second blood relationship, and dividing the blood relationship to be adjusted into a front-end blood relationship and a back-end blood relationship based on the relationship characteristics with the first blood relationship;
[0026] Adjust the front-segment blood relationship based on the first blood relationship to obtain the target front-segment blood relationship; adjust the back-segment blood relationship based on the target front-segment blood relationship to obtain the target back-segment blood relationship;
[0027] Based on the target's front-end blood relationship and the target's back-end blood relationship, an adjusted second blood relationship is obtained.
[0028] Preferably, in S2, based on the association relationship between the data identifiers and the data, determining the dependency and transfer relationship between the data identifiers includes:
[0029] Based on the blood relationship map, determine the flow relationship between data;
[0030] Based on the association relationship between data identifiers and data, combined with the flow relationship, the dependency and transmission relationship between data identifiers are determined.
[0031] Preferably, in said S3, the changes of the data source are monitored in real time, and when a change in the data source is detected, a data tag synchronization identification process is triggered to perform data identification on the changed data, including:
[0032] Monitor the current identification value of the data source in real time and compare the current identification value with the stored identification value. If there is a mismatch, it is determined that the data source has changed.
[0033] When a change in the data source is detected, the data tag synchronization recognition process is triggered to perform data identification and recognition on the changed data.
[0034] Preferably, in S4, according to the data change type and blood relationship map of the data source, combined with the data identifier recognition result, the target data identifier to be synchronized is determined, and the target data identifier is synchronously updated, including:
[0035] Based on the data change type of the data source, the type conversion characteristics are determined, and based on the blood relationship map, the state change characteristics are determined;
[0036] Determine the target data identifier of the data whose data source has changed from the data identifier identification result;
[0037] Based on the type conversion characteristics and state change characteristics, the target data identifier is synchronously updated.
[0038] Preferably, the synchronously updating the target data identifier based on the type conversion feature and the state change feature includes:
[0039] Parsing the target data identifier, and dividing the target data identifier into a plurality of identification fields according to the parsing result;
[0040] Acquire a first identification field related to the type conversion feature, and replace the first identification field based on the type conversion feature to obtain a first target identification field;
[0041] Acquire a second identification field related to the state change feature, and replace the second identification field based on the state change feature to obtain a second target identification field;
[0042] An initial identification field is formed based on the first target identification field, the second target identification field, and the unchanged identification field;
[0043] Based on the identifier display constraint rules, the initial identifier field is subjected to standardization verification. If the standardization verification fails, the identifier synchronization update fails, and identifier synchronization is performed again.
[0044] If the standard verification passes, the uniqueness verification of the initial identification field of the data is performed based on the uniqueness feature of the identification. If the uniqueness verification fails, a unique field is added to the initial identification field based on the identification construction rule to achieve synchronous update of the target data identification.
[0045] If the uniqueness verification is passed, the target data identifier is synchronously updated based on the initial identification field.
[0046] Preferably, in S5, based on the synchronous update result, obtaining the latest data identifier and the latest data status of the changed data, and verifying the consistency of the latest data identifier and the latest data status, includes:
[0047] Get the latest data identifier of the changed data from the synchronization update result;
[0048] Obtain the latest data status from the results of the data flow process;
[0049] Obtaining a state identifier from the latest data identifier, matching the state identifier with the latest data state to obtain a matching degree;
[0050] Determining whether the matching degree is greater than a preset matching degree;
[0051] If so, confirm that the latest data identifier is consistent with the latest data status;
[0052] Otherwise, it is determined that the latest data identifier is inconsistent with the latest data status.
[0053] Preferably, matching the state identifier with the latest data state to obtain a matching degree includes:
[0054] Based on the preset identification conversion rules, determine the actual standard status identification corresponding to the latest data status;
[0055] Based on a preset identifier conversion rule, the state identifier is standardized to obtain a marked standard state identifier;
[0056] The actual standard state identifier is matched with the marked standard state identifier to obtain a matching degree.
[0057] Compared with the prior art, the present invention has achieved the following beneficial effects:
[0058] By assigning a unique data tag to each piece of data in the data source when it is first connected to the lineage update system, and recording the association between the data identifier and the data, an identification basis is provided for the synchronization of industrial data identifiers. By analyzing the flow process of all data, a lineage relationship map between all data is established, and based on the association between the data identifier and the data, the dependency and transmission relationship between the data identifiers are determined, the changes in the data source are monitored in real time, and when a change in the data source is detected, the data tag synchronization identification process is triggered to identify the data identifier of the changed data. According to the data change type and lineage relationship map of the data source, combined with the data identifier identification results, the target data identifier that needs to be synchronized is determined, and the target data identifier is synchronized and updated, realizing real-time monitoring of data changes and synchronous update of data identifiers. Based on the synchronization update results, the latest data identifier and the latest data status of the changed data are obtained, and the consistency of the latest data identifier and the latest data status is verified to ensure the accuracy and consistency of the data throughout its life cycle, thereby ensuring a clear lineage update process, good update effect, and high update efficiency.
[0059] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures specifically pointed out in this application document.
[0060] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0062] Figure 1Flowchart of a method for synchronously identifying industrial data identifiers during a lineage update process according to an embodiment of the present invention;
[0063] Figure 2 This is a flow chart of recording the association relationship between data identifiers and data in an embodiment of the present invention;
[0064] Figure 3 The flowchart of performing data identification and recognition on changed data in an embodiment of the present invention. DETAILED DESCRIPTION
[0065] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0066] Example 1:
[0067] The embodiment of the present invention provides a method for synchronously identifying industrial data identifiers during lineage update. Figure 1 Shown, including:
[0068] S1: When a data source is first connected to the lineage update system, a unique data tag is assigned to each piece of data in the data source, and the association between the data tag and the data is recorded;
[0069] S2: By analyzing the flow of all data, a blood relationship map between all data is established, and based on the association between data identifiers and data, the dependency and transmission relationship between data identifiers is determined;
[0070] S3: Monitors changes in data sources in real time. When changes are detected, the data tag synchronization process is triggered to identify the changed data.
[0071] S4: Based on the data change type and blood relationship map of the data source, combined with the data identifier recognition result, determine the target data identifier that needs to be synchronized, and synchronize and update the target data identifier;
[0072] S5: Based on the synchronization update result, the latest data identifier and the latest data status of the changed data are obtained, and the consistency of the latest data identifier and the latest data status is verified.
[0073] In this embodiment, the association relationship between the data identifier and the data is unique.
[0074] In this embodiment, the flow process of all data includes the direction of data and data conversion, etc.
[0075] In this embodiment, the dependency and transfer relationship between data identifiers is determined according to the relationship between data.
[0076] The beneficial effects of the above design scheme are: by assigning a unique data tag to each piece of data in the data source when the data source is first connected to the lineage update system, and recording the association between the data identifier and the data, providing an identification basis for the synchronization of industrial data identifiers, by analyzing the flow process of all data, establishing a lineage relationship map between all data, and based on the association between the data identifier and the data, determining the dependency and transmission relationship between the data identifiers, monitoring the changes in the data source in real time, and when the data source is detected to have changed, triggering the data tag synchronization identification process, performing data identification identification on the changed data, according to the data change type and lineage relationship map of the data source, combined with the data identifier identification results, determining the target data identifier that needs to be synchronized, and synchronizing the target data identifier to achieve real-time monitoring of data changes, synchronously updating data identifiers, based on the synchronization update results, obtaining the latest data identifier and the latest data status of the changed data, and verifying the consistency of the latest data identifier and the latest data status, ensuring the accuracy and consistency of the data throughout its life cycle, thereby ensuring a clear lineage update process, good update effect, and high update efficiency.
[0077] Example 2:
[0078] Based on Example 1, the present invention provides a method for synchronously identifying industrial data identifiers during lineage update, such as Figure 2 As shown, in S1, when a data source is first connected to the lineage update system, a unique data tag is assigned to each piece of data in the data source, and the association relationship between the data tag and the data is recorded, including:
[0079] Obtain the data type and data status of each piece of data in the data source, and obtain the data characteristics of each piece of data;
[0080] Based on the data type and data state, an initial identifier is matched for the corresponding data, and based on the data characteristics, a unique identification mark is performed on the basis of the initial identifier to obtain a data identifier;
[0081] An association relationship between the data identifier and the data is established, and the association relationship is stored.
[0082] In this embodiment, the data status is, for example, data in an encryption, application, or storage state.
[0083] The beneficial effect of the above design scheme is: by assigning a unique data tag to each piece of data in the data source when the data source is first connected to the lineage update system, and recording the association between the data identifier and the data, it provides an identification basis for the synchronization of industrial data identifiers.
[0084] Example 3:
[0085] Based on Example 1, this embodiment of the present invention provides a method for synchronously identifying industrial data identifiers during a lineage update process. In S2, a lineage relationship map between all data is established by analyzing the flow process of all data, including:
[0086] Obtaining data identifiers of all data, determining data-related information between the data based on the data identifiers, and establishing a first blood relationship between the data based on the data-related information;
[0087] Obtain the data flow direction of the data from the flow process of all data, and establish a data flow structure diagram based on the data flow direction of all data, and based on the flow direction and flow node characteristics in the data flow structure diagram;
[0088] Based on the flow direction and flow node characteristics, the data flow structure diagram is split, and a data flow identifier is set for each branch from the split result;
[0089] Based on the data flow identifier, determining the data flow pattern of each piece of data, and determining the flow similarity of the data flow patterns between all the data;
[0090] Based on the flow similarity, a second blood relationship between the data is established;
[0091] Based on the first blood relationship, verify the second blood relationship to determine whether the second blood relationship is established on the basis of the first blood relationship;
[0092] If yes, establishing a blood relationship map between the data based on the second blood relationship;
[0093] Otherwise, based on the first blood relationship, the second blood relationship is adjusted, and a blood relationship map between the data is established based on the adjusted second blood relationship.
[0094] In this embodiment, the data-related information includes correlations of data attributes, data types, and the like.
[0095] In this embodiment, the more data flow identifiers that both pieces of data pass through, the higher the flow similarity.
[0096] In this embodiment, the first blood relationship is determined based on the characteristics of the data itself, and the second blood relationship is determined based on the characteristics of the data flow.
[0097] The beneficial effect of the above design scheme is: by starting from the characteristics of the data itself and the characteristics of the data flow, a blood relationship map between data is constructed to ensure the accuracy and comprehensiveness of the obtained blood relationship map, and provide a basis for the subsequent synchronous identification of industrial data labels.
[0098] Example 4:
[0099] Based on Example 3, this embodiment of the present invention provides a method for synchronously identifying industrial data identifiers during a lineage update process, wherein adjusting the second lineage relationship based on the first lineage relationship includes:
[0100] Obtaining the blood relationship to be adjusted that conflicts with the first blood relationship in the second blood relationship, and dividing the blood relationship to be adjusted into a front-end blood relationship and a back-end blood relationship based on the relationship characteristics with the first blood relationship;
[0101] Adjust the front-segment blood relationship based on the first blood relationship to obtain the target front-segment blood relationship; adjust the back-segment blood relationship based on the target front-segment blood relationship to obtain the target back-segment blood relationship;
[0102] Based on the target's front-end blood relationship and the target's back-end blood relationship, an adjusted second blood relationship is obtained.
[0103] In this embodiment, the blood relationship to be adjusted that is theoretically the same as the first blood relationship is the front-end blood relationship, and the blood relationship to be adjusted that is determined with the participation of the first blood relationship is the back-end blood relationship.
[0104] The beneficial effect of the above design scheme is: by dividing the blood relationship to be adjusted into the front blood relationship and the back blood relationship based on the relationship characteristics with the first blood relationship, adjusting the front blood relationship based on the first blood relationship to obtain the target front blood relationship, adjusting the back blood relationship based on the target front blood relationship to obtain the target back blood relationship, and obtaining the adjusted second blood relationship based on the target front blood relationship and the target back blood relationship, ensuring the accuracy of the adjusted second blood relationship, and providing a basis for the subsequent synchronous identification of industrial data identifiers.
[0105] Example 5:
[0106] Based on Example 1, this embodiment of the present invention provides a method for synchronously identifying industrial data identifiers during a lineage update process. In S2, based on the association relationship between the data identifiers and the data, the dependency and transfer relationship between the data identifiers is determined, including:
[0107] Based on the blood relationship map, determine the flow relationship between data;
[0108] Based on the association relationship between data identifiers and data, combined with the flow relationship, the dependency and transmission relationship between data identifiers are determined.
[0109] The beneficial effect of the above design scheme is: by determining the dependency and transmission relationship between data identifiers based on the association relationship between data identifiers and data, a basis is provided for the acquisition, identification and synchronization of subsequent identifiers.
[0110] Example 6:
[0111] Based on Example 1, the present invention provides a method for synchronously identifying industrial data identifiers during lineage update, such as Figure 3 As shown, in S3, the changes of the data source are monitored in real time, and when a change in the data source is detected, the data tag synchronization identification process is triggered to perform data identification on the changed data, including:
[0112] Monitor the current identification value of the data source in real time and compare the current identification value with the stored identification value. If there is a mismatch, it is determined that the data source has changed.
[0113] When a change in the data source is detected, the data tag synchronization recognition process is triggered to perform data identification and recognition on the changed data.
[0114] In this embodiment, the identification value of the data source is determined according to the data type, data status, etc.
[0115] The beneficial effects of the above design scheme are: achieving real-time monitoring of data changes, triggering the data tag synchronization identification process, and performing data identification and recognition on the changed data.
[0116] Example 7:
[0117] Based on Example 1, this embodiment of the present invention provides a method for synchronously identifying industrial data identifiers during a lineage update process. In S4, based on the data change type and lineage relationship map of the data source, combined with the data identifier identification result, a target data identifier to be synchronized is determined, and the target data identifier is synchronously updated, including:
[0118] Based on the data change type of the data source, the type conversion characteristics are determined, and based on the blood relationship map, the state change characteristics are determined;
[0119] Determine the target data identifier of the data whose data source has changed from the data identifier identification result;
[0120] Based on the type conversion characteristics and state change characteristics, the target data identifier is synchronously updated.
[0121] The beneficial effects of the above design scheme are: by determining the type of data change based on the data source change type, determining the state change characteristics based on the blood relationship map, and synchronously updating the target data identifier based on the type conversion characteristics and state change characteristics, the synchronous update of the data identifier is determined from both the data itself and the data identifier, thereby ensuring the accuracy of the data identifier update.
[0122] Example 8:
[0123] Based on Example 7, this embodiment of the present invention provides a method for synchronously identifying industrial data identifiers during a lineage update process, wherein the method synchronously updates the target data identifier based on the type conversion feature and the state change feature, including:
[0124] Parsing the target data identifier, and dividing the target data identifier into a plurality of identification fields according to the parsing result;
[0125] Acquire a first identification field related to the type conversion feature, and replace the first identification field based on the type conversion feature to obtain a first target identification field;
[0126] Acquire a second identification field related to the state change feature, and replace the second identification field based on the state change feature to obtain a second target identification field;
[0127] An initial identification field is formed based on the first target identification field, the second target identification field, and the unchanged identification field;
[0128] Based on the identifier display constraint rules, the initial identifier field is subjected to standardization verification. If the standardization verification fails, the identifier synchronization update fails, and identifier synchronization is performed again.
[0129] If the standard verification passes, the uniqueness verification of the initial identification field of the data is performed based on the uniqueness feature of the identification. If the uniqueness verification fails, a unique field is added to the initial identification field based on the identification construction rule to achieve synchronous update of the target data identification.
[0130] If the uniqueness verification is passed, the target data identifier is synchronously updated based on the initial identification field.
[0131] The beneficial effects of the above design scheme are: by identifying and updating the identifier from two aspects: type conversion characteristics and state change characteristics, and verifying the standardization and uniqueness after the update, the accuracy of the synchronous update of the target data identifier is guaranteed, and the accuracy and consistency of the data throughout the entire life cycle are achieved.
[0132] Example 9:
[0133] Based on Example 1, this embodiment of the present invention provides a method for synchronously identifying industrial data identifiers during a lineage update process. In S5, based on the synchronous update result, the latest data identifier and the latest data status of the changed data are obtained, and the consistency of the latest data identifier and the latest data status is verified, including:
[0134] Get the latest data identifier of the changed data from the synchronization update result;
[0135] Obtain the latest data status from the results of the data flow process;
[0136] Obtaining a state identifier from the latest data identifier, matching the state identifier with the latest data state to obtain a matching degree;
[0137] Determining whether the matching degree is greater than a preset matching degree;
[0138] If so, confirm that the latest data identifier is consistent with the latest data status;
[0139] Otherwise, it is determined that the latest data identifier is inconsistent with the latest data status.
[0140] The beneficial effect of the above design scheme is to ensure the accuracy and consistency of data throughout its entire life cycle by verifying the consistency between the latest data identifier and the latest data status.
[0141] Example 10:
[0142] Based on Example 9, this embodiment of the present invention provides a method for synchronously identifying industrial data identifiers during a lineage update process, wherein the state identifier is matched with the latest data state to obtain a matching degree, including:
[0143] Based on the preset identification conversion rules, determine the actual standard status identification corresponding to the latest data status;
[0144] Based on a preset identifier conversion rule, the state identifier is standardized to obtain a marked standard state identifier;
[0145] The actual standard state identifier is matched with the marked standard state identifier to obtain a matching degree.
[0146] In this embodiment, the preset identifier conversion rule is pre-set and used to determine a standard data identifier.
[0147] The beneficial effects of the above design scheme are: by determining the actual standard state identifier corresponding to the latest data state based on the preset identifier conversion rules, standardizing the state identifier to obtain the marked standard state identifier, matching the actual standard state identifier with the marked standard state identifier to obtain the matching degree, ensuring the accuracy of the matching degree between the determined state identifier and the latest data state, and providing an accurate data basis for consistency verification.
[0148] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of this application document and its equivalents, the present invention is intended to include these modifications and variations.
Claims
1. A method for synchronously identifying industrial data identifiers during lineage update, characterized in that: include: S1: When a data source is first connected to the lineage update system, a unique data tag is assigned to each piece of data in the data source, and the association between the data tag and the data is recorded; S2: By analyzing the flow of all data, a blood relationship map is established between all data. Based on the association between data identifiers and data, the dependency and transfer relationship between data identifiers is determined, including: Obtaining data identifiers of all data, determining data-related information between the data based on the data identifiers, and establishing a first blood relationship between the data based on the data-related information; Obtain the data flow direction of the data from the flow process of all data, establish a data flow structure diagram based on the data flow direction of all data, and determine the flow direction and flow node characteristics based on the data flow structure diagram; Based on the flow direction and flow node characteristics, the data flow structure diagram is split, and a data flow identifier is set for each branch from the split result; Based on the data flow identifier, determining the data flow pattern of each piece of data, and determining the flow similarity of the data flow patterns between all the data; Based on the flow similarity, a second blood relationship between the data is established; Based on the first blood relationship, verify the second blood relationship to determine whether the second blood relationship is established on the basis of the first blood relationship; If yes, establishing a blood relationship map between the data based on the second blood relationship; Otherwise, based on the first blood relationship, the second blood relationship is adjusted, and a blood relationship map between the data is established based on the adjusted second blood relationship; Based on the blood relationship map, determine the flow relationship between data; Based on the association between data identifiers and data, combined with the flow relationship, determine the dependency and transfer relationship between data identifiers; S3: Monitors changes in data sources in real time. When changes are detected, the data tag synchronization process is triggered to identify the changed data. S4: Based on the data change type and blood relationship map of the data source, combined with the data identifier recognition result, determine the target data identifier that needs to be synchronized, and synchronize and update the target data identifier; S5: Based on the synchronization update result, the latest data identifier and the latest data status of the changed data are obtained, and the consistency of the latest data identifier and the latest data status is verified.
2. The method for synchronously identifying industrial data identifiers during lineage update according to claim 1, characterized in that: In S1, when a data source is first connected to the lineage update system, a unique data tag is assigned to each piece of data in the data source, and the association between the data tag and the data is recorded, including: Obtain the data type and data status of each piece of data in the data source, and obtain the data characteristics of each piece of data; Based on the data type and data state, an initial identifier is matched for the corresponding data, and based on the data characteristics, a unique identification mark is performed on the basis of the initial identifier to obtain a data identifier; An association relationship between the data identifier and the data is established, and the association relationship is stored.
3. The method for synchronously identifying industrial data identifiers during lineage update according to claim 1, characterized in that: The adjusting the second blood relationship based on the first blood relationship includes: Obtaining the blood relationship to be adjusted that conflicts with the first blood relationship in the second blood relationship, and dividing the blood relationship to be adjusted into a front-end blood relationship and a back-end blood relationship based on the relationship characteristics with the first blood relationship; Adjust the front-segment blood relationship based on the first blood relationship to obtain the target front-segment blood relationship; adjust the back-segment blood relationship based on the target front-segment blood relationship to obtain the target back-segment blood relationship; Based on the target's front-end blood relationship and the target's back-end blood relationship, an adjusted second blood relationship is obtained.
4. The method for synchronously identifying industrial data identifiers during lineage update according to claim 1, characterized in that: In S3, changes in the data source are monitored in real time. When a change in the data source is detected, a data tag synchronization identification process is triggered to identify the changed data, including: Monitor the current identification value of the data source in real time and compare the current identification value with the stored identification value. If there is a mismatch, it is determined that the data source has changed. When a change in the data source is detected, the data tag synchronization recognition process is triggered to perform data identification and recognition on the changed data.
5. The method for synchronously identifying industrial data identifiers during lineage update according to claim 1, characterized in that: In S4, based on the data change type and blood relationship map of the data source, combined with the data identifier recognition result, the target data identifier to be synchronized is determined, and the target data identifier is synchronously updated, including: Based on the data change type of the data source, the type conversion characteristics are determined, and based on the blood relationship map, the state change characteristics are determined; Determine the target data identifier of the data whose data source has changed from the data identifier identification result; Based on the type conversion characteristics and state change characteristics, the target data identifier is synchronously updated.
6. The method for synchronously identifying industrial data identifiers during lineage update according to claim 5, characterized in that: The synchronously updating the target data identifier based on the type conversion feature and the state change feature includes: Parsing the target data identifier, and dividing the target data identifier into a plurality of identification fields according to the parsing result; Acquire a first identification field related to the type conversion feature, and replace the first identification field based on the type conversion feature to obtain a first target identification field; Acquire a second identification field related to the state change feature, and replace the second identification field based on the state change feature to obtain a second target identification field; An initial identification field is formed based on the first target identification field, the second target identification field, and the unchanged identification field; Based on the identifier display constraint rules, the initial identifier field is subjected to standardization verification. If the standardization verification fails, the identifier synchronization update fails, and identifier synchronization is performed again. If the standard verification passes, the uniqueness verification of the initial identification field of the data is performed based on the uniqueness feature of the identification. If the uniqueness verification fails, a unique field is added to the initial identification field based on the identification construction rule to achieve synchronous update of the target data identification. If the uniqueness verification is passed, the target data identifier is synchronously updated based on the initial identification field.
7. The method for synchronously identifying industrial data identifiers during lineage update according to claim 1, characterized in that: In S5, based on the synchronous update result, the latest data identifier and the latest data status of the changed data are obtained, and the consistency of the latest data identifier and the latest data status is verified, including: Get the latest data identifier of the changed data from the synchronization update result; Obtain the latest data status from the results of the data flow process; Obtaining a state identifier from the latest data identifier, matching the state identifier with the latest data state to obtain a matching degree; Determining whether the matching degree is greater than a preset matching degree; If so, confirm that the latest data identifier is consistent with the latest data status; Otherwise, it is determined that the latest data identifier is inconsistent with the latest data status.
8. The method for synchronously identifying industrial data identifiers during lineage update according to claim 7, characterized in that: The matching of the state identifier with the latest data state to obtain a matching degree includes: Based on the preset identification conversion rules, determine the actual standard status identification corresponding to the latest data status; Based on a preset identifier conversion rule, the state identifier is standardized to obtain a marked standard state identifier; The actual standard state identifier is matched with the marked standard state identifier to obtain a matching degree.
Citation Information
Patent Citations
Metadata management method, device, computer device, and storage medium
CN109241358A
Metadata-driven data inspection and version management method and device and electronic equipment
CN114218301A