A multi-source data master data identification and standardization management method for coal enterprises

CN122614818APending Publication Date: 2026-08-21HUADIAN COAL IND GRP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610760113.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-29
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

[0006]本发明的一个目的在于提出一种面向煤炭企业的多源数据主数据识别与标准化治理方法,本发明提出一种面向煤炭企业的多源数据主数据识别与标准化治理方法,旨在解决煤炭企业多业务系统中同一业务对象名称不一致、字段表达不统一、来源关系难追溯以及冲突字段难以可信决策的问题

Benefits of technology

本发明针对煤炭企业多源数据分散、来源关系复杂和字段出处难以追溯的问题,在采集阶段引入源域血缘锚定采集机制,对设备系统数据、生产系统数据、人员系统数据、物资系统数据和业务台账数据进行来源关联和血缘标记,形成血缘锚定原始数据集合,使后续字段解析、对象识别和冲突决策均能够保留数据来源、字段归属、版本变化和对象关联依据,从而提高多源主数据治理过程的可追溯性和数据基础完整性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122614818A_ABST
    Figure CN122614818A_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-source data master data identification and standardization management method for coal enterprises, comprising the following steps: step one: establish multi-source data acquisition index, form blood relation anchoring original data set;Step two: parse field, identify type, uniform format, form field analysis data set;Step three: entity segmentation, alias identification and attribute normalization, form candidate master data object;Step four: matching evaluation based on improved HGT model, form object matching evaluation result;Step five: associated clustering and merging entity, form merged master data object;Step six: multi-source credibility decision, form conflict field decision result;Step seven: coding, mapping and checking, form standard master data;Step eight: release and update standard master data, form management closed loop.The application uses improved HGT model, realizes the accurate identification and standardization management of multi-source master data of coal enterprises.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multi-source data master data identification and standardized governance technology, and in particular to a method for multi-source data master data identification and standardized governance for coal enterprises. Background Technology

[0002] Coal enterprises involve various types of data in their production and operation processes, including equipment management, production scheduling, personnel organization, material supply, maintenance, and business ledgers. This data is typically distributed across equipment systems, production systems, personnel systems, material systems, and various business ledgers. Differences in system development time, data standards, field naming rules, and maintenance methods can lead to inconsistencies in names, coding, missing attributes, and similar but different expressions for the same business object across different sources.

[0003] In existing master data governance processes, manual rules, field comparisons, or fixed coding rules are typically used to clean and merge multi-source data. While this approach can handle some data with uniform format and stable field structure, in the complex business scenarios of coal enterprises, abbreviations of equipment names, aliases of materials, personnel job changes, historical records in ledgers, and upstream and downstream production relationships are intertwined. Relying solely on name similarity or single field matching can easily lead to the omission of the same object or the incorrect merging of objects with similar names but different business meanings.

[0004] Meanwhile, existing master data standardization processes often focus more on result encoding and field mapping, while neglecting to consider data source lineage, field version changes, business ledger continuity, and the credibility of conflicting fields. When the same field has inconsistent values ​​in multiple systems, processing usually relies on manual judgment or fixed source priority, which makes it difficult to comprehensively reflect the source authority, time freshness, field completeness, and historical consistency, thus affecting the accuracy, traceability, and continuous governance capabilities of standard master data.

[0005] Therefore, how to provide a method for identifying and standardizing multi-source master data for coal enterprises is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose a method for multi-source master data identification and standardized governance in coal enterprises. This method aims to address issues such as inconsistent names of the same business object across multiple business systems within coal enterprises, inconsistent field representations, difficulty in tracing source relationships, and challenges in making reliable decisions regarding conflicting fields. The method combines a source domain lineage anchoring acquisition mechanism, an improved HGT master data identification model, and a business anchor point counter-evidence gating mechanism. It establishes a governance chain around equipment system data, production system data, personnel system data, material system data, and business ledger data, encompassing multi-source acquisition, field parsing, candidate object formation, object matching evaluation, and the release and updating of standard master data.

[0007] Building upon this foundation, a source domain lineage anchoring acquisition mechanism is used to associate and label multi-source data, ensuring that the original data retains source, field, version, and object attribution information before entering field parsing and object identification. Furthermore, an improved HGT master data identification model is used to identify heterogeneous relationships among candidate master data objects, incorporating source relationships, attribute relationships, lineage relationships, business scenario relationships, and upstream and downstream relationships among candidate master data objects into a unified evaluation process. Simultaneously, a business anchor point counter-evidence gating mechanism is used to enhance business anchor point relationships that support the identification of the same entity and suppress business counter-evidence relationships that are prone to erroneous merging, thereby generating more accurate object matching evaluation results.

[0008] This method can improve the accuracy of candidate master data object identification and entity merging in scenarios where coal enterprises have complex multi-source data structures, inconsistent field expressions, and overlapping business object relationships. Through multi-source credibility decision-making, unified coding, field mapping, rule verification, and feedback correction, it can achieve traceability, updability, and closed-loop governance of standard master data. It overcomes the problems of existing master data governance methods that rely on manual rules, have a high rate of erroneous merging, and lack sufficient basis for decision-making on conflicting fields. It has the technical effect of improving the accuracy, standardization, and continuous maintenance capabilities of master data governance in coal enterprises.

[0009] A method for identifying and standardizing master data from multiple sources in coal enterprises, according to an embodiment of the present invention, includes the following steps: Step 1: Establish a multi-source data collection index, collect and process multi-source data based on the source domain lineage anchoring collection mechanism, and perform source association and lineage marking on the collected multi-source data to form a lineage anchoring raw data set; Step 2: Perform field parsing, field type identification, data format unification, and abnormal format correction on the original bloodline anchoring data set to form a field parsing data set; Step 3: Perform entity segmentation, alias recognition, and attribute normalization on the field parsing data set to identify candidate expressions pointing to the same business object from different data sources and form candidate master data objects; Step 4: Based on the improved HGT master data identification model, heterogeneous relationship identification and object matching evaluation are performed on candidate master data objects to form object matching evaluation results. The improved HGT master data identification model includes a heterogeneous master data relationship construction layer, a type semantic embedding layer, an anchor point counter-evidence gating propagation layer, and a matching evaluation output layer. The anchor point counter-evidence gating propagation layer embeds a business anchor point counter-evidence gating mechanism. Step 5: Based on the object matching evaluation results, perform association clustering and entity merging on the candidate master data objects to form merged master data objects; Step 6: Perform multi-source credibility decision on conflicting fields in the merged master data object to form conflicting field decision results; Step 7: Based on the decision results of conflicting fields, perform unified encoding, field mapping, rule validation, and standardization governance on the merged master data objects to form standard master data; Step 8: Publish and continuously update the standard master data, and provide feedback and correction to the object matching assessment results and standard master data based on the master data change information, forming a closed loop for master data standardization governance.

[0010] Optionally, step one specifically includes: Identify the data sources and data objects to be addressed in coal enterprises, and establish a multi-source data collection index. The multi-source data includes: equipment system data, production system data, personnel system data, material system data, and business ledger data. Based on the multi-source data collection index, determine the data source, collection interface, business form and collection batch corresponding to each type of multi-source data, and process the multi-source data to obtain collection records; Based on the source domain lineage anchoring acquisition mechanism, source anchoring identifiers are generated according to data source, acquisition interface, business form and acquisition batch; field anchoring identifiers are generated according to original field name, field unit, field format and field belonging object; version anchoring identifiers are generated according to acquisition batch, update time, business effective status and historical change records; and object anchoring identifiers are generated according to the belonging relationship and business association relationship between data objects. Write the source anchoring identifier, field anchoring identifier, version anchoring identifier, and object anchoring identifier into the corresponding lineage marker field of the collection record, respectively, to form the source lineage marker, field lineage marker, version lineage marker, and object lineage marker; The source lineage marker, field lineage marker, version lineage marker, and object lineage marker are bound to the collection records to form a lineage-anchored original data set.

[0011] Optionally, step two specifically involves: The collected records in the original dataset for kinship anchoring are decomposed into fields to extract the original field names, original field values, field units, field formats, field sources, and kinship markers; Based on the original field name, original field value, field unit, field format, and lineage marker, the meaning of the field is parsed, and the field correspondence between the original field and the standard field is established; Based on the field correspondence, the original fields are identified by field type and divided into object identifier field, object name field, attribute description field, spatial location field, status record field, time record field, and business association field; Based on the field type identification results, the original field values ​​are standardized in terms of data format, and the encoding format, name format, time format, numerical format, unit of measurement format and status expression format in data from different sources are converted into a unified field expression; Correct any abnormal formats and abnormal field states that exist during the data format unification process. Abnormal formats and abnormal field states include field values ​​that are empty, field values ​​that contain redundant characters, inconsistent field units, inconsistent time expressions, inconsistent status expressions, and inconsistent encoding lengths. The unified field representation, field correspondence, field type identification results, and lineage markers are linked to form a field parsing data set.

[0012] Optionally, step three specifically includes: Based on the object identifier field, object name field, attribute description field and business association field in the field parsing data set, the field parsing data set is split into entities to obtain the entity splitting results; Based on the entity segmentation results, extract the object name, abbreviation, former name, ledger record name, and business description name from data from different sources, and perform alias recognition on the object name, abbreviation, former name, ledger record name, and business description name to obtain the alias recognition results; Based on the alias recognition results, establish a set of candidate representations of the same business object in data from different sources; Perform attribute normalization on the attribute description fields in the candidate expression set, converting attribute names, attribute units, attribute formats, and attribute value ranges into unified attribute expressions; By associating the candidate expression set with the unified attribute expression, candidate expressions pointing to the same business object in data from different sources are identified, forming candidate master data objects.

[0013] Optionally, step four specifically involves: Based on the object type, object name, alias expression, field attributes, source association, lineage relationship, and business association of candidate master data objects, a heterogeneous master data relationship graph is constructed through a heterogeneous master data relationship construction layer. Candidate master data objects are used as object nodes, and source systems, field attributes, business ledgers, production processes, spatial locations, responsible entities, and upstream and downstream business objects are used as association nodes. The source relationship, attribute relationship, lineage relationship, ledger record relationship, scenario relationship, location relationship, responsibility relationship, and upstream and downstream relationship between object nodes and association nodes are used as heterogeneous relationship edges. The heterogeneous relationship edges are marked with relationship direction and relationship strength to obtain the heterogeneous master data graph representation. The heterogeneous graph representation of master data is used as the input type semantic embedding layer. The object nodes, associated nodes and heterogeneous relationship edges are type-encoded respectively. The object name, alias expression, attribute expression, source expression, lineage expression and business description expression of the candidate master data objects are embedded. The type encoding results and the embedding results are concatenated, fused and linearly mapped to obtain the type semantic representation of the candidate master data objects. Type semantic representation is input into the anchor point counter-evidence gating propagation layer. The supporting and exclusive relationships between candidate master data objects are identified through the business anchor point counter-evidence gating mechanism. The relationships corresponding to object identity consistency, spatial location consistency, business scenario consistency, ledger record consistency, responsible entity consistency, and upstream and downstream relationship consistency are determined as business anchor point relationships. The relationships corresponding to object type conflict, location chain conflict, business scenario conflict, key field conflict, lineage conflict, and state sequence conflict are determined as business counter-evidence relationships. Anchor point enhancement weights are generated based on business anchor point relationships, and disproving weights are generated based on business disproving relationships. Anchor point enhancement weights and disproving weights are then gated and fused using a business anchor point disproving gating mechanism to obtain gating adjustment weights. In the process of improving the heterogeneous attention propagation of the HGT master data recognition model, type-related query representations, key representations, and value representations are generated based on the node type of the object node and the relationship type of the heterogeneous relation edge. The node attention weight and relation edge propagation weight are modified by gating adjustment weights, so that the semantic propagation of candidate master data objects with business anchor relationships is enhanced, and the semantic propagation of candidate master data objects with business counter-evidence relationships is reduced, thus obtaining the gating propagation semantic representation. The gating propagation semantic representation is input to the matching evaluation output layer, and the name similarity, attribute overlap, business scenario consistency, upstream and downstream correlation, lineage consistency and counter-evidence conflict degree between candidate master data objects are jointly evaluated to generate object matching confidence. The matching level, the same entity judgment result, and the counter-evidence suppression mark are determined based on the object matching confidence score, and the object matching confidence score, matching level, same entity judgment result, and counter-evidence suppression mark are used as the object matching evaluation results.

[0014] Optionally, the business anchor counter-evidence gating mechanism specifically includes: Based on the object identifier, object name, alias expression and key attributes of the candidate master data objects, extract the object identity anchor points, and determine whether there is object identity consistency among the candidate master data objects based on the object identity anchor points; Based on the spatial location field of the candidate master data object, the location record of the source system, and the location record of the business ledger, extract the spatial location anchor point, and determine whether there is spatial location consistency among the candidate master data objects based on the spatial location anchor point; Based on the production process, maintenance records, requisition records, job responsibility records, and business ledger records associated with the candidate master data objects, extract business scenario anchor points, and determine whether there is business scenario consistency among the candidate master data objects based on the business scenario anchor points. Based on the equipment subordination, production process chain, material requisition, personnel responsibility, and ledger record relationships among candidate master data objects, upstream and downstream relationship anchor points are extracted, and the consistency of upstream and downstream relationships among candidate master data objects is determined based on the upstream and downstream relationship anchor points. The consistency of object identity, spatial location, business scenario, and upstream and downstream relationship are integrated to generate business anchor relationships, and the anchor enhancement weight is calculated based on the number, strength, and source authority of the business anchor relationships. Based on the differences in object type, location chain, business scenario, key field, lineage, and state sequence among candidate master data objects, business counter-evidence relationships are extracted, and counter-evidence suppression weights are calculated based on the number of conflicts, conflict level, and source authority of the business counter-evidence relationships. Anchor point enhancement weights and counter-evidence suppression weights are fused to generate gating adjustment weights. These gating adjustment weights are used to increase the semantic propagation strength between candidate master data objects with business anchor point relationships and decrease the semantic propagation strength between candidate master data objects with business counter-evidence relationships. The node attention weights and relation edge propagation weights in the improved HGT master data recognition model are adjusted according to the gating adjustment weights to generate a gating propagation semantic representation.

[0015] Optionally, step five specifically includes: Based on the object matching evaluation results, extract the object matching confidence, matching level, same entity judgment result and counter-evidence suppression flag between candidate master data objects; Candidate master data objects are used as clustering nodes, and object matching confidence, matching level, same entity judgment result and counter-evidence suppression flag are used as clustering edge attributes to construct a candidate master data association graph. Based on the clustering edge attributes between candidate master data objects in the candidate master data association graph, the candidate master data objects are initially grouped to form master data candidate clusters; Relation propagation clustering is performed on candidate master data objects within the candidate master data clusters. The object matching confidence, matching level and counter-evidence suppression labels are propagated and updated in the candidate master data association graph to obtain master data clusters. Based on the master data clusters, the object identifier, object name, alias expression, attribute fields and business association fields in the candidate master data objects are merged to form merged master data objects.

[0016] Optionally, step six specifically includes: Identify fields in the merged master data object that have different sources, the same meaning, and inconsistent values, and form a set of conflicting fields; A multi-dimensional credibility evaluation is conducted on the candidate field values ​​corresponding to each conflicting field item in the conflicting field set. The evaluation dimensions include the source authority determined by the data source, business affiliation, field lineage and source system management authority; the time freshness determined by the update time, business effective time, version status and historical change records; the field completeness determined by the field completeness, format standardization, unit consistency and value validity; and the historical consistency determined by the consistent occurrence in historical data, business ledgers and related objects. The credibility score of a field is generated by integrating factors such as source authority, time freshness, field completeness, and historical consistency. Based on the field credibility score, the target field value, alternative field values, and conflict trace information of the conflict field are determined to form the conflict field decision result.

[0017] Optionally, step seven specifically includes: Based on the conflict field decision results, determine the target field value in the merged master data object, and retain the source association, lineage marker and conflict decision basis corresponding to the target field value; A unified master data code is generated based on the object category, business level, source relationship, and coding order of the merged master data object, and the unified master data code is bound to the merged master data object; Based on the field parsing data set, the merged master data object, and the decision results of conflicting fields, establish the field mapping relationship between the original fields, parsed fields, merged fields, and standard fields; Based on the field mapping relationship, perform integrity verification, uniqueness verification, format verification, reference relationship verification, and status validity verification on the merged master data object to obtain the master data object that passes the verification. Standardize the master data objects that pass the verification by encapsulating the unified master data code, target field values, field mapping relationships and verification results to form standard master data.

[0018] Optionally, step eight specifically includes: The standard master data is published to the data management platform of the coal enterprise, and the corresponding publication version, publication time, publication status and publication scope are generated. Identify the change types based on master data change information: new objects, field changes, object merging, object splitting, object deactivation, and version rollback; Extract the corresponding changed object, changed field, field value before change, field value after change, and change source based on the change type to form a master data change record; Based on the master data change record, the object matching evaluation results are fed back and corrected, and the object matching confidence, the same entity judgment result and the counter-evidence suppression mark between candidate master data objects are updated. Based on the feedback and correction of the object matching evaluation results, the standard master data is continuously updated, and the version status, lineage markers and release status of the standard master data are updated synchronously to form a closed loop of master data standardization governance.

[0019] The beneficial effects of this invention are: This invention addresses the problems of scattered multi-source data, complex source relationships, and difficulty in tracing the origin of fields in coal enterprises. It introduces a source domain lineage anchoring acquisition mechanism during the acquisition phase. This mechanism associates and marks the source of equipment system data, production system data, personnel system data, material system data, and business ledger data, forming a lineage anchored original data set. This ensures that subsequent field parsing, object identification, and conflict decision-making can retain the data source, field ownership, version changes, and object association basis, thereby improving the traceability and data integrity of the multi-source master data governance process.

[0020] In the process of candidate master data object identification, the improved HGT master data identification model constructs a heterogeneous master data relationship graph by considering candidate master data objects, source systems, field attributes, business ledgers, spatial locations, and upstream and downstream business objects. It then utilizes a type semantic embedding layer and a matching evaluation output layer to jointly evaluate name similarity, attribute overlap, business scenario consistency, upstream and downstream relevance, and lineage consistency. Compared to methods that rely solely on name or single field matching, this process can more accurately distinguish data with different names but pointing to the same object, as well as data with similar names but different business meanings, reducing the risk of missed identification and mismatch.

[0021] By leveraging the business anchor counter-evidence gating mechanism in the anchor counter-evidence gating propagation layer, the method can enhance semantic propagation between the same entity by utilizing consistency in object identity, spatial location, business scenario, and upstream and downstream relationships. At the same time, it can suppress error propagation by utilizing differences in object type, location chain, key field, and state sequence, thereby obtaining more reliable object matching evaluation results and improving the accuracy of subsequent association clustering and entity merging.

[0022] During the standardization governance phase, the multi-source credibility decision-making mechanism generates field credibility scores from dimensions such as source authority, time freshness, field completeness, and historical consistency. It determines the target field value, alternative field value, and conflict trace information for conflicting fields, ensuring that the generation of standard master data has a clear basis. Furthermore, by combining unified coding, field mapping, rule verification, release updates, and feedback correction, a closed loop for sustainable master data standardization governance can be formed, improving the standardization level, reliability, and dynamic update capability of coal enterprise master data. Attached Figure Description

[0023] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is an overall flowchart of a multi-source data master data identification and standardized governance method for coal enterprises proposed in this invention; Figure 2 This is a flowchart illustrating the working principle of an improved HGT model for multi-source data master data identification and standardized governance in coal enterprises, as proposed in this invention. Detailed Implementation

[0024] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0025] refer to Figure 1 and Figure 2 A method for identifying and standardizing master data from multiple sources in coal enterprises includes the following steps: Step 1: Establish a multi-source data collection index, collect and process multi-source data based on the source domain lineage anchoring collection mechanism, and perform source association and lineage marking on the collected multi-source data to form a lineage anchoring raw data set; Step 2: Perform field parsing, field type identification, data format unification, and abnormal format correction on the original bloodline anchoring data set to form a field parsing data set; Step 3: Perform entity segmentation, alias recognition, and attribute normalization on the field parsing data set to identify candidate expressions pointing to the same business object from different data sources and form candidate master data objects; Step 4: Based on the improved HGT master data identification model, heterogeneous relationship identification and object matching evaluation are performed on candidate master data objects to form object matching evaluation results. The improved HGT master data identification model includes a heterogeneous master data relationship construction layer, a type semantic embedding layer, an anchor point counter-evidence gating propagation layer, and a matching evaluation output layer. The anchor point counter-evidence gating propagation layer embeds a business anchor point counter-evidence gating mechanism. Step 5: Based on the object matching evaluation results, perform association clustering and entity merging on the candidate master data objects to form merged master data objects; Step 6: Perform multi-source credibility decision on conflicting fields in the merged master data object to form conflicting field decision results; Step 7: Based on the decision results of conflicting fields, perform unified encoding, field mapping, rule validation, and standardization governance on the merged master data objects to form standard master data; Step 8: Publish and continuously update the standard master data, and provide feedback and correction to the object matching assessment results and standard master data based on the master data change information, forming a closed loop for master data standardization governance.

[0026] In this embodiment, step one specifically includes: Identify the data sources and data objects to be addressed in coal enterprises, and establish a multi-source data collection index. The multi-source data includes: equipment system data, production system data, personnel system data, material system data, and business ledger data. Based on the multi-source data collection index, determine the data source, collection interface, business form and collection batch corresponding to each type of multi-source data, and process the multi-source data to obtain collection records; Based on the source domain lineage anchoring acquisition mechanism, source anchoring identifiers are generated according to data source, acquisition interface, business form and acquisition batch; field anchoring identifiers are generated according to original field name, field unit, field format and field belonging object; version anchoring identifiers are generated according to acquisition batch, update time, business effective status and historical change records; and object anchoring identifiers are generated according to the belonging relationship and business association relationship between data objects. Write the source anchoring identifier, field anchoring identifier, version anchoring identifier, and object anchoring identifier into the corresponding lineage marker field of the collection record, respectively, to form the source lineage marker, field lineage marker, version lineage marker, and object lineage marker; The source lineage marker, field lineage marker, version lineage marker, and object lineage marker are bound to the collection records to form a lineage-anchored original data set; In the specific implementation process, the multi-source data acquisition index establishes index items according to data source, data object, acquisition interface, business form and acquisition batch. Data source is used to distinguish equipment system, production system, personnel system, material system and business ledger. Data object is used to identify equipment object, production object, personnel object, material object and ledger object. Acquisition interface is used to record the data entry method. Business form is used to record the storage location of data in the source system. Acquisition batch is used to record the data ownership relationship in the same round of acquisition processing. During multi-source data acquisition and processing, the original records from different sources are read according to the multi-source data acquisition index, and each original record is assigned a source association identifier and an object association identifier. The source association identifier is used to record the system, interface or form from which the original record comes, and the object association identifier is used to record the business object category corresponding to the original record. Source lineage markers are generated based on data source, collection interface, and business form; field lineage markers are generated based on original field name, field unit, field format, and field belonging object; version lineage markers are generated based on collection batch, update time, business effective status, and historical change records; and object lineage markers are generated based on the belonging relationship between equipment objects, production objects, personnel objects, material objects, and ledger objects. The source lineage marker, field lineage marker, version lineage marker, and object lineage marker are written into the lineage marker field of the corresponding original record and encapsulated together with the original field values ​​into a lineage-anchored original data set. This makes the lineage-anchored original data set simultaneously contain the original data content, source association, field origin relationship, version inheritance relationship, and object ownership relationship, providing a traceable data foundation for field parsing, candidate master data object formation, and heterogeneous relationship identification.

[0027] In this embodiment, step two specifically includes: The collected records in the original dataset for kinship anchoring are decomposed into fields to extract the original field names, original field values, field units, field formats, field sources, and kinship markers; Based on the original field name, original field value, field unit, field format, and lineage marker, the meaning of the field is parsed, and the field correspondence between the original field and the standard field is established; Based on the field correspondence, the original fields are identified by field type and divided into object identifier field, object name field, attribute description field, spatial location field, status record field, time record field, and business association field; Based on the field type identification results, the original field values ​​are standardized in terms of data format, and the encoding format, name format, time format, numerical format, unit of measurement format and status expression format in data from different sources are converted into a unified field expression; Correct any abnormal formats and abnormal field states that exist during the data format unification process. Abnormal formats and abnormal field states include field values ​​that are empty, field values ​​that contain redundant characters, inconsistent field units, inconsistent time expressions, inconsistent status expressions, and inconsistent encoding lengths. The unified field representation, field correspondence, field type identification results, and lineage markers are linked to form a field parsing data set; In the specific implementation process, the collected records in the original data set of bloodline anchoring are decomposed according to record identifier, field identifier and bloodline mark, and the original field name, original field value, field unit, field format, field source and bloodline mark are extracted and written into the field parsing cache table, so that each field value retains the corresponding source path, field origin and version relationship; When parsing field meanings, name matching is performed between the original field name and the standard field name, and the business meaning of the field is determined by combining the field unit, field format, field source, and lineage marker. For fields with different sources but similar meanings, a field correspondence relationship is established between the original field and the standard field. For fields with the same name but different source objects, the corresponding standard field is distinguished according to the field source and object lineage marker to avoid the incorrect merging of fields from different business objects. When identifying field types, based on the field correspondence and the data form of the original field values, the fields are divided into object identifier fields, object name fields, attribute description fields, spatial location fields, status record fields, time record fields, and business association fields, so that the field parsing data set can accept the corresponding parsing rules according to the field type; When the data format is unified, the encoding format is length-padding, case-consistent, and separator-cleaning are performed; the name format is space-removing, full-width / half-width conversion is performed, and abbreviations are retained; the time format is converted to a unified time expression; the numerical format is removed from non-numerical characters and retains effective precision; the unit of measurement format is converted or marked according to standard units; and the status expression format is converted to a unified status enumeration value. When correcting abnormal formats and abnormal field states, generate null status flags for null fields, perform character cleanup for fields containing redundant characters, perform unit conversion for fields with inconsistent units, perform time format conversion for fields with inconsistent time expressions, perform state enumeration mapping for fields with inconsistent state expressions, and mark fields with inconsistent encoding lengths as padded or truncated, while retaining the field values ​​before correction and the correction rules. The unified field expression, field correspondence, field type identification results, and lineage markers are associated and encapsulated according to the collected record identifier to form a field parsing data set, which can simultaneously reflect the standard meaning of the field, the field type, the unified expression, and the source lineage.

[0028] In this embodiment, step three specifically includes: Based on the object identifier field, object name field, attribute description field and business association field in the field parsing data set, the field parsing data set is split into entities to obtain the entity splitting results; Based on the entity segmentation results, extract the object name, abbreviation, former name, ledger record name, and business description name from data from different sources, and perform alias recognition on the object name, abbreviation, former name, ledger record name, and business description name to obtain the alias recognition results; Based on the alias recognition results, establish a set of candidate representations of the same business object in data from different sources; Perform attribute normalization on the attribute description fields in the candidate expression set, converting attribute names, attribute units, attribute formats, and attribute value ranges into unified attribute expressions; By associating the candidate expression set with the unified attribute expression, candidate expressions pointing to the same business object in data from different sources are identified, and candidate master data objects are formed. In the specific implementation process, entity segmentation is based primarily on the object identifier field, with the object name field, attribute description field, and business association field serving as auxiliary segmentation criteria. When the field parsing data set contains equipment number, material code, personnel number, or ledger number, records with the same object identifier field are grouped into the same entity fragment. When the object identifier field is missing, has an inconsistent format, or has undergone historical changes, auxiliary merging is performed by combining name expression, key attributes, and business association pointers to obtain the entity segmentation result. During alias identification, the object names, abbreviations, former names, ledger record names, and business description names in the entity segmentation results are normalized. The normalization process includes removing redundant spaces, unifying full-width and half-width characters, cleaning up invalid symbols, unifying English capitalization, and retaining the abbreviation for coal business. The normalized name expressions are then compared based on character similarity, term overlap, abbreviation expansion relationship, and source lineage relationship. Different name expressions pointing to the same business object are grouped into the same alias identification result. The candidate expression set is established based on the alias recognition results. It unifies the object name, abbreviation, former name, ledger record name and business description name of the same business object in different source data, and retains the field source, source lineage marker, version lineage marker and object lineage marker corresponding to each candidate expression; Attribute normalization performs unified processing on the attribute description fields in the candidate expression set for attribute names, attribute units, attribute formats, and attribute value ranges; attribute name normalization maps synonymous attribute names to unified attribute names based on the standard attribute dictionary; attribute unit normalization splits the numerical part and unit part and determines the target unit based on the field type; attribute format normalization converts numerical, text, status, time, and encoded attributes into a unified expression form; attribute value range normalization determines whether the attribute value meets the allowed range, enumerated range, or business constraint range based on the object type and the standard attribute dictionary, and generates corresponding tags for out-of-bounds values, abnormal states, location-to-be-verified values, and encoded abnormal values. The normalized attribute description fields are encapsulated into a unified attribute expression, and the unified attribute expression is associated with the candidate expression set according to object identifier, object name, source lineage, version lineage and business relationship, forming a candidate master data object containing candidate name expression, unified attribute expression, source relationship, lineage relationship, exception flag and business relationship.

[0029] In this embodiment, step four specifically includes: Based on the object type, object name, alias expression, field attributes, source association, lineage relationship, and business association of candidate master data objects, a heterogeneous master data relationship graph is constructed through a heterogeneous master data relationship construction layer. Candidate master data objects are used as object nodes, and source systems, field attributes, business ledgers, production processes, spatial locations, responsible entities, and upstream and downstream business objects are used as association nodes. The source relationship, attribute relationship, lineage relationship, ledger record relationship, scenario relationship, location relationship, responsibility relationship, and upstream and downstream relationship between object nodes and association nodes are used as heterogeneous relationship edges. The heterogeneous relationship edges are marked with relationship direction and relationship strength to obtain the heterogeneous master data graph representation. The heterogeneous graph representation of master data is used as the input type semantic embedding layer. The object nodes, associated nodes and heterogeneous relationship edges are type-encoded respectively. The object name, alias expression, attribute expression, source expression, lineage expression and business description expression of the candidate master data objects are embedded. The type encoding results and the embedding results are concatenated, fused and linearly mapped to obtain the type semantic representation of the candidate master data objects. Type semantic representation is input into the anchor point counter-evidence gating propagation layer. The supporting and exclusive relationships between candidate master data objects are identified through the business anchor point counter-evidence gating mechanism. The relationships corresponding to object identity consistency, spatial location consistency, business scenario consistency, ledger record consistency, responsible entity consistency, and upstream and downstream relationship consistency are determined as business anchor point relationships. The relationships corresponding to object type conflict, location chain conflict, business scenario conflict, key field conflict, lineage conflict, and state sequence conflict are determined as business counter-evidence relationships. Anchor point enhancement weights are generated based on business anchor point relationships, and disproving weights are generated based on business disproving relationships. Anchor point enhancement weights and disproving weights are then gated and fused using a business anchor point disproving gating mechanism to obtain gating adjustment weights. In the process of improving the heterogeneous attention propagation of the HGT master data recognition model, type-related query representations, key representations, and value representations are generated based on the node type of the object node and the relationship type of the heterogeneous relation edge. The node attention weight and relation edge propagation weight are modified by gating adjustment weights, so that the semantic propagation of candidate master data objects with business anchor relationships is enhanced, and the semantic propagation of candidate master data objects with business counter-evidence relationships is reduced, thus obtaining the gating propagation semantic representation. The gating propagation semantic representation is input to the matching evaluation output layer, and the name similarity, attribute overlap, business scenario consistency, upstream and downstream correlation, lineage consistency and counter-evidence conflict degree between candidate master data objects are jointly evaluated to generate object matching confidence. The matching level, the judgment result of the same entity, and the counter-evidence suppression mark are determined based on the object matching confidence, and the object matching confidence, matching level, judgment result of the same entity, and counter-evidence suppression mark are used as the object matching evaluation results. In the specific implementation process, the heterogeneous master data relationship construction layer establishes object nodes with candidate master data objects as the center. Each candidate master data object carries object type, object name, alias expression, unified attribute expression, source lineage mark, version lineage mark, and object lineage mark. Source system, field attribute, business ledger, production process, spatial location, responsible entity, and upstream and downstream business objects serve as association nodes to represent data source, field attribute, business record, production scenario, spatial level, management responsibility, and business chain relationship. Heterogeneous relationship edges are generated according to the actual business relationships between candidate master data objects and associated nodes, including source relationship edges, attribute relationship edges, lineage relationship edges, ledger record relationship edges, scenario relationship edges, location relationship edges, responsibility relationship edges, and upstream and downstream relationship edges; the relationship direction marker is determined based on the relationship initiating object and the relationship pointing to object, and the relationship strength marker is determined based on the degree of field consistency, source authority, frequency of ledger occurrence, number of business associations, and degree of lineage consistency, thus obtaining the master data heterogeneous graph representation; The type semantic embedding layer sets type codes for object nodes, associated nodes, and heterogeneous relationship edges, and converts object type, source type, field type, scenario type, location type, responsibility type, and relationship type into dense vectors. The name expression, attribute expression, source expression, lineage expression, and business description expression of the candidate master data object are encoded separately and then concatenated and fused. During the concatenation and fusion, the first and last parts are concatenated in the order of object type code, source type code, field type code, relationship type code, name expression vector, attribute expression vector, source expression vector, lineage expression vector, and business description expression vector to form a concatenated semantic vector. When different vectors have inconsistent dimensions, vectors with insufficient dimensions are expanded to the target dimension by zero-value padding, and vectors with excessive dimensions are reduced to the target dimension by truncation while retaining the main expression dimension. During linear mapping, the concatenated semantic vector is multiplied by the linear mapping weight matrix and the bias vector is superimposed to obtain the mapped semantic vector. The mapped semantic vector is then subjected to non-linear activation, missing mask correction, and normalization to transform the combination relationship between source, field, lineage, and business description into a scale-uniform type semantic representation. Normalization is performed on the activation semantic vector. The feature mean and feature standard deviation are calculated. The feature mean is subtracted from each feature value and divided by the sum of the feature standard deviation and the stability coefficient. When the feature standard deviation is zero, the normalized semantic vector is set to a zero vector and bias compensation is performed through learnable translation coefficients. The normalized semantic vector is then scaled back through learnable scaling coefficients and learnable translation coefficients to obtain the type semantic representation of the candidate master data object. The anchor point disproving gating propagation layer takes type semantic representation as input, identifies supporting and exclusive relationships between candidate master data objects, extracts relationships that support the judgment of the same entity as business anchor point relationships, and extracts relationships that should not be merged as business disproving relationships; the anchor point enhancement weight and disproving suppression weight are generated by the business anchor point disproving gating mechanism and serve as the adjustment basis for heterogeneous attention propagation. In the heterogeneous attention propagation process, the improved HGT master data recognition model generates type-related query representations, key representations, and value representations based on the node type of the object node and the relationship type of the heterogeneous relation edges. First, the basic attention weights are calculated based on the query representations and key representations, and then the basic attention weights are corrected using gating adjustment weights. During the correction, the basic attention weights are multiplied by the gating adjustment weights, and the correction results on the adjacent edges of the same object node are normalized to obtain the corrected node attention weights. The gated propagation semantic representation is obtained by weighted aggregation of the value representations of adjacent nodes using the corrected node attention weights and relation edge propagation weights. During weighted aggregation, the value representations of adjacent nodes are multiplied by the corresponding corrected node attention weights and relation edge propagation weights, and the weighted value representations are summed to obtain the gated aggregated representation. The gated aggregated representation is then added to the original type semantic representation of the candidate master data object using a residual join method, and the addition result is normalized to obtain the gated propagation semantic representation. The matching evaluation output layer calculates name similarity, attribute overlap, business scenario consistency, upstream and downstream relevance, lineage consistency, and counter-evidence conflict degree based on gated propagation semantic representation. When generating object matching confidence, name similarity is weighted at 0.20, attribute overlap at 0.20, business scenario consistency at 0.18, upstream and downstream relevance at 0.17, and lineage consistency at 0.25, respectively, and then weighted and summed to obtain a positive matching score. The counter-evidence conflict degree is then converted into a suppression score with a suppression coefficient of 0.30, and the suppression score is subtracted from the positive matching score to obtain the object matching confidence. When the object matching confidence score is greater than or equal to 0.85, the corresponding candidate master data objects are determined to be the same entity; when the object matching confidence score is greater than or equal to 0.60 and less than 0.85, the corresponding candidate master data objects are determined to be a relationship to be confirmed; when the object matching confidence score is less than 0.60 or there is a strong counter-evidence suppression flag, the corresponding candidate master data objects are determined to be different entities; the matching level, the same entity judgment result, and the counter-evidence suppression flag are determined based on the object matching confidence score and encapsulated as the object matching evaluation result.

[0030] In this embodiment, the business anchor counter-evidence gating mechanism is specifically as follows: Based on the object identifier, object name, alias expression and key attributes of the candidate master data objects, extract the object identity anchor points, and determine whether there is object identity consistency among the candidate master data objects based on the object identity anchor points; Based on the spatial location field of the candidate master data object, the location record of the source system, and the location record of the business ledger, extract the spatial location anchor point, and determine whether there is spatial location consistency among the candidate master data objects based on the spatial location anchor point; Based on the production process, maintenance records, requisition records, job responsibility records, and business ledger records associated with the candidate master data objects, extract business scenario anchor points, and determine whether there is business scenario consistency among the candidate master data objects based on the business scenario anchor points. Based on the equipment subordination, production process chain, material requisition, personnel responsibility, and ledger record relationships among candidate master data objects, upstream and downstream relationship anchor points are extracted, and the consistency of upstream and downstream relationships among candidate master data objects is determined based on the upstream and downstream relationship anchor points. The consistency of object identity, spatial location, business scenario, and upstream and downstream relationship are integrated to generate business anchor relationships, and the anchor enhancement weight is calculated based on the number, strength, and source authority of the business anchor relationships. Based on the differences in object type, location chain, business scenario, key field, lineage, and state sequence among candidate master data objects, business counter-evidence relationships are extracted, and counter-evidence suppression weights are calculated based on the number of conflicts, conflict level, and source authority of the business counter-evidence relationships. Anchor point enhancement weights and counter-evidence suppression weights are fused to generate gating adjustment weights. These gating adjustment weights are used to increase the semantic propagation strength between candidate master data objects with business anchor point relationships and decrease the semantic propagation strength between candidate master data objects with business counter-evidence relationships. The node attention weights and relation edge propagation weights in the improved HGT master data recognition model are adjusted according to the gating adjustment weights to generate a gating propagation semantic representation; In the specific implementation process, the business anchor counter-evidence gating mechanism takes candidate master data object pairs as the processing unit, reads the object identifier, object name, alias expression, key attributes, spatial location field, source system location record, business ledger location record, production process, maintenance record, requisition record, job responsibility record and upstream and downstream business relationship of any two candidate master data objects, and organizes them into object identity information, spatial location information, business scenario information and upstream and downstream relationship information; Object identity anchors are extracted based on object identifiers, object names, aliases, and key attributes; spatial location anchors are extracted based on the hierarchical matching results of mines, mining areas, working faces, roadways, warehouses, and installation locations; business scenario anchors are extracted based on the co-occurrence of production processes, maintenance records, requisition records, job responsibility records, and business ledger records; upstream and downstream relationship anchors are extracted based on equipment subordination relationships, production process chain relationships, material requisition relationships, personnel responsibility relationships, and ledger record relationships. When generating business anchor relationships, object identity consistency, spatial location consistency, business scenario consistency, and upstream / downstream relationship consistency are converted into anchor scores. These scores are then assigned in a tiered manner: strong consistency is assigned 1.00, weak consistency is assigned 0.70, pending confirmation is assigned 0.40, and inconsistency is assigned 0. After standardizing the anchor scores, the anchor scores for object identity consistency are weighted at 0.35, spatial location consistency at 0.25, business scenario consistency at 0.20, and upstream / downstream relationship consistency at 0.20, and then weighted and summed to obtain the business anchor relationship strength. This ensures that object identity and key business relationships contribute more significantly to the strength of business anchor relationships. When calculating the anchor point enhancement weight, the strength of business anchor point relationships, the number of business anchor point relationships, and the authority of the source are converted to the same numerical range. The strength of business anchor point relationships is then weighted and summed with a fusion weight of 0.50, the number of business anchor point relationships with a fusion weight of 0.20, and the authority of the source with a fusion weight of 0.30 to obtain the anchor point enhancement weight. The strength of business anchor point relationships is used to characterize the consistency of anchor points, the number of business anchor point relationships is used to characterize the coverage of supporting evidence, and the authority of the source is used to characterize the credibility of the source of anchor point information. Business-related counter-evidence relationships are extracted based on differences in object type, position chain, business scenario, key field, lineage, and state sequence. When calculating the counter-evidence suppression weight, the number of conflicts, conflict level, and source authority are first converted to the same numerical range. Then, the number of conflicts is weighted and summed using a counter-evidence fusion weight of 0.25, the conflict level using a counter-evidence fusion weight of 0.45, and the source authority using a counter-evidence fusion weight of 0.30 to obtain the counter-evidence suppression weight. The number of conflicts is used to characterize the coverage of evidence opposing the merger, the conflict level is used to characterize the repulsion strength of the counter-evidence relationship, and the source authority is used to characterize the credibility of the counter-evidence information, enabling key field conflicts, strong position chain conflicts, and high-authority source conflicts to exert stronger suppression on semantic propagation. When generating the gating adjustment weights, the anchor point enhancement weight is used as a positive adjustment term, and the proof-of-contrast suppression weight is used as a negative adjustment term. The difference between the positive and negative adjustment terms is compressed using the Sigmoid function to obtain the gating adjustment weights within the defined interval. The gating adjustment weights increase with the increase of the anchor point enhancement weights and decrease with the increase of the proof-of-contrast suppression weights. When the anchor point enhancement weight is high and the proof-of-contrast suppression weight is low, the gating adjustment weights tend to enhance propagation; when the anchor point enhancement weight is low and the proof-of-contrast suppression weight is high, the gating adjustment weights tend to suppress propagation; when both anchor points and proof-of-contrast exist simultaneously, the gating adjustment weights are reduced according to the proof-of-contrast priority principle. The gating adjustment weights are applied to the basic node attention weights and source relationship edges, attribute relationship edges, lineage relationship edges, ledger record relationship edges, scene relationship edges, location relationship edges, responsibility relationship edges, and upstream and downstream relationship edges between candidate master data objects. When a relationship edge carries a business anchor relationship, the relationship edge propagation weight is increased; when a relationship edge carries a business counter-evidence relationship, the relationship edge propagation weight is decreased. By generating a gating propagation semantic representation in the above manner, the improved HGT master data recognition model enhances credible matching relationships and suppresses erroneous matching relationships during the object matching evaluation process.

[0031] In this embodiment, step five specifically includes: Based on the object matching evaluation results, extract the object matching confidence, matching level, same entity judgment result and counter-evidence suppression flag between candidate master data objects; Candidate master data objects are used as clustering nodes, and object matching confidence, matching level, same entity judgment result and counter-evidence suppression flag are used as clustering edge attributes to construct a candidate master data association graph. Based on the clustering edge attributes between candidate master data objects in the candidate master data association graph, the candidate master data objects are initially grouped to form master data candidate clusters; Relation propagation clustering is performed on candidate master data objects within the candidate master data clusters. The object matching confidence, matching level and counter-evidence suppression labels are propagated and updated in the candidate master data association graph to obtain master data clusters. Based on the master data clusters, the object identifier, object name, alias expression, attribute fields and business association fields in the candidate master data objects are merged to form merged master data objects; In the specific implementation process, the object matching confidence, matching level, same entity judgment result, and counter-evidence suppression mark between candidate master data objects are first read from the object matching evaluation results, and the candidate master data objects are written into the candidate master data association graph as cluster nodes. When two candidate master data objects have the same entity judgment result or a relationship to be confirmed, a cluster edge is established between the two cluster nodes. When two candidate master data objects are marked as different entities and have a strong counter-evidence suppression mark, no cluster edge in the entity merging direction is established, and the relationship between the two cluster nodes is recorded as an exclusion relationship. During initial grouping, the candidate master data association graph is traversed according to the clustering edge attributes. Candidate master data objects with consistent judgment results for the same entity, high confidence levels for object matching, and no strong counter-evidence suppression markers are grouped into the same master data candidate cluster. For candidate master data objects with medium confidence levels or weak counter-evidence suppression markers, a pending confirmation and merging state is generated. For candidate master data objects with strong counter-evidence suppression markers, they are either kept in independent grouping or assigned to an exclusion relationship set. During relational propagation clustering, the candidate master data clusters are used as the local propagation range. The object matching confidence, matching level and disproving suppression label in the candidate master data association graph are propagated and updated. High-confidence clustering edges enhance the merging tendency, and disproving suppression labels enhance the exclusivity relationship. When the same candidate master data object is connected to multiple candidate master data clusters and a disproving suppression label exists, the exclusivity relationship is retained first to avoid cross-cluster erroneous merging. After the propagation update is completed, a consistency check is performed on the candidate master data objects within the master data candidate cluster to check for incompatible conflicts in object type, key fields, spatial location, state sequence, and lineage. The master data candidate cluster that passes the consistency check is determined as the master data cluster. Candidate master data objects with strong conflicts are split out of the master data candidate cluster and re-form independent cluster nodes. When merging entities, object identifiers, object names, aliases, attribute fields, and business-related fields are merged by master data clusters. When multiple candidate values ​​exist for the same type of field, conflicting values ​​are not directly deleted during the entity merging stage. Instead, conflicting fields, field sources, lineage markers, and candidate field values ​​are retained together, enabling the merged master data object to handle conflicting field decision-making.

[0032] In this embodiment, step six specifically includes: Identify fields in the merged master data object that have different sources, the same meaning, and inconsistent values, and form a set of conflicting fields; A multi-dimensional credibility evaluation is conducted on the candidate field values ​​corresponding to each conflicting field item in the conflicting field set. The evaluation dimensions include the source authority determined by the data source, business affiliation, field lineage and source system management authority; the time freshness determined by the update time, business effective time, version status and historical change records; the field completeness determined by the field completeness, format standardization, unit consistency and value validity; and the historical consistency determined by the consistent occurrence in historical data, business ledgers and related objects. The credibility score of a field is generated by integrating factors such as source authority, time freshness, field completeness, and historical consistency. Based on the field credibility score, the target field value, alternative field values, and conflict trace information of the conflicting field are determined to form the conflicting field decision result; In the specific implementation process, conflict field identification takes the merged master data object as the processing unit, reads synonymous field items from different sources in the merged master data object, and writes field items with the same meaning but different values ​​into the conflict field set; field items with the same value but different sources are not treated as conflict fields, but are retained as the basis for historical consistency evaluation. In multidimensional credibility evaluation, source authority, time freshness, field completeness, and historical consistency are calculated for the candidate field values ​​corresponding to each conflicting field item. Source authority is determined based on data source, business affiliation, field lineage, and source system management permissions. Time freshness is determined based on update time, business effective time, version status, and historical change records. Field completeness is determined based on field integrity, format standardization, unit consistency, and value validity. Historical consistency is determined based on the consistent occurrence of candidate field values ​​in historical data, business ledgers, and related objects. When generating field credibility scores, source authority, time freshness, field completeness, and historical consistency are converted to a unified numerical range, and fusion weights are set according to field type. For object code, object type, and key location fields, source authority is weighted at 0.35, time freshness at 0.15, field completeness at 0.35, and historical consistency at 0.15, and the resulting score is obtained by weighted summation. For status record, responsible entity, and business activation fields, source authority is weighted at 0.25, time freshness at 0.35, and field completeness at 0.20. The field credibility score is obtained by weighting and summing the fusion weights and historical consistency using a fusion weight of 0.20. For the name, alias, and business description fields, the field credibility score is obtained by weighting and summing the source authority using a fusion weight of 0.20, time freshness using a fusion weight of 0.10, field completeness using a fusion weight of 0.30, and historical consistency using a fusion weight of 0.40. For other fields, the field credibility score is obtained by weighting and summing the source authority using a fusion weight of 0.30, time freshness using a fusion weight of 0.25, field completeness using a fusion weight of 0.25, and historical consistency using a fusion weight of 0.20. When a certain evaluation dimension is missing, the corresponding fusion weight of that evaluation dimension is reduced, and the fusion weight is renormalized according to the remaining evaluation dimensions to ensure that the field credibility score will not be abnormally offset due to a single missing dimension. When making decisions about conflicting fields, the field credibility scores of each candidate field value under the same conflicting field item are sorted. When the highest field credibility score is not lower than 0.80 and no strong conflict flag is triggered, the candidate field value corresponding to the highest score is determined as the target field value, and the remaining candidate field values ​​are retained as alternative field values. When the difference between the highest field credibility score and the second highest field credibility score is less than 0.05, the conflicting field is marked as a field to be confirmed and multiple candidate field values ​​are retained. When the field credibility scores of all candidate field values ​​are lower than 0.60, the conflicting field is marked as a low-confidence conflicting field and a manual review flag is retained. Conflict trace information includes conflict field name, candidate field value, target field value, alternative field value, field credibility score, evaluation dimension result, data source, lineage marker, decision time and decision basis; the target field value, alternative field value and conflict trace information are encapsulated into conflict field decision result.

[0033] In this embodiment, step seven specifically includes: Based on the conflict field decision results, determine the target field value in the merged master data object, and retain the source association, lineage marker and conflict decision basis corresponding to the target field value; A unified master data code is generated based on the object category, business level, source relationship, and coding order of the merged master data object, and the unified master data code is bound to the merged master data object; Based on the field parsing data set, the merged master data object, and the decision results of conflicting fields, establish the field mapping relationship between the original fields, parsed fields, merged fields, and standard fields; Based on the field mapping relationship, perform integrity verification, uniqueness verification, format verification, reference relationship verification, and status validity verification on the merged master data object to obtain the master data object that passes the verification. Standardize the master data objects that pass the verification by encapsulating the unified master data code, target field values, field mapping relationships and verification results to form standard master data; In the specific implementation process, for each merged master data object, the target field value, alternative field value, field credibility score and conflict trace information in the conflict field decision result are read. The target field value is written into the standard field position corresponding to the merged master data object, the alternative field value is retained in the candidate field record, and the source association, lineage mark and conflict decision basis corresponding to the target field value are synchronously bound to the standard field. The unified master data code is generated based on the object category, business level, source relationship, and coding order of the merged master data object. During code generation, the object category is converted into an object category code, the business level is converted into a level code, and the source relationship is converted into a source code. Then, the codes are concatenated in the order of object category code, level code, source code, sequence code, and check code. When the merged master data object already has a historical unified master data code, it is determined whether it belongs to the continuation of the same object based on the version lineage mark and master data change record. If it belongs to the continuation of the same object, the original unified master data code is retained and the version status is updated. The field mapping relationship is established according to the chain relationship from the original field to the parsed field, from the parsed field to the merged field, and from the merged field to the standard field. For each mapping relationship, the field source, mapping rules, value changes, conflict decision basis and lineage mark are recorded, so that the field mapping relationship can reflect the transformation path of the field from the original state to the standard state. Rule validation includes integrity validation, uniqueness validation, format validation, reference relationship validation, and status validity validation. Integrity validation is used to determine whether necessary fields are complete; uniqueness validation is used to determine whether the unified master data code and key object identifier are duplicated; format validation is used to determine whether the field format meets the standard requirements; reference relationship validation is used to determine whether there are valid references to the responsible entity, spatial location, upstream and downstream business objects, and business ledger records; and status validity validation is used to determine whether there are contradictions between the running status, the disabled status, the version status, and the release status. When performing standardized governance on the master data objects that have passed the verification, the unified master data code, target field value, field mapping relationship, verification result, lineage mark, alternative field value and conflict trace information are encapsulated into standard master data. The standard master data includes at least the master data code, master data name, object category, standard field set, alias set, business association set, lineage information set, field mapping relationship, verification status and version status.

[0034] In this embodiment, step eight specifically includes: The standard master data is published to the data management platform of the coal enterprise, and the corresponding publication version, publication time, publication status and publication scope are generated. Identify the change types based on master data change information: new objects, field changes, object merging, object splitting, object deactivation, and version rollback; Extract the corresponding changed object, changed field, field value before change, field value after change, and change source based on the change type to form a master data change record; Based on the master data change record, the object matching evaluation results are fed back and corrected, and the object matching confidence, the same entity judgment result and the counter-evidence suppression mark between candidate master data objects are updated. Based on the object matching evaluation results after feedback correction, the standard master data is continuously updated, and the version status, lineage mark and release status of the standard master data are updated synchronously to form a closed loop of master data standardization governance. In the specific implementation process, before the standard master data is released, the unified master data code, standard field set, field mapping relationship, verification status and version status are read. For the standard master data that passes the verification and has a valid version status, a release task is generated. The release task records the release version, release time, release status and release scope. The release scope is used to indicate the business systems, data interfaces or governance directories that the standard master data can be synchronized to. When standard master data is published to the data management platform of coal enterprises, it is written into the corresponding master data directory according to the object category, the unified master data code is used as the main index, and the master data name, standard field set, alias set, business association set and lineage information set are used as the published content. Field mapping relationship and conflict trace information are used as the traceability content. For data that needs to be synchronized to the business system, a distribution record is generated according to the publication scope, and the distribution object, distribution interface, distribution status and distribution feedback are recorded. Master data change information comes from newly collected data, business system modification records, manual review records, release feedback records, and historical version records; based on the change information, new objects, field changes, object merging, object splitting, object deactivation, and version rollback are identified, and master data change records are generated. When correcting the object matching evaluation results, update the object matching confidence, same entity judgment result, and counter-evidence suppression flag between candidate master data objects based on the master data change record; when object merging is confirmed to be valid, increase the corresponding object matching confidence; when object splitting or erroneous merging is confirmed, decrease the corresponding object matching confidence and strengthen the counter-evidence suppression flag; when field changes have a clear version succession relationship, update the lineage consistency and status sequence judgment to avoid normal business changes being misjudged as conflicts; When standard master data is continuously updated, the system determines whether to re-execute association clustering, entity merging, and conflict field decisions based on the object matching evaluation results after feedback correction. When only field changes occur and the object matching relationship remains unchanged, the corresponding standard field values, version status, and lineage markers are directly updated. When object merging or object splitting causes changes in object boundaries, the merged master data object or the split standard master data is regenerated, and the field mapping relationship is re-established. When an object is deactivated, the release status is updated, and when a version is rolled back, the target historical version is restored and rollback trace information is retained. By publishing standard master data, generating master data change records, providing feedback and correction on object matching evaluation results, continuously updating standard master data, updating version status, updating lineage markers, and updating publication status, a closed loop of master data standardization governance is formed, from multi-source data collection to standard master data publication and then to change feedback.

[0035] Example 1: To verify the feasibility of this invention in practice, it was applied to the data governance scenario of a coal enterprise. This enterprise simultaneously runs an equipment management system, a production scheduling system, a personnel management system, a materials management system, and multiple business ledgers. The data suffers from problems such as inconsistent equipment numbers, abbreviations and former names for equipment, inconsistent material specifications, delayed personnel job change records, and inconsistencies between maintenance ledgers and equipment system records. Traditional manual rule-based governance methods rely primarily on field name matching, fixed source priority, and manual review. These methods struggle to accurately distinguish data with similar names but different business objects, and easily overlook data with different names but actually pointing to the same business object. This leads to duplicate standard master data, insufficient decision-making basis for conflicting fields, and low reliability when sharing data across systems.

[0036] During implementation, a multi-source data acquisition index is first established for equipment system data, production system data, personnel system data, material system data, and business ledger data. Based on the source domain lineage anchoring acquisition mechanism, the acquired records are bound to source lineage, field lineage, version lineage, and object lineage to form a lineage anchoring raw data set. Next, the lineage anchoring raw data set undergoes field parsing, field type identification, format unification, and abnormal format correction to form a field parsing data set. Then, through entity segmentation, alias identification, and attribute normalization, candidate master data objects are formed. The candidate master data objects are input into the improved HGT master data identification model to construct a heterogeneous master data relationship diagram containing candidate master data objects, source systems, field attributes, business ledgers, production processes, spatial locations, responsible entities, and upstream and downstream business objects. The semantic propagation of object identity consistency, spatial location consistency, business scenario consistency, and upstream and downstream relationship consistency is enhanced through a business anchor counter-verification gating mechanism, suppressing the propagation of errors such as object type conflicts, key field conflicts, location chain conflicts, and state sequence conflicts, and finally generating object matching evaluation results. Based on the object matching evaluation results, association clustering, entity merging, multi-source credibility decision for conflicting fields, unified coding, field mapping, and rule verification are performed to generate standard master data and publish it to the data management platform.

[0037] To verify the practical effectiveness of the method of the present invention, a total of 10,000 records of equipment, personnel, materials, and business ledger objects from the enterprise's multi-source system were selected as governance samples. After manual review, 2,860 records were found to contain duplicate expressions, aliases, field conflicts, or version changes across systems. The method of the present invention was compared with traditional manual rule governance methods and name similarity matching governance methods. The statistical results are shown in Table 1 below.

[0038] Table 1. Comparison of the effects of multi-source master data identification and standardized governance in coal enterprises.

[0039] As shown in Table 1, the method of this invention significantly outperforms traditional manual rule-based governance and name similarity matching governance methods in terms of candidate master data object identification accuracy, automatic decision-making accuracy for conflict fields, traceability coverage of field sources, and pass rate of standard master data verification. The candidate master data object identification accuracy increased from 84.60% in the traditional manual rule-based governance method to 96.72%, indicating that this invention can more accurately identify data expressions pointing to the same business object across systems and ledgers. The rate of missed merging of the same object decreased from 9.80% to 1.92%, and the rate of incorrect merging of different objects decreased from 5.35% to 0.86%, indicating that this invention can simultaneously reduce the problems of missed identification of heteronymous objects and incorrect merging of heteronymous objects. The automatic decision-making accuracy for conflict fields increased from 78.40% to 94.58%, and the proportion of manual review intervention decreased from 42.30% to 12.45%, indicating that this invention can reduce reliance on human experience judgment.

[0040] Among the above metrics, the candidate master data object identification accuracy reflects whether business objects in multi-source data can be correctly identified; the same object omission merging rate reflects whether data with different names and codes but actually pointing to the same object are omitted; the different object mismerging rate reflects whether data with similar names but different business meanings are incorrectly merged; the conflict field automatic decision accuracy reflects the reliability of target field value selection when the values ​​of multiple source fields are inconsistent; the field source traceability coverage reflects whether standard master data can be traced back to the source system, field lineage, version lineage, and object lineage; and the master data change feedback update accuracy reflects the continuous maintenance capability of standard master data after business changes.

[0041] The reason why this invention achieves the above-mentioned effects is that the source domain lineage anchoring acquisition mechanism preserves the source, field, version, and object ownership relationship of multi-source data before entering the governance process, providing a traceable basis for field parsing, object identification, and conflict decision-making; the improved HGT master data identification model incorporates the source system, field attributes, business ledger, production process, spatial location, responsible entity, and upstream and downstream business objects into the heterogeneous master data relationship graph, so that object matching evaluation no longer depends on a single name or a single field; the business anchor counter-evidence gating mechanism further enhances the semantic propagation between candidate master data objects with consistent object identity, spatial location, business scenario, and upstream and downstream relationships, and suppresses the error propagation of conflicts in object type, key fields, location chain, and state sequence, thereby improving the reliability of object matching evaluation results.

[0042] As can be seen from this embodiment, the present invention can achieve accurate identification of multi-source master data objects, reliable decision-making on conflicting fields, and continuous updating of standard master data in the complex environment of coal enterprises where multiple systems, multiple ledgers, and multiple fields coexist. At the same time, by forming a closed loop of master data standardization governance through unified coding, field mapping, rule verification, and feedback correction, it improves the accuracy, traceability, automation, and continuous maintenance capabilities of master data governance in coal enterprises, and has good engineering application value.

[0043] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for multi-source data master data identification and standardized governance for coal enterprises, characterized in that, Includes the following steps: Step 1: Establish a multi-source data collection index, collect and process multi-source data based on the source domain lineage anchoring collection mechanism, and perform source association and lineage marking on the collected multi-source data to form a lineage anchoring raw data set; Step 2: Perform field parsing, field type identification, data format unification, and abnormal format correction on the original bloodline anchoring data set to form a field parsing data set; Step 3: Perform entity segmentation, alias recognition, and attribute normalization on the field parsing data set to identify candidate expressions pointing to the same business object from different data sources and form candidate master data objects; Step 4: Based on the improved HGT master data identification model, heterogeneous relationship identification and object matching evaluation are performed on candidate master data objects to form object matching evaluation results. The improved HGT master data identification model includes a heterogeneous master data relationship construction layer, a type semantic embedding layer, an anchor point counter-evidence gating propagation layer, and a matching evaluation output layer. The anchor point counter-evidence gating propagation layer embeds a business anchor point counter-evidence gating mechanism. Step 5: Based on the object matching evaluation results, perform association clustering and entity merging on the candidate master data objects to form merged master data objects; Step 6: Perform multi-source credibility decision on conflicting fields in the merged master data object to generate conflicting field decision results; Step 7: Based on the decision results of conflicting fields, perform unified encoding, field mapping, rule validation, and standardization governance on the merged master data objects to form standard master data; Step 8: Publish and continuously update the standard master data, and provide feedback and correction to the object matching assessment results and standard master data based on the master data change information, forming a closed loop for master data standardization governance.

2. The method for multi-source data master data identification and standardized governance for coal enterprises according to claim 1, characterized in that, Step one specifically involves: Identify the data sources and data objects to be addressed in coal enterprises, and establish a multi-source data collection index. The multi-source data includes: equipment system data, production system data, personnel system data, material system data, and business ledger data. Based on the multi-source data collection index, determine the data source, collection interface, business form, and collection batch corresponding to each type of multi-source data, and process the multi-source data to obtain collection records; Based on the source domain lineage anchoring acquisition mechanism, source anchoring identifiers are generated according to data source, acquisition interface, business form and acquisition batch; field anchoring identifiers are generated according to original field name, field unit, field format and field belonging object; version anchoring identifiers are generated according to acquisition batch, update time, business effective status and historical change records; and object anchoring identifiers are generated according to the belonging relationship and business association relationship between data objects. Write the source anchoring identifier, field anchoring identifier, version anchoring identifier, and object anchoring identifier into the corresponding lineage marker field of the collection record, respectively, to form the source lineage marker, field lineage marker, version lineage marker, and object lineage marker; The source lineage marker, field lineage marker, version lineage marker, and object lineage marker are bound to the collection records to form a lineage-anchored original data set.

3. The method for multi-source data master data identification and standardized governance for coal enterprises according to claim 1, characterized in that, Step two specifically involves: The collected records in the original kinship anchoring dataset are decomposed into fields to extract the original field names, original field values, field units, field formats, field sources, and kinship markers; Based on the original field name, original field value, field unit, field format, and lineage marker, the meaning of the field is parsed, and the field correspondence between the original field and the standard field is established; Based on the field correspondence, the original fields are identified by field type and divided into object identifier field, object name field, attribute description field, spatial location field, status record field, time record field and business association field; Based on the field type identification results, the original field values ​​are standardized in terms of data format, and the encoding format, name format, time format, numerical format, unit of measurement format and status expression format in data from different sources are converted into a unified field expression; Correct any abnormal formats and abnormal field states that exist during the data format unification process. Abnormal formats and abnormal field states include field values ​​that are empty, field values ​​that contain redundant characters, inconsistent field units, inconsistent time expressions, inconsistent status expressions, and inconsistent encoding lengths. The unified field representation, field correspondence, field type identification results, and lineage markers are linked to form a field parsing data set.

4. The method for multi-source data master data identification and standardized governance for coal enterprises according to claim 1, characterized in that, Step three specifically involves: Based on the object identifier field, object name field, attribute description field and business association field in the field parsing data set, the field parsing data set is split into entities to obtain the entity splitting results; Based on the entity segmentation results, extract the object name, abbreviation, former name, ledger record name, and business description name from data from different sources, and perform alias recognition on the object name, abbreviation, former name, ledger record name, and business description name to obtain the alias recognition results; Based on the alias recognition results, establish a set of candidate representations of the same business object in data from different sources; Perform attribute normalization on the attribute description fields in the candidate expression set, converting attribute names, attribute units, attribute formats, and attribute value ranges into unified attribute expressions; By associating the candidate expression set with the unified attribute expression, candidate expressions pointing to the same business object in data from different sources are identified, forming candidate master data objects.

5. The method for multi-source data master data identification and standardized governance for coal enterprises according to claim 1, characterized in that, Step four specifically involves: Based on the object type, object name, alias expression, field attributes, source association, lineage relationship, and business association of candidate master data objects, a heterogeneous master data relationship graph is constructed through a heterogeneous master data relationship construction layer. Candidate master data objects are used as object nodes, and source systems, field attributes, business ledgers, production processes, spatial locations, responsible entities, and upstream and downstream business objects are used as association nodes. The source relationship, attribute relationship, lineage relationship, ledger record relationship, scenario relationship, location relationship, responsibility relationship, and upstream and downstream relationship between object nodes and association nodes are used as heterogeneous relationship edges. The heterogeneous relationship edges are marked with relationship direction and relationship strength to obtain the heterogeneous master data graph representation. The heterogeneous graph representation of master data is used as the input type semantic embedding layer. The object nodes, associated nodes and heterogeneous relationship edges are type-encoded respectively. The object name, alias expression, attribute expression, source expression, lineage expression and business description expression of the candidate master data objects are embedded. The type encoding results and the embedding results are concatenated, fused and linearly mapped to obtain the type semantic representation of the candidate master data objects. Type semantic representation is input into the anchor point counter-evidence gating propagation layer. The supporting and exclusive relationships between candidate master data objects are identified through the business anchor point counter-evidence gating mechanism. The relationships corresponding to object identity consistency, spatial location consistency, business scenario consistency, ledger record consistency, responsible entity consistency, and upstream and downstream relationship consistency are determined as business anchor point relationships. The relationships corresponding to object type conflict, location chain conflict, business scenario conflict, key field conflict, lineage conflict, and state sequence conflict are determined as business counter-evidence relationships. Anchor point enhancement weights are generated based on business anchor point relationships, and disprovenance suppression weights are generated based on business disprovenance relationships. Anchor point enhancement weights and disprovenance suppression weights are then gated and fused through a business anchor point disprovenance gating mechanism to obtain gating adjustment weights. In the process of improving the heterogeneous attention propagation of the HGT master data recognition model, type-related query representations, key representations, and value representations are generated based on the node type of the object node and the relationship type of the heterogeneous relation edge. The node attention weight and relation edge propagation weight are modified by gating adjustment weights, so that the semantic propagation of candidate master data objects with business anchor relationships is enhanced, and the semantic propagation of candidate master data objects with business counter-evidence relationships is reduced, thus obtaining the gating propagation semantic representation. The gating propagation semantic representation is input to the matching evaluation output layer, and the name similarity, attribute overlap, business scenario consistency, upstream and downstream correlation, lineage consistency and counter-evidence conflict degree between candidate master data objects are jointly evaluated to generate object matching confidence. The matching level, the same entity judgment result, and the counter-evidence suppression mark are determined based on the object matching confidence level, and the object matching confidence level, matching level, same entity judgment result, and counter-evidence suppression mark are used as the object matching evaluation results.

6. The method for multi-source data master data identification and standardized governance for coal enterprises according to claim 5, characterized in that, The specific business anchor counter-evidence gating mechanism is as follows: Based on the object identifier, object name, alias expression and key attributes of the candidate master data objects, extract the object identity anchor points, and determine whether there is object identity consistency among the candidate master data objects based on the object identity anchor points; Based on the spatial location field of the candidate master data object, the location record of the source system, and the location record of the business ledger, extract the spatial location anchor point, and determine whether there is spatial location consistency among the candidate master data objects based on the spatial location anchor point; Based on the production process, maintenance records, requisition records, job responsibility records and business ledger records associated with the candidate master data objects, extract business scenario anchor points, and determine whether there is business scenario consistency among the candidate master data objects based on the business scenario anchor points. Based on the equipment subordination, production process chain, material requisition, personnel responsibility, and ledger record relationships among candidate master data objects, upstream and downstream relationship anchor points are extracted, and the consistency of upstream and downstream relationships among candidate master data objects is determined based on the upstream and downstream relationship anchor points. The consistency of object identity, spatial location, business scenario, and upstream and downstream relationship are integrated to generate business anchor relationships, and the anchor enhancement weight is calculated based on the number, strength, and source authority of the business anchor relationships. Based on the differences in object type, location chain, business scenario, key field, lineage, and state sequence among candidate master data objects, business counter-evidence relationships are extracted, and counter-evidence suppression weights are calculated based on the number of conflicts, conflict level, and source authority of the business counter-evidence relationships. Anchor point enhancement weights and counter-evidence suppression weights are fused to generate gating adjustment weights. These gating adjustment weights are used to increase the semantic propagation strength between candidate master data objects with business anchor point relationships and decrease the semantic propagation strength between candidate master data objects with business counter-evidence relationships. The node attention weights and relation edge propagation weights in the improved HGT master data recognition model are adjusted according to the gating adjustment weights to generate a gating propagation semantic representation.

7. The method for multi-source data master data identification and standardized governance for coal enterprises according to claim 1, characterized in that, Step five specifically involves: Based on the object matching evaluation results, extract the object matching confidence, matching level, same entity judgment result and counter-evidence suppression mark between candidate master data objects; Candidate master data objects are used as clustering nodes, and object matching confidence, matching level, same entity judgment result and counter-evidence suppression label are used as clustering edge attributes to construct a candidate master data association graph. Based on the clustering edge attributes between candidate master data objects in the candidate master data association graph, the candidate master data objects are initially grouped to form master data candidate clusters; Relation propagation clustering is performed on candidate master data objects within the candidate master data clusters. The object matching confidence, matching level and counter-evidence suppression labels are propagated and updated in the candidate master data association graph to obtain master data clusters. Based on the master data clusters, the object identifier, object name, alias expression, attribute fields and business association fields in the candidate master data objects are merged to form merged master data objects.

8. The method for multi-source data master data identification and standardized governance for coal enterprises according to claim 1, characterized in that, Step six specifically involves: Identify fields in the merged master data object that have different sources, the same field meaning, and inconsistent field values, and form a set of conflicting fields; A multi-dimensional credibility evaluation is conducted on the candidate field values ​​corresponding to each conflicting field item in the conflicting field set. The evaluation dimensions include the source authority determined by the data source, business affiliation, field lineage and source system management authority; the time freshness determined by the update time, business effective time, version status and historical change records; the field completeness determined by the field completeness, format standardization, unit consistency and value validity; and the historical consistency determined by the consistent occurrence in historical data, business ledgers and related objects. The credibility score of a field is generated by integrating factors such as source authority, time freshness, field completeness, and historical consistency. Based on the field credibility score, the target field value, alternative field values, and conflict trace information of the conflict field are determined to form the conflict field decision result.

9. A method for multi-source data master data identification and standardized governance for coal enterprises according to claim 1, characterized in that, Step seven specifically involves: Based on the conflict field decision results, determine the target field value in the merged master data object, and retain the source association, lineage marker and conflict decision basis corresponding to the target field value; A unified master data code is generated based on the object category, business level, source relationship, and coding order of the merged master data object, and the unified master data code is bound to the merged master data object; Based on the field parsing data set, the merged master data object, and the decision results of conflicting fields, establish the field mapping relationship between the original fields, parsed fields, merged fields, and standard fields; Based on the field mapping relationship, perform integrity verification, uniqueness verification, format verification, reference relationship verification, and status validity verification on the merged master data object to obtain the master data object that passes the verification. Standardize the master data objects that pass the verification by encapsulating the unified master data code, target field values, field mapping relationships and verification results to form standard master data.

10. A method for multi-source data master data identification and standardized governance for coal enterprises according to claim 1, characterized in that, Step eight specifically involves: The standard master data is published to the data management platform of the coal enterprise, and the corresponding publication version, publication time, publication status and publication scope are generated. Identify the change types based on master data change information: new objects, field changes, object merging, object splitting, object deactivation, and version rollback; Extract the corresponding changed object, changed field, field value before change, field value after change, and change source based on the change type to form a master data change record; Based on the master data change record, the object matching evaluation results are fed back and corrected, and the object matching confidence, the same entity judgment result and the counter-evidence suppression flag between candidate master data objects are updated. Based on the feedback and correction of the object matching evaluation results, the standard master data is continuously updated, and the version status, lineage markers and release status of the standard master data are updated synchronously to form a closed loop of master data standardization governance.