Data dynamic matching processing method and system for supply chain collaboration
Patent Information
- Application Number
- CN202610046604.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-14
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-01-14
AI Technical Summary
[0004]有鉴于此,本申请实施例提供了一种供应链协同的数据动态匹配处理方法及系统,以解决现有技术存在的多源异构数据难以动态对齐、接口版本演进适配差、匹配结果缺乏可审计追溯的问题
通过接收来自至少两个协同方的数据接入契约,基于数据接入契约建立契约目录并记录契约版本与字段约束,数据接入契约包括同步接口契约与异步事件契约;基于数据接入契约对接入数据执行分层规范化处理,以得到规范化记录;针对协同方数据模型与平台规范模型生成映射计划,并将映射计划编译为可执行的映射拓扑;基于映射拓扑对规范化记录执行字段映射以得到映射后记录;针对待对齐实体从映射后记录中生成候选实体对集合,生成候选实体对集合包括基于标准标识的候选生成、基于规则分桶的候选生成以及基于向量索引近邻召回的候选生成;对候选实体对集合执行多证据匹配评分,以输出通过图一致性校验的匹配链接及置信区间,多证据匹配评分包括基于字段级相似特征的概率评分、基于向量相似度的语义证据融合以及基于关系约束的图一致性校验;基于匹配链接生成匹配合约,并为匹配合约生成证据包,将匹配合约发布至下游协同链路以供按统一链接键进行查询或订阅。本申请能够提升跨方数据匹配准确性、增强版本适配性、降低人工维护成。
Smart Images

Figure CN122065042B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method and system for dynamic data matching and processing in supply chain collaboration. Background Technology
[0002] In supply chain collaboration scenarios, multiple entities, including brands, suppliers, logistics and warehousing, and distributors and retailers, need to share master data and event data in stages such as order placement, shipment, arrival, warehousing, and reconciliation to support cross-enterprise business collaboration and process traceability. Existing technologies employ two main approaches: one uses electronic data interchange or interface connections to transmit business documents and master data, and manually configures field mapping relationships to achieve data conversion between heterogeneous systems; the other uses master data synchronization or event standards to standardize descriptions of goods, locations, enterprises, and business events, and aggregates and displays collaborative data through a centralized platform.
[0003] However, the existing solutions mentioned above typically rely on static mapping rules and fixed interface versions, which are difficult to adapt to the reality of continuous addition of collaborating parties, frequent field evolution, and multiple versions running in parallel. This results in high mapping maintenance costs and difficulty in maintaining consistency in the long term. At the same time, the same entity may have multiple encodings, multiple aliases, and inconsistent granularity in different systems, lacking an effective dynamic alignment mechanism. This makes it difficult to stably associate events with a unified entity, even though events can be collected. In addition, the matching results often lack auditable evidence chains and version traceability capabilities, making it difficult to support dispute arbitration and compliance audits. Furthermore, the availability of collaborative matching is further limited when multiple parties are unwilling to share plaintext data. Summary of the Invention
[0004] In view of this, embodiments of this application provide a method and system for dynamic data matching in supply chain collaboration to solve the problems of difficulty in dynamically aligning multi-source heterogeneous data, poor adaptation to interface version evolution, and lack of auditable and traceable matching results in the prior art.
[0005] A first aspect of this application provides a method for dynamic data matching in supply chain collaboration, comprising: receiving data access contracts from at least two collaborating parties; establishing a contract catalog based on the data access contracts and recording contract versions and field constraints, wherein the data access contracts include synchronous interface contracts and asynchronous event contracts; performing hierarchical normalization processing on the access data based on the data access contracts to obtain normalized records; generating a mapping plan for the collaborating party data model and the platform specification model, and compiling the mapping plan into an executable mapping topology; performing field mapping on the normalized records based on the mapping topology to obtain mapped records; generating a candidate entity pair set from the mapped records for entities to be aligned, wherein the candidate entity pair set includes candidate generation based on standard identifiers, candidate generation based on rule bucketing, and candidate generation based on vector index nearest neighbor recall; and performing multi-evidence matching scoring on the candidate entity pair set to output a pass / fail result. Figure 1 Consistency verification involves matching links and confidence intervals; multi-evidence matching scoring includes probability scoring based on field-level similarity features, semantic evidence fusion based on vector similarity, and scoring based on relational constraints. Figure 1 Consistency verification; generate a matching contract based on the matching link, generate an evidence package for the matching contract, and publish the matching contract to the downstream collaborative link for querying or subscription by pressing the unified link key.
[0006] A second aspect of this application provides a data dynamic matching and processing system for supply chain collaboration, comprising: a receiving module, configured to receive data access contracts from at least two collaborating parties, establish a contract catalog based on the data access contracts, and record contract versions and field constraints, wherein the data access contracts include synchronous interface contracts and asynchronous event contracts; a processing module, configured to perform hierarchical normalization processing on the access data based on the data access contracts to obtain normalized records; a mapping module, configured to generate a mapping plan for the collaborating party data model and the platform specification model, and compile the mapping plan into an executable mapping topology; and perform field mapping on the normalized records based on the mapping topology to obtain mapped records; a generation module, configured to generate a candidate entity pair set from the mapped records for entities to be aligned, wherein the generation of the candidate entity pair set includes candidate generation based on standard identifiers, candidate generation based on rule bucketing, and candidate generation based on vector index nearest neighbor recall; and a scoring module, configured to perform multi-evidence matching scoring on the candidate entity pair set to output matching links and confidence intervals that pass consistency checks, wherein the multi-evidence matching scoring includes probability scoring based on field-level similarity features, semantic evidence fusion based on vector similarity, and scoring based on relational constraints. Figure 1 Consistency verification; the publishing module is used to generate a matching contract based on the matching link, generate an evidence package for the matching contract, and publish the matching contract to the downstream collaborative link for querying or subscription by pressing the unified link key.
[0007] A third aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.
[0008] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: By receiving data access contracts from at least two collaborating parties, a contract catalog is established based on the data access contracts, recording contract versions and field constraints. Data access contracts include synchronous interface contracts and asynchronous event contracts. Layered normalization processing is performed on the accessed data based on the data access contracts to obtain normalized records. A mapping plan is generated for the collaborating party's data model and the platform's specification model, and the mapping plan is compiled into an executable mapping topology. Field mapping is performed on the normalized records based on the mapping topology to obtain mapped records. For entities to be aligned, a candidate entity pair set is generated from the mapped records. The candidate entity pair set includes candidate generation based on standard identifiers, candidate generation based on rule bucketing, and candidate generation based on vector index nearest neighbor recall. Multi-evidence matching scoring is performed on the candidate entity pair set to output a pass / fail result. Figure 1 Consistency verification involves matching links and confidence intervals; multi-evidence matching scoring includes probability scoring based on field-level similarity features, semantic evidence fusion based on vector similarity, and scoring based on relational constraints. Figure 1 Consistency verification; generating a matching contract based on the matching link, and generating an evidence package for the matching contract, then publishing the matching contract to downstream collaborative links for querying or subscription by a unified link key. This application can improve the accuracy of cross-party data matching, enhance version compatibility, and reduce manual maintenance costs. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is a flowchart illustrating the dynamic data matching and processing method for supply chain collaboration provided in an embodiment of this application. Figure 2 This is a schematic diagram of the structure of the dynamic data matching and processing system for supply chain collaboration provided in this application embodiment; Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0011] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0012] In existing technologies, supply chain collaboration typically achieves cross-enterprise data transmission and aggregation through electronic data interchange, interface integration, or centralized data platforms, supplemented by master data standards or event standards to standardize the description of goods, locations, enterprises, and business events. These solutions often rely on pre-configured field mappings and fixed-version interfaces, or focus on event collection and display, making it difficult to achieve stable data consistency and entity association in situations involving multiple collaborating parties and parallel system evolution.
[0013] Against this backdrop, the prominent problems with existing technologies are: multi-source heterogeneous data lack a dynamic alignment mechanism for sustainable evolution; changes in interface and field versions lead to high mapping maintenance costs and easy failures; the same entity has multiple codes, multiple aliases, and granularity differences in different systems, making it difficult to stably associate event data with a unified entity; at the same time, the matching results lack auditable and traceable evidence chains and version basis, making it difficult to meet the needs of dispute arbitration and compliance auditing.
[0014] To address the aforementioned issues, this application provides a method and system for dynamic data matching in supply chain collaboration: A contract directory is established by receiving synchronous interface contracts and asynchronous event contracts, and version and field constraint governance is implemented; based on the contracts, hierarchical normalization processing is performed on the access data to form normalized records; a mapping plan is generated for the collaborating party's data model and the platform's specification model, and compiled into an executable mapping topology to perform field mapping on the normalized records to obtain mapped records; based on this, a candidate entity pair set is generated using a combination of standard identifiers, rule-based bucketing, and vector index nearest neighbor recall, and field-level probability scoring, semantic evidence fusion, and relation-constraint-based... Figure 1 Consistency verification performs multi-evidence matching scoring, outputs the matching links and confidence intervals that pass the verification; further, it generates matching contracts and evidence packages containing mapping versions and evidence summaries, and publishes them to downstream collaborative links for querying or subscription.
[0015] Through the above technical solutions, this application can improve the accuracy and stability of dynamic data matching across collaborating parties, enhance the adaptability to interface and field version evolution, reduce manual maintenance and mapping adjustment costs, and enable the matching results to have evidentiary and traceable audit support capabilities.
[0016] The technical solution of this application will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0017] Figure 1 This is a flowchart illustrating the dynamic data matching and processing method for supply chain collaboration provided in an embodiment of this application. Figure 1 As shown, the data dynamic matching processing method for supply chain collaboration may specifically include: S101, Receive data access contracts from at least two collaborating parties, establish a contract directory based on the data access contracts and record the contract version and field constraints, the data access contracts include synchronous interface contracts and asynchronous event contracts; S102, Perform hierarchical normalization processing on the access data based on the data access contract to obtain normalized records; S103: Generate a mapping plan for the collaborating party's data model and the platform's specification model, and compile the mapping plan into an executable mapping topology; perform field mapping on the normalized records based on the mapping topology to obtain the mapped records; S104, Generate a set of candidate entity pairs from the mapped records for the entities to be aligned. The generation of the candidate entity pair set includes candidate generation based on standard identifiers, candidate generation based on rule-based bucketing, and candidate generation based on vector index nearest neighbor recall. S105, Perform multi-evidence matching scoring on the candidate entity pair set to output the pass / fail results. Figure 1 Consistency verification involves matching links and confidence intervals; multi-evidence matching scoring includes probability scoring based on field-level similarity features, semantic evidence fusion based on vector similarity, and scoring based on relational constraints. Figure 1 Consistency check; S106, Generate a matching contract based on the matching link, generate an evidence package for the matching contract, and publish the matching contract to the downstream collaborative link for querying or subscription by pressing the unified link key.
[0018] In some embodiments, receiving data access contracts from at least two collaborating parties, establishing a contract catalog based on the data access contracts, and recording contract versions and field constraints include: Obtain the synchronous interface contracts and asynchronous event contracts submitted by each collaborating party, and perform unified contract parsing on the synchronous interface contracts and asynchronous event contracts to generate contract metadata containing a list of fields, field types, required constraints, enumeration field constraints, and data format constraints; Contract version identifiers are generated based on contract metadata, and a contract version relationship graph is established to record the compatibility relationships and changes between different contract versions; An executable set of contract verification rules is generated based on the field constraints in the contract metadata, and the set of contract verification rules is associated with and stored with the corresponding contract version identifier. When the access data arrives, a consistency check is performed on the access data based on the contract verification rule set, and the verification result is bound and recorded with the source collaborator identifier and contract version identifier of the access data for subsequent hierarchical normalization processing and mapping topology calls.
[0019] Specifically, the data access contract describes the accessible data channels provided by the collaborating party and their data structure constraints. The data access contract includes at least a synchronous interface contract and an asynchronous event contract. The synchronous interface contract describes the data interface constraints for request-response interactions, including at least the interface path, request field structure, response field structure, and field-level constraints. The asynchronous event contract describes the event channel constraints for publish-subscribe interactions, including at least the event topic identifier, event payload field structure, event trigger type, and field-level constraints.
[0020] In some examples, the supply chain collaboration platform receives data access contracts from at least two collaborating parties. Specifically, the platform configures a contract submission entry point for each collaborating party, which can be at least one of three methods: upload from the management end, interface submission, or automatic collection through the collaboration gateway. Taking "Manufacturer Collaborator A" and "Logistics Collaborator B" as an example, Collaborator A submits a synchronous interface contract to describe the field structure of the "Order Line Query Interface" and "Product Master Data Query Interface"; Collaborator B submits an asynchronous event contract to describe the event payload structure of the "Shipping Event Theme" and "Arrival Event Theme".
[0021] After receiving the aforementioned contracts, the collaborative platform performs unified contract parsing on both synchronous interface contracts and asynchronous event contracts, generating contract metadata. Unified contract parsing means that regardless of whether the contract originates from a synchronous interface or an asynchronous event, it is parsed and normalized according to a unified field description model, resulting in a unified field list and set of field-level constraints. This avoids subsequent processing steps requiring separate adaptation for different channel types. The field list enumerates the set of fields allowed by the contract and their hierarchical paths; the field type identifies the data type category of the field, such as string, number, boolean, array, or object; required constraints indicate whether a field must appear in a specific scenario; enumeration field constraints indicate the enumeration set or enumeration mapping table of allowed values for the field; and data format constraints indicate the format rules of the field, such as date format, timestamp format, encoding rules, or regular expression format. For nested fields, unified parsing converts them into field items with hierarchical paths and records parent-child relationships for subsequent validation and indexing.
[0022] In some examples, the collaboration platform generates contract version identifiers based on contract metadata and establishes a contract version relationship graph to record compatibility relationships and changes between different contract versions. Specifically, the collaboration platform calculates a version identifier for each received contract metadata and writes it as a node in the version relationship graph into the contract directory. If the same collaborating party submits a new contract on the same channel, the collaboration platform establishes a version-derived edge between the new version identifier and the previous version identifier, and records compatibility and changes on that edge.
[0023] For example, compatibility is used to characterize whether the new version maintains a resolvable and verifiable relationship with the data access of the old version, including at least three states: backward compatibility, partial compatibility, and incompatibility; change difference is used to characterize field-level changes, including at least field addition, field deletion, field type change, required constraint change, enumeration field change, and data format constraint change.
[0024] In some examples, the collaboration platform generates an executable set of contract verification rules based on the field constraints in the contract metadata, and stores the set of contract verification rules in association with the corresponding contract version identifier. The executable set of contract verification rules refers to the set of verification rules transformed from field constraints. It can determine the existence, type, value range, and format of fields item by item when the incoming data arrives, and output the verification conclusion and exception details.
[0025] For example, the contract validation rule set includes at least mandatory field validation rules, field type validation rules, enumeration field validation rules, and data format validation rules. Mandatory field validation rules determine whether mandatory fields are missing or empty; field type validation rules determine whether field values can be parsed into the contract-specified type; enumeration field validation rules determine whether field values fall within an allowed enumeration set or can be mapped to a target enumeration; and data format validation rules determine whether field values conform to preset format constraints. The collaborative platform binds the rule set with the contract version identifier and stores it in the contract directory, so that any subsequent access data only needs to carry the contract version identifier to locate the corresponding rule set.
[0026] In some examples, consistency verification refers to judging the structural and field consistency of the access data according to the contract verification rule set, outputting a verification pass status or a verification failure status, and outputting a list of abnormal fields and exception types when verification fails. For example, exception types can include missing exceptions, type exceptions, enumeration exceptions, and format exceptions. The collaboration platform generates an access record identifier for each piece of access data, and binds the record source collaborator identifier, contract version identifier, and verification result to the access record identifier.
[0027] When the subsequent hierarchical normalization process reads the binding record, the processing strategy can be determined based on the verification results: data that passes verification directly enters the normalization process; data with non-critical field anomalies enters the "normalization process with anomaly markers," retaining the anomaly field markers for subsequent mapping topology selection of degradation rules; data with missing critical fields or anomaly types of critical fields enters the "pending completion / pending arbitration queue," and the field constraints and enumeration domain constraints in the contract are referenced when generating completion candidates. When the subsequent mapping topology calls the binding record, it can select the matching mapping topology version based on the contract version identifier and trigger mapping assertion verification or rollback strategies based on the list of anomaly fields.
[0028] This embodiment achieves unified contract governance for synchronous interfaces and asynchronous events of multiple collaborators through the collaborative work of the aforementioned contract catalog, contract version relationship diagram, and contract verification rule set. This ensures that the access data has a locatable version basis and an executable field constraint verification foundation before entering the hierarchical normalization processing and mapping topology. As a result, it reduces the manual maintenance cost of contract adaptation and field verification in the case of continuous evolution of collaborator contracts and parallel development of multiple versions, and improves the stability and consistency of subsequent data processing links.
[0029] In some embodiments, hierarchical normalization processing is performed on access data based on a data access contract to obtain normalized records, including: Based on the source collaborator identifier and contract version identifier of the access data, a data access contract matching the access data is determined, and a set of standardized rules corresponding to the data access contract is generated; Based on a set of standardized rules, the access data is normalized in terms of standard identifier, field type and unit, and enumeration field alignment, so as to generate intermediate records carrying a unified identifier representation. When a key field is missing in an intermediate record or the value of a key field does not meet the field constraints, generate key field completion candidates and mark the key field completion candidates with confidence information. The intermediate record, which has undergone normalization and alignment and carries the contract version identifier and confidence information, is encapsulated into a normalized record.
[0030] Specifically, "layered normalization processing" refers to performing layered processing on the identifier representation, field type and unit of measurement, and enumeration value caliber of the accessed data according to the field constraints defined in the contract and the normalization rules preset by the platform. It also generates traceable completion candidates and confidence information when key fields are missing or do not meet constraints. The "normalization rule set" is an executable set of rules formed by combining the field constraints, enumeration domain constraints, data format constraints in the data access contract, and the platform-side normalization strategy. It includes at least identifier normalization rules, type and unit normalization rules, enumeration domain alignment rules, and key field detection and completion trigger rules.
[0031] In some examples, after the access data arrives, the collaboration platform first determines the data access contract matching the access data based on the source collaborator identifier and contract version identifier, and generates a set of standardized rules corresponding to the data access contract. Specifically, the collaboration platform locates the corresponding data access contract in the contract directory using "source collaborator identifier + channel identifier + contract version identifier", and reads the contract's field list, field types, required constraints, enumerated field constraints, and data format constraints.
[0032] Building upon this, the platform further merges contract field constraints with its own normalization strategy to form a set of standardized rules for that contract version. The platform-side normalization strategy may include a unified identifier encoding strategy, a numeric field type priority strategy, a unit conversion benchmark, an enumerated field mapping table version, and a key field priority list. Through this method, records accessed by the same collaborator under different contract versions will be matched with different sets of normalization rules, thus avoiding the mixing of normalization logic when multiple versions are running concurrently.
[0033] In some examples, the collaboration platform performs standard identifier normalization, field type and unit normalization, and enumeration field alignment on the access data based on a set of normalization rules to generate intermediate records carrying a unified identifier representation. Specifically, standard identifier normalization is used to convert entity-related identifier fields in the access data into a unified identifier representation on the platform side. Identifier fields may include at least one of product identifiers, location identifiers, enterprise identifiers, batch identifiers, or logistics unit identifiers.
[0034] For identifiers that meet the standard structure, the platform performs structure parsing and verification according to the identifier normalization rules, and converts the verified identifiers into a unified identifier representation. For identifiers that do not meet the standard structure but can be matched in the identifier mapping table maintained by the platform, the platform looks up the corresponding unified identifier representation through the mapping table. For identifiers that cannot be parsed and cannot be matched, the platform retains the original identifier in the intermediate record and marks it as "to be completed", providing input for the subsequent generation of completion candidates.
[0035] Field type and unit normalization ensures that field values match the field types specified in the contract and unifies the same physical quantity fields to the platform's base unit. The platform performs type conversion on field values according to the type and unit normalization rules, performs unit conversion on fields containing unit information, and records the normalized values and unit base markers in the intermediate record. Enumeration field alignment converts enumeration fields such as status codes, transportation methods, and packaging levels defined by collaborators into a unified enumeration set on the platform side. The platform performs mapping on the original enumeration values according to the enumeration field alignment rules. If mapping is successful, the target enumeration value is written and the version of the mapping table used is recorded; if mapping fails, the original enumeration value is recorded in the intermediate record and the enumeration exception type is marked for subsequent generation of completion candidates or triggering manual arbitration.
[0036] The following is a specific example. Collaborator A is a manufacturer. Its synchronous interface pushes a "Product Master Data Change Record," where the product identifier field is A_ITEM_001, the packaging level field is CASE, and the weight field is 12.5kg. Collaborator B is a logistics provider. Its asynchronous event channel publishes a "Shipping Event," where the shipping location field is WH-01, the transportation mode field is TRUCK, and the event payload also includes an order line identifier field of SO123-10. The collaboration platform, based on the source collaborator identifier and contract version identifier, locates the V2 synchronous interface contract of collaborator A and the E1 event contract of collaborator B, respectively, and generates the corresponding set of normalized rules.
[0037] For the record of collaborator A, the platform maps A_ITEM_001 to the unified product identifier U_ITEM_789 through the identifier mapping table during the identifier normalization stage; during the type and unit normalization stage, the weight field is normalized to the base unit g and the unit base mark is recorded; during the enumeration field alignment stage, CASE is mapped to the platform's unified packaging level enumeration PACK_L2. For the shipping event of collaborator B, the platform performs location identifier mapping on WH-01 to obtain the unified location identifier U_LOC_301 during the identifier normalization stage; during the enumeration field alignment stage, if TRUCK can be mapped in the current enumeration mapping table, it is written into the unified transportation mode enumeration TRANS_ROAD; otherwise, it is marked as an enumeration exception and enters the candidate completion process. After the above processing is completed, the platform generates intermediate records carrying unified identifier representations as input for subsequent mapping topology and entity candidate generation.
[0038] In some examples, when a key field is missing in an intermediate record or the value of a key field does not meet the field constraints, the collaboration platform generates key field completion candidates and marks the key field completion candidates with confidence information. Key fields refer to the set of fields marked as required by the field constraints corresponding to the contract version identifier or set by the platform as required for entity alignment, such as at least one of the following: uniform product identifier, uniform location identifier, packaging level enumeration, event timestamp, or order line identifier.
[0039] The platform performs key field checks on intermediate records. If a key field is found to be missing or its value does not meet the field constraints, it triggers the generation of completion candidates. Completion candidate generation can be based on at least the following criteria: rule inference based on field context, field derivation relationships recorded in historical mapping plan versions, association constraints in the collaborative knowledge graph, and nearest neighbor candidates based on the enumerated domain mapping table. The platform generates confidence information for each completion candidate. Confidence information can be determined by a combination of candidate source type, candidate hit consistency, and the number of candidate conflicts, and the candidate and confidence information are bound to the intermediate record.
[0040] Furthermore, the collaboration platform encapsulates the intermediate records, which have undergone normalization and alignment processing and carry contract version identifiers and confidence information, into normalized records. Specifically, a normalized record includes at least: the source collaborator identifier, the contract version identifier, a unified identifier representation set, a set of normalized field values, enumerated alignment results, anomaly marker set, and key field completion candidates and confidence information. During encapsulation, the platform retains the association information between the original fields and the normalized fields to support subsequent evidence package generation and audit traceability; simultaneously, it retains anomaly type markers for fields marked as anomalous, enabling subsequent mapping topologies to select applicable conversion rules or degradation processing strategies based on the anomaly type. The encapsulated normalized record is written to a normalized storage area or cache queue and outputs a record index to the subsequent mapping topology execution module to ensure that subsequent field mappings can be correctly scheduled according to the contract version and record status.
[0041] This embodiment generates a standardized rule set based on the data access contract and performs hierarchical execution of identifier normalization, type and unit normalization, enumeration field alignment, and key field completion candidate generation. This enables data from multiple collaborators and multiple versions to form a standardized record with unified identifier representation and traceable confidence information. This reduces the manual maintenance cost of heterogeneous data cleaning and caliber alignment, improves the adaptability and stability of subsequent mapping topology execution and entity matching links to data version evolution and field missing scenarios, and provides a consistent field basis and anomaly marking basis for the construction of subsequent matching evidence packages.
[0042] In some embodiments, a mapping plan is generated for the collaborating party's data model and the platform specification model, and the mapping plan is compiled into an executable mapping topology, including: Based on the field constraints of the data access contract and the field definitions of the platform specification model, the set of fields to be mapped is determined, and the field correspondence between the set of fields to be mapped and the collaborator's data model and the platform specification model is generated. Based on the field correspondence, a mapping plan carrying conditional predicates and conversion rules is generated. The conversion rules include one or more of the following: field type conversion, unit conversion, and enumeration value mapping. Perform topology compilation on the mapping plan to construct a mapping topology consisting of field generation nodes and dependent edges, and determine the execution order of field mapping based on the dependent edges; Generate an assertion verification rule set associated with the contract version for the mapping topology, and perform version marking and release control on the mapping topology. Release control includes labeling the applicable version range, canary activation, and binding rollback strategies.
[0043] Specifically, the "collaborator data model" refers to the set of field structures and semantic definitions defined by the synchronous interface contract or asynchronous event contract of the collaborator; the "platform specification model" refers to the set of standard field structures and semantic definitions used by the collaborative platform to uniformly carry master data and event data, which is used to support subsequent candidate generation, multi-evidence matching scoring and matching contract generation.
[0044] The "mapping plan" is a configuration object that provides a structured representation of field correspondences, mapping conditions, and transformation rules. It describes how to convert collaborator fields into platform fields. The "mapping topology" is an executable dependency graph formed by compiling the mapping plan. It uses field generation nodes as the basic unit and represents the sequential constraints of field generation through dependency edges, thereby achieving scheduled execution according to dependency order.
[0045] "Conditional predicates" are used to limit the applicable conditions of mapping rules, such as enabling the corresponding rule when the record type, event type, or a certain enumeration value meets preset conditions. "Assertion validation rule set" is used to perform constraint validation on the generated target fields during the field mapping process. It is derived from the field constraints of the data access contract and the field definitions of the platform specification model, and is associated with the contract version.
[0046] In some examples, the collaboration platform determines the set of fields to be mapped based on the field constraints of the data access contract and the field definitions of the platform specification model, and generates the field correspondence between the data model of the collaborating party and the platform specification model. Specifically, the collaboration platform first reads the corresponding data access contract from the contract directory based on the source collaborating party identifier and the contract version identifier, and extracts the field list, field types, required constraints, enumeration domain constraints, and data format constraints; at the same time, it extracts the target field list, target field types, target enumeration domains, and target caliber constraints from the platform specification model.
[0047] Based on this, the collaborative platform determines the set of fields to be mapped. The set of fields to be mapped includes at least one of the key fields marked as required or referenced by subsequent entity alignment links in the platform specification model, such as the unified entity identifier field, time field, location field, packaging level field, and event type field.
[0048] Subsequently, the collaborative platform generates field mapping relationships for the set of fields to be mapped. Field mapping relationships can include at least one of three scenarios: one-to-one mapping, one-to-many derivation, and many-to-one fusion. One-to-one mapping is used for direct renaming or path transformation; one-to-many derivation is used to derive multiple target fields from one source field; and many-to-one fusion is used to merge multiple source fields to generate one target field, such as merging "province, city, district, detailed address" to generate the platform-side "standardized address field".
[0049] The following example illustrates this. Collaborator A is the manufacturer, and its synchronization interface outputs product master data fields including item_id, gross_weight, and pack_level. Collaborator B is the logistics provider, and its event payload fields include ship_from, ship_to, transport_mode, and event_time. The platform specification model requires output fields to include a unified product identifier U_ITEM, a unified location identifier U_LOC, a unified transportation mode enumeration U_TRANS, a unified timestamp U_TIME, and a packaging level enumeration U_PACK. Based on this, the collaborative platform determines the set of fields to be mapped as U_ITEM, U_LOC, U_TRANS, U_TIME, and U_PACK, and generates field mapping relationships. For example, item_id corresponds to U_ITEM, gross_weight corresponds to the weight field and requires unit conversion, pack_level corresponds to U_PACK and requires enumeration mapping, ship_from and ship_to correspond to different role fields of U_LOC, and event_time corresponds to U_TIME.
[0050] In some examples, the collaboration platform generates a mapping plan carrying conditional predicates and transformation rules based on the field correspondence. Specifically, the collaboration platform generates one or more mapping rules for each target field, and each mapping rule includes at least a source field reference, a target field write, a conditional predicate, and a transformation rule.
[0051] The conditional predicate is used to distinguish the field definitions under different record types or event types. For example, when the event type is "shipping event," `ship_from` is used as the shipping location, and when the event type is "arrival event," `ship_to` is used as the arrival location. Conversion rules are used to perform standardization processing on the source field values before writing them, including at least one or more of the following: field type conversion, unit conversion, and enumeration value mapping. Field type conversion converts string values to numeric types or converts time strings to timestamps; unit conversion converts units such as kg and lb used by collaborators to the platform's standard units; and enumeration value mapping maps the values of collaborators' enumeration fields to the platform's unified enumeration field values and records the mapping table version.
[0052] Taking the above example, gross_weight is converted from the string "12.5kg" into a numerical value and then written into the platform weight field after being converted into the base unit; pack_level is mapped from CASE to the unified enumeration corresponding to U_PACK; transport_mode is mapped to TRANS_ROAD in U_TRANS if it is TRUCK, and to TRANS_AIR if it is AIR.
[0053] In some examples, the collaboration platform performs topology compilation on the mapping plan, constructing a mapping topology composed of field generation nodes and dependency edges, and determines the execution order of field mapping based on the dependency edges. Specifically, the collaboration platform converts each target field generation rule in the mapping plan into a field generation node. A field generation node includes at least a set of input fields, an output target field, applicable condition predicates, and a set of transformation rules. The collaboration platform analyzes the dependencies between different field generation nodes and generates dependency edges. Dependency edges indicate the relationship where the output field of one field generation node is referenced by another field generation node as an input field, thus forming a directed acyclic mapping topology. The execution order of the mapping topology is determined by the topological order of the dependency edges to ensure that the intermediate fields that are depended upon are generated first, and then the derived fields that depend on those intermediate fields are generated.
[0054] For example, the platform can first execute the "Identifier Unification Node" to generate a unified product identifier U_ITEM, and then execute the "Packaging Level Derivation Node" to generate U_PACK based on U_ITEM and pack_level; the platform can first execute the "Time Parsing Node" to generate U_TIME, and then execute the "Event Window Derivation Node" to generate a time bucket field for subsequent candidate generation based on U_TIME. Through topology compilation, the collaborative platform transforms the originally static field mapping relationship into a schedulable, verifiable, and traceable execution graph structure.
[0055] In some examples, the collaboration platform generates an assertion validation rule set associated with the contract version for the mapping topology, and performs version marking and release control on the mapping topology. Specifically, the assertion validation rule set is derived from two parts: one part comes from the field constraints of the data access contract, which is used to ensure that the mapping result does not violate the type and format boundaries declared by the collaboration party's contract; the other part comes from the field definitions of the platform specification model, which is used to ensure that the target fields after mapping satisfy the mandatory constraints, type constraints, and enumeration domain constraints of the platform side.
[0056] Assertion validation rules can include at least one of the following: target field existence assertion, target field type assertion, target field enumeration domain assertion, and cross-field consistency assertion. The collaboration platform binds the assertion validation rule set to the contract version identifier and generates a mapping topology version tag for the mapping topology. Release control manages the scope of mapping topologies across different contract versions and collaboration parties, and includes at least the following: applicable version range labeling, canary deployment, and rollback strategy binding.
[0057] The applicable version range label is used to limit which contract version identifiers a certain mapping topology version is applicable to; the gray-scale enablement is used to gradually enable the new mapping topology according to a preset ratio or according to a subset of collaborators when it is launched, and to monitor mapping anomalies and assertion failures; the rollback policy binding is used to automatically switch to the previous mapping topology version and record the rollback reason and triggering conditions when the anomaly rate exceeds the threshold or assertion failures occur in a set.
[0058] Continuing with the example above, collaborator B upgrades the event contract from E1 to E2. In E2, the `transport_mode` enumeration field is expanded to include the value `RAIL`, and a new field `carrier_code` is added. Upon receiving the new contract, the collaborating platform first records the differences between E1 and E2 in the contract version relationship graph. Then, based on the platform specification model, it recalculates the set of fields to be mapped and their correspondences, generating a new mapping plan that includes "carrier_code participating in transportation mode inference," and compiles it to obtain the new mapping topology version T2. The collaborating platform generates an assertion verification rule set for T2 to ensure that the output of the `U_TRANS` field must fall within the platform enumeration field.
[0059] Subsequently, the platform implemented a phased rollout of T2: T2 was initially enabled on some logistics network points or some carrier data. If it was found that the failure rate of U_TRANS assertions increased due to the lack of coverage of RAIL mapping rules, a rollback strategy was triggered to switch back to T1. After the enumerated mapping from RAIL to TRANS_RAIL was subsequently completed, it was rolled out again in a phased rollout. Through the above process, the mapping topology can iterate with the contract version while maintaining the controllable release and traceability of the mapping link.
[0060] This embodiment generates field correspondences based on contract field constraints and platform specification model field definitions, constructs a mapping plan carrying conditional predicates and transformation rules, and compiles the mapping plan into a mapping topology with dependency order. At the same time, it constrains and governs the mapping topology with assertion verification rule sets and versioned release control, enabling multi-version data models of collaborating parties to form an executable and rollbackable field mapping link on the platform side. This reduces the maintenance cost of mapping rules, enhances the adaptability to contract version changes and differences in caliber, and provides a consistent field output basis and auditable version basis for subsequent candidate entity pair generation and multi-evidence matching scoring.
[0061] In some embodiments, performing field mapping on normalized records based on mapping topology to obtain mapped records includes: Obtain the contract version identifier corresponding to the normalized record, and select the mapping topology that matches the contract version identifier based on the contract version identifier; The field generation nodes are scheduled sequentially according to the execution order of the mapping topology. Field extraction, condition predicate determination and transformation rule processing are performed on the source fields in the normalized records to generate target fields corresponding to the platform's standard model. During the execution of the field generation node, the assertion verification rule set associated with the mapping topology is called to perform constraint verification on the generated target field, and field-level mapping exception information is recorded when the verification fails. The validated target field is associated and encapsulated with the field-level mapping exception information and the version tag of the mapping topology to generate the mapped record.
[0062] Specifically, after generating standardized records, the supply chain collaboration platform further performs field mapping based on the mapping topology to obtain mapped records. "Mapped records" refer to structured data objects that conform to the platform's standardized model field structure. They carry at least a contract version identifier, a mapping topology version marker, and field-level mapping anomaly information, which are used for subsequent candidate entity pair generation and multi-evidence matching scoring.
[0063] In some examples, the collaboration platform first obtains the contract version identifier corresponding to the normalized record, and then selects a mapping topology that matches the contract version identifier. Specifically, the platform reads the source collaborator identifier and the contract version identifier from the normalized record, and searches the mapping topology directory for mapping topology versions whose applicable version range includes the contract version identifier. When multiple available mapping topology versions exist, the platform prioritizes the mapping topology version that is enabled and whose gray-scale strategy is applied, and writes the selection result into the context of this mapping session.
[0064] Furthermore, the platform sequentially schedules field generation nodes according to the execution order of the mapping topology. These nodes perform field extraction, conditional predicate determination, and transformation rule processing on the source fields in the normalized records to generate target fields corresponding to the platform's standardized model. The field generation nodes describe the generation logic of the target fields; their inputs are source fields or intermediate fields from the normalized records, and their output is the target field of the platform's standardized model.
[0065] When scheduling each field generation node, the platform first extracts the field value from the normalized record according to the input field path configured for the field generation node, and then determines whether the condition predicate is satisfied. When the condition predicate is satisfied, the transformation rule is called to process the field value and write it to the target field. When the condition predicate is not satisfied, the node is skipped or a default placeholder is written and the reason for skipping is recorded.
[0066] For example, the conversion rules may include at least one of field type conversion, unit conversion, and enumeration value mapping. Taking the "shipment event" of logistics collaborator B as an example, the normalized record contains a unified location identifier U_LOC_301, the original transportation mode enumeration TRUCK, and the event time field 2026-01-13T10:15:30. When the conditional predicate of the event type "shipment event" is satisfied, the field generation node configured in the mapping topology writes U_LOC_301 into the platform field "shipment location identifier", converts TRUCK to the unified transportation mode enumeration TRANS_ROAD through enumeration mapping and writes it into the platform field "transportation mode", and converts the time field to the platform unified timestamp field U_TIME.
[0067] In some examples, during the execution of the field generation node, the platform invokes the assertion validation rule set associated with the mapping topology to perform constraint validation on the generated target field, and records field-level mapping exception information when the validation fails. The assertion validation rule set includes at least the target field existence assertion, type assertion, and enumeration field assertion.
[0068] Immediately after each field generation node outputs the target field, the platform executes the corresponding assertion. If the assertion succeeds, the target field is marked as available; if the assertion fails, field-level mapping exception information is recorded. This information includes at least the path of the exception field, the exception type, the assertion trigger identifier, and a summary of the input field. Using the example above, if the TRUCK fails to map to the platform's unified enumeration domain, causing the transportation mode field enumeration domain assertion to fail, the platform records the enumeration exception and retains the original enumeration value and mapping table version information. This information can be used as uncertain evidence or to trigger a rollback strategy during the subsequent multi-evidence matching and scoring stage.
[0069] In some examples, the platform associates and encapsulates the validated target fields with field-level mapping anomaly information and the version tag of the mapping topology to generate a post-mapping record. Specifically, during encapsulation, the platform writes the mapping topology version tag into the header of the post-mapping record, writes all target fields into the field structure corresponding to the platform's standardized model, and attaches field-level mapping anomaly information to the post-mapping record in the form of an anomaly list. At the same time, it retains the mapping link index from the standardized record to the target field, enabling subsequent evidence packages to trace back to the contract version, mapping topology version, and assertion validation results used.
[0070] This embodiment selects a matching mapping topology based on the contract version identifier, schedules field generation nodes according to the topology order, and outputs a mapped record with anomaly markers by combining the assertion verification rule set. This enables multi-collaborator, multi-version data to stably generate structured output that conforms to the platform specification model on the platform side, thereby reducing the mapping failure rate caused by differences in field definitions, enhancing the traceability and controllable release capability of the mapping process, and providing a consistent data foundation for subsequent entity candidate generation and matching evidence construction.
[0071] In some embodiments, generating a set of candidate entity pairs from the mapped record for the entities to be aligned includes: Based on the unified identifier in the mapped record, candidate generation based on the standard identifier is performed on the entities to be aligned to obtain the first set of candidate entity pairs; Based on at least one key attribute in the mapped record, perform rule-based bucketing to generate candidate pairs for the entities to be aligned, so as to obtain a set of second candidate entity pairs; Entity semantic vectors are generated based on the text-type attributes in the mapped records, and nearest neighbor retrieval is performed in the vector index to obtain a third set of candidate entity pairs; The first set of candidate entity pairs, the second set of candidate entity pairs, and the third set of candidate entity pairs are deduplicated and merged. Candidate source tags are added to the merged candidate entity pairs to obtain the candidate entity pair set.
[0072] Specifically, after obtaining and recording the mapping, the supply chain collaboration platform generates a set of candidate entity pairs for the entities to be aligned, which serve as candidate inputs for subsequent multi-evidence matching and scoring. Entities to be aligned refer to entity objects that require establishing identity associations across collaborating parties, including at least one of the following: product entity, enterprise entity, location entity, batch entity, or logistics unit entity. Candidate entity pairs are paired objects consisting of the entity identifiers of the two collaborating parties, used to represent candidate associations that "may be the same entity."
[0073] Key attributes are a set of attribute fields that distinguish entities. They can be preset by the platform's specification model and bound to the entity type. Examples include the specifications, brand, and packaging level of a product entity; the country, province, city, district, and postal code of a location entity; and the registration number and tax ID of a company entity. Rule binning refers to discretizing or grouping key attributes and assigning records to different bins to limit the range of candidate comparisons.
[0074] In some examples, the platform performs candidate generation based on standard identifiers for entities to be aligned, using the unified identifier representation in the mapped record, to obtain a first set of candidate entity pairs. Specifically, the platform maintains a standard identifier index for different entity types, with the unified identifier representation as the key and the list of entity identifiers as the value. When a usable unified identifier representation exists in the mapped record, the platform directly locates the set of entities with the same key in the standard identifier index and pairs the current collaborating entity identifier with the matched entity identifiers of other collaborating parties to generate a first set of candidate entity pairs. To avoid role confusion, the platform simultaneously validates the entity type and entity role fields when generating candidates; for example, it distinguishes between the shipping location and the destination for location entities, and between the shipper and the carrier for enterprise entities.
[0075] Furthermore, the platform performs rule-based bucketing candidate generation on the entities to be aligned based on at least one key attribute in the mapped record to obtain a second set of candidate entity pairs. Specifically, the platform configures a bucketing rule set for each entity type. The bucketing rule set includes at least one bucket key construction method, and the bucket key can be generated from the normalized value, truncated value, or combined value of the key attribute. The platform generates bucket keys according to the bucketing rule set based on the key attributes in the mapped record, and retrieves the set of candidate entities under the same bucket key in the bucketing index, thereby generating the second set of candidate entity pairs. To improve bucketing stability, the platform can introduce fault-tolerance strategies for the bucket keys, such as taking the main specification segment for specification fields, taking the province / city level for address fields, and discretizing numerical fields by interval, and recording the bucketing rule version in the bucketing index for subsequent traceability.
[0076] Furthermore, the platform generates entity semantic vectors based on the textual attributes in the mapped records and performs nearest neighbor retrieval in the vector index to obtain a third set of candidate entity pairs. Specifically, the platform pre-defines a set of textual attribute fields, such as product name, company name, address description, and goods description, and standardizes and concatenates these textual attributes to generate text representations. The platform vectorizes these text representations to obtain entity semantic vectors and performs nearest neighbor retrieval in the vector index using entity type, language tags, and time window tags as retrieval filters to obtain a list of candidate entities with similarity exceeding a preset threshold. The platform then pairs the current entity with candidate entities to generate a third set of candidate entity pairs. To control the size of the candidates, the platform can limit the number of results returned for each query and attach an initial similarity value to the nearest neighbor results as a priori feature for subsequent scoring.
[0077] Furthermore, the platform performs deduplication and merging on the first, second, and third candidate entity pair sets, and adds candidate source tags to the merged candidate entity pairs to obtain the candidate entity pair set. Specifically, the platform merges the three types of candidates using "entity type + collaborating party entity identifier pair" as the deduplication key, and performs union annotation on candidate sources with the same deduplication key. At the same time, the platform can generate an initial candidate level according to the source priority. For example, candidates that hit both the standard identifier and the rule bucket are given a higher priority, while candidates that only hit the vector nearest neighbor are given a pending verification priority, and this priority is written into the candidate source tag.
[0078] The following example illustrates this. Collaborator A reports a product entity record. After mapping, the record contains a unified product identifier U_ITEM_789, a packaging level U_PACK, and the product name "IndustrialCleaner500ml". Collaborator C reports a product entity record. After mapping, the product identifier in this record is an alias code and lacks a unified product identifier, but the name is "Industrial Cleaner 500 ml". The platform directly matches the same U_ITEM_789 candidate for Collaborator A's record using the standard identifier index, generating the first candidate. For Collaborator C's record, it uses rule-based bucketing with a bucket key constructed from "brand + specification range + packaging level" to match candidates in the same bucket, generating the second candidate. Simultaneously, it generates semantic vectors for the names of the two records and matches them in the nearest neighbor search of the vector index, generating the third candidate. The platform finally deduplicates and merges the three types of candidates, annotating the candidate source with "standard identifier + bucket + vector", and outputs a set of candidate entity pairs for subsequent multi-evidence matching and scoring.
[0079] This embodiment improves candidate recall coverage while maintaining a controllable candidate size by combining standard identifier candidates, rule-based binning candidates, and vector nearest neighbor candidates through a candidate generation mechanism. It also provides candidate source markers and prior information for subsequent matching and scoring, thereby reducing the computational cost of full comparison and reducing missed recalls caused by missing keys, aliases, and differences in caliber.
[0080] In some embodiments, multi-evidence matching scoring is performed on the candidate entity pair set to output a pass / fail score. Figure 1 The consistency check includes matching links and confidence intervals, including: Extract field-level similarity features of each candidate entity pair in the candidate entity pair set, and calculate the probability score of the candidate entity pair based on the field-level similarity features; For at least some candidate entity pairs in the candidate entity pair set, semantic vectors are generated based on entity attributes and vector similarity is calculated. The vector similarity is then fused with the probability score as semantic evidence to obtain a comprehensive score for the candidate entity pairs. A collaborative knowledge relationship graph is constructed based on relational constraints, and a comprehensive score is applied to candidate entity pairs. Figure 1 Consistency checks are performed to eliminate candidate entity pairs that do not meet relational constraints. Based on Figure 1 The comprehensive score of candidate entity pairs in the consistency check generates matching links, and the confidence interval corresponding to the matching links is determined based on the comprehensive score.
[0081] Specifically, after obtaining the set of candidate entity pairs, the supply chain collaboration platform performs multi-evidence matching scoring on the set of candidate entity pairs to output the pass / fail result. Figure 1 Consistency verification includes matching links and confidence intervals. Matching links refer to the identity associations between entities across collaborating parties, which can be represented as structured objects of "entity identifier pair + entity type + link direction or role".
[0082] Confidence intervals characterize the range of credibility of a matching link. They can be represented by upper and lower bounds or confidence level ranges and are linked to the comprehensive score and source of evidence. Field-level similarity features are a set of comparable features extracted from the structured fields of two entities, including at least one or more of the following: identifier consistency features, enumeration consistency features, numerical difference features, and character similarity features.
[0083] Semantic evidence refers to auxiliary matching evidence provided by unstructured features such as text semantic vector similarity. Relationship constraints refer to the structural consistency conditions that should be met in supply chain business relationships, such as the association constraints between order lines and goods, shipping locations and carrier routes, companies and tax numbers, and locations and administrative levels. A collaborative knowledge relationship graph is a graph structure with entities as nodes and business relationships as edges, where nodes and edges can carry source collaborators, time windows, and version tags.
[0084] In some examples, the platform extracts field-level similarity features from each candidate entity pair in the candidate entity pair set and calculates a probability score for the candidate entity pairs based on these field-level similarity features. Specifically, the platform configures feature extraction templates for different entity types. Taking a product entity as an example, field-level similarity features may include whether the unified identifier is consistent, whether the brand field is consistent, the difference range of the specification field, whether the packaging level enumeration is consistent, and whether the weight field difference falls within a preset threshold, etc. Taking a company entity as an example, it may include the consistency of registration number or tax number, name similarity, consistency of country / region, etc. Taking a location entity as an example, it may include the consistency of administrative division level, consistency of postal code prefix, consistency of latitude and longitude grid, etc. The platform inputs the extracted field-level similarity features into the probability scoring model to output a probability score. The probability scoring model can be a probability mapping based on feature weights or based on a trained probability classifier, outputting a probability score in the range of 0 to 1, and recording the top-contributing field features as an evidence summary.
[0085] In some examples, for at least a portion of the candidate entity pairs in the candidate entity pair set, the platform generates semantic vectors based on entity attributes and calculates vector similarity. This vector similarity is then fused with the probability score as semantic evidence to obtain a comprehensive score for the candidate entity pairs. Specifically, the platform generates semantic vectors for the textual attributes of the entities on both sides of a candidate entity pair. These textual attributes may include at least one of the following: product name, company name, address description, or goods description. The platform calculates the similarity between the semantic vectors on both sides as vector similarity and fuses this similarity with the probability score. The fusion method may include weighted fusion, gated fusion, or segmented fusion: when the probability score is in the intermediate uncertainty range, the weight of semantic evidence is increased to distinguish between naming or synonym scenarios; when the probability score is already in the high confidence range, semantic evidence is used for consistency confirmation rather than as the primary decision-maker. The platform writes the fused comprehensive score along with the candidate source tag into the scoring record of the candidate entity pairs for subsequent evaluation. Figure 1 Consistency checks provide input.
[0086] Let's continue with the previous example. Collaborator A's product name is "IndustrialCleaner500ml", and Collaborator C's product name is "Industrial Cleaner 500 ml". The two lack the same unified identifier in their structured fields, resulting in a probability score in the middle range. After the platform generates semantic vectors, it obtains a high vector similarity. After fusion, the overall score is improved, and "semantic evidence enhancement" is written into the evidence summary.
[0087] In some examples, the platform constructs a collaborative knowledge graph based on relational constraints and performs a comprehensive scoring of candidate entity pairs. Figure 1 Consistency checks are performed to eliminate candidate entity pairs that do not meet relationship constraints. Specifically, the platform extracts entity nodes and business relationship edges from the mapped records and historical matching contracts, constructs a collaborative knowledge relationship graph, and adds candidate entity pairs as "optional link edges" to the graph. The platform pre-sets relationship constraint rules, such as that product entities associated with the same order line should maintain consistent links on different collaborating parties; the combination of the place of shipment and the place of arrival under the same carrier should not conflict; and the same enterprise tax number should not correspond to multiple mutually exclusive enterprise entities.
[0088] The platform consistently selects candidate entity pairs based on their comprehensive scores, from highest to lowest, and checks for violations of relationship constraints when adding each candidate link. If a violation is found, the candidate is removed or downgraded to a pending arbitration status. Extending the above example, if accepting a candidate product link would result in the same order line being associated with two mutually exclusive product entities on the platform, the platform would reject the candidate link based on order line relationship constraints, even if its comprehensive score is high, thus ensuring global consistency.
[0089] In some examples, the platform is based on... Figure 1 The platform generates matching links based on the comprehensive score of candidate entity pairs in the consistency check, and determines the confidence interval corresponding to the matching links based on the comprehensive score. Specifically, the platform will... Figure 1 Candidate entity pairs for consistency verification are solidified into matching links, and each matching link is bound to a comprehensive score, candidate source marker, evidence summary, and... Figure 1 Consistency verification records. The platform determines the confidence interval based on the comprehensive score and evidence coverage. For example, matching links with a high comprehensive score and that match both the standard identifier and structured features are assigned a high confidence interval; matching links with a medium comprehensive score but strong semantic evidence and that do not trigger relationship conflicts are assigned a medium confidence interval; matching links with a comprehensive score at the boundary and with field-level anomaly markers are assigned a low confidence interval and marked as requiring review.
[0090] This embodiment fuses field-level probability scoring with semantic evidence to form a comprehensive score, and then applies relational constraints to the collaborative knowledge graph for enforcement. Figure 1 Consistency verification can output globally consistent matching links and confidence intervals in scenarios with multiple collaborators, multiple heterogeneous sources, and differences in aliasing, thereby improving the reliability of matching results and reducing the error rate of subsequent data association caused by conflicting links. At the same time, it provides traceable scoring and verification basis for subsequent matching contract generation and evidence package accumulation.
[0091] In some embodiments, a matching contract is generated based on the matching link, and an evidence package is generated for the matching contract. The matching contract is then published to downstream collaborative links for querying or subscription by a unified link key, including: Assign a unified link key to each matching link that passes the consistency check, and generate a matching contract containing the unified link key and a list of participant entity identifiers; Based on the mapping plan version used when generating matching links, field-level evidence summaries, and Figure 1 The consistency verification records are used to construct an evidence package, which is then associated with and stored in the matching contract. The evidence package also includes input data version markers and revocable policies. Encapsulate the matching contract as a matching contract change event and publish it to the downstream collaborative link, or write the matching contract into a unified query interface so that downstream users can execute queries by pressing a unified link key; When the revocable policy is triggered, a revocation contract for the unified link key is generated and published to the downstream collaborative link.
[0092] Specifically, the supply chain collaboration platform outputs through... Figure 1After verifying the consistency of the matching links and confidence intervals, the matching results are further solidified into matching contracts that can be stably consumed downstream. An auditable evidence package is generated for the matching contracts, and then the matching contracts are published to downstream collaborative links for querying or subscription using a unified link key. A matching contract is a structured encapsulation object of a matching link, used to reference the same entity across collaborative parties using a unified key.
[0093] A unified link key is a globally unique identifier assigned by the platform to each matching link. It can be generated from a combination of entity type, participant entity identifiers, and a time window marker, and deduplication rules ensure that the same matching link yields a stable key value across multiple calculations. An evidence package encapsulates the evidence set from the matching process, and includes at least the mapping plan version, field-level evidence summaries, and... Figure 1 Consistency verification records, input data version tags, and reversibility policies are used for subsequent audit traceability, dispute verification, and reversal compensation.
[0094] In some examples, the platform assigns a unified link key to each matching link that passes the consistency check and generates a matching contract containing the unified link key and a list of participant entity identifiers. Specifically, the platform extracts the entity type, participant identifier, and list of participant entity identifiers for each matching link, and normalizes and sorts the list of participant entity identifiers according to a preset sorting rule to avoid key-value drift. Based on this, the platform generates a unified link key and writes the unified link key into the contract header of the matching contract. The contract content of the matching contract includes at least the unified link key, entity type, list of participant entity identifiers, matching confidence interval, matching effective time, and contract version number.
[0095] For example, taking a product entity as an example, the product entity identifier A_ITEM_001 of collaborating party A and the product entity identifier C_SKU_77 of collaborating party C are connected through... Figure 1 After consistency verification, a matching link is formed. The platform assigns a unified link key UKEY_ITEM_20260113_0001 to the link and generates a matching contract, writing A_ITEM_001 and C_SKU_77 as the list of participating entity identifiers.
[0096] In some examples, the platform bases its decisions on the version of the mapping plan used when generating matching links, field-level evidence summaries, and... Figure 1 Consistency verification records are used to construct an evidence package, which is then associated and stored with the matching contract. Specifically, the platform reads the mapping plan version or mapping topology version tag from the mapping link, the field-level evidence summary and comprehensive score from the scoring link, and the consistency link... Figure 1 Consistency verification records and summaries of removal reasons are compiled, and the above information is packaged into the main body of the evidence package.
[0097] Field-level evidence summaries may include consensus conclusions of fields with high contribution, enumeration mapping hit information, semantic vector similarity intervals, and field-level anomaly information summaries. Figure 1 The consistency verification record may include applicable relational constraint identifiers, a summary of the verification successful path, and conflict detection results. The evidence package also includes input data version tags and a revocable policy. The input data version tags are used to identify the access data version and contract version on which the evidence is based, and include at least the source collaborator identifier, contract version identifier, access record identifier, and mapped record version tag.
[0098] A revocable policy defines the conditions under which a matched contract can be revoked and provides compensation actions after revocation. It includes at least revocation trigger conditions, revocation priority, and revocation propagation scope. Revocation trigger conditions may include contract version rollback, changes to the enumeration mapping table version causing the assertion failure rate to exceed a threshold, and subsequent... Figure 1 Consistency checks revealed conflicting links and corrective appeals initiated by collaborating parties. The platform associates and stores evidence packages with matching contracts using a unified link key, enabling downstream users to retrieve both the contract and evidence summary when querying the unified link key.
[0099] In some examples, the platform encapsulates the matching contract into a matching contract change event and publishes it to the downstream collaborative link, or writes the matching contract into a unified query interface for downstream systems to execute queries using a unified link key. Specifically, the platform generates a contract change event for the matching contract, and the event payload includes at least a unified link key, a list of participating entity identifiers, a contract version number, a change type marker, and an evidence package index. The downstream collaborative link can be a message subscription channel or a collaborative data bus. After subscribing, downstream systems can establish a local mapping cache of "unified link key to local entity identifier" for rapid association of business events such as orders, shipments, and arrivals.
[0100] On the other hand, the platform can also write the matching contract into the contract storage area corresponding to the unified query interface, so that downstream systems can directly query the contract and evidence package summary through the unified link key when they cannot subscribe or need to trace back. Taking the warehousing and distribution system as an example, after receiving the contract change event UKEY_ITEM_20260113_0001, it associates the unified link key with the local product master data, so that when the inbound event arrives, it can directly associate with the same product entity by pressing the unified link key.
[0101] In some examples, when a revocable policy is triggered, the platform generates a revocation contract for the unified link key and publishes the revocation contract to the downstream collaborative links. Specifically, when the platform detects that the revocation trigger condition is met, it generates a revocation contract. The revocation contract includes at least the unified link key, the version number of the revoked contract, a revocation reason flag, the revocation effective time, and an alternative link suggestion or isolation flag. The platform encapsulates the revocation contract as a revocation event and publishes it to the downstream collaborative links to drive the downstream system to revoke the corresponding local association cache; at the same time, it updates the contract status in the unified query interface, retaining the evidence package before revocation for audit traceability. Taking the above example, if it is subsequently found that collaborator C's C_SKU_77 is confused with another product due to code reuse, triggering a relationship constraint conflict and meeting the revocation condition, the platform generates a revocation contract for UKEY_ITEM_20260113_0001 and issues it, so that the downstream system can stop the erroneous association in time.
[0102] This embodiment assigns a unified link key to the matching link and generates a matching contract to map versions and scoring evidence. Figure 1 The consistency verification records are used to construct evidence packages and store them in association. They are then distributed downstream via event publishing and query interfaces and support revocable policies. This can enhance the referability, traceability and governance of matching results, thereby reducing the integration cost of cross-system entity association and improving the controllability of collaborative links in scenarios of dispute verification and change of standards.
[0103] The following are system embodiments of this application, which can be used to execute the method embodiments of this application. For details not disclosed in the system embodiments of this application, please refer to the method embodiments of this application.
[0104] Figure 2 This is a schematic diagram of the structure of the dynamic data matching and processing system for supply chain collaboration provided in this application embodiment. For example... Figure 2 As shown, the data dynamic matching and processing system for supply chain collaboration includes: The receiving module 201 is used to receive data access contracts from at least two collaborating parties, establish a contract directory based on the data access contracts and record the contract version and field constraints. The data access contracts include synchronous interface contracts and asynchronous event contracts. Processing module 202 is used to perform hierarchical normalization processing on access data based on the data access contract to obtain normalized records; The mapping module 203 is used to generate a mapping plan for the collaborating party's data model and the platform specification model, and compile the mapping plan into an executable mapping topology; based on the mapping topology, it performs field mapping on the normalized records to obtain the mapped records; The generation module 204 is used to generate a set of candidate entity pairs from the mapped records for the entities to be aligned. The generated set of candidate entity pairs includes candidate generation based on standard identifiers, candidate generation based on rule-based bucketing, and candidate generation based on vector index nearest neighbor recall. Scoring module 205 is used to perform multi-evidence matching scoring on the candidate entity pair set to output matching links and confidence intervals that pass consistency checks. Multi-evidence matching scoring includes probability scoring based on field-level similarity features, semantic evidence fusion based on vector similarity, and scoring based on relational constraints. Figure 1 Consistency check; The publishing module 206 is used to generate a matching contract based on the matching link, generate an evidence package for the matching contract, and publish the matching contract to the downstream collaborative link for querying or subscription by pressing the unified link key.
[0105] Figure 3 This is a schematic diagram of the electronic device 3 provided in an embodiment of this application. Figure 3 As shown, the electronic device 3 of this embodiment includes a processor 301, a memory 302, and a computer program 303 stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program 303, it implements the steps in the various method embodiments described above. Alternatively, when the processor 301 executes the computer program 303, it implements the functions of each module / unit in the system embodiments described above.
[0106] Electronic device 3 can be a desktop computer, laptop, handheld computer, cloud server, or other electronic device. Electronic device 3 may include, but is not limited to, processor 301 and memory 302. Those skilled in the art will understand that... Figure 3 This is merely an example of electronic device 3 and does not constitute a limitation on electronic device 3. It may include more or fewer components than shown, or different components.
[0107] The processor 301 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0108] The memory 302 can be an internal storage unit of the electronic device 3, such as a hard disk or memory of the electronic device 3. The memory 302 can also be an external storage device of the electronic device 3, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 3. The memory 302 can also include both internal and external storage units of the electronic device 3. The memory 302 is used to store computer programs and other programs and data required by the electronic device.
[0109] The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium (e.g., a computer-readable storage medium). Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by a processor, it can implement the steps of the various method embodiments described above. The computer program can include computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable storage medium can include: any entity or device capable of carrying computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0110] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for dynamic data matching and processing in supply chain collaboration, characterized in that, include: Receive data access contracts from at least two collaborating parties, establish a contract directory based on the data access contracts and record the contract versions and field constraints, the data access contracts include synchronous interface contracts and asynchronous event contracts; Based on the data access contract, hierarchical normalization processing is performed on the access data to obtain normalized records; A mapping plan is generated for the collaborating party's data model and the platform's specification model, and the mapping plan is compiled into an executable mapping topology; based on the mapping topology, field mapping is performed on the normalized records to obtain the mapped records; For entities to be aligned, a set of candidate entity pairs is generated from the mapped records. The generated set of candidate entity pairs includes candidate generation based on standard identifiers, candidate generation based on rule-based bucketing, and candidate generation based on vector index nearest neighbor recall. The candidate entity pair set is subjected to multi-evidence matching scoring to output matching links and confidence intervals that pass graph consistency verification. The multi-evidence matching scoring includes probability scoring based on field-level similarity features, semantic evidence fusion based on vector similarity, and graph consistency verification based on relational constraints. A matching contract is generated based on the matching link, and an evidence package is generated for the matching contract. The matching contract is published to the downstream collaborative link for querying or subscription by pressing the unified link key. The step of receiving data access contracts from at least two collaborating parties, establishing a contract catalog based on the data access contracts, and recording contract versions and field constraints includes: Obtain the synchronous interface contract and asynchronous event contract submitted by each collaborating party, and perform unified contract parsing on the synchronous interface contract and the asynchronous event contract to generate contract metadata containing a list of fields, field types, required constraints, enumeration domain constraints, and data format constraints; Based on the contract metadata, a contract version identifier is generated, and a contract version relationship diagram is established to record the compatibility relationships and changes between different contract versions; An executable contract verification rule set is generated based on the field constraints in the contract metadata, and the contract verification rule set is associated with and stored with the corresponding contract version identifier; When the access data arrives, a consistency check is performed on the access data based on the contract verification rule set, and the verification result is bound and recorded with the source collaborator identifier and contract version identifier of the access data for subsequent hierarchical normalization processing and mapping topology calls; The step of generating a mapping plan for the collaborating party's data model and the platform specification model, and compiling the mapping plan into an executable mapping topology, includes: Based on the field constraints of the data access contract and the field definitions of the platform specification model, a set of fields to be mapped is determined, and the field correspondence between the set of fields to be mapped and the collaborator data model and the platform specification model is generated. Based on the field correspondence, a mapping plan carrying conditional predicates and conversion rules is generated. The conversion rules include one or more of the following: field type conversion, unit conversion, and enumeration value mapping. Perform topology compilation on the mapping plan to construct a mapping topology consisting of field generation nodes and dependency edges, and determine the execution order of field mapping based on the dependency edges; Generate an assertion verification rule set associated with the contract version for the mapping topology, and perform version marking and release control on the mapping topology. The release control includes labeling the applicable version range, canary activation, and binding the rollback strategy. The step of performing field mapping on the normalized records based on the mapping topology to obtain mapped records includes: Obtain the contract version identifier corresponding to the normalized record, and select a mapping topology that matches the contract version identifier based on the contract version identifier; The field generation nodes are scheduled sequentially according to the execution order of the mapping topology, and the source fields in the normalized records are processed by field extraction, condition predicate determination and transformation rule processing to generate target fields corresponding to the platform standard model. During the execution of the field generation node, the assertion verification rule set associated with the mapping topology is called to perform constraint verification on the generated target field, and field-level mapping exception information is recorded when the verification fails. The target field that passes the verification is associated and encapsulated with the field-level mapping exception information and the version tag of the mapping topology to generate the mapped record.
2. The method according to claim 1, characterized in that, The step of performing hierarchical normalization processing on the access data based on the data access contract to obtain normalized records includes: Based on the source collaborator identifier and contract version identifier of the access data, a data access contract matching the access data is determined, and a set of normalized rules corresponding to the data access contract is generated; Based on the set of normalized rules, the access data is normalized in terms of standard identifier, field type and unit, and enumeration field alignment, so as to generate intermediate records carrying a unified identifier representation. When a key field is missing in the intermediate record or the value of the key field does not meet the field constraint, a key field completion candidate is generated and confidence information is marked on the key field completion candidate. The intermediate record, which has undergone normalization and alignment and carries the contract version identifier and the confidence information, is encapsulated into the normalized record.
3. The method according to claim 1, characterized in that, The step of generating a set of candidate entity pairs from the mapped records for the entities to be aligned includes: Based on the unified identifier in the mapped record, candidate generation based on standard identifier is performed on the entities to be aligned to obtain a first set of candidate entity pairs; Based on at least one key attribute in the mapped record, perform rule-based bucketing to generate candidate pairs for the entities to be aligned, so as to obtain a second set of candidate entity pairs. Entity semantic vectors are generated based on the text class attributes in the mapped records, and nearest neighbor retrieval is performed in the vector index to obtain a third candidate entity pair set; The first set of candidate entity pairs, the second set of candidate entity pairs, and the third set of candidate entity pairs are deduplicated and merged, and candidate source tags are added to the merged candidate entity pairs to obtain the set of candidate entity pairs.
4. The method according to claim 1, characterized in that, The step of performing multi-evidence matching scoring on the candidate entity pair set to output matching links and confidence intervals that pass graph consistency verification includes: Extract field-level similarity features of each candidate entity pair in the candidate entity pair set, and calculate the probability score of the candidate entity pair based on the field-level similarity features; For at least some candidate entity pairs in the candidate entity pair set, semantic vectors are generated based on entity attributes and vector similarity is calculated. The vector similarity is then fused with the probability score as semantic evidence to obtain a comprehensive score for the candidate entity pairs. A collaborative knowledge relationship graph is constructed based on relational constraints, and graph consistency verification is performed on the comprehensive score of the candidate entity pairs to eliminate candidate entity pairs that do not satisfy the relational constraints. A matching link is generated based on the comprehensive score of the candidate entity pairs that pass the graph consistency check, and a confidence interval corresponding to the matching link is determined based on the comprehensive score.
5. The method according to claim 1, characterized in that, The process of generating a matching contract based on the matching link, generating an evidence package for the matching contract, and publishing the matching contract to downstream collaborative links for querying or subscribing by a unified link key includes: Assign a unified link key to each of the matching links that pass the consistency check, and generate a matching contract containing the unified link key and a list of participant entity identifiers; An evidence package is constructed based on the mapping plan version used when generating the matching link, the field-level evidence digest, and the graph consistency verification record, and the evidence package is associated with and stored with the matching contract. The evidence package also includes an input data version tag and a revocable policy. The matching contract can be encapsulated as a matching contract change event and published to the downstream collaborative link, or the matching contract can be written into a unified query interface for downstream users to perform queries by pressing the unified link key. When the revocable policy is triggered, a revocation contract is generated for the unified link key, and the revocation contract is published to the downstream collaborative link.
6. A dynamic data matching and processing system for supply chain collaboration, characterized in that, include: A receiving module is used to receive data access contracts from at least two collaborating parties, establish a contract directory based on the data access contracts, and record the contract version and field constraints. The data access contracts include synchronous interface contracts and asynchronous event contracts. The processing module is used to perform hierarchical normalization processing on the access data based on the data access contract to obtain normalized records; The mapping module is used to generate a mapping plan for the collaborating party's data model and the platform's specification model, and compile the mapping plan into an executable mapping topology; based on the mapping topology, it performs field mapping on the normalized records to obtain the mapped records; The generation module is used to generate a set of candidate entity pairs from the mapped records for the entities to be aligned. The generated set of candidate entity pairs includes candidate generation based on standard identifiers, candidate generation based on rule-based bucketing, and candidate generation based on vector index nearest neighbor recall. The scoring module is used to perform multi-evidence matching scoring on the candidate entity pair set to output matching links and confidence intervals that pass the consistency check. The multi-evidence matching scoring includes probability scoring based on field-level similarity features, semantic evidence fusion based on vector similarity, and graph consistency check based on relational constraints. The publishing module is used to generate a matching contract based on the matching link, generate an evidence package for the matching contract, and publish the matching contract to the downstream collaborative link for querying or subscription by pressing the unified link key; The receiving module is used to acquire synchronous interface contracts and asynchronous event contracts submitted by each collaborating party, and perform unified contract parsing on the synchronous interface contracts and asynchronous event contracts to generate contract metadata containing a field list, field types, required constraints, enumeration domain constraints, and data format constraints; generate contract version identifiers based on the contract metadata, and establish a contract version relationship graph to record the compatibility relationships and change differences between different contract versions; generate an executable contract verification rule set for the field constraints in the contract metadata, and associate the contract verification rule set with the corresponding contract version identifier for storage; when access data arrives, perform consistency verification on the access data based on the contract verification rule set, and bind and record the verification result with the source collaborating party identifier and contract version identifier of the access data for subsequent hierarchical normalization processing and mapping topology calls; The mapping module is used to determine the set of fields to be mapped based on the field constraints of the data access contract and the field definitions of the platform specification model, and to generate the field correspondence between the set of fields to be mapped and the collaborator data model and the platform specification model; based on the field correspondence, a mapping plan carrying conditional predicates and conversion rules is generated, wherein the conversion rules include one or more of field type conversion, unit conversion, and enumeration value mapping; topology compilation is performed on the mapping plan to construct a mapping topology composed of field generation nodes and dependency edges, and the execution order of field mapping is determined based on the dependency edges; an assertion verification rule set associated with the contract version is generated for the mapping topology, and version marking and release control are performed on the mapping topology, wherein the release control includes applicable version range marking, gray-scale activation, and rollback strategy binding; The mapping module is further configured to obtain the contract version identifier corresponding to the normalized record, and select a mapping topology matching the contract version identifier based on the contract version identifier; sequentially schedule field generation nodes according to the execution order of the mapping topology, and perform field extraction, conditional predicate judgment, and transformation rule processing on the source fields in the normalized record to generate target fields corresponding to the platform specification model; during the execution of the field generation node, call the assertion verification rule set associated with the mapping topology to perform constraint verification on the generated target fields, and record field-level mapping exception information when the verification fails; associate and encapsulate the verified target fields with the field-level mapping exception information and the version tag of the mapping topology to generate the mapped record.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
A fast docking insurance company contract platform for small and medium-sized generations
CN108985949A
Supply chain order placing management method and system based on big data analysis
CN120706760A