Embedded intelligent capture component-driven data archiving processing method and system

CN122527082APending Publication Date: 2026-08-07JIANGSU JIYANG SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU JIYANG SOFTWARE CO LTD
Filing Date
2026-07-10
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0007]本发明的一个目的在于提出一种嵌入式智能捕获组件驱动的数据归档处理方法及系统,针对现有技术难以兼容单轨制原生档案数据精准捕获与双轨制受控实体关联信息同步全量捕获、异构业务系统字段语义不统一导致捕获与映射不准、以及归档依赖被动推送而不可控且缺少合规证据链固化的问题,提出了如下技术方案:在业务系统运行环境中部署嵌入式智能捕获组件并生成捕获配置,基于变更数据捕获形成变更事件流并提取字段差分,按事务关联聚合生成事务簇并建立因果链,基于多视图字段特征通过对比学习模型实现字段语义对齐并输出映射置信分值,按归档模式生成归档数据集并在双轨制下事件驱动补采实体物理信息、权属凭证信息和管护记录信息;进一步对映射置信分值分组执行条件共形预测校准并以双阈值路由自动归档、纠偏补采或人工复核,自动归档时生成含数据指纹、因果链、对齐结果、置信度与合规校验记录的封装包并签名与时间戳固化入库

Benefits of technology

1、实现单轨制与双轨制归档的一体化兼容与全量关联捕获:基于业务与档案的双向内核,通过归档模式参数驱动事件监听范围与生成规则,在单轨制下精准捕获档案原生变更数据,在双轨制下基于受控实体标识(如针对地图空间数据归档领域为空间要素ID,针对实物/文物数据归档领域为实物编号)事件驱动补采多类型电子档案的受控实体物理信息(如空间数据的坐标系参数、实物数据的材质尺寸等)、权属凭证信息和管护记录信息并合并归档,避免因归档模式差异导致的数据范围缺失与关联口径不一致。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122527082A_ABST
    Figure CN122527082A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of informationization and electronic archives management, and discloses a data archiving processing method and system driven by an embedded intelligent capture component; a component reads archiving mode parameters and standard fields to generate capture configuration, executes changed data capture and extracts field difference, forms a transaction cluster according to a transaction and establishes a cause-effect chain; field semantic alignment is completed based on multi-view feature comparison learning, and a confidence score is output; an archiving data set is generated according to single-track or double-track; under the double-track, event-driven entity physical information, ownership certificate information and management and protection record information are supplemented; the confidence score is grouped, conformally calibrated and routed by double thresholds; when automatically archiving, data fingerprints, cause-effect chains, verification records and signature time stamps are encapsulated and solidified into a database.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information technology and electronic records management, and in particular to a data archiving and processing method and system driven by an embedded intelligent capture component. Background Technology

[0002] With the advancement of digital and smart museum construction, archival management is gradually shifting from paper-based ledgers to information systems, forming multi-source business data encompassing basic information of controlled entities, acquisition and ownership certificates, protection, restoration and maintenance records, entry and exit records, and inspection data. To meet regulatory requirements and long-term preservation needs, the industry typically builds archival management platforms or electronic document archiving systems to aggregate business system data according to established archiving standards, forming an archival data foundation. Existing archiving methods mainly include: business systems actively pushing archived data via interfaces, or archiving platforms periodically extracting and batch-processing data for cleaning and transformation via interfaces. Some scenarios also employ message queue relays, data exchange platforms, and change capture based on database logs to achieve incremental data collection. In terms of field mapping, it largely relies on pre-configured data dictionaries, fixed rule mapping tables, or manual standardization, combined with metadata standards to achieve structured archiving. Regarding compliance and evidence preservation, some systems introduce capabilities such as integrity verification, electronic signatures, and timestamps for result solidification and auditing.

[0003] However, existing technologies still have the following shortcomings: 1. The single-track independent archiving system and the dual-track entity and archive combined archiving system differ significantly in data scope, capture objects and related criteria. Existing archiving links often cannot simultaneously achieve accurate capture of the original archive data and synchronous full capture of entity-related information.

[0004] 2. In heterogeneous business system environments, data structures, field semantics, and interface forms are not uniform. Relying on manual organization or fixed rule mapping can easily lead to incomplete capture, inaccurate field identification, and unstable associations, resulting in insufficient adaptability and accuracy.

[0005] 3. Archiving relies on passive pushes or periodic extractions from business systems, making the archiving process uncontrollable and prone to omissions, delays, or misalignments. Furthermore, it lacks automatic verification of compliance elements and the solidification of traceable evidence chains at the moment of capture, resulting in insufficient credibility and auditability of the archiving results.

[0006] Therefore, there is a need for a data archiving and processing method and system that can overcome the shortcomings of the existing technologies. Summary of the Invention

[0007] One objective of this invention is to propose a data archiving processing method and system driven by an embedded intelligent capture component. Addressing the problems of existing technologies such as difficulty in accurately capturing native archive data in a single-track system and synchronously capturing all associated information of controlled entities in a dual-track system, inaccurate capture and mapping due to inconsistent field semantics in heterogeneous business systems, and uncontrollable archiving relying on passive push and lacking solidified compliance evidence chains, the following technical solution is proposed: An embedded intelligent capture component is deployed in the business system operating environment to generate capture configurations. A change event stream is formed based on change data capture, and field differences are extracted. Transaction clusters are generated by aggregating transaction associations and establishing causal chains. Based on multi-view field features, a contrastive learning model is used to achieve field semantic alignment and output mapping confidence scores. An archive dataset is generated according to the archiving mode, and under the dual-track system, event-driven supplementary collection of entity physical information, ownership certificate information, and maintenance record information is performed. Furthermore, conditional conformal prediction calibration is performed on the mapping confidence scores in groups, and automatic archiving, correction supplementary collection, or manual review is performed using dual-threshold routing. During automatic archiving, a package containing data fingerprints, causal chains, alignment results, confidence levels, and compliance verification records is generated, signed, and timestamped before being stored in the database. This invention has the technical advantages of automatic cross-system adaptation, controllable and monitorable archiving process, verifiable, traceable and reliably encapsulated archiving results.

[0008] This invention provides a data archiving processing method driven by an embedded intelligent capture component, comprising: S1, deploying an embedded intelligent capture component, reading archiving mode parameters and a set of archiving standard fields, generating a capture configuration, including data source connection information, event listening range, and rules and parameters related to field differential extraction, transaction association, data grouping, and dual-threshold routing; S2, performing change data capture on the data source of the business system according to the capture configuration, forming a change event stream, and generating a change record stream according to the field differential extraction rules; S3, aggregating the change record stream according to the transaction association rules to obtain a transaction cluster, and establishing a causal chain of change records within the transaction cluster to generate transaction cluster data; S4, generating multi-view field features based on the field differential of the causal chain, inputting a multi-view field comparison learning model to output field semantic vectors, and calculating the similarity with the standard field semantic vectors corresponding to the archiving standard field set to obtain field semantic alignment results, including field mapping relationships and mapping confidence scores; S5, converting the transaction cluster data into structured data under the archiving standard field set according to the field mapping relationship to obtain basic archived data, with the basic archived data used in single-track systems. As an archived dataset, in the dual-track system, the controlled entity identifier is determined based on the basic archived data. After event-driven supplementary collection of the controlled entity's physical information, ownership certificate information, and management record information associated with the controlled entity identifier, the archived dataset is generated. The mapping confidence scores of the archived datasets are collected to form a mapping confidence score set. S6: The archived dataset is grouped according to the grouping rules. For each group, conformal prediction calibration based on the mapping confidence score set is performed to obtain the calibrated confidence level. The calibrated confidence level is compared with the dual threshold parameters to generate a routing decision and obtain the routed archived data. S7: After performing event-driven correction supplementary collection on the data of the correction supplementary collection and updating the archived dataset, the process returns to S6. For the data of the manually reviewed route, a review task is output and associated with transaction cluster data, field semantic alignment results, and calibrated confidence levels. For the data of the automatically archived route, an archive package is generated, which includes the data fingerprint, causal chain, field semantic alignment results, calibrated confidence level, and verification records generated according to the preset compliance rule set of the archived dataset. The archive package is electronically signed, timestamped, and containerized before being written to the archive storage system.

[0009] Optionally, S1 includes: The embedded intelligent capture component is started in the operating environment of the business system; The archive mode parameters are obtained through the preset configuration interface, and the set of archive standard fields is obtained from the preset configuration library. The event monitoring scope is determined based on the archiving mode parameters. When the archiving mode parameters indicate single-track archiving, the event monitoring scope is limited to the data source objects and event types corresponding to the original archive data. When the archiving mode parameters indicate dual-track archiving, the event monitoring scope further includes the data source objects and event types corresponding to the physical information of the controlled entity, ownership certificate information, and management record information. Field difference extraction rules are generated based on the archiving standard field set. These rules are used to at least limit the range of fields to be extracted and the record format for field differences. Transaction association rules are generated based on the archiving mode parameters. These transaction association rules are used to at least limit the extraction location of the transaction identifier and the time window for transaction aggregation. Grouping rules are generated based on archiving mode parameters. These grouping rules are at least used to limit the data grouping method based on data source system identifier, event type, and field category. Based on the archiving mode parameters, dual threshold parameters are generated. The dual threshold parameters include a first threshold and a second threshold used to distinguish between automatically archived routes, correction and supplementary acquisition routes, and manually reviewed routes. The data source connection information, event listening range, field difference extraction rules, transaction association rules, grouping rules, and dual threshold parameters are written into the capture configuration and persisted to obtain the capture configuration.

[0010] Optionally, S2 includes: A connection is established with the data source of the business system based on the data source connection information in the capture configuration, and change monitoring is enabled for the data source according to the event monitoring range in the capture configuration to continuously acquire change events matching the event monitoring range and form a change event stream. For each change event in the change event stream, event metadata is parsed to obtain the event metadata, which includes the data source system identifier, data source object identifier, operation type, occurrence time, and association identifier used for transaction aggregation. According to the field difference extraction rules in the capture configuration, the set of fields that have changed is extracted from the change event, and field differences are generated for each field in the set of fields, where each field difference includes a field identifier, the field value before the change, the field value after the change, and field value pattern information. The event metadata and the field differences are combined according to a preset record format to generate a change record containing the event metadata and the field difference list. The change record is output to a preset change record stream channel or change record queue to form a change record stream.

[0011] Optionally, S3 includes: Transaction identifiers for transaction aggregation are extracted from change records in the change record stream. These transaction identifiers are determined from association identifiers in the event metadata according to transaction association rules in the capture configuration. A time window for transaction aggregation is set according to the transaction association rules in the capture configuration, and change records with the same transaction identifier and occurrence time falling within the same time window are aggregated into the same transaction cluster to form multiple transaction clusters. For each transaction cluster, the change records within the cluster are sorted according to their occurrence time to obtain a sorting result. A reference relationship is established between adjacent change records based on the sorting result, and each change record is set to reference its previous change record, forming a causal chain corresponding to the transaction cluster. The transaction clusters and the causal chains are associated and stored to generate transaction cluster data containing the causal chains.

[0012] Optionally, S4 includes: For the causal chain in the transaction cluster data, traverse each field difference in the causal chain and generate multi-view field features. The generation of multi-view field features includes: extracting the field identifier corresponding to the field difference to generate field name features. Based on the field values ​​before and after the change in the field difference, the data type, length, value range, and unit information of the field values ​​are determined to generate field value pattern features; Extract the field identifier and field value pattern information of adjacent fields from the field difference list that is in the same change record as the field difference to generate adjacent field context features; Extract the data source system identifier, data source object identifier, operation type, and occurrence time from the event metadata of the change record corresponding to the field difference to generate event metadata context features; The field name feature, field value pattern feature, adjacent field context feature, and event metadata context feature are combined to obtain the multi-view field feature corresponding to the field difference; The multi-view field features are fed into a trained multi-view field contrast learning model to output a field semantic vector. Obtain the standard field semantic vectors corresponding to the archived standard field set from the pre-established standard field semantic vector library; Calculate the similarity between the semantic vector of a field and the semantic vector of each standard field, and determine the field mapping relationship based on the principle of maximizing similarity; The maximum similarity value is determined as the mapping confidence score, and a field semantic alignment result is generated, which includes the field mapping relationship and the mapping confidence score.

[0013] Optionally, S5 includes: Based on the field mapping relationships in the field semantic alignment results, field transformation processing is performed on each field difference list in the transaction cluster data. This field transformation process includes replacing field identifiers with the corresponding archive standard field identifiers from the archive standard field set, and performing format and unit normalization on the changed field values ​​in the field differences. The archive standard field identifiers obtained through the field transformation process and their corresponding changed field values ​​are then aggregated. In cases where multiple changed field values ​​appear for the same archive standard field identifier within the same transaction cluster, the chronological order is determined based on the causal chain in the transaction cluster data, and the changed field value with the latest chronological order is selected to generate structured data under the archive standard field set to obtain the basic archive. Data; when the archiving mode parameter indicates single-track archiving, the basic archiving data is determined as the archiving dataset; when the archiving mode parameter indicates dual-track archiving, the field values ​​corresponding to the preset archiving standard fields are read from the basic archiving data to determine the controlled entity identifier, and event-driven association full data acquisition is triggered based on the controlled entity identifier to obtain the controlled entity physical information, ownership certificate information, and maintenance record information associated with the controlled entity identifier. The controlled entity physical information, ownership certificate information, and maintenance record information are converted into structured data under the archiving standard field set and then merged with the basic archiving data to obtain the archiving dataset; at the same time, the mapping confidence scores corresponding to each archiving standard field identifier in the archiving dataset are collected to form a mapping confidence score set.

[0014] Optionally, S6 includes: Based on the grouping rules in the capture configuration, the data source system identifier, event type, and field category are extracted from each structured data in the archived dataset. The dataset is then divided into at least one data group according to the principle that the data source system identifier, event type, and field category are consistent. For each data group, the corresponding mapping confidence score set is read, and grouping conditional conformal prediction calibration is performed on the mapping confidence score set according to a preset calibration coverage parameter to obtain the calibrated confidence level for that data group. For each piece of structured data, its calibrated confidence level is compared with the dual threshold parameters in the capture configuration. An automatic archiving route is generated when the calibrated confidence level is not less than the first threshold; a correction and supplementary acquisition route is generated when the calibrated confidence level is less than the first threshold but not less than the second threshold; and a manually reviewed route is generated when the calibrated confidence level is less than the second threshold. The automatic archiving route, the correction and supplementary acquisition route, and the manually reviewed route are aggregated to form a routing decision. The archived dataset is then labeled according to the routing decision to generate routed archived data.

[0015] Optionally, the S7 includes: For archived data following a route where the routing decision is to correct and supplement the route, a supplementary acquisition request is generated based on missing or insufficiently confident archiving standard fields and corresponding controlled entity identifiers in the archived data. An event-driven corrective supplementary acquisition is then performed using an embedded intelligent capture component to obtain the supplementary acquisition result. The supplementary acquisition result is converted into structured data under the archiving standard field set and merged with the archived dataset to update the archived dataset. Simultaneously, the mapping confidence score set is updated, and the process returns to step S6. For archived data following a route where the routing decision is to manually review the route, a review task is generated. This review task is at least associated with the transaction cluster data and fields corresponding to the archived data following the route. The semantic alignment results and calibrated confidence scores are output to the verification terminal. For the archived data after the routing decision is an automatic archiving route, an archive package is generated. The generation of the archive package includes calculating the data fingerprint of the archived dataset, writing the causal chain, field semantic alignment results, and calibrated confidence scores in the transaction cluster data into the archive package, performing element integrity verification according to a preset set of compliance rules and generating verification records, writing the verification records into the archive package, performing electronic signature and timestamp solidification on the archive package, and containerizing the archive package in a preset electronic file container format. The containerized archive package is then written into the archive storage system.

[0016] Optionally, the training of the multi-view field comparison learning model includes: obtaining historical change records from at least two business systems, and generating change records containing event metadata and a field difference list according to step S2; The change records are aggregated according to the transaction association rules of S3 to form transaction cluster data, and a causal chain is established within each transaction cluster. For each field difference in the causal chain, generate multi-view field features according to S4; Based on the archive standard field set and the preset field semantic generation rules, a field mapping pseudo-annotation set is constructed. The field semantic generation rules include field name feature similarity constraints, field value pattern feature consistency constraints, adjacent field context feature constraints, and event metadata context feature constraints. The field mapping pseudo-annotation set is used to indicate the correspondence between multi-view field features and archive standard fields in the archive standard field set. Positive sample pairs are constructed based on the field mapping pseudo-label set. The two multi-view field features in the positive sample pair are mapped to the same archiving standard field, and the positive sample pair includes positive sample pairs between multi-view field features from different business systems. Negative sample pairs are constructed based on the field mapping pseudo-label set. The two multi-view field features in the negative sample pairs are mapped to different archiving standard fields. The negative sample pairs include negative sample pairs with similar field name features and different field value pattern features. The negative sample pairs are subjected to hard negative sample screening, which includes using a multi-view field contrast learning model to be trained to calculate the similarity of the negative sample pairs and selecting negative sample pairs with a similarity greater than a preset similarity threshold as hard negative sample pairs. The multi-view field contrast learning model is iteratively trained using a contrastive learning loss function. The contrastive learning loss function is used to increase the similarity of the field semantic vectors of positive sample pairs and decrease the similarity of the field semantic vectors of negative sample pairs until a preset convergence condition is met or a preset iteration threshold is reached, thereby obtaining the trained multi-view field contrast learning model. Based on the trained multi-view field comparison learning model, the multi-view field features mapped to the same archive standard field in the field mapping pseudo-annotation set are used to generate corresponding field semantic vectors, and the field semantic vectors are aggregated to generate standard field semantic vectors corresponding to the archive standard field.

[0017] On the other hand, the present invention also provides a data archiving and processing system driven by an embedded intelligent capture component, comprising: a capture configuration module for reading archiving mode parameters and a set of archiving standard fields to generate a capture configuration; a change capture module for performing change data capture on the data source of the business system according to the capture configuration, forming a change event stream, and generating a change record stream according to field difference extraction rules; a transaction processing module for aggregating the change record stream according to transaction association rules to obtain a transaction cluster, and establishing a causal chain of change records within the transaction cluster to generate transaction cluster data; a field alignment module for generating multi-view field features based on field difference of the causal chain, inputting a multi-view field comparison learning model to output a field semantic vector, calculating the similarity with the standard field semantic vector corresponding to the archiving standard field set, and obtaining a field semantic alignment result, including a field mapping relationship and a mapping confidence score; and an archive generation module for converting the transaction cluster data into structured data under the archiving standard field set according to the field mapping relationship to obtain an archive dataset, and aggregating the mapping confidence scores to form a mapping. The system comprises a set of confidence scores, where, in a dual-track system, controlled entity identifiers are determined based on the archived dataset, and event-driven supplementary collection of physical information, ownership certificate information, and maintenance record information of the controlled entity associated with the identifier is performed before merging and updating the archived dataset; a calibration routing module is used to group the archived dataset according to grouping rules, perform conformal prediction calibration on each group based on the mapped confidence score set to obtain a calibrated confidence level, compare it with dual threshold parameters to generate a routing decision, and output automatic archived routes, correction supplementary collection routes, or manual review routes; and an archive execution module is used to generate an archive package for the data of the automatic archived routes and write it to the archive storage system. The archive package contains the data fingerprint, causal chain, field semantic alignment results, calibrated confidence level, and compliance verification records of the archived dataset, and performs electronic signature, timestamp solidification, and containerization encapsulation. For the data of the correction supplementary collection routes, event-driven correction supplementary collection is performed and the archived dataset is updated. For the data of the manual review routes, a review task is output and associated with transaction cluster data, field semantic alignment results, and calibrated confidence level.

[0018] The beneficial effects of this invention are: 1. Achieve integrated compatibility and full-data correlation capture for single-track and dual-track archiving: Based on the bidirectional kernel of business and archives, the event listening range and generation rules are driven by the archiving mode parameters. Under the single-track system, the original change data of archives is accurately captured. Under the dual-track system, the controlled entity identifier (e.g., spatial element ID for map spatial data archiving, and physical number for physical / cultural relic data archiving) is used to supplement the controlled entity physical information (e.g., coordinate system parameters of spatial data, material size of physical data, etc.), ownership certificate information and maintenance record information of multiple types of electronic archives and merge them for archiving, so as to avoid the loss of data range and inconsistency of correlation due to the difference in archiving mode.

[0019] 2. Improve the data capture adaptability and field mapping accuracy in heterogeneous business system environments: By organizing change records through transaction clusters and causal chains, multi-view features are constructed based on field differences, including field names, field value patterns, adjacent field contexts, and event metadata contexts. Contrastive learning is used to achieve field semantic alignment and confidence score output, reducing reliance on manual processing and fixed rule mapping. This allows for flexible compatibility with different business front-ends such as map spatial GIS systems and cultural relic / physical asset management systems, reducing problems such as incomplete capture, inaccurate identification, and unstable association.

[0020] 3. Enhance the controllability, compliance, and auditability of the archiving process: The mapping confidence score is calibrated by conformal prediction based on grouping conditions, and automatic archiving, correction and supplementary collection, and manual review are performed using dual-threshold gating routing to form a supervised closed-loop archiving control; at the same time, during automatic archiving, data fingerprints, causal chains, field alignment results, calibration confidence and compliance verification records are encapsulated and electronically signed and timestamped to form verifiable and traceable archiving credentials. Attached Figure Description

[0021] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a data archiving processing method driven by an embedded intelligent capture component proposed in this invention. Detailed Implementation

[0022] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0023] refer to Figure 1An embedded intelligent capture component-driven data archiving processing method includes: S1, deploying the embedded intelligent capture component, reading archiving mode parameters and archiving standard field set, generating capture configuration, including data source connection information, event listening range, and rules and parameters related to field differential extraction, transaction association, data grouping, and dual threshold routing; S2, performing change data capture on the data source of the business system according to the capture configuration, forming a change event stream, and generating a change record stream according to the field differential extraction rules; S3, aggregating the change record stream according to the transaction association rules to obtain a transaction cluster, and establishing a causal chain of change records within the transaction cluster to generate transaction cluster data; S4, generating multi-view field features based on the causal chain's field differential, inputting the multi-view field comparison learning model to output field semantic vectors, and calculating the similarity with the standard field semantic vectors corresponding to the archiving standard field set to obtain field semantic alignment results, including field mapping relationships and mapping confidence scores; S5, converting the transaction cluster data into structured data under the archiving standard field set according to the field mapping relationship to obtain basic archived data. In single-track systems, the basic archived data is used as... For the archived dataset, in the dual-track system, the controlled entity identifier is determined based on the basic archived data. After event-driven supplementary collection of the controlled entity's physical information, ownership certificate information, and maintenance record information associated with the controlled entity identifier, the data is merged to generate the archived dataset. The mapping confidence scores of the archived dataset are collected to form a mapping confidence score set. S6: The archived dataset is grouped according to the grouping rules. For each group, conformal prediction calibration based on the mapping confidence score set is performed to obtain the calibrated confidence level. The calibrated confidence level is compared with the dual threshold parameters to generate a routing decision and obtain the routed archived data. S7: After performing event-driven correction supplementary collection on the data of the correction supplementary collection and updating the archived dataset, the process returns to S6. For the data of the manually reviewed route, a review task is output and associated with transaction cluster data, field semantic alignment results, and calibrated confidence levels. For the data of the automatically archived route, an archive package is generated, which includes the data fingerprint, causal chain, field semantic alignment results, calibrated confidence levels, and verification records generated according to the preset compliance rule set of the archived dataset. The archive package is electronically signed, timestamped, and containerized before being written to the archive storage system.

[0024] In this specific embodiment, S1 includes: This is accomplished by starting an embedded intelligent capture component in the runtime environment of the business system. The embedded intelligent capture component runs in-process mode, bound to the application process of the business system, and reads the local controlled configuration file at startup to obtain data source connection information. The data source connection information includes database type, host address, port, database name, username, authentication credentials, and connection pool parameters. The connection pool parameters are set to a maximum of 20 connections and a connection timeout of 30 seconds, so that the embedded intelligent capture component can maintain a continuous available connection to the data source and have real-time monitoring capabilities throughout the lifecycle of the business process. The embedded intelligent capture component simultaneously launches a preset configuration interface client and initiates a pull request to the configuration interface at a fixed address to obtain the archive mode parameters. The archive mode parameters are defined as a structured parameter set including an archive mode identifier, a monitoring object set identifier, a transaction aggregation time window, a grouping dimension switch, and dual threshold parameters. The archive mode identifier is set to one of two values, either single-track or dual-track, and is explicitly set to dual-track in this embodiment. After the embedded intelligent capture component completes the retrieval of the archive mode parameters, it reads the archive standard field set from the preset configuration library. The archive standard field set consists of multiple archive standard field records, and each archive standard field record includes an archive standard field identifier, a Chinese name of the field, a field category identifier, a field data type, a unit identifier, and a required field mark. The preset configuration library stores the data in a versioned manner and returns the currently effective version number, so that subsequent capture and conversion are performed under the same version of the archive caliber. Based on the archiving mode parameters, the embedded intelligent capture component determines the event listening range and generates a listening list. The event listening range is defined as a set of binary pairs consisting of a data source object identifier and an event type. When the archiving mode identifier is dual-track, the event listening range includes both the data source object and the event type corresponding to the original archive data, and additionally includes the data source object corresponding to the controlled entity physical information, the data source object corresponding to the ownership certificate information, and the data source object corresponding to the management record information. The event types are set to three categories: insertion event, update event, and deletion event, enabling the embedded intelligent capture component to achieve unified listening on the same source side for the capture objects of dual-track archiving. The embedded intelligent capture component generates field difference extraction rules based on the archived standard field set. These rules are set to a field whitelist extraction mechanism and drive the extraction granularity with field category identifiers in the archived standard field set. Specifically, archived standard field identifiers marked as mandatory are added to the mandatory extraction list, while those marked as optional are added to the candidate extraction list. Primary key field identifiers and foreign key field identifiers used for association are added to the association extraction list. The record format for field difference is fixed as a four-element structure containing field identifier, field value before change, field value after change, and field value pattern information. The field value pattern information is set to consist of four sub-items: data type, length, value range, and unit identifier. The unit identifier is normalized using a built-in unit dictionary. This dictionary contains six unit identifiers: “mm”, “cm”, “m”, “g”, “kg”, and “℃”, along with their standardized names. The unit normalization result is determined using a full-match rule, ensuring that the subsequent extraction range and record format for field difference are strictly limited at the configuration level. The embedded intelligent capture component generates transaction association rules based on the archive mode parameters. These rules are configured to read the association identifier field from the event metadata of the change event and use it as the transaction identifier. In this embodiment, the association identifier field is fixedly named tx_id, and its value is derived from the transaction number of the database change commit. Simultaneously, the transaction aggregation time window is set to... ,in This parameter represents the time window parameter that aggregates change records under the same transaction identifier that occur within the same 5-second window into the same transaction cluster, thereby clarifying the boundaries of transaction aggregation at the configuration level. The embedded intelligent capture component generates grouping rules based on archiving mode parameters. These grouping rules are configured to group data according to a consistent three-dimensional structure: data source system identifier, event type, and field category, and encode this three-dimensional combination as a grouping key. The grouping key satisfies the formula... ,in Indicates the grouping key. Indicates the system identifier of the data source. Indicates the event type, This indicates the field category identifier, enabling subsequent calibration and routing of the mapping confidence score to be performed separately under the statistical distributions of different sources, different events, and different field categories; The embedded intelligent capture component generates dual threshold parameters based on the archiving mode parameters, the dual threshold parameters including a first threshold. With the second threshold ,in Used to define the lower confidence limit of automatically archived routes and Used to define the upper limit of confidence for manually reviewed routes, thereby providing fixed and auditable gating parameters for subsequent routing decisions; The embedded intelligent capture component writes data source connection information, event listening range, field differential extraction rules, transaction association rules, grouping rules, and dual threshold parameters into the capture configuration and persists it. The capture configuration is stored in a JSON structured document and includes a configuration version number, generation time, and configuration verification digest. The configuration verification digest is obtained by performing SHA-256 calculation on the full text of the capture configuration and is bound to the version number and stored in the local persistent directory, so that the capture configuration has consistency verification capability and can be traced and referenced in subsequent operations.

[0025] In this specific embodiment, S2 includes: The embedded intelligent capture component executes the capture based on the persistent capture configuration. First, the embedded intelligent capture component uses the data source connection information in the capture configuration to establish a long connection to the business system data source and initializes the connection pool, so that the maximum number of connections in the connection pool is constant at 20 and the connection timeout is constant at 30s. In this embodiment, the data source is determined to be a MySQL database and its change capture method is determined to be change data capture based on transaction log. Specifically, the database side enables binary logs and sets the log format to ROW and the row mirroring to FULL, so that the update event contains both the field value before the change and the field value after the change and can meet the input conditions for field difference extraction. The embedded intelligent capture component enables change monitoring based on the event monitoring scope in the capture configuration. The event monitoring scope is implemented as a joint filter of table-level whitelist and event type whitelist. The table-level whitelist is obtained by parsing the data source object identifier set in the capture configuration and corresponds one-to-one with the database name and table name in the database. The event type whitelist fixedly includes three categories: insert events, update events and delete events and is bound to the MySQL row event type mapping. When the embedded intelligent capture component starts monitoring, it reads the binary log file name and position offset from the locally persisted monitoring position and establishes a subscription cursor accordingly. It consumes binary log events sequentially in a single thread to form a change event stream and ensures that events within the same transaction enter the processing pipeline in log order. For each change event in the change event stream, the embedded intelligent capture component parses and obtains the event metadata. The event metadata permanently includes the data source system identifier, data source object identifier, operation type, occurrence time, and association identifier used for transaction aggregation. The data source system identifier is generated by a system identifier constant pre-written in the capture configuration and is consistent with the grouping rules in step S1. The meanings are consistent. The data source object identifier is generated from the database table to which the current event belongs and has the same meaning as the data source object identifier in the event listening scope of step S1. The operation type is fixedly encoded as one of the three values: INSERT, UPDATE, and DELETE. The occurrence time is taken from the timestamp in the database event header and uniformly converted to UTC millisecond timestamp. The association identifier is extracted from the transaction number field of the transaction to which the event belongs and the field name is fixed as tx_id and consistent with the transaction association rules in step S1. At the same time, the embedded intelligent capture component generates a unique event identifier for each change event to support idempotent processing. The unique event identifier is obtained by concatenating the binary log file name, position offset, and row sequence number triple and is transmitted along with the event metadata. The embedded intelligent capture component performs field difference extraction on change events according to the field difference extraction rules in the capture configuration. Specifically, in the INSERT event, the field value before the change is set to null and the field value after the change is taken from its own mirror. In the DELETE event, the field value after the change is set to null and the field value before the change is taken from its own mirror. In the UPDATE event, the field values ​​before and after the change in the row mirror are compared field by field and field difference is generated only for the changed fields. At the same time, the corresponding fields in the forced extraction list and the associated extraction list in the field difference extraction rules are forcibly merged into the difference set to ensure that the required elements and associated elements can be completely reconstructed in the change record. Each field difference includes four items: field identifier, field value before change, field value after change, and field value pattern information. The field identifier is taken from the database column name and serves as the input source for the field name feature in subsequent step S4. The field value pattern information consists of data type, length, value range, and unit identifier. The data type and length are read from the database data dictionary. The value range is generated according to the data type rules and is determined by one of three rules: minimum and maximum values ​​for integer types, effective number of digits range for floating-point types, and maximum length range for character types. The unit identifier is obtained by reading the unit code in the database column comment and performing a full match and normalization with the unit dictionary in step S1. When there is no unit code in the column comment, the unit identifier is set to an empty string. The embedded intelligent capture component combines event metadata and field differences according to a preset record format to generate change records and outputs them to a preset change record stream channel to form a change record stream. In this embodiment, the change record stream channel is determined to be a Kafka topic with a fixed topic name of change_record_stream and uses the transaction identifier tx_id as the message key to maintain the order of messages within the same transaction. At the same time, each change record is serialized and written in the form of a structured object, and its logical structure satisfies the formula. ,in Indicates the sequence number is Change records and This represents the record sequence number generated in a monotonically increasing manner according to the monitoring site. The event metadata representing the change record includes the data source system identifier, data source object identifier, operation type, occurrence time, associated identifier tx_id, and event unique identifier. The list represents the field difference list of the change record, and the list elements are field difference objects containing field identifiers, field values ​​before the change, field values ​​after the change, and field value pattern information. After Kafka confirms the successful write, the embedded intelligent capture component writes the corresponding binary log position as the new listening position back to the local persistent storage and performs duplicate consumption deduplication with the unique event identifier. Thus, even after the component restarts abnormally, it can still continue to generate a continuous stream of change events and change records from the last confirmed position and maintain the traceability and consistency of the change records.

[0026] In this specific embodiment, S3 includes: The transaction processing module performs streaming aggregation on the change record stream. As the sole consumer of the Kafka topic `change_record_stream`, the transaction processing module pulls each change record in the order of arrival. And parse event metadata The association identifier tx_id in the middle is used to determine the transaction identifier. The transaction identifier is consistent with the transaction association rule in step S1 and is used to aggregate change records belonging to the same database transaction into the same transaction cluster. The transaction processing module maintains an aggregated state table partitioned by time window for each transaction identifier in memory and uses RocksDB as the local state storage to enable restart recovery. The time window is taken from the transaction aggregated time window in step S1. Furthermore, in this embodiment, the UTC millisecond timestamp is aligned to a fixed 5-second hour boundary, thereby identifying change records with the same transaction identifier whose occurrence time falls within the same window as the same transaction cluster. The definition of a transaction cluster satisfies the formula: ; in Indicates the transaction identifier is And the window start time is transaction clusters, Indicates from event metadata The transaction identifier tx_id read from the middle, tx_id This represents the function that retrieves the transaction identifier from the event metadata. Indicates the sequence number is Change records and The record number is consistent with that in step S2. This represents the event metadata of the change record. Represents event metadata The occurrence time is in milliseconds. Indicates the start time of the aligned window in milliseconds. This represents the transaction aggregation time window, which is set to 5 seconds in this implementation. When writing change records to the aggregate state of the corresponding transaction cluster, the transaction processing module simultaneously writes a deduplication index with the event unique identifier as the key and discards duplicate change records when duplicate event unique identifiers are found, thereby ensuring the idempotency consistency of the record set within the transaction cluster. When the window of a transaction cluster closes and the current processing time exceeds the window closing time. And when the fixed delay tolerance time is 2 seconds, the transaction processing module closes the transaction cluster and sorts all change records within the transaction cluster, with the sorting primary key being the occurrence time. The order of changes within a transaction cluster is determined by ascending order of events and, when events occur at the same time, by using the lexicographical order of the unique event identifier as the stability secondary key. Based on the aforementioned order, the transaction processing module establishes reference relationships and generates causal chains between adjacent change records. Specifically, it sets the first change record in the sequence as the head node of the causal chain and sets its predecessor reference to null. The predecessor reference of the change record is set to the first in the sequence. The event unique identifier of each change record is written to the causal chain node table along with the previous record, so that the causal chain can restore the order of operations within the transaction in a single-chain structure where "each record references its previous record". The transaction processing module encapsulates the closed transaction cluster, its included list of change records, sorting results, causal chain node table, window start and end times, and transaction identifier into transaction cluster data and writes it to the transaction cluster repository `transaction_cluster_store`. In this embodiment, the transaction cluster repository is persisted using a relational table structure and uses the transaction identifier `tx_id` and window start time as its identifiers. A composite primary key is formed, and the table stores the unique event identifier of the head node of the causal chain to support subsequent steps of traversing the entire causal chain by the head node and extracting field differences.

[0027] In this specific embodiment, S4 includes: The field alignment module performs semantic field alignment on the transaction cluster data. The field alignment module first reads the composite primary key tx_id and the window start time from the transaction_cluster_store. The transaction cluster data is analyzed, and the unique event identifier of the head node of the causal chain is located. Then, the chain is traversed node by node along the predecessor reference direction to obtain the sequence of change records arranged in deterministic order within the transaction cluster. And iterate through each change record Time difference list of its fields Each field in the dataset is differentially analyzed to generate multi-view field features one by one; The multi-view field features are composed of four parts: field name features, field value pattern features, adjacent field context features, and event metadata context features, which are concatenated in a fixed order. All four parts are fixed-length vectors. The field name features are obtained by performing lowercase normalization and word segmentation on the field identifiers in the field difference to obtain up to 8 tags, which are then mapped to a tag index sequence and input into the field name encoder. The field name encoder is a 2-layer Transformer encoding network with 4 attention heads in each layer and uses a hidden layer representation with a dimension of 128 and a tag embedding vector with a dimension of 64. The field name encoder outputs the hidden layer vector with the first tag position and projects it through a linear layer to obtain the field name features with a dimension of 128. The field value pattern feature is determined by the field value pattern information in the field difference, and the field value pattern information includes four sub-items: data type, length, value range, and unit identifier. Specifically, the data type is mapped to 16 discrete IDs through a data type dictionary and embedded into a 16-dimensional vector; the length is mapped to 10 discrete IDs through length bucketing rules and embedded into an 8-dimensional vector; the value range is mapped to 10 discrete IDs through range bucketing rules and embedded into an 8-dimensional vector; and the unit identifier is determined through the following steps... The unit dictionary is mapped to discrete numbers and embedded as 8-dimensional vectors. The four types of embedded vectors are concatenated and input into a two-layer fully connected network to obtain field value pattern features with a dimension of 128. The number of channels in the two fully connected networks are as follows: And the ReLU activation function is used; Adjacent field context features are defined by changes to the same change record that differ from the field. The field difference list is determined and the selection rule for adjacent fields is fixed as follows: after sorting the data table column number corresponding to the data source object identifier in ascending order, the two adjacent field identifiers on the left and the two adjacent field identifiers on the right are taken. If there are less than 4, empty identifiers are used to fill in the gaps and the empty identifiers are mapped to all-zero vectors. Then, the 4 adjacent field identifiers are respectively processed by the field name encoder that shares parameters with the field name feature to obtain 4 128-dimensional vectors, and the average of the 4 vectors is taken to form a 128-dimensional adjacent field context feature. Event metadata context features are derived from event metadata. The data source system identifier, data source object identifier, operation type, and occurrence time are generated. The data source system identifier is embedded as a 16-dimensional vector through a system identifier dictionary, and its semantics correspond to the grouping key. Consistent with the data source object identifier, which is embedded as a 32-dimensional vector using an object identifier dictionary, and whose semantics are consistent with the data source object identifier in the event listening scope, the operation type is embedded as an 8-dimensional vector using an operation type dictionary, with values ​​limited to one of three types: INSERT, UPDATE, and DELETE. The occurrence time is bucketed by hour using UTC millisecond timestamps and embedded as a 16-dimensional vector. The four types of embedded vectors are concatenated and input into a two-layer fully connected network to obtain a 128-dimensional event metadata context feature. The number of channels in the two fully connected networks are as follows: And the ReLU activation function is used; The field alignment module concatenates the four 128-dimensional feature vectors in a fixed order into a 512-dimensional input vector, which is then input into the fusion encoder of the multi-view field contrast learning model to obtain the field semantic vector. The fusion encoder is a two-layer fully connected network with the number of channels sequentially as follows: The ReLU activation function is employed, and the encoder output is fused as a 256-dimensional field semantic vector for execution. Normalization is used to eliminate differences in modulus length, thereby ensuring consistency in similarity calculation. The field alignment module loads a set of standard field semantic vectors with the same version number as the archived standard field set read in step S1 from the standard field semantic vector library and constructs a matrix index in memory to support batch similarity calculation. Then, it calculates the cosine similarity between each field semantic vector and all standard field semantic vectors and determines the field mapping relationship and mapping confidence score based on the maximum similarity principle. The cosine similarity satisfies the formula: ; in Indicates the first The semantic vector of the field corresponding to the difference of the field and the field of the field are related to the field semantic vector of the field difference. The similarity between the semantic vectors of the standard fields corresponding to the archived standard fields. Indicates the first The semantic vector of the field differences is and its dimension is . Indicates the first The standard field semantic vector of each archive standard field has a dimension of . This represents the vector transpose operation. The second norm of a vector; The field alignment module will enable Standard field index for obtaining the maximum value The field mapping relationship for the field difference is determined. When multiple standard field indexes are tied for the largest, the unique mapping result is determined according to the rule of the smallest lexicographical order of the archived standard field identifier. The maximum similarity value is determined as the mapping confidence score to form the field semantic alignment result. The field semantic alignment result records the mapping relationship from the field identifier to the archived standard field identifier at the field difference granularity and associates it with the mapping confidence score. It is written to field_alignment_store along with the transaction cluster data.

[0028] In this specific embodiment, S5 includes: The archive generation module performs the operation, and the archive mode parameter is set to dual-track. The archive generation module reads the current transaction identifier (tx_id) and window start time from the transaction_cluster_store. The corresponding transaction cluster data is obtained and the causal chain and its sequentially arranged change record sequence are obtained. At the same time, the field semantic alignment results corresponding one-to-one with the transaction cluster data are read from field_alignment_store to obtain the field mapping relationship and mapping confidence score of each field difference. The archive generation module performs field conversion processing on each field difference list in the transaction cluster data according to the field mapping relationship. The field conversion processing includes replacing the field identifier in the field difference with the archive standard field identifier in the archive standard field set, and performing format normalization and unit normalization on the changed field values ​​in the field difference. The format normalization is fixed to unify the character values ​​to UTF-8 encoding and remove leading and trailing whitespace, and unify the date and time values ​​to UTC ISO8601 format strings while retaining millisecond precision, and converting the numeric strings to decimal. The data is parsed as a numeric type and retained to two decimal places. Unit normalization is fixed to convert length values ​​with unit identifiers "mm", "cm", and "m" to "m" with conversion factors of 0.001, 0.01, and 1 respectively. Mass values ​​with unit identifiers "g" and "kg" are converted to "kg" with conversion factors of 0.001 and 1 respectively. Temperature values ​​with unit identifiers "℃" are kept as "℃" without conversion. At the same time, when the unit identifier is missing in the field difference, the unit normalization process is set to not be executed and the empty unit identifier is recorded in the field value mode information. The archive generation module collects the archive standard field identifiers and their corresponding modified field values ​​obtained through field transformation and generates structured data under the archive standard field set as the basic archive data. Where multiple modified field values ​​appear for the same archive standard field identifier within the same transaction cluster, the chronological order is determined based on a causal chain, and the most recent modified field value is selected. The determination of "most recent" satisfies the following formula: ; in Indicates the archiving standard field identifier, This indicates that the archiving standard field in the basic archived data is identified as... The final value, Indicates the first in the causal chain The change record identifies the archiving standard fields. The resulting changed field values, This represents the index of change records that maximizes the occurrence time. The operator that takes the maximum value of the argument variable. Indicates the archiving standard field identifier within this transaction cluster. An index collection of records showing changes to field values ​​after modification. Indicates the first The occurrence time of each change record is taken from the event metadata and is a UTC millisecond timestamp; After generating basic archive data, the archive generation module reads the field value corresponding to the preset archive standard field identifier "controlled entity identifier" to determine the controlled entity identifier. It then generates an event-driven association full supplementary collection request based on the controlled entity identifier and writes it into the internal topic supplement_request_stream. The controlled entity identifier is used as the message key to ensure that supplementary collection requests with the same controlled entity identifier are processed serially. The request body of the supplementary collection request always includes the controlled entity identifier, the supplementary collection category set, and the request timestamp. The supplementary collection category set always includes three categories: controlled entity physical information, ownership certificate information, and management record information. The supplementary data collection executor subscribes to the supplement_request_stream and, upon receiving a supplementary data collection request, performs a consistent snapshot query on the data source by capturing the data source connection information in the configuration. Specifically, the physical information of the controlled entity is obtained by querying the entity information table with the controlled entity identifier as the primary key and retrieving the record corresponding to the latest update timestamp; the ownership certificate information is obtained by querying the ownership certificate table with the controlled entity identifier as the foreign key and returning all records in descending order of certificate effective time; and the maintenance record information is obtained by querying the maintenance record table with the controlled entity identifier as the foreign key and returning all records in descending order of record time. The supplementary data collection executor generates a data structure consistent with step S2 for each field of the three types of query results and reuses the field mapping relationship formed in step S4 to complete the replacement of archiving standard field identifiers and format normalization and unit normalization, thereby converting them into structured data under the archiving standard field set. The archive generation module merges the three types of structured data obtained from the supplementary collection with the basic archive data to generate an archive dataset. The merging rules are fixed so that the standard field identifier with the same name in the basic archive data has a coverage priority and the supplementary collection structured data is used to fill the fields that have not been assigned values ​​in the basic archive data. At the same time, the ownership certificate information and maintenance record information are written into the archive dataset in a descending sequence table structure according to their business time field. While generating the archive dataset, the archive generation module gathers the mapping confidence scores corresponding to each archive standard field identifier in the archive dataset to form a mapping confidence score set. The mapping confidence score of each archive standard field identifier is taken from the mapping confidence score corresponding to the field difference that produces the final value of the field. When the final value of the field comes from supplementary collection of structured data, it is taken from the mapping confidence score of the corresponding field difference during the supplementary collection and transformation process.

[0029] In this specific embodiment, S6 includes: The calibration routing module is executed after the archive generation module outputs the archived dataset and the set of mapped confidence scores. The calibration routing module first writes group metadata to each piece of structured data in the archived dataset and extracts the data source system identifier for each piece of data. Event Type With field category identifier The data source system identifier The event metadata is taken from step S2 and kept consistent with the basic archived data that triggered the supplementary acquisition in the supplementary acquisition and merging scenario. Event type The operation type is taken from the tail change record of the causal chain in step S3 and is fixedly inherited in the supplementary collection and merging scenario. Field category identifier The field category identifiers in the archiving standard field set read in step S1 are determined, and different field category identifiers are assigned to the basic archived data, controlled entity physical information, ownership certificate information and maintenance record information to ensure that the categories can be distinguished. The calibration routing module will select modules that meet the same grouping key according to the grouping rules in step S1. The structured data is divided into the same data group, and an independent calibrator instance is established for each data group. The calibrator instance is defined as a grouped conditional conformal prediction calibrator, and its calibration coverage parameter is fixed. In this embodiment, Write the capture configuration for auditing purposes; For each piece of structured data, the calibration routing module reads the mapping confidence scores of all archived standard field identifiers from the mapping confidence score set corresponding to that structured data and takes the minimum value as the mapping confidence score of that structured data. This allows subsequent routes to use the weakest field alignment result as a lower bound for confidence. For each group key The grouping condition conformal prediction calibrator loads the set of calibration samples corresponding to the grouping key from the calibration sample library and constructs a set of inconsistencies in scores. The calibration sample library is generated from historical manual review results and only includes samples that have been manually reviewed and determined to have "correct field mapping relationships," thus ensuring the authenticity of the calibration samples. The calibration sample library is configured for each grouping key. Fixed storage recently Each calibration sample contains a mapped confidence score, which is then converted into a non-consistent score. The inconsistency score It is defined as 1 minus the mapping confidence score of the calibration sample; The mapping confidence score of the grouped conditional conformal prediction calibrator to the current structured data Calculate the calibrated confidence level And Write back the routing metadata of this structured data, the calibrated confidence level. Experience using conformal prediction - Value is calculated and satisfies: ; in Indicates the grouping key is Time-based confidence score of mapping The output of the calibration function, This represents the calibrated confidence level and its value range is [value missing]. Indicates the grouping key and is identified by the data source system. Event Type With field category identifier composition, This represents the mapping confidence score of the current structured data, taken as the minimum value of the mapping confidence score corresponding to each archived standard field identifier within the structured data. Indicates the grouping key is The set of inconsistent scores and by The non-consistency scores of each calibration sample constitute the total score. Represents a set The Middle The inconsistency score of each calibration sample, symbol Represents the cardinality operation for sets. Indicates from set Select from those that meet the condition that the inconsistency score is not less than the current inconsistency threshold. A subset formed by the elements of Indicates the grouping key is The number of calibration samples is constant at 2000 in this embodiment; The calibration routing module will calibrate the confidence level. The first threshold is compared with the dual threshold parameters written in step S1 to capture the configuration and generate a routing decision, wherein the second threshold is... Used to generate automatically archived routes and when The structured data is then labeled as an automatic archiving route, with a second threshold. Used to generate manually reviewed routes and when The structured data is then marked as a route for manual review, and at the same time... The structured data is then labeled as a correction and supplementary data acquisition route; The calibration routing module will automatically archive routes, correct and supplement routes, and manually verified routes, along with their annotation results and grouping keys. Mapping confidence score With calibrated confidence level The archived data after routing is written together and output to route_result_store so that step S7 can trigger the generation of archive encapsulation, event-driven correction and supplementary data collection, or manual review tasks according to the routing type.

[0030] In this specific embodiment, S7 includes: The archiving execution module performs corrective and supplementary data collection, manual review, or automatic archiving processing on the archived data output by route_result_store according to the routing decision. When the routing decision is a correction and supplementary acquisition route, the archiving execution module reads the controlled entity identifiers from the archived data following that route and scans the archived dataset to locate missing and insufficient confidence fields. The missing fields are defined as a set of archiving standard field identifiers that are marked as required and have empty values ​​in the archived dataset. The insufficient confidence fields are defined as the calibrated confidence levels in the archived dataset. Less than the first threshold Furthermore, the corresponding field mapping is a set of archiving standard field identifiers with a confidence score less than 0.85. The archiving execution module merges the missing fields and fields with insufficient confidence into a set of fields to be supplemented and generates a supplementary collection request. The supplementary collection request always includes a controlled entity identifier, a set of fields to be supplemented, a set of supplementary collection categories, and a request timestamp. The supplementary collection category set always includes three categories: controlled entity physical information, ownership certificate information, and management record information. The archiving execution module writes the supplementary collection request into the `supplement_request_stream` and it is subscribed to and processed by the supplementary collection executor. The server performs a consistency snapshot query on the data source based on the controlled entity identifier and returns only the fields corresponding to the set of fields to be supplemented. Then, it reuses the field conversion processing of step S5 to complete the replacement of the archive standard field identifier, format normalization and unit normalization, and merges it with the original archive dataset to update the archive dataset. At the same time, it updates the mapping confidence score set with the mapping confidence score corresponding to the field difference that produces the final value in the supplementary collection result, and writes the updated archive dataset and mapping confidence score set back to route_result_store, thereby triggering the return to step S6 for recalibration and routing. When the routing decision is to manually review the route, the archiving execution module generates a review task and writes it into the review_task_queue. The review task is fixedly associated with the transaction cluster data, causal chain, field semantic alignment result, mapping confidence score set, and calibrated confidence level of the archived data after the route. Grouping keys And a routing decision reason field, which is fixedly written as "calibrated confidence level is less than the second threshold". "The verification terminal will then display the field mapping relationships that require manual confirmation and the source of the field values." When the routing decision is automatic archiving, the archiving execution module generates an archive package and writes it to the archive storage system. The archive package adopts a ZIP container structure and always contains the following files: archive_dataset.json, transaction_cluster.json, causal_chain.json, field_alignment.json, calibrated_confidence.json, compliance_check.json, manifest.json, signature.bin, and timestamp.tst. The archiving execution module calculates data fingerprints for the archived dataset and writes the data fingerprints to manifest.json. The data fingerprints satisfy the formula. ,in This represents a data fingerprint, which is a hexadecimal string of length 64. SHA256() represents the SHA-256 cryptographic hash function. This indicates the byte sequence obtained by performing normalized serialization on the archived dataset, wherein the normalized serialization is fixed to output archive_dataset.json in UTF-8 encoding, sort the keys according to the lexicographical order of the archived standard field identifiers, remove all semantically meaningless whitespace characters, and unify the numerical values ​​into decimal text representation; The archiving execution module performs element integrity verification based on a preset set of compliance rules and generates verification records written to compliance_check.json. In this embodiment, the preset set of compliance rules includes rules for verifying the non-empty nature of required fields, rules for verifying the consistency of field data types, rules for verifying the validity of unit identifiers, rules for verifying the consistency of controlled entity identifiers, and rules for verifying the traversability of causal chains. Each rule outputs a rule identifier, verification time, verification input summary, verification result, and failure reason fields, thereby providing auditable compliance evidence within the archive package. The archiving execution module performs a digital signature on manifest.json and writes the signature result to signature.bin. The digital signature is completed using the RSA private key corresponding to the X.509 certificate, with a fixed key length of 2048 bits and using the SHA256withRSA signature algorithm. Subsequently, the archiving execution module submits the hash of manifest.json to the timestamp server and obtains a timestamp token in RFC3161 format, which is written to timestamp.tst to achieve time solidification after signing. After the archive execution module completes the ZIP containerization, it writes the archive package to the archive storage system using an object key consisting of the archive number and the archive time, and simultaneously writes it to the index table to record the controlled entity identifier, transaction identifier tx_id, and window start time. Data fingerprint calibrated confidence level The signature certificate serial number enables subsequent signature verification, validation, traceability, and auditing.

[0031] In this specific embodiment, the training of the multi-view field comparison learning model and the generation of the standard field semantic vector library are completed on an offline training server and are consistent with the model structure called in step S4. The training data consists of historical change records from two business systems, with each business system corresponding to a different data source system identifier. Historical change records are converted into change records containing event metadata and a field difference list according to step S2 and written to the historical change record database. Then, according to the transaction association rules in step S3, they are identified by transaction identifiers. With time window Data is aggregated to form transaction clusters and causal chains are established within the transaction clusters to ensure that the training samples contain field differences, adjacent field contexts, and event metadata contexts consistent with those in the online dataset. For each transaction cluster of data, the training pipeline traverses each field difference along the causal chain and generates multi-view field features according to step S4. The multi-view field features are obtained by concatenating field name features, field value pattern features, adjacent field context features and event metadata context features in a fixed order. Their feature construction, word segmentation rules, dictionary mapping rules, embedding dimensions and encoding network structure are consistent with step S4, thereby ensuring that the input distribution is consistent between the training stage and the inference stage. The training pipeline constructs a field mapping pseudo-annotation set based on the archived standard field set and field semantic generation rules. The field semantic generation rules generate pseudo-annotations with deterministic constraints and do not rely on manual interaction. The field name feature similarity constraint calculates the Jaccard similarity of the tag set after field identifier word segmentation and requires that the similarity is not less than 0.60. The field value pattern feature consistency constraint requires that the data type number is consistent and the unit identifier is consistent, and the difference between the length bucket number is not greater than 1 and the difference between the value range bucket number is not greater than 1. The adjacent field context feature constraint requires that at least one of the four adjacent fields of the difference of the field satisfies the field name feature similarity constraint and the field value pattern feature consistency constraint. The event metadata context feature constraint requires that the data source object identifier is mapped to the same object category identifier and the operation type belongs to the same category set. The object category identifier is given by the static mapping table in the data source object identifier dictionary and the mapping table is fixed at the beginning of training and bound to the capture configuration version number. The training pipeline calculates the satisfaction of four types of constraints for each multi-view field feature and each archive standard field, and uses the sum of the satisfaction numbers as the pseudo-annotation score. The pseudo-annotation score is required to be no less than 3 points before it is written into the field mapping pseudo-annotation set. At the same time, when multiple archive standard fields have the same highest score, the unique correspondence is determined according to the rule of the smallest lexicographical order of the archive standard field identifier. This makes the field mapping pseudo-annotation set unique and reproducible under a given input. The training pipeline constructs positive and negative sample pairs based on the field mapping pseudo-label set. A positive sample pair consists of two multi-view field features mapped to the same archiving standard field, and the two multi-view field features come from different data source system identifiers. To enhance cross-system alignment capabilities, negative sample pairs are composed of two multi-view field features mapped to different archiving standard fields, and one type of negative sample pair satisfies the field name feature similarity constraint while not satisfying the field value pattern feature consistency constraint, so as to form negative samples with similar field names but different semantics. The training pipeline performs hard negative sample filtering on negative sample pairs, and the filtering method is fixed to use the multi-view field contrast learning model to be trained to forward compute the field semantic vector of the negative sample pairs and calculate the cosine similarity. Negative sample pairs with a cosine similarity greater than 0.70 are selected as hard negative sample pairs and added to the training batch to improve the learning strength of the discrimination boundary. The multi-view field comparison learning model is trained end-to-end according to the network structure in step S4, with parameters initialized using a Xavier uniform distribution. Training uses the Adam optimizer with a fixed learning rate. And the weight decay is fixed at The training batch size is fixed at 256, and the number of positive sample pairs and the number of hard negative sample pairs in each batch are fixed at 128. The number of training rounds is fixed at 50, and the entire set of pseudo-labeled samples is traversed once in each round. The training loss function employs a contrastive learning loss function and is applied to the field semantic vector. After normalization, the contrastive learning loss function is calculated to satisfy the following formula: ; in This represents the contrastive learning loss for a single batch. This represents the number of semantic vectors of the fields involved in the comparison in a single batch, and in this implementation, it is taken as 256. Indexes representing the semantic vectors of fields. Representation and Index The semantic vector index of another field is compared with the corresponding field semantic vector. Indicates that the index is The field semantic vector and the multi-view field comparison learning model for the first The forward output of the multi-view field features is consistent with the definition in step S4. Indicates and The field semantic vectors that constitute positive sample pairs are guaranteed to map to the same archived standard field by the field mapping pseudo-annotation set. Represents the semantic vector of the field With field semantic vector cosine similarity, Represents the temperature coefficient and is taken as in this embodiment. Represents an exponential function; After 50 rounds of training, the training pipeline freezes the parameters of the multi-view field contrastive learning model and generates a standard field semantic vector library. The generation method involves forward calculating the semantic vectors of all multi-view field features mapped to the same archived standard field from the field mapping pseudo-annotation set, grouping them according to the archived standard field identifier, and then taking the arithmetic mean of the semantic vectors within the group corresponding to each archived standard field identifier to obtain the standard field semantic vector for that archived standard field. This process is then repeated. Normalization is performed to maintain consistency with the similarity calculation in step S4. Subsequently, the standard field semantic vector, along with the archived standard field identifier and the archived standard field set version number, is written into the standard field semantic vector library and the version number is used as the index key for loading and calling in step S4.

[0032] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

[0033] This invention employs a combination of embedded event-driven capture, self-supervised field semantic alignment, and confidence-gated routing algorithms to transform archiving capabilities from "passive push from business systems" to "continuous, automatic, and supervised proactive acquisition from the data source side." Specifically, the event stream and field differencing generated by change data capture preserve the original change facts of the archives, while the transaction clusters and causal chains formed by transaction aggregation restore the timing and dependencies of business operations, thus providing a more accurate foundation for single-track archiving. Building upon this, the multi-view field features generated by field differencing achieve cross-system field semantic alignment through comparative learning, reducing mapping errors caused by differences in field naming and structure between heterogeneous business systems, and quantifying alignment reliability with mapping confidence scores. Subsequently, the confidence scores are calibrated through conformal prediction based on grouping conditions, and routing is performed using dual thresholds, forming a closed-loop control system for automatic archiving, correction and supplementary data collection, and manual review. This mechanism reduces omissions, delays, and caliber drift, and during automatic archiving, data fingerprints, causal chains, semantic alignment results, and verification records are reliably encapsulated and solidified, enhancing the verifiability and auditability of the archiving results.

[0034] This invention addresses the aforementioned technical issues through algorithmic improvements: First, it expands traditional incremental acquisition at the record level to dual-granularity capture at the event and field levels, introducing transaction clusters and causal chain organization changes. This allows field semantic judgment to not only rely on static field names but also leverage intra-transaction context and temporal relationships to enhance stability and traceability. Second, it improves field mapping from single-rule matching to multi-view comparative learning representation, jointly modeling field name features, field value pattern features, adjacent field context features, and event metadata context features. Hierarchical alignment can be superimposed to further constrain mapping consistency across different object levels, thus better adapting to heterogeneous systems and improving alignment accuracy. Third, it improves direct threshold judgment of similarity scores to grouping condition conformal prediction calibration plus dual-threshold gating. Uncertainty is calibrated within groups of different source systems, event types, and field categories, making routing decisions more interpretable and controllable. Furthermore, it continuously improves the integrity and reliability of archived data through corrective acquisition paths.

Claims

1. A data archiving and processing method driven by an embedded intelligent capture component, characterized in that, include: S1. Deploy the embedded intelligent capture component, read the archive mode parameters and archive standard field set, and generate the capture configuration, including data source connection information, event listening range, and rules and parameters related to field differential extraction, transaction association, data grouping, and dual threshold routing. S2. Based on the capture configuration, perform change data capture on the data source of the business system to form a change event stream, and generate a change record stream according to the field difference extraction rules; S3. Aggregate the change record stream according to the transaction association rules to obtain a transaction cluster, and establish a causal chain of change records within the transaction cluster to generate transaction cluster data; S4. Generate multi-view field features based on causal chain field difference, input the multi-view field comparison learning model to output field semantic vectors, and calculate the similarity with the standard field semantic vectors corresponding to the archived standard field set to obtain field semantic alignment results, including field mapping relationships and mapping confidence scores; S5. Convert transaction cluster data into structured data under the archived standard field set according to field mapping relationships to obtain basic archived data. In single-track system, the basic archived data is used as the archived dataset. In dual-track system, the controlled entity identifier is determined based on the basic archived data, and the physical information, ownership certificate information and maintenance record information of the controlled entity associated with the controlled entity identifier are collected in an event-driven manner and then merged to generate the archived dataset. The mapping confidence scores of the archived dataset are collected to form a mapping confidence score set; S6. Group the archived dataset according to grouping rules, perform grouping condition conformal prediction calibration on each group based on the mapping confidence score set to obtain calibrated confidence, and compare the calibrated confidence with dual threshold parameters to generate routing decisions to obtain the routed archived data; S7. After performing event-driven correction and supplementary data collection on the data of the corrected and supplemented routes and updating the archived dataset, return to S6. Output the review task for the data of the manually reviewed routes and associate it with transaction cluster data, field semantic alignment results and calibrated confidence. Generate an archive package for the data of the automatically archived routes, which includes the data fingerprint, causal chain, field semantic alignment results, calibrated confidence and verification records generated according to the preset compliance rule set of the archived dataset. After electronically signing, timestamping and containerizing the archive package, write it to the archive storage system.

2. The data archiving and processing method driven by the embedded intelligent capture component according to claim 1, characterized in that, S1 includes: The embedded intelligent capture component is started in the operating environment of the business system; The archive mode parameters are obtained through the preset configuration interface, and the set of archive standard fields is obtained from the preset configuration library. The event monitoring scope is determined based on the archiving mode parameters. When the archiving mode parameters indicate single-track archiving, the event monitoring scope is limited to the data source objects and event types corresponding to the original archive data. When the archiving mode parameters indicate dual-track archiving, the event monitoring scope further includes the data source objects and event types corresponding to the physical information of the controlled entity, ownership certificate information, and management record information. Field difference extraction rules are generated based on the archiving standard field set. These rules are used to at least limit the range of fields to be extracted and the record format for field differences. Transaction association rules are generated based on the archiving mode parameters. These transaction association rules are used to at least limit the extraction location of the transaction identifier and the time window for transaction aggregation. Grouping rules are generated based on archiving mode parameters. These grouping rules are at least used to limit the data grouping method based on data source system identifier, event type, and field category. Based on the archiving mode parameters, dual threshold parameters are generated. The dual threshold parameters include a first threshold and a second threshold used to distinguish between automatically archived routes, correction and supplementary acquisition routes, and manually reviewed routes. The data source connection information, event listening range, field difference extraction rules, transaction association rules, grouping rules, and dual threshold parameters are written into the capture configuration and persisted to obtain the capture configuration.

3. The data archiving and processing method driven by the embedded intelligent capture component according to claim 1, characterized in that, S2 includes: A connection is established with the data source of the business system based on the data source connection information in the capture configuration, and change listening is enabled for the data source based on the event listening range in the capture configuration, so as to continuously acquire change events that match the event listening range and form a change event stream; For each change event in the change event stream, event metadata is parsed to obtain the event metadata, which includes the data source system identifier, data source object identifier, operation type, occurrence time, and association identifier used for transaction aggregation; Based on the field difference extraction rules in the capture configuration, the set of fields that have changed is extracted from the change event, and field differences are generated for each field in the set of fields. Each field difference includes a field identifier, the field value before the change, the field value after the change, and field value pattern information. The event metadata and the field differences are combined according to a preset record format to generate a change record containing event metadata and a list of field differences; The change records are output to a preset change record stream channel or change record queue to form a change record stream.

4. The data archiving and processing method driven by the embedded intelligent capture component according to claim 1, characterized in that, S3 includes: extracting transaction identifiers for transaction aggregation from the change records in the change record stream, wherein the transaction identifiers are determined from the association identifiers in the event metadata according to the transaction association rules in the capture configuration; setting a time window for transaction aggregation according to the transaction association rules in the capture configuration, and aggregating change records with the same transaction identifier and occurrence time falling within the same time window into the same transaction cluster to form multiple transaction clusters; for each transaction cluster, sorting the change records within the transaction cluster according to the occurrence time of the change records to obtain a sorting result; establishing a reference relationship between adjacent change records according to the sorting result, and setting each change record to reference its previous change record to form a causal chain corresponding to the transaction cluster; storing the transaction clusters and the causal chains in association to generate transaction cluster data containing the causal chains.

5. The data archiving and processing method driven by the embedded intelligent capture component according to claim 1, characterized in that, S4 includes: For the causal chain in the transaction cluster data, traverse each field difference in the causal chain and generate multi-view field features. The generation of multi-view field features includes: extracting the field identifier corresponding to the field difference to generate field name features. Based on the field values ​​before and after the change in the field difference, the data type, length, value range, and unit information of the field values ​​are determined to generate field value pattern features; Extract the field identifier and field value pattern information of adjacent fields from the field difference list that is in the same change record as the field difference to generate adjacent field context features; Extract the data source system identifier, data source object identifier, operation type, and occurrence time from the event metadata of the change record corresponding to the field difference to generate event metadata context features; The field name feature, field value pattern feature, adjacent field context feature, and event metadata context feature are combined to obtain the multi-view field feature corresponding to the field difference; The multi-view field features are fed into a trained multi-view field contrast learning model to output a field semantic vector. Obtain the standard field semantic vectors corresponding to the archived standard field set from the pre-established standard field semantic vector library; Calculate the similarity between the semantic vector of a field and the semantic vector of each standard field, and determine the field mapping relationship based on the principle of maximizing similarity; The maximum similarity value is determined as the mapping confidence score, and a field semantic alignment result is generated, which includes the field mapping relationship and the mapping confidence score.

6. The data archiving and processing method driven by the embedded intelligent capture component according to claim 1, characterized in that, S5 includes: based on the field mapping relationship in the field semantic alignment result, performing field transformation processing on each field difference list in the transaction cluster data. The field transformation processing includes replacing the field identifier with the corresponding archive standard field identifier in the archive standard field set, and performing format normalization and unit normalization processing on the changed field values ​​in the field difference. The archive standard field identifiers obtained through the field transformation processing and their corresponding changed field values ​​are then aggregated. In cases where multiple changed field values ​​occur for the same archive standard field identifier within the same transaction cluster, the time order is determined based on the causal chain in the transaction cluster data, and the changed field value with the latest time order is selected to generate structured data under the archive standard field set to obtain the base... The system first compiles basic archive data. When the archive mode parameter indicates single-track archive, the basic archive data is determined as the archive dataset. When the archive mode parameter indicates dual-track archive, the field values ​​corresponding to the preset archive standard fields are read from the basic archive data to determine the controlled entity identifier. Based on the controlled entity identifier, an event-driven association full-volume supplementary collection is triggered to obtain the controlled entity physical information, ownership certificate information, and maintenance record information associated with the controlled entity identifier. The controlled entity physical information, ownership certificate information, and maintenance record information are converted into structured data under the archive standard field set and then merged with the basic archive data to obtain the archive dataset. At the same time, the mapping confidence scores corresponding to each archive standard field identifier in the archive dataset are collected to form a mapping confidence score set.

7. The data archiving and processing method driven by the embedded intelligent capture component according to claim 1, characterized in that, S6 includes: Based on the grouping rules in the capture configuration, extract the data source system identifier, event type, and field category for each structured data in the archived dataset, and divide it into at least one data group according to the principle that the data source system identifier, event type, and field category are consistent; For each data group, read the set of mapped confidence scores corresponding to the data group, and perform grouped conditional conformal prediction calibration on the set of mapped confidence scores according to the preset calibration coverage parameter to obtain the calibrated confidence level corresponding to the data group; For each piece of structured data, its calibrated confidence level is compared with the dual threshold parameters in the capture configuration. An automatic archiving route is generated when the calibrated confidence level is not less than the first threshold, a correction and supplementary acquisition route is generated when the calibrated confidence level is less than the first threshold but not less than the second threshold, and a manually reviewed route is generated when the calibrated confidence level is less than the second threshold. The automatic archiving route, the correction and supplementary acquisition route, and the manually reviewed route are aggregated to form a routing decision, and the archived dataset is labeled according to the routing decision to generate the routed archived data.

8. The data archiving and processing method driven by the embedded intelligent capture component according to claim 1, characterized in that, S7 includes: For archived data following a route where the routing decision is to correct and supplement the route, a supplementary acquisition request is generated based on missing or insufficiently confident archiving standard fields and corresponding controlled entity identifiers in the archived data. An event-driven corrective supplementary acquisition is then performed using an embedded intelligent capture component to obtain the supplementary acquisition result. The supplementary acquisition result is converted into structured data under the archiving standard field set and merged with the archived dataset to update the archived dataset. Simultaneously, the mapping confidence score set is updated, and the process returns to step S6. For archived data following a route where the routing decision is to manually review the route, a review task is generated. This review task is at least associated with the transaction cluster data and fields corresponding to the archived data following the route. The semantic alignment results and calibrated confidence scores are output to the verification terminal. For the archived data after the routing decision is an automatic archiving route, an archive package is generated. The generation of the archive package includes calculating the data fingerprint of the archived dataset, writing the causal chain, field semantic alignment results, and calibrated confidence scores in the transaction cluster data into the archive package, performing element integrity verification according to a preset set of compliance rules and generating verification records, writing the verification records into the archive package, performing electronic signature and timestamp solidification on the archive package, and containerizing the archive package in a preset electronic file container format. The containerized archive package is then written into the archive storage system.

9. The data archiving and processing method driven by the embedded intelligent capture component according to claim 1, characterized in that, The training of the multi-view field comparison learning model includes: obtaining historical change records from at least two business systems, and generating change records containing event metadata and a field difference list according to step S2; The change records are aggregated according to the transaction association rules of S3 to form transaction cluster data, and a causal chain is established within each transaction cluster. For each field difference in the causal chain, generate multi-view field features according to S4; Based on the archive standard field set and the preset field semantic generation rules, a field mapping pseudo-annotation set is constructed. The field semantic generation rules include field name feature similarity constraints, field value pattern feature consistency constraints, adjacent field context feature constraints, and event metadata context feature constraints. The field mapping pseudo-annotation set is used to indicate the correspondence between multi-view field features and archive standard fields in the archive standard field set. Positive sample pairs are constructed based on the field mapping pseudo-label set. The two multi-view field features in the positive sample pair are mapped to the same archiving standard field, and the positive sample pair includes positive sample pairs between multi-view field features from different business systems. Negative sample pairs are constructed based on the field mapping pseudo-label set. The two multi-view field features in the negative sample pairs are mapped to different archiving standard fields. The negative sample pairs include negative sample pairs with similar field name features and different field value pattern features. The negative sample pairs are subjected to hard negative sample screening, which includes using a multi-view field contrast learning model to be trained to calculate the similarity of the negative sample pairs and selecting negative sample pairs with a similarity greater than a preset similarity threshold as hard negative sample pairs. The multi-view field contrast learning model is iteratively trained using a contrastive learning loss function. The contrastive learning loss function is used to increase the similarity of the field semantic vectors of positive sample pairs and decrease the similarity of the field semantic vectors of negative sample pairs until a preset convergence condition is met or a preset iteration threshold is reached, thereby obtaining the trained multi-view field contrast learning model. Based on the trained multi-view field comparison learning model, the multi-view field features mapped to the same archive standard field in the field mapping pseudo-annotation set are used to generate corresponding field semantic vectors, and the field semantic vectors are aggregated to generate standard field semantic vectors corresponding to the archive standard field.

10. A data archiving and processing system driven by an embedded intelligent capture component, used to execute the data archiving and processing method driven by an embedded intelligent capture component as described in any one of claims 1 to 8, characterized in that, include: The capture configuration module is used to read the archive mode parameters and the archive standard field set to generate the capture configuration; The change capture module is used to capture change data of the data source of the business system according to the capture configuration, form a change event stream, and generate a change record stream according to the field difference extraction rules. The transaction processing module is used to aggregate the change record stream according to transaction association rules to obtain a transaction cluster, and to establish a causal chain of change records within the transaction cluster to generate transaction cluster data; The field alignment module is used to generate multi-view field features based on causal chain field difference. It takes the multi-view field comparison learning model as input and outputs the field semantic vector. It calculates the similarity with the standard field semantic vector corresponding to the archived standard field set to obtain the field semantic alignment result, including the field mapping relationship and mapping confidence score. The archive generation module is used to convert transaction cluster data into structured data under the archive standard field set according to the field mapping relationship to obtain the archive dataset, and to collect the mapping confidence scores to form the mapping confidence score set. In the dual-track system, the controlled entity identifier is determined based on the archive dataset, and the controlled entity physical information, ownership certificate information and maintenance record information associated with the controlled entity identifier are collected in an event-driven manner and then merged to update the archive dataset. The calibration routing module is used to group the archived dataset according to the grouping rules, perform grouping conditional conformal prediction calibration on each group based on the mapping confidence score set to obtain the calibrated confidence level, and compare it with the dual threshold parameters to generate routing decisions, outputting automatic archived routes, correction and supplementary sampling routes, or manually verified routes. The archiving execution module is used to generate an archive package for the data of the automatically archived route and write it to the archive storage system. The archive package contains the data fingerprint, causal chain, field semantic alignment result, calibrated confidence level and compliance verification record of the archived dataset, and performs electronic signature, timestamp solidification and containerization encapsulation. It performs event-driven correction and supplementary acquisition of the data of the correction and supplementary acquisition route and updates the archived dataset. It outputs a review task for the data of the manually reviewed route and associates it with transaction cluster data, field semantic alignment result and calibrated confidence level.