Automatic archive filing method based on multi-source heterogeneous data fusion
By combining structural diagram modeling and the PHATE algorithm, a multi-source heterogeneous data fusion method was developed, which solved the problem of inconsistent archival data integration and achieved a high-precision and automated archiving process, applicable to government, finance and other fields.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-02
- Publication Date
- 2026-03-17
AI Technical Summary
Existing technologies struggle to effectively integrate structured, semi-structured, and unstructured archival data, lack deep structural representation, have limited accuracy in similarity measurement, insufficient robustness in embedding representation, and lack adaptive mechanisms in label prediction, resulting in unstable and low-accuracy archiving results.
By employing structural graph modeling, slice Gromov-Wasserstein distance calculation, PHATE potential thermal diffusion embedding, structural perturbation correction, and label prediction methods, multi-source heterogeneous archive data is processed in a unified manner to achieve intelligent semantic classification and automatic path writing.
It improves the accuracy and automation of archiving, ensures data consistency and accuracy, reduces the need for manual intervention, and is suitable for multimodal archival management systems in government and finance.
Smart Images

Figure CN121681884A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data management and intelligent information processing, and particularly relates to an archive automatic filing method based on multi-source heterogeneous data fusion. BACKGROUND
[0002] With the large-scale deployment of information systems and the growth of document management needs, enterprises, institutions and public sectors have accumulated a large amount of archive data with diverse sources and complex structures. These archive data not only include structured information such as tables, database entries, but also cover a large amount of semi-structured documents, scanned images, text records and other unstructured data. How to efficiently and accurately classify and manage these heterogeneous archive data has become an important research topic in the field of archive management.
[0003] In the prior art, the archive automatic filing system usually adopts a processing method based on keyword matching, template rules, shallow text classification or traditional machine learning algorithms. Although these methods have certain practicability in specific scenarios, they still have the following main deficiencies when facing the filing task of multi-source heterogeneous archives: Limited modal fusion capability: the existing technology mainly processes single data type, and it is difficult to effectively integrate structured, semi-structured and unstructured archive data, and lacks a unified modeling framework. Lack of depth in structure representation: traditional methods usually rely on flat field extraction or shallow semantic analysis, and cannot fully extract the field level relationship and topological structure contained in the archive data, resulting in insufficient utilization of structure information. Limited precision of similarity measurement: in the similarity analysis between structure graphs, simple edit distance, label matching and other methods are often used, which cannot accurately capture the deep alignment relationship between complex structures. Insufficient robustness of embedding representation: most filing methods do not consider the potential bias caused by data structure disturbance, lack of embedding consistency constraints, and are prone to unstable filing results or distorted semantic expression. Lack of adaptive mechanism in label prediction: the filing label mainly depends on static training model or fixed matching rule, which cannot dynamically respond to the changes in the distribution of filing data structure, and the filing accuracy and generalization ability are limited. Insufficient automatic filing path configuration.
[0004] Therefore, how to provide an archive automatic filing method based on multi-source heterogeneous data fusion is a problem that those skilled in the art need to solve. SUMMARY
[0005] One purpose of the present application is to propose an archive automatic filing method based on multi-source heterogeneous data fusion, which adopts structural graph modeling, slice Gromov-Wasserstein distance calculation, PHATE potential heat diffusion embedding, structural disturbance correction and label prediction technology path, uniformly processes structured, semi-structured and unstructured archive data, realizes intelligent semantic classification and path automatic writing of archives, and has the advantages of high filing accuracy, high automation degree of processing flow and strong multi-format compatibility.
[0006] According to an embodiment of the present application, an archive automatic filing method based on multi-source heterogeneous data fusion comprises the following steps: S1, obtaining structured archive data, semi-structured archive data and unstructured archive data from at least two preset data sources, performing format conversion on each type of archive data, and generating a unified representation data set; S2, performing structure modeling operation on each archive data in the unified representation data set, extracting field level and topological features, constructing corresponding structure graph, and forming structure graph set; S3, based on slice Gromov-Wasserstein distance, calculating the structural similarity between any two structure graphs in the structure graph set, and generating a structure similarity matrix; S4, taking the structure similarity matrix as input, applying PHATE algorithm to simulate heat diffusion process, and extracting semantic embedding representation of the corresponding archive in the potential space; S5, performing disturbance measurement on the semantic embedding representation, calculating the structural deviation of each embedding vector relative to the original structure graph; when the deviation exceeds the preset disturbance measurement threshold, adjusting the embedding value based on the structural deviation control function to generate an embedding vector set that meets the structural consistency; S6, predicting the filing label of each archive data based on the distribution mode in the embedding vector set; S7, writing the archive data corresponding to the filing label into the specified filing category path to complete automatic filing.
[0007] The application establishes a complete process from data acquisition, structure modeling, similarity measurement, semantic embedding, structure consistency adjustment to automatic label prediction and path writing, effectively solves the problems of various data formats, inconsistent structures and low archiving efficiency in existing archive management systems. The method measures the similarity between structure graphs through slice Gromov-Wasserstein distance, and generates high-quality semantic embedding representation by combining the PHATE algorithm to simulate the heat diffusion process. Through structure deviation control and adjustment of embedding vectors, the consistency and accuracy of data are ensured, and finally the whole process of automatic archiving is realized through archiving label prediction and path automatic writing. The method significantly improves the archiving efficiency and accuracy, reduces the demand for manual intervention, has wide practical application value, and has important popularization significance in government, financial and other multi-modal archive management systems.
[0008] Optionally, the S1 specifically comprises: S11, extracting archive data from two or more than two preset data sources respectively, the archive data including structured archive data, semi-structured archive data and unstructured archive data, each type of archive data being collected in original format; S12, for structured archive data, field renaming, type unification and coding specification processing are carried out based on field mapping template, and a data item set with unified field definition is generated; S13, for semi-structured archive data, the data node structure is parsed, the key-value pair relationship with identification is extracted, and it is converted into two-dimensional field representation, and it is aligned with structured data format through mapping relationship table; S14, for unstructured archive data, a preset text parsing rule and an entity recognition model are used to extract key entities and superordinate and subordinate semantic structures in the text, and a structured fragment representation with entity relationship as a unit is constructed; S15, the above three types of processed archive data are respectively mapped to the field set defined in the unified data format template, the missing fields are filled with default placeholders, the redundant fields are discarded, and a unified representation data set is constructed.
[0009] By format conversion and unified processing of multi-source heterogeneous archive data, the application ensures that the archive data from different data sources can be standardized and represented according to a unified template. In specific implementation, the structured data is renamed, type-unified and coding-normalized through a field mapping template, ensuring the consistency and comparability of the data; the semi-structured data is converted into a two-dimensional field representation aligned with the format of the structured data by analyzing the node structure and extracting key-value pairs; for unstructured data, a text parsing rule and an entity recognition model are used to extract key entities and their relationships in the text and convert them into a structured fragment representation. Finally, all data are mapped to a unified format, missing fields are filled and redundant fields are discarded to generate a consistent data set. This method effectively solves the problems of format difference and information inconsistency of heterogeneous data sources, and improves the accuracy and efficiency of data integration.
[0010] Optionally, S2 specifically includes: S21, performing a field parsing operation on each piece of archive data in the unified representation data set to extract field name, field value and field path identification in the original data structure, and constituting a field record sequence; S22, identifying the inclusion relationship between fields according to the hierarchical number in the field path identification, establishing a hierarchical connection relationship between fields in a depth-first order, and generating a field hierarchical structure table; S23, constructing a field node set according to the field hierarchical structure table, each field being a node in the graph, and setting node attributes including field name, type mark and hierarchical number; S24, extracting two types of connection edges for the field node set, respectively: Hierarchical edge: if field A is the direct superior of field B, a directed edge from field A to field B is established; Sequential edge: if field E and field F appear adjacent to each other in the same hierarchical structure, a sequential edge from field E to field F is established; S25, constructing a structure graph corresponding to the archive data according to the field node set and the above connection edges, defined as graph , wherein, ; represents the field node set, is a hierarchical edge set, is a sequential edge set, is the total number of fields; S26, all structure graphs constructed by the archive data are summarized to form a structure graph set.
[0011] By detailed field parsing on the archive data in the unified representation data set, the field name, field value and field path identification of each data can be accurately extracted and constructed into a field record sequence. On this basis, according to the hierarchical number of the field path, the containing relationship between fields is automatically identified, and a field hierarchical structure table is established in a depth-first order to form a clear field hierarchy. Each field is regarded as a node in the graph, and the node attributes include the field name, type mark and hierarchical number. A field node set is further generated. Then, the hierarchical edges and sequential edges are extracted to construct the relationship graph between data fields, so as to accurately represent the structural characteristics of each archive data. Finally, by aggregating all the structure graphs constructed by the archive data, a structure graph set is formed, which provides high-quality structured data for subsequent similarity calculation and archiving label prediction.
[0012] Optionally, the S3 specifically comprises: S31, performing node path distance calculation on each structure graph in the structure graph set to obtain the shortest path length between any two field nodes, and constructing an internal distance matrix of the structure graph; for the structure graph , a distance matrix is generated, wherein represents the path length from node to node ; similarly, for the structure graph , a distance matrix is constructed. S32, mapping the structure graph and the structure graph to a unified structure comparison space to construct a structure difference tensor , wherein each tensor element represents the path difference degree between the cross-graph node pair: , wherein , . S33, selecting a direction set on the unit sphere, and projecting the tensor along each direction to compress the high-dimensional structure difference into a one-dimensional slice distance sample sequence; each group of projections corresponds to generate two one-dimensional sample sets and , which respectively represent the structure difference expression of the structure graph and in the direction . S34, constructing a cumulative distribution function for each pair of slice samples , which is defined as:
[0013] ; where, is an indicator function, is a distance variable, S35, in each direction , the Wasserstein-1 distance between the two slice cumulative distribution functions is calculated, defined as: S36, the structural difference distance in all directions is averaged to obtain the structure graph and the slice Gromov-Wasserstein distance of the structure graph ; S37, for any two structure graphs in the structure graph set, repeat the node path distance calculation, structure difference tensor construction, direction slice projection, slice distribution construction and Wasserstein distance calculation operations to generate the slice Gromov-Wasserstein distance between all structure graph pairs, and fill the results into the structure similarity matrix where, .
[0014] By performing node path distance calculation on each structure graph in the structure graph set, and constructing the distance matrix between nodes, the present application realizes accurate measurement of the internal distance of the structure graph. On this basis, based on the slice Gromov-Wasserstein distance calculation method, the structure graph is mapped to a unified structure comparison space, a structure difference tensor is constructed, and the path difference between the node pairs across the graph is quantified. Further, a set of directions is selected on the unit sphere, a projection operation of high-dimensional structure difference is performed, which is compressed into a one-dimensional slice sample sequence, and a cumulative distribution function is constructed for each pair of slice samples, and the Wasserstein-1 distance is calculated. Through multi-directional structure difference calculation and averaging, the slice Gromov-Wasserstein distance between the structure graphs is finally obtained. This method effectively solves the structure similarity calculation problem of multi-source heterogeneous data, realizes the structure alignment and similarity measurement between different data sources, and provides high-precision basic data support for the archiving label prediction and path writing of the archive data.
[0015] Optionally, the S4 specifically comprises: S41, the slice Gromov-Wasserstein distance calculation results between all structure graphs constructed by the archive data in the structure graph set are taken as input to construct a structure similarity matrix where, represents the first structure graph and the first similarity between two structure graphs; S42, apply PHATE algorithm to structure similarity matrix for latent heat diffusion modeling, construct Markov transition matrix , each element is defined as follows: ; wherein, is a direction guide factor, when the structure graph is a semantic subgraph of the structure graph , set , otherwise , used to guide the diffusion to enhance in the semantic master-slave direction; S43, for the structure graph , calculate the average path length of its field node set , set the number of structure adaptive diffusion time steps , the calculation method is as follows: ; wherein, is a diffusion control factor, usually taking an empirical value range [1, 10], used to adjust the response sensitivity of diffusion time to structure depth; S44, for each structure graph , execute times Markov diffusion, that is, calculate the power of the transition matrix, obtain the propagation probability distribution of the th structure graph in the semantic diffusion field ; S45, convert the diffusion probability distribution of all structure graphs into logarithmic heat diffusion potential, calculate the latent heat diffusion distance matrix between structure graph pairs , each element is defined as: ; wherein, represents the stable transition probability from the structure graph to the structure graph , represents the diffusion potential mapping, used to construct the global structure semantic relationship; S46, use the metric-preserving multidimensional scaling embedding method to perform low-dimensional mapping with the heat diffusion distance matrix as input, construct the latent embedding matrix , wherein the th row represents the latent semantic embedding representation of the structure graph ; S47, construct structure perturbation preserving regularization term , which is used to keep the consistency of the heat diffusion distance between embedding vector pairs in the embedding process, is defined as follows: ; wherein, is the Euclidean distance between the structure graph and the structure graph , which is used to measure the degree of keeping between the semantic expression and the heat diffusion relationship after mapping; S48, by jointly optimizing the metric keeping objective function and the structure disturbance keeping regular term, obtaining the embedding matrix , taking each vector in the matrix as the semantic embedding representation of the profile data represented by the structure graph in the latent space.
[0016] The application combines slice Gromov-Wasserstein distance and PHATE algorithm to propose a new method to optimize automatic archiving of multi-source heterogeneous data. In specific implementation, first, the slice Gromov-Wasserstein distance calculation results between all structure graphs are taken as input to construct a structure similarity matrix. The PHATE algorithm is applied to latent heat diffusion modeling of the structure similarity matrix, and a direction guiding factor is combined to enhance the diffusion process in the semantic master-slave relationship. The direction guiding factor is set according to the superior-inferior relationship of the structure graph, so that the diffusion effect of the subgraph is amplified, thereby effectively transmitting hierarchical information. For the adaptive diffusion time of the structure graph, the system calculates the diffusion time step according to the average path length of each graph to ensure that the response sensitivity of the diffusion process to the structure depth is adjustable, thereby avoiding the problems of over-diffusion or under-diffusion. To keep the consistency of the structure in the embedding space, the system constructs a regular term, which is used to keep the consistency of the heat diffusion distance between embedding vectors in the low-dimensional mapping process. The regular term is combined with the objective function to ensure that the embedding vector can accurately reflect the structure and semantics of the original data. Through these technologies, the application greatly improves the automation processing precision and efficiency of the archived data, solves the problems of inaccurate structure alignment and incomplete data fusion in the traditional method, and improves the robustness and accuracy of the profile data processing.
[0017] Optionally, the S5 specifically comprises: S51, for each semantic embedding vector corresponding to the profile data, based on the field node pair set in the original structure graph, the average diffusion distance of the node pair in the heat diffusion space is calculated, and the Euclidean distance between the corresponding embedding vector pairs in the semantic embedding space is extracted synchronously; S52, for each field node pair, calculate the absolute difference between the diffusion distance between the field node pair in the heat diffusion space and the Euclidean distance between the corresponding embedding vectors in the semantic embedding space, and average the differences of all field node pairs to generate a structural deviation value of the semantic embedding vector relative to the original structure graph; S53, compare the structural deviation value with the structural disturbance metric threshold, the structural disturbance metric threshold is a preset fixed real value, representing the maximum allowed structural consistency error range;When the structural deviation value is greater than the structural disturbance metric threshold, it is determined that the current embedding vector does not maintain the original structural consistency; S54, for each semantic embedding vector with a structural deviation value greater than the disturbance metric threshold, perform an embedding adjustment operation: construct a structural deviation correction vector in the direction of the distance gradient between the current embedding vector and its corresponding heat diffusion distribution of the original structure graph, and the length of the correction vector is proportional to the structural deviation value; S55, add the structural deviation correction vector to the original semantic embedding vector to generate a corrected embedding vector;Repeat the vector correction operation until the structural deviation value is not greater than the structural disturbance metric threshold; S56, combine all the corrected semantic embedding vectors and the original embedding vectors that meet the structural consistency requirement into an embedding vector set as input data for subsequent archival label prediction operations, and each vector in the embedding vector set maintains the structural consistency characteristics of its corresponding structure graph.
[0018] The application introduces a structural deviation metric and a correction mechanism, and proposes a semantic embedding vector consistency maintenance method for multi-source heterogeneous archival data. In the implementation process, first, calculate the difference between the diffusion distance between the semantic embedding vector of each archival data and the field node pair in the original structure graph and the Euclidean distance, generate a structural deviation value by comparing these differences. Then, compare the structural deviation value with the preset structural disturbance metric threshold, if the structural deviation value exceeds the threshold, it means that the current embedding vector does not maintain the original structural consistency, at this time, according to the distance gradient between the heat diffusion distribution and the current embedding vector, a structural deviation correction vector is constructed, the length of the correction vector is proportional to the deviation value, through multiple iterations of correction, until the deviation value is lower than the threshold, ensuring the structural consistency of each embedding vector. In the final correction process, all embedding vectors that meet the consistency requirement are combined into a vector set as input data for subsequent archival label prediction.
[0019] Optionally, the S6 specifically comprises: S61, perform a normalization operation on all embedding vectors in the embedding vector set, scale the numerical value of each embedding vector in each dimension to a fixed interval, and construct a uniform scale feature space; S62. The distribution structure in the set of embedded vectors is modeled using the density estimation method. The kernel density function is used to fit the local sample density around each embedded vector in the embedding space to form a continuous embedding density distribution mapping. S63. Perform cluster-level structure recognition operation on the embedded density distribution map, and use graph-based partitioning or spectral clustering methods to divide the density concentration region and generate clusters. Each cluster represents a potential archive category candidate region. S64. Calculate the mean vector of the embedded vectors in each candidate region of the archive category as the category center representation, and construct a category label reference table. Each entry in the reference table consists of the category center vector and a predefined archive label. S65. For each embedded vector corresponding to the archive data to be archived, calculate the distance relationship between the embedded vector and all category center vectors, and select the archive label corresponding to the category center vector with the smallest distance as the predicted label of the embedded vector. S66. Backfill all predicted archive tags according to the archive data index to generate an archive tag sequence.
[0020] This invention achieves efficient archival label prediction for multi-source heterogeneous data by constructing a unified-scale feature space and density estimation method. In specific implementation, a normalization operation is performed on all embedding vectors in the embedding vector set, scaling the value of each embedding vector to a fixed interval to construct a unified feature space, ensuring data consistency. Next, a kernel density function is used to model the embedding space, fitting the local sample density around each embedding vector to form a continuous embedding density distribution mapping. Density concentration regions in the embedding space are identified based on graph partitioning or spectral clustering methods, generating multiple clusters, each corresponding to a potential archival category candidate region. By calculating the mean vector of the embedding vectors in each category candidate region, category centers are determined, and a category label reference table is constructed, containing category center vectors and predefined archival labels. For archival data to be archived, the distance between its embedding vector and all category center vectors is calculated, and the label corresponding to the category center vector with the smallest distance is selected as the prediction result. All predicted archival labels are backfilled according to the archival data index, generating a complete archival label sequence. This method improves the accuracy and automation level of data archiving, avoiding the subjective intervention and errors in label classification in traditional manual archiving.
[0021] Optionally, S7 specifically includes: S71. Receive the archive tag sequence output by the archive tag prediction module, where each tag in the archive tag sequence corresponds to a unique archive data index; S72. Parse the mapping table between archive tags and predefined archive category paths in the system to obtain the archive category path corresponding to the archive tag; S73. For each piece of archival data, retrieve the archival tag in the archival tag sequence based on the unique index of the archival data, and determine the target archival category path of the archival data according to the mapping table; S74. Write the archive data into the data storage directory or database table structure under the target archive category path. When writing, keep the data format, field naming and hierarchical relationship consistent with the original data to complete the archive storage of the archive data. S75. Perform archiving tag parsing, archiving path matching and data writing operations on all archival data, generate archiving completion identifier, and establish an archiving index table for the archived data. The archiving index table records the correspondence between the archival data index, archiving tag and archiving category path. S76. Perform integrity verification on the archived data to confirm that all data under the archived category path is stored accurately. If an exception occurs during the archiving process, record the exception log. Data that has not been archived is rewritten to the specified category path through the supplementary recording mechanism until all archived data is automatically archived.
[0022] This invention significantly improves data processing efficiency and accuracy by establishing an automated archiving process. In its implementation, the system first receives a tag sequence output from the archiving tag prediction module and maps each tag to an archival data index. The system then parses the tags according to a pre-defined archiving category path mapping table, determines the target archiving path for each piece of archival data, and writes the data to the corresponding directory, ensuring consistency in format, field naming, and hierarchical relationships with the original data. Subsequently, the system executes the complete archiving process, including archiving tag parsing, path matching, data writing, and generating an archiving completion identifier, establishing an archiving index table to record the index and path relationships of the data. Finally, the system performs integrity verification on the archived data; if any anomalies are found, it automatically adds incomplete archiving data. This method achieves fully automated archiving, reduces the need for manual intervention, ensures the accuracy and consistency of archived data, and improves the system's automation level and reliability.
[0023] The beneficial effects of this invention are: This invention transforms structured, semi-structured, and unstructured archival data into a unified structural diagram representation through unified format conversion and structural modeling. It also combines the Gromov-Wasserstein distance of slices to achieve accurate similarity measurement between structural diagrams, effectively improving the fusion and alignment capabilities of multi-source heterogeneous archives at the structural level. This solves the problems of single archiving basis and insufficient structural matching accuracy in existing technologies.
[0024] This invention employs an improved PHATE latent thermal diffusion embedding mechanism, which combines semantic direction guidance, structural adaptive diffusion time adjustment, and thermal diffusion regularization constraints. This mechanism can map the similarity relationships between structural graphs to a latent space that maintains structural consistency, thereby enhancing the distinguishability of semantic relationships between archives and improving the stability and interpretability of the embedded expression.
[0025] This invention introduces a structural perturbation measurement mechanism, which corrects the structural deviation of the potential embedding vector in one direction and optimizes the deviation value iteratively by combining the perturbation control function. This ensures that the embedding representation maintains consistency with the original structure graph while expressing semantics, and overcomes the problem of traditional embedding methods easily losing structural details. Attached Figure Description
[0026] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0027] Fig. 1 This is a flowchart of an automatic archiving method based on multi-source heterogeneous data fusion proposed in this invention; Fig. 2 This is a flowchart of the structural graph similarity calculation based on slice Gromov-Wasserstein distance proposed in this invention; Fig. 3 This is a flowchart of the semantic space construction process based on the PHATE latent thermal diffusion embedding algorithm proposed in this invention. Detailed Implementation
[0028] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0029] refer to Figs. 1-3 An automatic archiving method based on multi-source heterogeneous data fusion includes the following steps: S1. Obtain structured archive data, semi-structured archive data and unstructured archive data from at least two preset data sources, convert the format of each type of archive data, and generate a data set with a unified representation; S2. Perform structural modeling operations on each file in the unified data set, extract field hierarchy and topological features, construct the corresponding structural graph, and form a set of structural graphs. S3. Based on the slice Gromov-Wasserstein distance, calculate the structural similarity between any two structural graphs in the set of structural graphs and generate a structural similarity matrix. S4. Using the structural similarity matrix as input, apply the PHATE algorithm to simulate the thermal diffusion process and extract the semantic embedding representation of the corresponding file in the latent space. S5. Perform perturbation measurement on the semantic embedding representation and calculate the structural deviation of each embedding vector relative to the original structure graph. When the deviation exceeds the preset perturbation measurement threshold, adjust the embedding value based on the structural deviation control function to generate a set of embedding vectors that meet the structural consistency. S6. Based on the distribution pattern in the set of embedded vectors, predict the archival tag for each piece of archival data; S7. Write the archive data corresponding to the archive tag to the specified archive category path to complete automatic archiving.
[0030] In this embodiment, S1 specifically includes: S11. Extract archival data from two or more preset data sources respectively. The archival data includes structured archival data, semi-structured archival data and unstructured archival data. Each type of archival data is collected in its original format. S12. For structured archive data, based on the field mapping template, perform field renaming, type unification and encoding standardization processing to generate a set of data items with unified field definitions; S13. For semi-structured archive data, parse its data node structure, extract key-value pairs with distinctive characteristics, and convert them into two-dimensional field representations, aligning them with structured data formats through a mapping relationship table; S14. For unstructured archive data, use preset text parsing rules and entity recognition models to extract key entities and hierarchical semantic structures in the text, and construct a structured fragment representation based on entity relationships. S15. Map the three types of processed archive data to the field set defined in the unified data format template, fill missing fields with default placeholders, discard redundant fields, and construct a unified data set.
[0031] In this embodiment, S2 specifically includes: S21. Perform field parsing operations on each file data in the uniformly represented data set to extract the field name, field value, and path identifier of the field in the original data structure to form a field record sequence. S22. Based on the hierarchical number in the field path identifier, identify the inclusion relationship between fields, establish the hierarchical connection relationship between fields in depth-first order, and generate a field hierarchy structure table. S23. Construct a set of field nodes based on the field hierarchy structure table. Each field is a node in the graph. Set the node attributes, including field name, type label and hierarchy number. S24. Extract two types of connection edges from the field node set: Hierarchical edge: If field A is the direct parent of field B, then a directed edge is created from field A to field B; Sequential edge: If field E and field F appear adjacent to each other in the same hierarchical structure, then a sequential edge is created from field E to field F; S25. Based on the set of field nodes and the aforementioned connecting edges, construct a structural graph corresponding to the archive data, defined as a graph. ,in, ; Represents a collection of field nodes. For a set of hierarchical edges, For a set of ordered edges, Total number of fields; S26. Summarize the structure diagrams constructed from all the archive data to form a set of structure diagrams.
[0032] In this embodiment, S3 specifically includes: S31. Calculate the node path distance for each structural graph in the structural graph set, obtain the shortest path length between any two field nodes, and construct the internal distance matrix of the structural graph; for structural graphs... Generate distance matrix ,in Represents a node To the node The path length; similarly, for the structure graph Construct the distance matrix ; S32, Structure diagram With structural diagram Mapping to a unified structural comparison space, constructing a structural difference tensor Each tensor element represents the path dissimilarity between cross-graph node pairs: ,in, , ; S33. Select a set of directions on a unit sphere. For tensors Along each direction A projection operation is performed to compress the high-dimensional structural differences into a one-dimensional slice distance sample sequence; each projection generates two one-dimensional sample sets. and , respectively representing the structure diagram and In direction Structural differences are expressed below; S34. For each pair of slice samples Construct the cumulative distribution function, defined as:
[0033] ; in, For indicator functions, For distance variables, , ; S35, in each direction Above, calculate the Wasserstein-1 distance between the cumulative distribution functions of two slices, defined as: ; S36, Distance of structural differences in all directions Calculate the average value to obtain the structure diagram. With structural diagram The Gromov-Wasserstein distance of the slice; S37. For any two structural graphs in the set of structural graphs, repeatedly perform the operations of node path distance calculation, structural difference tensor construction, orientation slice projection, slice distribution construction, and Wasserstein distance calculation to generate the slice Gromov-Wasserstein distance between all structural graph pairs, and fill the results into the structural similarity matrix. ,in .
[0034] In this embodiment, S4 specifically includes: S41. Using the Gromov-Wasserstein distance calculation results between the slices of the structural graphs constructed from all the archive data in the structural graph set as input, construct a structural similarity matrix. ,in Indicates the first The first structural diagram and the first Similarity between structural diagrams; S42. Apply the PHATE algorithm to model the potential thermal diffusion of the structural similarity matrix and construct the Markov transition matrix. Each element is defined as follows: ; in, Indicates the directional guiding factor, when the structure diagram It is a structural diagram When setting the semantic lower-level graph, ,otherwise This is used to guide diffusion in a semantically master-slave direction; S43, Regarding the structural diagram Calculate the average path length of its field node set. Set the adaptive diffusion time step of the structure The calculation method is as follows: ; in, As a diffusion control factor, it is usually taken as an empirical value range [1, 10], which is used to adjust the sensitivity of the diffusion time to the structural depth; S44. For each structural diagram ,implement Sub-Markov diffusion, i.e., calculating the power of the transition matrix. , obtained the Probability distribution of propagation of structural graphs in semantic diffusion field ; S45. Convert the diffusion probability distribution of all structure diagrams into logarithmic thermal diffusion potential, and calculate the potential thermal diffusion distance matrix between structure diagram pairs. Each element is defined as: ; in, Representing the structure diagram Diffusion to structural diagram The stable transition probability, Represents a diffusion potential mapping, used to construct global structural semantic relations; S46, using the thermal diffusion distance matrix Using the metric-preserving multidimensional scaling embedding method as input, a low-dimensional mapping is performed to construct the latent embedding matrix. , of which OK Representation of structure diagram Latent semantic embedding representation; S47. Constructing a structural perturbation-preserving regularity term This is used to maintain the consistency of thermal diffusion distance between embedded vector pairs during the embedding process, and is defined as follows: ; in, Structure diagram embedded in space With structural diagram The Euclidean distance is used to measure the degree of preservation between the semantic representation and the thermal diffusion relationship after mapping; S48. Obtain the embedding matrix by jointly optimizing the objective function and the structural perturbation-preserving regularization term. , each vector in the matrix As a structural diagram The semantic embedding representation of the archive data in the latent space.
[0035] In this embodiment, S5 specifically includes: S51. For each file data corresponding to the semantic embedding vector, based on the set of field node pairs in the original structure graph, calculate the average diffusion distance of the node pairs in the thermal diffusion space, and simultaneously extract the Euclidean distance between the corresponding embedding vector pairs in the semantic embedding space. S52. For each field node pair, calculate the absolute difference between the diffusion distance between the field node pairs in the thermal diffusion space and the Euclidean distance between the corresponding embedding vectors in the semantic embedding space, and average the differences of all field node pairs to generate the structural deviation value of the semantic embedding vector relative to the original structure graph. S53. Compare the structural deviation value with the structural disturbance measurement threshold. The structural disturbance measurement threshold is a preset fixed real value, representing the maximum allowable range of structural consistency error. When the structural deviation value is greater than the structural disturbance measurement threshold, it is determined that the current embedded vector does not maintain the original structural consistency. S54. For each semantic embedding vector whose structural deviation value is greater than the perturbation metric threshold, perform an embedding adjustment operation: construct a structural deviation correction vector with the distance gradient between the current embedding vector and the thermal diffusion distribution of its corresponding original structural diagram as the direction. The magnitude of the correction vector is proportional to the structural deviation value. S55. Add the structural deviation correction vector to the original semantic embedding vector to generate the corrected embedding vector; repeat the vector correction operation until the structural deviation value is not greater than the structural disturbance measurement threshold. S56. Merge all the corrected semantic embedding vectors with the original embedding vectors that meet the structural consistency requirements into an embedding vector set, which serves as the input data for subsequent archive label prediction operations. Each vector in the embedding vector set maintains the structural consistency features of its corresponding structure graph.
[0036] In this embodiment, S6 specifically includes: S61. Perform normalization on all embedded vectors in the set of embedded vectors, scaling the value of each embedded vector in each dimension to a fixed range, and constructing a feature space of uniform scale. S62. The distribution structure in the set of embedded vectors is modeled using the density estimation method. The kernel density function is used to fit the local sample density around each embedded vector in the embedding space to form a continuous embedding density distribution mapping. S63. Perform cluster-level structure recognition operation on the embedded density distribution map, and use graph-based partitioning or spectral clustering methods to divide the density concentration region and generate clusters. Each cluster represents a potential archive category candidate region. S64. Calculate the mean vector of the embedded vectors in each candidate region of the archive category as the category center representation, and construct a category label reference table. Each entry in the reference table consists of the category center vector and a predefined archive label. S65. For each embedded vector corresponding to the archive data to be archived, calculate the distance relationship between the embedded vector and all category center vectors, and select the archive label corresponding to the category center vector with the smallest distance as the predicted label of the embedded vector. S66. Backfill all predicted archive tags according to the archive data index to generate an archive tag sequence.
[0037] In this embodiment, S7 specifically includes: S71. Receive the archive tag sequence output by the archive tag prediction module, where each tag in the archive tag sequence corresponds to a unique archive data index; S72. Parse the mapping table between archive tags and predefined archive category paths in the system to obtain the archive category path corresponding to the archive tag; S73. For each piece of archival data, retrieve the archival tag in the archival tag sequence based on the unique index of the archival data, and determine the target archival category path of the archival data according to the mapping table; S74. Write the archive data into the data storage directory or database table structure under the target archive category path. When writing, keep the data format, field naming and hierarchical relationship consistent with the original data to complete the archive storage of the archive data. S75. Perform archiving tag parsing, archiving path matching and data writing operations on all archival data, generate archiving completion identifier, and establish an archiving index table for the archived data. The archiving index table records the correspondence between the archival data index, archiving tag and archiving category path. S76. Perform integrity verification on the archived data to confirm that all data under the archived category path is stored accurately. If an exception occurs during the archiving process, record the exception log. Data that has not been archived is rewritten to the specified category path through the supplementary recording mechanism until all archived data is automatically archived.
[0038] Example 1: To verify the technical advantages of the automatic archiving method based on multi-source heterogeneous data fusion, this invention was applied to the data resource management center of a large state-owned enterprise. This center is responsible for the archiving management of data generated by various business systems within the group, including the ERP production system, financial management platform, email system, and OA collaborative office system. The business data comes from diverse sources, with formats including structured database tables, semi-structured XML and JSON messages, unstructured scanned PDFs, and images. Approximately 40,000 new documents are added each quarter. The archiving process involves significant manual intervention and inconsistent structures, leading to difficulties in data retrieval, inaccurate labeling, and poor historical archiving integrity.
[0039] From March to June 2025, the center used a specific quarter as a sample to embed the method of this invention into the archiving module of an existing archiving system. First, raw archival data from the four major business platforms were aggregated to the data fusion middleware layer via interface scheduling. The first step of the method involved renaming structured fields, standardizing types, and encoding standards for the Excel data exported from the ERP system database tables and the financial system. Email body texts and XML / JSON format files from the email system underwent semi-structured parsing, extracting key nodes such as sender, recipient, subject, and body fragments, and aligning them to a standard structure using a mapping table. Contract and meeting minutes images scanned from the OA system were used to extract entities using OCR text parsing rules and merged into structured fragments according to document type. After multiple steps of cleaning and missing field completion, the field completion rate of the unified data template reached 98.5%.
[0040] The technical features are reflected in the fact that for each data entry, a structure diagram is built based on field paths and hierarchical relationships, and node attributes include field names, data types, and hierarchical numbers. For all structure diagrams, Gromov-Wasserstein distance is used to calculate structural similarity, automatically generating a similarity matrix. Based on the archival class target signature system of real business, it accurately reflects the semantic structural similarity of document types such as contracts, invoices, expense reports, and emails in various business scenarios.
[0041] Based on the business archiving process, the PHATE latent thermal diffusion embedding algorithm is used to model the diffusion relationships between structure graphs. The diffusion step size is automatically set to simulate the hierarchical association between the main contract, supplementary agreements, payment vouchers, and approval forms in real business processes. The system automatically calculates the semantic distribution in the embedding space and introduces a perturbation consistency metric for certain special documents (such as abnormal process emails and scanned image archives) to correct structural deviations in real time, ensuring that all embedding vectors maintain more than 95% consistency with the original structure.
[0042] During the archive label prediction phase, the system normalizes the embedding vectors and automatically identifies the archive category using spectral clustering. The Euclidean distance between the embedding vector and the category center is used to determine the label assignment. The classification results are automatically mapped to the archive path rule table, and the system automatically writes the archive file paths and updates the data index. After archiving, the system automatically performs integrity checks and mis-archive reviews, and all results are automatically written to the system log.
[0043] During the experiment, the project team set up a control group. In the third quarter of 2024, the traditional archiving method of manual rule-based and keyword matching was used throughout, while in the third quarter of 2025, the archiving scheme of this invention was fully applied. The experimental data for both groups are as follows.
[0044] Table 1 Comparison of the effects of traditional manual archiving and the automatic archiving method of this invention.
[0045] As shown in Table 1, the average archiving time decreased from 7.8 minutes to 1.1 minutes, improving archiving efficiency by more than 7 times. The proportion of manual intervention decreased from 51.3% to 5.7%, with most steps becoming unattended. Archiving label accuracy increased from 83.1% to 97.5%, archiving completeness (field rate) increased from 89.2% to 99.1%, and the error rate decreased to less than 1%. Archiving retrieval time was significantly shortened, and the number of business complaints after archiving was greatly reduced, leading to increased satisfaction among business departments. Label consistency checks revealed that the automated archiving results of this invention maintained a label consistency rate of over 98% in multiple batches of audits, while manual archiving often resulted in label confusion due to subjective judgment and unfamiliarity with the process.
[0046] During implementation, the system was also specifically tested in the quarterly invoice archiving process of the ERP finance module. In large-scale invoice archiving tasks (1500 invoices per batch), the automatic archiving process can automatically identify key fields such as "invoice number," "invoice amount," and "invoice date" based on structural modeling. It also automatically archives invoices to the correct contract item based on the structural similarity between the contract number and project node, without manual intervention. Verification data shows that the accuracy rate of automatically archived invoice labels reached 99.4%, while the accuracy rate of manually archived batches was only 87.7%.
[0047] To further analyze the universality and practical promotion value of this invention, in June 2025, the project team conducted a retrospective analysis of historical archive records, and the average performance of each archiving task from the fourth quarter of 2023 to the second quarter of 2025 was statistically analyzed as follows: Table 2 Summary of Multi-Quarter Archiving Performance
[0048] Statistics show that after adopting this invention, the average archiving time has remained at 1.0 to 1.2 minutes. The accuracy and completeness are higher than the traditional process. The logical consistency and structural matching of the archived data have been significantly improved, and the proportion of duplicate and incorrect archives has been greatly reduced.
[0049] In terms of practical application, this invention not only greatly improves the automation and efficiency of archiving, but also ensures scientific tag allocation, consistent structural expression, and high retrieval efficiency in the face of multi-source, heterogeneous, and complex business data scenarios, significantly reducing manual intervention and error risks, and realizing intelligent archiving and standardized management of the entire archival process.
[0050] In actual operation feedback, staff at the archives center generally reported that after the system went live, the traditional tedious processes such as archive distribution, label review, classification confirmation, and correction of erroneous archives were largely automated. Manual checks on individual data flagged by the system were now only required. The automated archiving system's support for multi-source structural modeling and semantic embedding significantly alleviated the manual burden of large-scale data archiving tasks. Especially during peak periods of quarterly closing, auditing, and annual report archiving, the system's response speed and accuracy to new archiving tasks were highly praised by business departments and management.
[0051] Therefore, this invention can provide a smart archiving solution that is structurally unified, semantically consistent, highly efficient, and requires minimal human intervention, addressing the key challenges faced by enterprises and institutions in multi-source archiving. It has strong promotional and application value.
[0052] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. An automatic archiving method based on multi-source heterogeneous data fusion, characterized in that, The method comprises the following steps: S1, obtaining structured archive data, semi-structured archive data and unstructured archive data from at least two preset data sources, performing format conversion on each type of archive data, and generating a uniformly represented data set; S2, performing structure modeling operations on each piece of archive data in the uniformly represented data set, extracting field hierarchy and topological features, constructing a corresponding structure graph, and forming a structure graph set; S3, calculating the structural similarity between any two structure graphs in the structure graph set based on the slice Gromov-Wasserstein distance, and generating a structure similarity matrix; S4, inputting the structure similarity matrix, applying the PHATE algorithm to simulate the heat diffusion process, and extracting the semantic embedding representation of the corresponding archive in the latent space; S5, performing perturbation measurement on the semantic embedding representation, calculating the structural deviation of each embedding vector relative to the original structure graph, and when the deviation exceeds a preset perturbation measurement threshold, adjusting the embedding value based on the structural deviation control function to generate an embedding vector set that satisfies the structural consistency; S6, predicting the archiving label of each piece of archive data based on the distribution pattern in the embedding vector set; S7, writing the archive data corresponding to the archiving label into a specified archiving category path to complete automatic archiving.
2. The automatic archiving method based on multi-source heterogeneous data fusion according to claim 1, characterized in that, The S1 specifically comprises: S11, extracting archive data from two or more preset data sources respectively, wherein the archive data comprises structured archive data, semi-structured archive data and unstructured archive data, and each type of archive data is collected in the original format; S12, for structured archive data, performing field renaming, type unification and coding specification processing based on a field mapping template to generate a data item set with uniform field definitions; S13, for semi-structured archive data, parsing its data node structure, extracting key-value pair relationships with identifiers, and converting them into two-dimensional field representations, and aligning them with structured data formats through a mapping relationship table; S14, for unstructured archive data, using a preset text parsing rule and an entity recognition model to extract key entities and superordinate and subordinate semantic structures in the text, and constructing a structured fragment representation based on entity relationships; S15, mapping the processed archive data to the field set defined in the unified data format template, filling the missing fields with default placeholders, discarding redundant fields, and constructing a uniformly represented data set.
3. The method of claim 1, wherein, The S2 specifically comprises: S21, performing field parsing operations on each piece of archive data in the uniformly represented data set, extracting field names, field values, and field path identifiers in the original data structure, and constructing a field record sequence; S22, identifying the inclusion relationship between fields according to the hierarchical numbers in the field path identifiers, establishing hierarchical connection relationships between fields in a depth-first order, and generating a field hierarchy table; S23, constructing a field node set according to the field hierarchy table, each field being a node in the graph, and setting node attributes including field name, type mark and hierarchical number; S24, extracting two types of connection edges from the field node set, respectively: Hierarchy edge: if field A is the direct superior of field B, then a directed edge is created from field A to field B; Sequential edge: if field E and field F appear adjacent in the same hierarchy, then a sequential edge is created from field E to field F; S25, constructing a structure graph corresponding to the archive data according to the field node set and the connection edge, defined as graph , wherein ; the field node set is represented by the hierarchical edge set is represented by the sequential edge set is represented by the total number of fields is represented by S26, all the structure diagrams of the archive data are summarized to form a structure diagram set.
4. The automatic archiving method based on multi-source heterogeneous data fusion according to claim 1, characterized in that, The S3 specifically includes: S31, performing node path distance calculation on each structure graph in the structure graph set, obtaining the shortest path length between any two field nodes, and constructing an internal distance matrix of the structure graph; for the structure graph , a distance matrix is generated , wherein denotes the path length from node to node ; similarly, for the structure graph , a distance matrix is constructed S32, construct a structure graph with the structure graph mapping to a unified structure comparison space, construct a structure difference tensor where each tensor element represents a degree of path difference across a pair of graph nodes: where, , ; S33, selecting a direction set on the unit sphere , structural difference tensor , along each direction , projection operation is performed to compress the high-dimensional structural difference into a one-dimensional slice distance sample sequence; each group of projections corresponds to generate two one-dimensional sample sets , and , respectively represent the structural graph , and , the structural difference expression under the direction S34, for each pair of slice samples constructing a cumulative distribution function; S35、in each direction In this case, the Wasserstein-1 distance between the two slice cumulative distribution functions is calculated. S36, structure difference distance in all directions average, get structure map with structure map slice Gromov-Wasserstein distance; S37、for any two structure graphs in the structure graph set, repeat the execution of node path distance calculation, structure difference tensor construction, direction slice projection, slice distribution construction and Wasserstein distance calculation operations, generate the slice Gromov-Wasserstein distance between all structure graph pairs, and fill the results into the structure similarity matrix wherein .
5. The method of claim 1, wherein the method is characterized by, The S4 specifically includes: S41. Using the Gromov-Wasserstein distance calculation results between the slices of the structural graphs constructed from all the archive data in the structural graph set as input, construct a structural similarity matrix. ,in Indicates the first The first structural diagram and the first Similarity between structural diagrams; S42, apply the PHATE algorithm to the structural similarity matrix for latent heat diffusion modeling, construct a Markov transition matrix ; S43, for a structure diagram , calculate the average path length of its field node set , set the structure adaptive diffusion time step number ; S44, for each structure graph , execute times Markov diffusion, i.e. calculate the power of the transition matrix , obtain the propagation probability distribution of the th structure graph in the semantic diffusion field ; S45, the diffusion probability distribution of all structure diagrams is converted into a logarithmic thermal diffusion potential, and a potential thermal diffusion distance matrix between structure diagram pairs is calculated; S46, using the low-dimensional mapping method of metric preserving multidimensional scaling embedding method with the heat diffusion distance matrix as input, constructing a latent embedding matrix wherein the first row represents the latent semantic embedding representation of the structure diagram ; S47, constructing a structure disturbance preserving regularization term In the embedding process, the thermal diffusion distance consistency between the embedding vector pairs is maintained; S48, obtain the embedding matrix by keeping the objective function and the structure perturbation keeping regular term through joint optimization measurement , each vector in the matrix as a structure diagram the semantic embedding representation of the profile data in the latent space.
6. The automatic archiving method based on multi-source heterogeneous data fusion according to claim 1, characterized in that, The S5 specifically includes: S51, for each semantic embedding vector corresponding to the archive data, based on the field node pair set in the original structure diagram, the average diffusion distance of the node pair in the thermal diffusion space is calculated, and the Euclidean distance between the corresponding embedding vector pair in the semantic embedding space is extracted synchronously; S52, for each field node pair, the absolute difference between the diffusion distance between the field node pairs in the thermal diffusion space and the Euclidean distance between the corresponding embedding vectors in the semantic embedding space is calculated, and the average of the difference values of all field node pairs is calculated to generate a structure deviation value of the semantic embedding vector relative to the original structure diagram; S53, compare the structure deviation value with the structure disturbance metric threshold value, the structure disturbance metric threshold value is a preset fixed real value, representing the maximum allowed structure consistency error range; when the structure deviation value is greater than the structure disturbance metric threshold value, it is determined that the current embedding vector does not maintain the original structure consistency; S54, for each semantic embedding vector with a structure deviation value greater than the disturbance metric threshold value, perform embedding adjustment operation: construct a structure deviation correction vector with the distance gradient between the current embedding vector and the thermal diffusion distribution of its corresponding original structure diagram as the direction, and the module length of the correction vector is proportional to the structure deviation value; S55, add the structure deviation correction vector to the original semantic embedding vector to generate the corrected embedding vector; repeat the vector correction operation until the structure deviation value is not greater than the structure disturbance metric threshold value; S56, merge all the corrected semantic embedding vectors and the original embedding vectors that meet the structure consistency requirement into an embedding vector set.
7. The method of claim 1, wherein the method further comprises: The S6 specifically includes: S61, perform normalization operation on all embedding vectors in the embedding vector set, scale the numerical value of each embedding vector in each dimension to a fixed interval, and construct a uniform scale feature space; S62, use density estimation method to model the distribution structure in the embedding vector set, use kernel density function to fit the local sample density around each embedding vector in the embedding space, and form a continuous embedding density distribution mapping; S63, perform cluster level structure recognition operation on the embedding density distribution mapping, divide the density concentrated area using the graph partitioning or spectral clustering method, and generate clustering clusters, each clustering cluster represents a potential archiving category candidate area; S64, calculate the mean vector of the embedding vectors in each archiving category candidate area as the category center representation, and construct a category label reference table, each entry in the reference table consists of a category center vector and a predefined archiving label. S65, for each embedding vector corresponding to the to-be-archived archive data, calculate the distance relationship between the embedding vector and all category center vectors, and select the archived label corresponding to the category center vector with the minimum distance as the predicted label of the embedding vector; S66, backfill all predicted archived labels according to the archive data index to generate an archived label sequence.
8. The automatic archiving method based on multi-source heterogeneous data fusion according to claim 1, characterized in that, The S7 specifically comprises: S71, receiving the archived label sequence output by the archived label prediction module, each label in the archived label sequence corresponding to a unique archive data index; S72, parsing the archived label and the pre-defined archived category path mapping table in the system to obtain the archived category path corresponding to the archived label; S73, for each archive data, retrieving the archived label in the archived label sequence according to the unique index of the archive data, and determining the target archived category path of the archive data according to the mapping table; S74, writing the archive data into the data storage directory or the database table structure under the target archived category path, keeping the data format, field naming and hierarchical relationship consistent with the original data during writing, and completing the archived storage of the archive data; S75, performing the archived label analysis, archived path matching and data writing operations on all archive data to generate an archived completion identifier, establishing an archived index table for the archived data, and recording the corresponding relationship among the archive data index, the archived label and the archived category path in the archived index table; S76, performing integrity check on the archived archive data to confirm that all data under the archived category path are accurately stored, recording abnormal logs when the archived process is abnormal, and re-writing the uncompleted archived data into the specified category path through the supplement recording mechanism until all archive data are automatically archived.