A document management method and system for logistics transportation

By building a multi-dimensional correlation network and PageRank algorithm, the problems of data islands and abnormal documents in traditional document management systems are solved, and efficient, accurate and real-time collaboration of logistics document management is achieved.

CN119990944BActive Publication Date: 2025-07-18SHANGHAI WINLINK NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510457250.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-18
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

Traditional document management systems are difficult to effectively capture the dynamic relationship between multiple types of documents, resulting in frequent data islands and abnormal documents, which cannot support high-time and high-precision logistics coordination needs.

Method used

By building a multi-dimensional correlation network, the correlation formula is used to quantify the correlation and abnormal correlation between documents, and in-depth sorting and incremental reconstruction are combined with the PageRank algorithm to achieve deep semantic fusion across business modules and accurate identification of abnormal documents.

Benefits of technology

It significantly improves the accuracy and coordination efficiency of logistics document management, realizes accurate identification and dynamic optimization of abnormal documents, and ensures the real-time and robustness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
Patent Text Reader

Abstract

The present invention relates to the technical field of intelligent data management, and particularly relates to a document management method and system for logistics transportation. The method includes the following steps: S1: Obtain document data of N different types, extract the association information between the N different types of documents, and based on the association information, obtain the association degree between the N different types of documents; the association information refers to the exact match degree of data fields and the business process coupling relationship between different documents; the association degree is calculated through an association degree formula; S2: Based on the association degree, obtain an abnormal association degree; the abnormal association degree is obtained through topological structure analysis. Through a multi-dimensional association model and dynamic weight adjustment, the present invention accurately identifies abnormal documents and blocks risks, hierarchically sorts core nodes in combination with the PageRank algorithm, and quickly embeds new data using an incremental strategy, realizing cross-system collaborative optimization and improving the real-time performance and robustness of logistics document management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent data management, and particularly to a document management method and system for logistics transportation. Background Art

[0002] As a core link of the modern supply chain, logistics transportation involves multiple business processes such as procurement, warehousing, transportation, and settlement. The types of documents generated during the process are complex and closely related. Traditional document management systems mostly rely on manual entry and static storage, and it is difficult to effectively capture the dynamic association relationships between multiple types of documents. Especially in scenarios of cross-system and multi-node collaboration, document information is scattered in different business modules, forming data islands, resulting in frequent problems such as misalignment of key fields and breakage of process timing dependencies. At the same time, due to imperfect data cleaning mechanisms or abnormal operations, abnormal documents such as "phantom documents" (still being referenced after physical deletion) and "zombie documents" (being misused after logical deletion) often appear, further exacerbating information chaos. Although existing technologies can achieve basic data storage and retrieval, they lack the ability to quantitatively analyze the complex coupling relationships between documents, cannot identify abnormal interactions or dynamically optimize the association network in real time, and are difficult to support the high-efficiency and high-precision logistics collaboration requirements. With the accelerating digital transformation of the logistics industry, there is an urgent need for a management solution that can deeply integrate business processes and data intelligence to systematically solve the problems of association governance and risk prevention and control of heterogeneous documents. Summary of the Invention

[0003] In order to overcome the shortcoming of difficult collaboration of multi-source heterogeneous documents, the present invention provides a document management method and system for logistics transportation.

[0004] The technical solution is as follows: A document management method for logistics transportation, comprising the following steps:

[0005] S1: Obtain N different types of document data, extract the association information between the N different types of documents, and based on the association information, obtain the association degree between the N different types of documents; the association information refers to the exact match degree of data fields and the business process coupling relationship between different documents; the association degree is calculated through an association degree formula;

[0006] S2: Based on the association degree, obtain an abnormal association degree; the abnormal association degree refers to quantifying the interaction density and abnormal pattern characteristics of documents with phantom documents, zombie documents, and fragmented documents through topological structure analysis; the phantom document refers to a document that has been physically deleted but is still referenced by multiple nodes; the zombie document is a document that is still misused after logical deletion; the fragmented document is a document that remains uncleaned correctly after a deletion failure; the abnormal association degree formula is as follows,

[0007]

[0008] Among them, is the abnormal correlation degree, is the abnormal weight type, is the document and the type The number of interactions between abnormal documents, is the total number of interactions of the document ; They are respectively abnormal type ghost documents, zombie documents, and fragmented documents;

[0009] S3: Based on the correlation degree, perform a depth sorting on N different types of documents. The depth sorting refers to sorting in descending order according to the magnitude of the correlation degree;

[0010] S4: Based on the depth sorting result, perform a correlation degree marking on the newly generated documents, trigger an incremental document reconstruction according to the updated network topology, and synchronously update the global sorting.

[0011] Preferably, the obtaining of the data of N different types of documents, extracting the correlation information between N different types of documents, and obtaining the correlation degree between N different types of documents based on the correlation information includes:

[0012] Sort N different types of documents in sequence according to the business process;

[0013] Based on the sorting result, extract the data fields of N different types of documents for matching;

[0014] Based on the sorting result and the matching result, calculate the correlation degree between N different types of documents by using the correlation degree formula.

[0015] Preferably, the calculating of the correlation degree between N different types of documents by using the correlation degree formula based on the sorting result and the matching result includes: The basic correlation degree formula is as follows,

[0016]

[0017] Among them, is the basic correlation degree, , are the weight coefficients, is the field matching degree, is the process coupling degree.

[0018] Preferably, the calculating of the correlation degree between N different types of documents by using the correlation degree formula based on the sorting result and the matching result includes: The field matching degree formula is as follows,

[0019]

[0020] Among them, is the field matching degree, is the weight of the th field, is the th value of the field of the document, is the th value of the field of the document, is the matching function; the process coupling formula is as follows,

[0021]

[0022] wherein, is the process coupling degree, , , are the weight coefficients, is the temporal dependence strength, is the number of state transfers, is the shared resource ratio.

[0023] Preferably, obtaining the abnormal correlation degree based on the said correlation degree includes: the final correlation degree formula is as follows,

[0024]

[0025] wherein, is the final correlation degree, is the basic correlation degree, is the abnormal correlation degree of the document, is the abnormal correlation degree of the document;

[0026] If there is no abnormal interaction between two documents, the final correlation degree is equal to the basic value;

[0027] If there is an abnormal correlation in any document, the correlation degree is reduced according to the interaction density ratio.

[0028] Preferably, performing a depth sorting on N different types of documents based on the said correlation degree, where the depth sorting means sorting in descending order according to the size of the correlation degree, including:

[0029] For each document, accumulate the final correlation degree between the document and the remaining documents to form a global correlation strength index, which is used to quantify the influence of the document in the overall network;

[0030] Model the document and the document association relationship as a weighted graph structure, calculate the importance score of each node based on the PageRank algorithm, and divide the depth levels according to the score, the higher the importance, the higher the level;

[0031] When adding a new document, through local association calculation and incremental update strategy, efficiently adjust the association strength and importance score of affected nodes; when the score fluctuation exceeds the preset threshold, automatically mark it as a deeply changed node and output the updated sorting result.

[0032] Preferably, the modeling of documents and document association relationships as a weighted graph structure, calculating the importance score of each node based on the PageRank algorithm, and dividing the depth levels according to the score, the higher the importance, the higher the level, including:

[0033] The weighted graph takes documents as nodes, and the final association degree between two documents is used as the edge weight to construct a directed network;

[0034] Iteratively calculate the node importance based on the PageRank algorithm: each node distributes the current node score to adjacent nodes according to the proportion of the out-edge weight, and superimposes the damping factor to balance the random jump probability until the score converges;

[0035] Finally, divide the node PageRank values into discrete depth levels according to the quantiles. The higher the value, the deeper the level, reflecting the core degree in the global association network.

[0036] Preferably, when adding a new document, through local association calculation and incremental update strategy, efficiently adjust the association strength and importance score of affected nodes, including:

[0037] When adding a new document, only calculate the association degree between the new document and the existing documents, and lock the set of directly associated nodes;

[0038] Through the incremental update strategy, insert the newly added association edges into the graph structure, only trigger the re-accumulation of the global association strength of relevant nodes, and iteratively update the PageRank score based on the local subgraph;

[0039] Constrain the influence range of random jumps through the damping factor. When the node score changes exceed the threshold, expand and update the neighborhood layer by layer until the change amount of the node score is less than the set convergence threshold and stop the iteration.

[0040] Preferably, based on the depth sorting result, mark the association degree of the newly generated document, trigger incremental document reconstruction according to the updated network topology, and synchronously update the global sorting, including:

[0041] When marking the association degree of the new document based on the depth sorting result, preferentially select core nodes at the same level or higher levels to establish strong associations, and generate association edges through the dynamic weight distribution rule.

[0042] Preferably, a document management system for logistics transportation includes:

[0043] The data association analysis module extracts the matching degree and process coupling relationship between documents based on the process sequence and field matching rules, and dynamically calculates the basic association degree; the coupling degree integrates the time series, status and resource characteristics to form a comprehensive evaluation;

[0044] The anomaly detection module identifies the interaction density of abnormal documents through the topological network, quantifies the anomaly characteristics and corrects the association degree;

[0045] The deep sorting module models the document as a weighted directed graph, calculates the core degree of nodes to divide levels; when a new document is added, the node scores are locally updated to control the influence range;

[0046] The dynamic reconstruction module preferentially binds nodes with high core levels, assigns weights according to the level differences, triggers incremental reconstruction to update the topology, and ensures the real-time performance of sorting and the network stability; realizes anomaly isolation and dynamic balance maintenance.

[0047] Beneficial effects: By constructing a multi-dimensional association network and a dynamic optimization mechanism, the present invention significantly improves the accuracy and collaborative efficiency of logistics document management. First, based on the double quantitative analysis of field matching degree and process coupling degree, a dynamic association model between heterogeneous documents is established, breaking through the dependence of traditional systems on static data islands and realizing deep semantic fusion across business modules. Second, through topological structure analysis and anomaly pattern recognition, risk sources such as ghost documents and zombie documents are accurately located, and the association weights are dynamically adjusted in combination with the interaction density characteristics to effectively block the propagation path of abnormal data. Furthermore, using the global association strength index and weighted graph modeling technology, the core nodes of the document network are hierarchically sorted, and the influence of nodes is quantified through the PageRank algorithm to ensure that key business documents are preferentially reached in process collaboration. On this basis, the incremental update strategy combined with the local subgraph iteration mechanism realizes the rapid embedding of new documents and the adaptive adjustment of the network topology, avoiding the waste of resources in full-scale calculation. Finally, a closed-loop optimization system is formed, and through dynamic weight allocation and incremental reconstruction, the robustness of the association network is continuously improved, enabling the document system to have self-adaptive anomaly immunity and real-time collaborative response capabilities. Brief Description of the Drawings

[0048] Figure 1 It is a flowchart of the document management method for logistics transportation of the present invention;

[0049] Figure 2 It is a schematic structural diagram of the document management system for logistics transportation of the present invention. Detailed Embodiments

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0051] Embodiment 1: A document management method for logistics transportation, as Figure 1 shown, includes the following steps:

[0052] S1: Obtain N different types of document data, extract the association information between the N different types of documents, and based on the association information, obtain the association degree between the N different types of documents; the association information refers to the exact match degree of data fields between different documents and the coupling relationship of business processes; the association degree is calculated through the association degree formula;

[0053] S2: Based on the association degree, obtain the abnormal association degree; the abnormal association degree refers to quantifying the interaction density and abnormal pattern characteristics of documents with ghost documents, zombie documents, and fragmented documents through topological structure analysis; the ghost document refers to a document that has been physically deleted but is still referenced by multiple nodes; the zombie document is a document that is still misused after logical deletion; the fragmented document is a document that remains uncleaned correctly after a failed deletion; the abnormal association degree formula is as follows,

[0054]

[0055] Wherein, is the abnormal association degree, is the abnormal weight type, is the document and type the number of interactions of the abnormal document, is the document the total number of interactions, are the abnormal types of ghost documents, zombie documents, and fragmented documents respectively;

[0056] For further explanation, the abnormal association degree formula is as follows,

[0057]

[0058] Objective: Quantify the interaction risk between the document and abnormal documents (such as ghost documents , zombie documents , fragmented documents ).

[0059] Numerator: Statistic The number of interactions with a certain type of abnormal document , multiply it by its risk weight (e.g., the weight of zombie documents is higher).

[0060] Denominator: The total number of interactions , used for normalization (the proportion of abnormal interactions).

[0061] Essence of the formula: If a document frequently interacts with high-risk abnormal documents (e.g., (C_k (D_i )) / (C_total (D_i) ) is high and ω_G is large), then the abnormal correlation degree A_int increases significantly.

[0062] Example:

[0063] Purchase order The total number of interactions is 20 times, among which:

[0064] Interact with ghost documents ( ), 5 times ( = 0.6),

[0065] Interact with zombie documents ( ), 2 times ( = 0.3),

[0066] Interact with fragmented documents ( ), 1 time ( = 0.1).

[0067] Calculation: = 0.6 * 5 / 20 + 0.3 * 2 / 20 + 0.1 * 1 / 20 = 0.15 + 0.03 + 0.005 = 0.185;

[0068] S3: Based on the correlation degree, perform a deep sorting on N different types of documents. The deep sorting means sorting in descending order according to the size of the correlation degree;

[0069] S4: Based on the deep sorting result, perform a correlation degree marking on the newly generated documents, trigger an incremental document reconstruction according to the updated network topology, and synchronously update the global sorting.

[0070] Obtain the data of N different types of documents, extract the correlation information between N different types of documents, and based on the correlation information, obtain the correlation degree between N different types of documents, including:

[0071] Sort N different types of documents in sequence according to the business process;

[0072] Based on the sorting result, extract the data fields of N different types of documents for matching;

[0073] Based on the sorting result and the matching result, calculate the correlation degree between N different types of documents by using the correlation formula.

[0074] Based on the sorting result and the matching result, calculate the correlation degree between N different types of documents by using the correlation formula, including: the basic correlation formula is as follows,

[0075]

[0076] Among them, is the basic correlation degree, and are the weight coefficients, is the field matching degree, is the process coupling degree.

[0077] For further illustration, assume that in an enterprise's supply chain system, the correlation analysis between a purchase order ( ) and a goods receipt note ( ).

[0078] I. Calculation of the field matching degree ( ).

[0079] Logic: Compare the common fields of the two documents (such as order number, supplier name, material code).

[0080] If the order numbers are exactly the same, get 1 point; for partial matching (such as the last 4 digits are the same), get 0.5 points; other fields are scored according to the similarity ratio (such as a 0.3-point match for the supplier name).

[0081] Assume that the field weights are evenly distributed, and the total matching degree = 0.1 + 0.33 = 0.43.

[0082] II. Calculation of the process coupling degree ( ).

[0083] Logic: Analyze the business process dependency relationship.

[0084] If the goods receipt note must be triggered after the purchase order is generated (temporal dependency), the coupling strength = 0.8;

[0085] If the two documents share the same warehouse resource (resource dependency), the coupling strength = 0.2;

[0086] Assume that the sub-weights = 0.7, = 0.3, then = 0.7×0.8 + 0.3×0.2 = 0.62.

[0087] III. Calculation of the comprehensive correlation degree.

[0088] Parameter: Let = 0.6 (emphasis on field matching), = 0.4 (emphasis on process coupling).

[0089] Result: = 0.6 × 0.43 + 0.4 × 0.62 = 0.506.

[0090] Precise association: Quantify the association strength between two documents (0.506). When it is higher than the threshold (such as 0.5), the system automatically marks it as a "strong association pair";

[0091] Process optimization: Identify the timing dependence (0.8) of the purchase - warehousing process, trigger the automatic verification rule, and reduce manual checking;

[0092] Abnormal warning: If the field matching degree is lower than the threshold (assuming the best threshold for field matching is 0.2), but the process coupling degree is higher than the threshold (assuming the best threshold for process coupling is 0.8), prompt the risk of "document information misalignment".

[0093] Based on the sorting result and the matching result, use the association formula to calculate the association degree between N different types of documents, including: The field matching formula is as follows,

[0094]

[0095] Among them, is the field matching degree, is the weight of the th field, is the value of the th field of document is the value of the th field of document is the th field, is the matching function; The process coupling formula is as follows,

[0096]

[0097] Among them, is the process coupling degree, , , are the weight coefficients, is the timing dependence strength, is the number of state transmissions, is the proportion of shared resources.

[0098] Furthermore, the matching function includes the exact field formula, the numerical field formula, and the text field formula,

[0099] The exact field formula is as follows,

[0100]

[0101] The value is 0 or 1, applicable scenarios: Boolean, enumeration, or unique identifier fields;

[0102] Example: All order status fields are "signed for" → = 1; if one is "in transit" → = 0;

[0103] The formula for numerical fields is as follows,

[0104]

[0105] Applicable scenarios: Numerical fields that need to quantify the degree of difference;

[0106] Characteristic: When completely equal = 1, the greater the difference approaches 0; sensitive to outliers (e.g., 1 vs 100 → = 0.01);

[0107] Example: The weight field values are 80 kg and 100 kg respectively → = 1 - 20 / 100 = 0.8;

[0108] The formula for text fields is as follows,

[0109]

[0110] Applicable scenarios: Semantic similarity matching for text descriptions, addresses, etc.;

[0111] Prerequisite: The text needs to be converted into a numerical vector;

[0112] Characteristic: The value range is [-1, 1], usually taking the absolute value or normalizing to [0, 1]; ignoring the text length and focusing on direction similarity;

[0113] Example: "Zhangjiang Road, Pudong New Area" vs "Shanghai Zhangjiang High-Tech Park" → After vector quantization calculation, we get = 0.72;

[0114] The formula for field matching degree is as follows,

[0115]

[0116] The numerator: Traverse the common fields of the two documents (such as order number, supplier), and for each field, multiply its weight according to the matching degree (calculated by the δ function) and then accumulate.

[0117] Denominator: The sum of all weights, used for normalization to ensure the result is in the range of [0, 1].

[0118] Example:

[0119] The purchase order ( ) and the goods receipt note ( ) have 3 fields in total: order number (weight 0.5), material code (weight 0.3), and supplier (weight 0.2).

[0120] The order number is completely matched ( = 1), the material code is partially matched ( = 0.5), and the supplier is not matched ( = 0).

[0121] Calculation: = 0.5×1 + 0.3×0.5 + 0.2×00.5 + 0.3 + 0.2 = 0.65 / 1 = 0.65

[0122] The formula for process coupling degree is as follows

[0123]

[0124] Time - sequence dependence ( ): If the process has a strict sequence (such as a goods receipt note can only be created after a purchase order is generated), then the strength is high (example: = 0.8).

[0125] State transfer ( ): The number of times of state linkage between documents (such as the purchase order "paid" triggers the goods receipt note "awaiting receipt"). The more times, the higher the coupling degree.

[0126] Shared resources ( ): The proportion of two documents sharing the same resource (such as a warehouse, a capital account) (example: sharing the same warehouse accounts for 50%, = 0.5).

[0127] Weights ( , , ): Reflect the importance of different factors (example: = 0.6 emphasizes time - sequence, = 0.3 emphasizes state, = 0.1 emphasizes resources).

[0128] Example:

[0129] The process coupling between the purchase order and the goods receipt note.

[0130] Time - sequence dependence strength = 0.8, the number of state transfer times =Second (corresponding to =0.3), the proportion of shared warehouses =0.5.

[0131] Calculation: =0.6×0.8 + 0.3×2 + 0.1×0.5 = 0.48 + 0.6 + 0.05 = 1.13. Note: If the sum of weights exceeds 1, normalization is required.

[0132] Total correlation degree of two documents = + , if =0.7, =0.3, then R = 0.7×0.65 + 0.3×1.13 = 0.794. When it is higher than the threshold of 0.7, it is determined as a strong correlation and the automatic review process is triggered.

[0133] Anomaly detection: If the field matching degree is high but the process coupling degree is low, a "data island" risk is prompted; otherwise, a "process breakpoint" is warned.

[0134] Based on the above-mentioned correlation degree, an abnormal correlation degree is obtained, including: The final correlation degree formula is as follows,

[0135]

[0136] Wherein, is the final correlation degree, is the basic correlation degree, is the document 's abnormal correlation degree, is the document 's abnormal correlation degree;

[0137] If there is no abnormal interaction between the two documents, the final correlation degree is equal to the basic value;

[0138] If there is an abnormal correlation in any one of the documents, the correlation degree is reduced according to the interaction density ratio.

[0139] For further explanation, the final correlation degree formula is as follows,

[0140]

[0141] Internal logic:

[0142] Basic correlation degree correction: If there is no abnormal interaction between the two documents ( =0), = ; if there is an anomaly, the correlation degree is reduced according to the ratio of the average anomaly degree of the two documents.

[0143] Risk suppression mechanism: The higher the abnormal interaction density (such as = 0.3), the more obvious the final correlation decay (the multiplier approaches 0.7).

[0144] Example:

[0145] Purchase order and the inbound order The basic correlation = 0.8,

[0146] Abnormal correlation = 0.185,

[0147] Abnormal correlation = 0.05.

[0148] Calculation: = 0.8 * (1 - 0.185 + 0.052) = 0.8 * (1 - 0.1175) = 0.8 * 0.8825 = 0.706,

[0149] Risk dynamic adjustment: The correlation of normal documents remains stable (such as = 0.8);

[0150] If the document frequently correlates with zombie documents ( = 0.3), then = 0.8 * 0.85 = 0.68, and manual review is triggered when it is lower than the threshold of 0.7.

[0151] Abnormal location:

[0152] High The interaction of zombie documents will significantly reduce the correlation, directly exposing the "zombie body" risk nodes in the supply chain.

[0153] Based on the said correlation, perform a deep sorting on N different types of documents. The deep sorting refers to sorting in descending order according to the size of the correlation, including:

[0154] For each document, accumulate the final correlation between the document and the remaining documents to form a global correlation strength index, which is used to quantify the influence of the document in the overall network;

[0155] Model the document and the document correlation relationship as a weighted graph structure, calculate the importance scores of each node based on the PageRank algorithm, and divide the deep levels according to the scores. The higher the importance, the higher the level;

[0156] When adding a new document, the association strength and importance score of affected nodes are efficiently adjusted through local association calculation and incremental update strategy; when the score fluctuation exceeds the preset threshold, it is automatically marked as a deeply changed node, and the updated sorting result is output.

[0157] For further explanation, the influence of documents is quantified through global association strength (accumulative association degree), and after constructing a weighted graph, the importance of nodes is evaluated based on PageRank and hierarchical levels are divided; when adding a new document, only the association network is updated locally, and the score is dynamically adjusted through incremental calculation. When the fluctuation exceeds the threshold, deeply changed nodes are marked to achieve efficient and adaptive maintenance of the network structure. Map the document association to a dynamic graph model, dynamically divide hierarchical levels (such as dividing core levels by quantiles) according to the importance score of nodes (such as PageRank value), and adjust the hierarchical structure in real time through incremental update strategy to balance calculation efficiency and accuracy.

[0158] Model the document and its association relationship as a weighted graph structure, calculate the importance score of each node based on the PageRank algorithm, and divide the depth levels according to the score. The higher the importance, the higher the level, including:

[0159] The weighted graph takes documents as nodes, and the final association degree between two documents

[0160] is used as the edge weight to construct a directed network;

[0161] Iteratively calculate the importance of nodes based on the PageRank algorithm: each node distributes the current score of the node to adjacent nodes according to the proportion of the out-edge weight, and adds a damping factor to balance the random jump probability until the score converges;

[0162] For further explanation, the weighted graph takes documents as nodes, and the final association degree from to Directed edge weights (ignored if the association degree is lower than the threshold) are used to construct a directed network. Based on the PageRank algorithm, the importance of nodes is iteratively calculated: The scores of each node are initialized to 1 / N (N is the total number of nodes). In each round of iteration, the nodes distribute the current score to adjacent nodes according to the proportion of the out-edge weights, and the damping factor (usually taken as 0.85) is superimposed to simulate the random jump probability until the change in node scores reaches the maximum number of iterations (such as 100 times), and convergence is determined. Finally, the PageRank values of the nodes are divided into discrete depth levels (such as 4 levels) by the quartile method. The highest level (such as Top25%) corresponds to the nodes with the largest PageRank values, representing their core hub status in the global network. For example, the key documents that frequently drive multi-level interactions in the supply chain.

[0163] When adding a new document, through local association calculation and incremental update strategy, the association strength and importance score of the affected nodes are efficiently adjusted, including:

[0164] When adding a new document, only calculate the association degree between the new document and the existing documents, and lock the set of directly associated nodes;

[0165] Through the incremental update strategy, insert the newly added associated edges into the graph structure, only trigger the re-accumulation of the global association strength of the relevant nodes, and iteratively update the PageRank scores based on the local subgraph;

[0166] By constraining the influence range of random jumps with the damping factor, when the change in node scores exceeds the threshold, expand and update the neighborhood layer by layer until the change amount of node scores is less than the set convergence threshold and stop the iteration.

[0167] For further explanation, when adding a new document, only calculate its final association degree with the existing documents (such as similarity or causal strength). When the association degree exceeds the preset threshold (assuming the optimal threshold of the association degree is >0.3), it is regarded as a highly associated node; through the incremental update strategy, insert the new document and its associated edges into the original graph, only re-accumulate the global association strength (such as the sum of in-degree weights) of the affected nodes (the newly added nodes and the direct neighbors of the newly added nodes), and iteratively update the PageRank scores within the local subgraph (such as the 2-layer neighborhood): Only adjust the node scores within the subgraph in each round, distribute according to the edge weight ratio and superimpose the damping factor (such as 0.85) to limit the score diffusion range; when the change amount Δ of the node scores exceeds the threshold (assuming the threshold of the change amount of node scores is Δ> )), then expand and update to the outer neighborhood (such as the 3-layer) until the score changes of the entire graph are all lower than the threshold or reach the maximum number of iterations (such as 20 times) to ensure efficient convergence.

[0168] Based on the depth sorting result, perform correlation marking on the newly generated documents, trigger incremental document reconstruction according to the updated network topology, and synchronously update the global sorting, including:

[0169] When performing correlation marking on the new documents based on the depth sorting result, preferentially select the core nodes at the same level or higher levels to establish strong correlations, and generate correlation edges through the dynamic weight distribution rule.

[0170] For further illustration, based on the depth level (divided by the quartiles of the PageRank value), the new documents preferentially establish correlations with the nodes at the same level or higher levels (such as the Top50%), dynamically allocate weights according to the level difference (such as a 30% weight decay for each level difference), and generate strong correlation edges; if the number of newly added edges ≥ 3 or the total weight change > 10%, trigger incremental reconstruction: only recalculate the PageRank values of the affected nodes (including the newly added documents, the directly associated nodes of the newly added documents, and the direct neighbors of these nodes), and update the sorting through local iteration (such as a damping factor of 0.85 and a convergence threshold of 1e-5) to ensure synchronous adjustment of the global level and avoid recalculating the entire graph.

[0171] Embodiment 2: On the basis of Embodiment 1, a document management system for logistics transportation, as Figure 2 shown, includes:

[0172] A data association analysis module, which extracts the matching degree and process coupling relationship between documents based on the process sequence and field matching rules, and dynamically calculates the basic correlation degree; the coupling degree integrates the time sequence, status, and resource characteristics to form a comprehensive evaluation;

[0173] An anomaly detection module, which identifies the interaction density of abnormal documents through the topological network, quantifies the abnormal characteristics, and corrects the correlation degree;

[0174] A depth sorting module, which models the documents as a weighted directed graph, calculates the core degree of the nodes to divide the levels; when adding new documents, locally update the node scores to control the influence range;

[0175] A dynamic reconstruction module, which preferentially binds the nodes with high core levels, allocates weights according to the level difference, triggers incremental reconstruction to update the topology, and ensures the real-time nature of sorting and the stability of the network; realizes anomaly isolation and dynamic balance maintenance.

[0176] The above has introduced the present application in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A document management method for logistics transportation, characterized in that, It includes the following steps: S1: Obtain N different types of document data, extract the association information between the N different types of documents, and based on the association information, obtain the association degree between the N different types of documents; the association information refers to the exact match degree of data fields between different documents and the coupling relationship of business processes; the association degree is calculated through the association degree formula; S2: Based on the association degree, obtain the abnormal association degree; the abnormal association degree refers to quantifying the interaction density and abnormal pattern characteristics between documents and ghost documents, zombie documents, and fragmented documents through topological structure analysis; the ghost document refers to a document that has been physically deleted but is still referenced by multiple nodes; the zombie document is a document that is still misused after logical deletion; the fragmented document is a document that remains uncleared after a failed deletion; the abnormal association degree formula is as follows, Among them, is the anomaly correlation degree, is the anomaly weight type, is the document and the type is the number of interactions between the abnormal document is the document is the total number of interactions, are respectively the abnormal type ghost document, zombie document, fragmented document; S3: Based on the association degree, perform a deep sorting on the N different types of documents, and the deep sorting means sorting from large to small according to the size of the association degree; S4: Based on the deep sorting result, perform an association degree marking on the newly generated document, trigger incremental document reconstruction based on the updated network topology, and synchronously update the global sorting.

2. The document management method for logistics transportation according to claim 1, characterized in that, The obtaining of the N different types of document data, extracting the association information between the N different types of documents, and based on the association information, obtaining the association degree between the N different types of documents includes: Sort the N different types of documents in sequence according to the business process; Based on the sorting result, extract the data fields of the N different types of documents for matching; Based on the sorting result and the matching result, calculate the association degree between the N different types of documents using the association degree formula.

3. A document management method for logistics transportation according to claim 2, characterized in that, The calculating of the association degree between the N different types of documents using the association degree formula based on the sorting result and the matching result includes: The basic association degree formula is as follows, Among them, is the basic correlation degree, , are the weight coefficients, is the field matching degree, is the process coupling degree.

4. A document management method for logistics transportation according to claim 3, characterized in that, The calculating of the association degree between the N different types of documents using the association degree formula based on the sorting result and the matching result includes: The field matching degree formula is as follows, Among them, is the field matching degree, is the weight of the th field, is the value of the th field of the document ; is the value of the th field of the document ; is the matching function; the process coupling formula is as follows, Among them, is the process coupling degree, , , are weight coefficients, is the timing dependence intensity, is the number of state transfers, is the shared resource ratio.

5. A document management method for logistics transportation according to claim 1, characterized in that, The obtaining of the abnormal association degree based on the association degree includes: The final association degree formula is as follows, Among them, is the final correlation degree, is the basic correlation degree, is the abnormal correlation degree of the document, is the abnormal correlation degree of the document; If there is no abnormal interaction between two documents, the final association degree is equal to the base value; If there is an abnormal association in any document, the association degree is reduced according to the interaction density ratio.

6. A document management method for logistics transportation according to claim 1, characterized in that, The performing of a deep sorting on the N different types of documents based on the association degree, and the deep sorting means sorting from large to small according to the size of the association degree includes: For each document, accumulate the final association degree between the document and the remaining documents to form a global association strength index, which is used to quantify the influence of the document in the overall network; Model the document and the document association relationship as a weighted graph structure, calculate the importance scores of each node based on the PageRank algorithm, and divide the depth levels according to the scores, the higher the importance, the higher the level; When adding a new document, through local association calculation and incremental update strategy, efficiently adjust the association strength and importance scores of the affected nodes; when the score fluctuation exceeds the preset threshold, automatically mark it as a deep change node and output the updated sorting result.

7. A document management method for logistics transportation according to claim 6, characterized in that, Modeling the documents and their associated relationships as a weighted graph structure, calculating the importance scores of each node based on the PageRank algorithm, and dividing the depth levels according to the scores from high to low, where the higher the importance, the higher the level, including: The weighted graph uses documents as nodes, and the final correlation degree between two documents is used as the edge weight to construct a directed network; Iteratively calculating the node importance based on the PageRank algorithm: each node distributes the current score of the node to adjacent nodes according to the weight ratio of the outgoing edges, and superimposes the damping factor to balance the random jump probability until the score converges; Finally, dividing the node PageRank values into discrete depth levels according to quantiles. The higher the value, the deeper the level, reflecting the core degree in the global association network.

8. A document management method for logistics transportation according to claim 6, characterized in that, When adding a new document, efficiently adjusting the association strength and importance score of the affected nodes through local association calculation and incremental update strategy, including: When adding a new document, only calculate the association degree between the new document and the existing documents, and lock the set of directly associated nodes; Through the incremental update strategy, insert the newly added association edge into the graph structure, only trigger the re-accumulation of the global association strength of the relevant nodes, and iteratively update the PageRank score based on the local subgraph; Constrain the influence range of random jumps through the damping factor. When the change in the node score exceeds the threshold, expand and update the neighborhood layer by layer until the change amount of the node score is less than the set convergence threshold and stop the iteration.

9. A document management method for logistics transportation according to claim 1, characterized in that, Based on the depth sorting result, perform an association degree marking on the newly generated document, trigger incremental document reconstruction according to the updated network topology, and synchronously update the global sorting, including: When performing an association degree marking on a new document based on the depth sorting result, preferentially select core nodes at the same level or higher levels to establish strong associations, and generate association edges through the dynamic weight distribution rule.

10. A document management system for logistics transportation, which is used to implement the method for document management for logistics transportation according to any one of claims 1-9, characterized in that, Including: A data association analysis module that extracts the matching degree and process coupling relationship between documents based on the process sequence and field matching rules, and dynamically calculates the basic association degree; The coupling degree integrates time series, status, and resource characteristics to form a comprehensive evaluation; An anomaly detection module that identifies the interaction density of abnormal documents through the topological network, quantifies the abnormal characteristics, and corrects the association degree; A depth sorting module that models the document as a weighted directed graph, calculates the core degree of the nodes to divide the levels; when adding a new document, locally update the node scores and control the influence range; A dynamic reconstruction module that preferentially binds high-core-level nodes, assigns weights according to the level difference, triggers incremental reconstruction to update the topology, and ensures the real-time performance of sorting and the stability of the network; Realize anomaly isolation and dynamic balance maintenance.

Citation Information

Patent Citations

  • Abnormal medical insurance document identification method and device, computer equipment and storage medium

    CN111340638A

  • Multi-source heterogeneous data treatment method and system

    CN115718779A