Document full life cycle closed loop management method and system based on multi-source data fusion
Patent Information
- Application Number
- CN202610844543.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-11
- Publication Date
- 2026-09-25
AI Technical Summary
[0006]本发明的目的在于提供一种,以解决上述背景技术中提出的现有的问题
1、本发明首先通过在多源异构文档解析过程中引入行业专属关键字实时校验、分级预警以及语义词典支撑的多维特征提取与加权融合,解决了现有文档识别结果事后校验、不同来源数据语义口径不统一及字段冲突难以处理的问题,实现了多源文档数据的前置风险拦截、统一语义表达和可信融合输出。
Smart Images

Figure CN122817428A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of document data processing technology, specifically to a method and system for closed-loop management and control of the entire document lifecycle based on multi-source data fusion. Background Technology
[0002] With the continuous increase in the number of documents in government affairs, enterprise office work, and project management scenarios, the sources, formats, and field expressions of documents are becoming increasingly complex. Conventional document management methods mostly rely on single-point entry, manual verification, and segmented archiving, which makes it difficult to detect identification errors, missing fields, semantic inconsistencies, and permission risks in a timely manner. This leads to macro-level problems such as insufficient document data credibility, difficulty in process traceability, and fragmented security control.
[0003] Existing technologies typically employ optical character recognition to extract text from image-based documents, use template rules to complete fields, control document flow through a workflow engine, and combine role-based access control, log recording, and symmetric encryption to achieve basic management and security protection.
[0004] However, the above solutions still have the following problems: the identification results are mostly verified after the fact, and risky content and missing required fields are easy to flow into subsequent stages; the parsing results of documents from different sources are isolated from each other, lack a unified semantic standard, and it is difficult to handle conflicts in the same field; the general workflow can only record the process status, and it is difficult to accurately trace the problem data back to the responsible node and automatically trigger rectification; the access control and encryption strategies are mostly set independently and cannot be dynamically adjusted according to the document semantics and node status.
[0005] Therefore, this invention provides a method and system for closed-loop management of the entire document lifecycle based on multi-source data fusion. Summary of the Invention
[0006] The purpose of this invention is to provide a solution to the existing problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides the following technical solution: a closed-loop management method for the entire lifecycle of documents based on multi-source data fusion, comprising the following steps: S1. Parse multi-source heterogeneous format documents to extract raw data, perform real-time verification and hierarchical warning based on industry-specific keyword library during optical character recognition, and clean and standardize the raw data to obtain raw text data; S2. Construct and dynamically update an industry semantic dictionary that includes core semantic entities, semantic association rules, and synonym mappings; S3. Based on the original text data and the industry semantic dictionary, perform multi-dimensional semantic feature extraction to obtain a set of core semantic feature vectors; S4. Based on the document source credibility and data integrity, assign differentiated weights to the core semantic feature vector set and perform weighted fusion. Correct semantic conflicts through conflict resolution rules and encapsulate the fusion result into standardized structured data output. S5. Assign a unique node identifier to each flow node and establish an association mapping with document data and operation logs. Verify the standardized structured data and locate the problem node, trigger linkage rectification, and form a closed loop through secondary verification. S6. Based on the node identifier state switching operation permission, perform hierarchical encryption according to the industry semantic tags in the standardized structured data.
[0008] A further improvement of this invention is that the real-time verification and graded early warning process includes: S11. In the process of optical character recognition, a string matching algorithm based on a partial matching table is used to scan the recognition results segment by segment and compare them with the industry-specific keyword library. The keywords in the keyword library are divided into three categories: risk keywords, sensitive keywords, and required keywords, and are stored separately. S12. When a risky keyword is identified, a Level 1 warning is triggered, the identification process is paused, and the information is sent to the reviewer for manual confirmation. When a sensitive keyword is identified, a Level 2 warning is triggered, the location of the sensitive content is marked, and the location mark is passed to the hierarchical encryption step. When no required keyword is identified, a Level 3 warning is triggered, and the information is sent to the data entry personnel for a reminder to supplement the information. All warning information is automatically stored in the warning log table and bound to the document identifier and the node identifier.
[0009] A further improvement of this invention is that the process of constructing and dynamically updating the industry semantic dictionary includes: S21. Abstractly define the core business elements of the industry as core semantic entities, establish semantic association rules between semantic entities to define the structural relationship between entities, and establish synonym mapping to unify and normalize semantically similar business terms. S22. Use web crawling technology to capture the latest industry business standards, connect to the interfaces of professional business systems, import professional organization standard documents, receive customized standards uploaded by users, and automatically update dictionary content and weights according to a preset cycle based on semantic deviations reported by users.
[0010] A further improvement of this invention is that the process of extracting the multidimensional semantic features includes: The algorithm calculates word weights using the inverse document frequency (IVF) algorithm and transforms words into low-dimensional dense vectors using a word vector model. For core business fields in the industry, a dual extraction logic of format matching and semantic matching is employed. First, regular expressions are used to match the field format with the industry semantic dictionary to verify semantic legitimacy. Second, cosine similarity is used to verify the rationality of relationships between contextually related fields. Documents lacking core business fields or with semantic conflicts reaching a preset threshold are extracted for abnormal features and marked as requiring manual intervention. Redundant features are removed using variance analysis to obtain the core semantic feature vector set. The word vector model is trained using a labeled dataset formed by collecting multi-source document data from the target industry. Gradient descent is used to minimize semantic fusion error and dynamically adjust the learning rate. The final model is obtained after verifying semantic matching accuracy, fused data completeness, and conflict resolution accuracy using a test set.
[0011] A further improvement of this invention is that the process of assigning differentiated weights to perform weighted fusion includes: S41. Determine the source credibility based on the document acquisition method and assign a corresponding credibility weight; determine the data integrity based on the completeness of document fields and assign a corresponding integrity weight; the comprehensive weight of each source document is obtained by multiplying the credibility weight by the integrity weight. S42. Perform weighted fusion on the core semantic feature vector set of multiple source documents based on the comprehensive weight; the standardized structured data includes document basic information blocks, preprocessed original data blocks, semantic feature data blocks and fused structured data blocks, wherein the fused structured data blocks include business field values, semantic conflict records and fusion credibility scores.
[0012] A further improvement of this invention is that the allocation of a unique node identifier and the verification and location of the problem node include: Each circulation node is assigned a node identifier using a universally unique identification code algorithm. A mapping table is established in the underlying database to associate the node identifier with document data, operation logs, and permission information. During document circulation, a new node identifier is automatically generated while the parent node identifier is retained to form a node link. A dual verification mechanism of rule engine and semantic verification is used to verify the standardized structured data field by field and node by node. When a problem is found, a data fingerprint is generated using a hash algorithm and matched with the corresponding node identifier. The problem location is located by combining the content offset. A rectification notice is sent to the responsible personnel through a message queue. After rectification, a second verification is triggered. If the verification fails, the system locks the document and triggers a reminder again.
[0013] A further improvement of this invention is that the switching operation permissions and the execution of hierarchical encryption include: Based on the changes in the node identifier status, operation permissions are automatically assigned: input nodes are assigned editing permissions, review nodes are assigned approval permissions, and archive nodes are assigned read-only permissions; industry semantic tags in the standardized structured data and the locations of sensitive content marked in the hierarchical warning are read, and high-strength symmetric encryption is enabled for core documents or documents with sensitive tags and operation logs, while basic-strength symmetric encryption is enabled for ordinary documents.
[0014] A further improvement of this invention is that the conflict resolution rules include a time-priority rule, a permission-priority rule, and an official document-priority rule, which are executed sequentially. The time-priority rule uses the latest timestamp data as the standard, the permission-priority rule overwrites low-permission role data with high-permission role data, and the official document-priority rule uses data obtained from official channels as the standard. When the number of conflicts reaches a preset threshold, forced fusion is stopped and the conflict is marked as requiring manual intervention. All resolution processes are recorded in the semantic conflict record field.
[0015] A further improvement of this invention is that the closed-loop linkage in step S5 is supported by four core data tables at the bottom layer, namely, the node document association table, the operation log table, the problem record table, and the linkage rule table; the linkage rules are preset in the workflow engine's underlying configuration to establish the correspondence between nodes, responsible personnel, and reminder methods, and can be dynamically adjusted through the configuration interface; the underlying layer records the reminder message sending status, and automatically triggers a secondary reminder when it is not read; the system lock is only released after the secondary verification is passed, and the entire feedback path is automatically recorded and bound to the node identifier.
[0016] On the other hand, the present invention provides a closed-loop management system for the entire lifecycle of documents based on multi-source data fusion, comprising: The data preprocessing module is used to parse multi-source heterogeneous format documents to extract raw data, perform real-time verification and hierarchical warning based on an industry-specific keyword library during optical character recognition, and clean and standardize the raw data to obtain raw text data. The semantic dictionary module is used to build and dynamically update an industry semantic dictionary that includes core semantic entities, semantic association rules, and synonym mappings. The feature extraction module is used to perform multi-dimensional semantic feature extraction based on the original text data and the industry semantic dictionary to obtain a set of core semantic feature vectors; The data fusion module is used to assign differentiated weights to the core semantic feature vector set based on the credibility of the document source and the integrity of the data, perform weighted fusion, correct semantic conflicts through conflict resolution rules, and encapsulate the fusion result into standardized structured data output. The node management module is used to assign a unique node identifier to each flow node and establish an association mapping with document data and operation logs. It verifies the standardized structured data, locates problem nodes, and triggers linkage rectification to form a closed loop after secondary verification. The permission encryption module is used to switch operation permissions based on the node identifier status and to perform hierarchical encryption according to the industry semantic tags in the standardized structured data.
[0017] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention firstly solves the problems of post-verification of document recognition results, inconsistent semantic interpretation of data from different sources, and difficulty in handling field conflicts by introducing industry-specific keywords for real-time verification, hierarchical early warning, and multi-dimensional feature extraction and weighted fusion supported by semantic dictionaries during the parsing of multi-source heterogeneous documents. It achieves pre-emptive risk interception, unified semantic expression, and reliable fusion output of multi-source document data.
[0018] 2. By assigning a unique node identifier to each flow node of the document and establishing a correlation mapping between node identifiers, document data, operation logs, problem records and linkage rules, the problems of difficult accurate traceability of problem data, disconnection of rectification feedback links and static permission encryption strategies in the existing process management are solved. This enables automatic location of problem nodes, closed-loop linkage of rectification, dynamic switching of permissions and hierarchical security control of documents. Attached Figure Description
[0019] Figure 1 This is a flowchart of a document lifecycle closed-loop management method based on multi-source data fusion, according to the present invention. Detailed Implementation
[0020] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0021] The term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three cases: A exists alone, A and B exist simultaneously, and B exists alone.
[0022] Example 1 Figure 1 This embodiment presents a flowchart of a document lifecycle closed-loop management method based on multi-source data fusion, with the following steps: This embodiment is applied to an office document management platform. The main body is a server cluster deployed on the intranet. The source document types include portable document format files, office document files, spreadsheet files, and image files generated by scanning devices. The platform needs to process multiple documents of the same business matter from different business systems, different upload channels, and different data entry personnel, and aggregate them into a reliable, traceable, and transferable archive result.
[0023] In the parsing and extraction stage, the traditional approach is to have various format files output as plain text by a separate tool and then directly input them into the database. However, this approach has the following two problems: First, the text obtained from image documents after optical character recognition often contains typos and layout noise. Verification is only performed after recognition, meaning it's discovered retrospectively. If inappropriate expressions or missing key identifiers are found, the problem has already entered subsequent processes, making remediation costly. Second, the parsing results for different formats are isolated, making cross-referencing under a unified semantic standard difficult.
[0024] To address the aforementioned issues, this embodiment incorporates real-time pre-processing verification and tiered early warning during optical character recognition. This includes a pre-configured industry-specific keyword library in the system backend, categorizing keywords into three types: risky, sensitive, and required. As the recognition engine outputs text segment by segment, it simultaneously scans against this library. If a risky keyword is detected, a level one warning is triggered, pausing the recognition process and prompting manual confirmation from reviewers. If a sensitive keyword is detected, a level two warning is triggered, recording the start and end positions of the sensitive content in the document and passing these positions to subsequent encryption steps. If no required keyword is detected, a level three warning is triggered, prompting data entry personnel to supplement the data.
[0025] Through the above methods, the original reactive discovery is transformed into proactive interception. All warning entries are automatically written to the warning log table and bound to document and node identifiers. After verification, the extracted results undergo cleaning and standardization: the cleaning process removes spaces, line breaks, and special characters, and uses an industry dictionary to correct typos; the standardization process standardizes dates to year-month-day format, amounts to numbers plus units, and department names to standardized abbreviations, thus obtaining the original text data.
[0026] However, general language models struggle to understand specific expressions used in government affairs, such as determining whether an approval department and a competent business authority refer to the same entity. Therefore, this embodiment subsequently constructs an industry-specific semantic dictionary as a supporting layer, abstracting core semantic entities such as approval departments, archive numbers, and verification standards. It establishes semantic association rules, such as archive numbers consisting of year, department code, and serial number, and uses synonym mapping to unify review, approval, and signature into the same operational meaning. This dictionary is not a one-time creation but rather developed by crawling the latest industry standards, connecting to interfaces of professional business systems, importing standard documents from professional institutions, and receiving customized standards uploaded by users. It is automatically updated monthly with semantic deviations based on user feedback, ensuring accurate semantic understanding even when business rules change. User-reported semantic deviations are recorded as feedback entries, each containing the corrected word, the user-specified correct semantic entity, and the number of feedback entries. When the dictionary is updated, the mapping weight of entries in the synonym mapping that are fed back to incorrect entities is reduced. The reduction amount is proportional to the cumulative number of feedbacks for that mapping entry. When the mapping weight falls below a preset retention threshold, the mapping entry is removed from the dictionary. For correct semantic entities specified by the user, the corresponding mapping weight is increased, with the increased weight not exceeding the upper weight limit of 1. In this embodiment, the retention threshold is preferably set to 0.1, the upper weight limit is set to 1, and the reduction ratio is preferably set to 0.05. That is, for each cumulative feedback, the corresponding mapping weight is reduced by 0.05. The reason for setting the ratio to 0.05 is that if the ratio is too high, the dictionary will be severely disturbed by individual false feedbacks, while if the ratio is too low, the dictionary will be slow to respond to persistent deviations. 0.05 strikes a balance between stability and response speed.
[0027] After obtaining the original text data and the industry semantic dictionary, it is necessary to convert the scattered text into a computable and comparable vector representation. Therefore, this embodiment performs multi-dimensional semantic feature extraction: the word frequency inverse document frequency algorithm is used to measure the importance of words in a single document, the word vector model is used to map words into low-dimensional dense vectors to capture semantic associations, the format of core business fields such as archive number is first matched with regular expressions and then the semantic legality is verified with a dictionary, the cosine similarity is used to verify the rationality of the relationship between contextual fields, and redundant terms are eliminated through variance analysis. Finally, a set of core semantic feature vectors is obtained. This batch of vectors provides a unified comparison benchmark for the next step of eliminating multi-source conflicts.
[0028] Multiple documents from different sources often give contradictory values for the same business field. For example, one document might show an approval status of "approved," while another shows "under review." Simply taking the union or overwriting the last document will inevitably distort the archiving results. Therefore, this embodiment assigns differentiated weights to the core semantic feature vector set and performs weighted fusion. This includes determining the source credibility based on the document acquisition method and assigning a credibility weight, and determining data integrity based on the field completeness and assigning an integrity weight. The comprehensive weight of each source document is obtained by multiplying the two weights. Based on this, multi-source vectors are weighted and fused. The process can be represented as follows: In the formula, This represents the fused semantic vector. Indicates the first A set of core semantic feature vectors of the source document. Indicates the first The source credibility weight of each document. Indicates the first Data integrity weight of the source document This indicates the total number of source documents participating in the fusion. This is a very small positive number, used to prevent the denominator from being zero when the sum of the weights of all documents approaches zero. In this embodiment... Preferred selection This value is far lower than the overall weight of any normal document, ensuring that it does not disturb the normal fusion result and guarantees computational robustness in extreme cases. The reason for using multiplication instead of addition to construct the overall weight in the above process is that multiplication allows the overall weight to approach zero simultaneously when any dimension approaches zero, thus automatically removing documents with unreliable sources or severely missing fields from the dominant position; conversely, addition may allow a single high-scoring dimension to mask the defects of another dimension. After fusion, the result is packaged into standardized structured data output, containing basic document information blocks, preprocessed raw data blocks, semantic feature data blocks, and fused structured data blocks. The fused structured data blocks record business field values, semantic conflict records, and fusion credibility scores for subsequent circulation and archiving.
[0029] To ensure document traceability throughout the workflow, this embodiment assigns a unique node identifier to each document at every stage of the workflow, including creation, entry, review, completion, and archiving. An association mapping is established in the underlying database between node identifiers and document data, operation logs, and permission information. When a document enters a new node, a new node identifier is automatically generated while retaining the parent node identifier, thus forming a node chain. The verification process employs a dual mechanism of rule engine and semantic verification, scanning standardized structured data field by field and node by node. Once a problem is detected, a data fingerprint is generated using a hash algorithm to look up the node identifier, and the problem location is then determined by combining it with content offset. A rectification notification is then pushed to the responsible personnel via a message queue. After rectification, a second verification is triggered; if it fails, the document is locked and a reminder is issued again, thus forming a closed loop of problem detection, location marking, reminder triggering, and rectification feedback.
[0030] Finally, this embodiment, based on the node identifier status switching operation permissions, grants editing permissions to input nodes, approval permissions to review nodes, and read-only permissions to archive nodes. It also reads the industry semantic tags in the standardized structured data and the sensitive content locations of the aforementioned secondary warning records, and enables high-strength symmetric encryption for core documents or documents containing sensitive tags and operation logs, while enabling basic-strength symmetric encryption for ordinary documents.
[0031] Example 2 This embodiment further defines each step based on Embodiment 1, and is applied to document processing scenarios involving overlapping engineering acceptance and financial settlement within the same business platform. Such scenarios involve multiple key business fields such as acceptance amount, approval status, and archive number, and the document sources span both official interfaces and manual uploads, resulting in a high probability of conflicts and placing higher demands on the rationality of parameter values.
[0032] Regarding real-time verification and tiered early warning, this embodiment employs a string matching algorithm based on a partial matching table to scan the recognition results segment by segment. When a match fails, this algorithm uses the constructed partial matching table to skip unnecessary backtracking, ensuring that the time cost of segment-by-segment scanning increases linearly with text length, rather than multiplying with the size of the keyword database. Therefore, even when the keyword database expands to thousands of entries, the recognition stage can still maintain real-time performance. In the keyword database, risk keywords cover illegal expressions and sensitive words, sensitive keywords cover core confidential identifiers and privacy information, and required keywords cover document numbers and approval identifiers. These three categories are stored together, and the database supports manual addition, deletion, modification, and setting of matching thresholds for exact or fuzzy matching in the background.
[0033] The warning classification and handling follow the method of Example 1: Level 1 warnings suspend the process and push it to the reviewer for manual confirmation; Level 2 warnings record the start and end positions of sensitive content and pass them to the encryption stage; Level 3 warnings record missing keyword types and push them to the data entry personnel for supplementation; All warning entries, along with warning type, keyword content, identification location, warning time, handling personnel and handling results, are written into the warning log table and bound to document identifiers and node identifiers.
[0034] In terms of multidimensional semantic feature extraction, the inverse document frequency (IVF) algorithm is used to measure the importance of a word within a single document relative to the entire corpus. Its calculation can be expressed as: In the formula, Words In the document The weights in Words In the document The number of times it appears in This indicates the total number of documents in the corpus. This indicates that the corpus contains vocabulary. The number of documents, For extremely small positive numbers, this embodiment preferably takes... This minimal positive number takes into account both boundary conditions: when a word appears in all documents, it makes... equal This can avoid anomalies in several denominators; when a word is a newly added word that has not yet been included in any document, it can... Even when the denominator is zero, the operation can still continue, which is a fault-tolerant handling for insufficient sampling and zero denominator. Logarithmic compression of the inverse document frequency term is used here because the impact of document frequency changes on word discrimination ability decreases over time. Logarithmic transformation can suppress the weight of high-frequency general words and amplify the weight of rare specific words, thus better matching the sparse distribution of key fields in government affairs. After the word weights are determined, they are transformed into low-dimensional dense vectors by the word vector model, with 128 dimensions being the preferred choice. This dimension is calibrated by gradually increasing the dimension of the training corpus and observing the convergence of semantic matching accuracy. Too low a dimension is insufficient to bear the semantic association between government terms, easily leading to confusion of near-synonyms, while too high a dimension introduces redundant components and significantly increases the computational burden of subsequent fusion; 128 dimensions strike a balance between accuracy convergence and computational overhead. For core business fields such as archive number, this embodiment first uses regular expressions to match the format of the three sub-items: year, department code, and serial number, and then uses an industry semantic dictionary to verify semantic validity, ensuring that key field extraction is not misaligned.
[0035] To verify the reasonableness of the relationships between context-related fields, this embodiment calculates the cosine similarity between the related field vectors, which can be expressed as: In the formula, This represents the similarity value between two related field vectors. and These represent the semantic vectors of the two related fields being compared, such as approval date and effective date, department name and person in charge. Represents the dot product of two vectors. and Let them represent the magnitudes of the two vectors respectively. For extremely small positive numbers, this embodiment preferably takes... This value is used to prevent the denominator from being zero when the vector of a certain associated field is zero, causing the magnitude to be zero. This value is less than the common magnitude of word vector components, so it does not affect normal similarity calculation. Cosine similarity is chosen instead of Euclidean distance to measure the reasonableness of the association because the degree of association between semantic vectors is mainly reflected in directional consistency rather than length difference. Cosine similarity only describes the directional angle, which can avoid interference from differences in vector magnitudes in association judgment. When the similarity value is lower than the preset association threshold, the system determines that the context association is biased and includes the corresponding field in the anomaly list.
[0036] Furthermore, for documents that lack core business fields or whose semantic conflicts reach a preset threshold, this embodiment extracts anomalies and marks them as requiring manual intervention. The preset threshold for the number of semantic conflicts is preferably three. This threshold is determined by statistically analyzing the correspondence between the number of conflicts in historical documents and the necessity of manual review. When the number of conflicts is less than three, they are mostly occasional differences in expression and can be automatically handled by the resolution rules. Once the number of conflicts reaches or exceeds three, it indicates that there is a systemic contradiction between documents. Continuing to force fusion will introduce unreliable results, and it is more prudent to switch to manual intervention.
[0037] It should be noted that the feature extraction stage and the weighted fusion stage share the same conflict counter, which accumulates the count based on the same business field: the contextual correlation deviations discovered by the cosine similarity verification in the feature extraction stage and the field discrepancies processed by the conflict resolution rules in the weighted fusion stage are both counted in the same count. They are not two independent thresholds, but rather a unified determination of the number of conflicts for the same business field. Therefore, the preset thresholds mentioned in the two places in the claims point to the same value, namely three.
[0038] Finally, analysis of variance is used to remove redundant items with low relevance to document management business, retaining core items to reduce the computational burden of subsequent fusion. Specifically, all candidate feature items are grouped according to their respective document management business categories. For each feature item, the ratio of between-group variance to within-group variance is calculated as the discriminant statistic. The larger the statistic, the more significant the difference in the value of the feature item among different business categories and the higher its relevance to document management business. A discriminant threshold is set, and feature items with a statistic below the threshold are judged as redundant items and removed. The threshold is determined by gradually increasing the value of the labeled dataset and observing the marginal contribution of the retained feature items to the semantic matching accuracy. In this embodiment, 1.5 is preferred because a threshold that is too low will make it difficult to filter out noisy features, while a threshold that is too high will mistakenly delete feature items with certain discriminative ability, resulting in information loss. 1.5 strikes a balance between noise removal and retention of effective information.
[0039] Regarding weighted fusion and conflict resolution, this embodiment determines source credibility based on the document acquisition method. Documents obtained from official channels or core interfaces are preferably weighted at 0.8, while documents manually uploaded by users or obtained through image recognition are preferably weighted at 0.2. This ratio is determined by retrospectively analyzing the accuracy differences between the two types of sources in historical archive results. Given that the accuracy of official sources is significantly higher than that of manually uploaded documents, the former is given several times the weight of the latter, making the fusion result closer to a credible source. Correspondingly, data integrity is determined based on field completeness: the integrity weight of complete documents with no missing fields is preferably 0.7, while the integrity weight of documents with missing fields is preferably 0.3. This is because although incomplete documents can still provide partial information, they are not suitable for dominating the fusion process; therefore, their contribution is preserved while their impact is reduced.
[0040] The overall weight is obtained by multiplying the credibility weight and the completeness weight. When multiple source documents disagree on the same field, the conflict resolution rule pool is triggered. In this embodiment, the resolution rules are executed in sequence: timeliness priority, permission priority, and official document priority, including: The timeliness priority rule uses the latest timestamp data as the standard, the permission priority rule prioritizes data submitted by higher-permission roles, such as approvers, overriding data submitted by lower-permission roles, such as data entry personnel, and the official document priority rule prioritizes data obtained from official channels. The three rules are judged in order of priority until the disagreement is resolved. All resolution processes, along with conflict fields, rules used, and resolution results, are recorded in the semantic conflict record field of the merged structured data block. When the number of conflicts reaches the aforementioned threshold, forced fusion is stopped and the data is marked as awaiting manual intervention, directly connecting to the closed-loop linkage mechanism. The standardized structured data output from the fusion is encapsulated in a unified lightweight data exchange format: the document basic information block records the document identifier, document type, creation time, source department, uploader, and encryption level. The document identifier uses the format of year plus department code plus serial number, and the encryption level is divided into three levels, each corresponding to a different symmetric encryption strength. The preprocessed raw data block records the cleaned full text, the tabular data stored in a two-dimensional array, and the text cleaned by optical character recognition. The semantic feature data block records the core vocabulary vectors, word frequency inverse document frequency weights, and industry semantic tags. The fused structured data block records business field values, semantic conflict records, and fusion credibility scores.
[0041] In terms of node tracing and closed-loop linkage, this embodiment uses a universally unique identification code algorithm to generate a node identifier for each flow node. The identifier is written as node type code plus timestamp plus random number. For example, the node identifier for creating a node is written as CREATE plus a 20-digit timestamp plus a 6-digit random number. After the node identifier is bound to the document identifier, it is stored in the node-document association table. When the document flows, a new node identifier is automatically generated and the parent node identifier is retained to form a node link. Each node identifier corresponds to a unique operation permission, operator, and operation time, which are synchronously written to the operation log table. In the verification stage, a rule engine combined with semantic verification scans each field and node. When a field is missing, semantic conflict, or permission violation is found, a hash algorithm is used to generate a data fingerprint that uniquely corresponds to the content fragment. The node-document association table is queried to find the node identifier where the problem is located. Then, the content offset recorded during document parsing, i.e., the start and end positions of the text paragraph, the table row and column numbers, or the image coordinate range, is combined to generate a description of the problem location and write it to the problem record table.
[0042] The feedback path is pre-configured in the workflow engine's underlying settings, pre-determining the correspondence between nodes, responsible personnel, and reminder methods, and supports dynamic adjustment via a configuration interface. After the problem location is marked, the system retrieves the responsible personnel and reminder methods based on the node identifier in the problem record table, sends a reminder containing the problem location, problem description, rectification requirements, and rectification deadline via the message middleware, and generates a task to be rectified on the front end. Furthermore, the underlying system records the sending status of reminder messages. If the responsible personnel do not read the reminder before the rectification deadline, a second reminder is automatically triggered, with a preferred interval of one hour. This interval is determined by statistically analyzing the average response time of responsible personnel to rectification tasks: too short an interval will cause repeated disturbances, while too long an interval will delay rectification timeliness; one hour provides a reasonable processing window while ensuring the continuity of reminders. After the responsible personnel rectify the issue, the system automatically obtains the rectification data and re-triggers the verification: if the second verification passes, the status of the issue record table is updated to "rectified" and the system is allowed to proceed to the next node; if it fails, the system marks the location of the new issue and reminds the user again. The system lock is only lifted after the second verification passes. The issue discovery time, marked location, reminder sending time, rectification time, and verification result of the entire feedback path are all recorded and bound to the node identifier and document identifier.
[0043] The aforementioned closed loop is supported by four core data tables, including: a node-document association table that stores node identifier, document identifier, parent node identifier, node type, operation time, and operator identifier; an operation log table that stores log identifier, node identifier, operation content, operator identifier, operation time, and operation source address; a problem record table that stores problem identifier, document identifier, node identifier, problem description, problem location, rectification status, and rectification deadline; and a linkage rule table that stores rule identifier, node identifier, responsible personnel identifier, reminder method, and rectification deadline.
[0044] Regarding permission switching and hierarchical encryption, this embodiment automatically assigns permissions based on changes in node identifier status: input nodes are granted editing permissions, review nodes are granted approval permissions, and archive nodes are granted read-only permissions. In the encryption process, industry semantic tags and sensitive content locations in secondary warning records are read from standardized structured data. When a semantic tag matches a core term related to finance or confidentiality, or when a document contains sensitive markers, high-strength symmetric encryption is enabled for its content and corresponding operation logs, i.e., an advanced encryption standard with a 256-bit key length is used. Basic-strength symmetric encryption is enabled for ordinary documents and their operation logs, i.e., an advanced encryption standard with a 128-bit key length is used. This ensures security while avoiding excessive encryption overhead on ordinary documents.
[0045] Example 3 This embodiment presents a system for implementing the above method and the training and optimization process of its semantic fusion model. The system is deployed on a government cloud platform and consists of six deeply coupled functional modules, operating in the same environment as in Embodiment 1.
[0046] The data preprocessing module is responsible for parsing multi-source heterogeneous document formats to extract raw data. Portable document format files are processed by a text parsing library to extract text and layout information, office document files are processed by a document parsing library to extract paragraphs and tables, spreadsheet files are processed by a data processing library to extract cell content, and image files are processed by optical character recognition to extract text. During recognition, real-time verification and hierarchical early warning are performed based on an industry-specific keyword library. The raw text data is then cleaned and standardized before being output. The output of this module also serves as the input for subsequent modules, realizing the connection between pre-emptive prevention and control and the main process.
[0047] The semantic dictionary module is responsible for building an industry-specific semantic dictionary containing core semantic entities, semantic association rules, and synonym mappings, and updates it dynamically on a monthly basis, providing semantic support for the feature extraction module. The feature extraction module receives raw text data and the industry semantic dictionary, performs word frequency inverse document frequency calculation, word vector transformation, dual extraction of regularization and semantics, cosine similarity association verification, and variance analysis for redundancy reduction, and outputs a set of core semantic feature vectors.
[0048] The data fusion module assigns differentiated weights based on the credibility of the source and the integrity of the data to perform weighted fusion. After the discrepancies are corrected by conflict resolution rules, the data is packaged into standardized structured data output.
[0049] The node management module assigns a unique node identifier to each workflow node and establishes a mapping between it and document data and operation logs. It verifies standardized structured data, locates problematic nodes, and triggers coordinated rectification, forming a closed loop through secondary verification. The access control module switches operation permissions based on the node identifier status and performs hierarchical encryption based on industry semantic tags in the standardized structured data. It is particularly important to note that when this system is deployed on the government intranet, its external access points should be configured with identity authentication and access control, and interface calls between modules should also carry token verification to prevent unauthorized access to standardized structured data and operation logs without authentication; this security measure is indispensable during actual system operation.
[0050] The semantic fusion model upon which the feature extraction and data fusion modules rely requires training, validation, and optimization before deployment. The training dataset consists of over 100,000 source documents collected from the target industry over the past three years, covering portable document formats, office documents, spreadsheets, and image documents. It is divided into training, validation, and test sets in a 7:2:1 ratio. Before partitioning, the datasets are cleaned and standardized, and semantic labels and conflict cases are added to create a labeled dataset. This partitioning ratio is chosen because a larger proportion of the training set allows the model to fully learn the semantic relationships of government terminology, while the validation and test sets each occupy a certain proportion, enabling monitoring of generalization ability during training and providing independent evaluation results after training.
[0051] The semantic fusion model consists of a word vector layer, a feature aggregation layer, and a classification output layer connected in sequence. The word vector layer, as described above, is the trained word vector model that maps each word in the document to a 128-dimensional dense vector. The feature aggregation layer receives the output from the word vector layer and performs a weighted average of all word vectors in the document using the inverse document frequency weights of each word as coefficients, resulting in a 128-dimensional document-level semantic vector. This layer contains no trainable parameters; its function is to reduce the variable-length word vector sequence to a fixed-length document representation. The classification output layer is a fully connected layer with an input dimension of 128 dimensions and an output dimension equal to the total number of semantic label categories. This layer transforms the document-level semantic vectors linearly using a normalized exponential function into predicted values for each category. ,in That is, the first The sample at the th The predicted values for each category satisfy the condition that the sum of the predicted values for all categories for that sample is one. Therefore, the parameters of the word vector layer and the feature aggregation layer, as well as the connection weights of the classification output layer, together constitute the trainable parameters of the model, and are uniformly optimized using cross-entropy loss as the objective in the training process described below.
[0052] During model initialization, the word vector dimension was set to 128 dimensions, the window size to 5, the maximum number of iterations to 100, and the initial learning rate to 0.025. For word frequency-inverse document frequency initialization, the minimum word frequency was set to 3, and stop words were removed according to the industry stop word list. Since the semantic relationships in government documents are mostly concentrated within a few adjacent words, a window that is too small would break the relationships, while a window that is too large would introduce interference from irrelevant words; therefore, a window size of 5 was chosen. The minimum word frequency was set to 3 because words appearing less than three times are mostly occasional noise or individual typos, and filtering them out avoids low-frequency noise disturbing the word vector space. During training, the training set was input into the initialized model, and the semantic fusion error was minimized using the gradient descent algorithm. The error function used was cross-entropy loss, which can be expressed as: In the formula, This represents the average loss over a batch. This indicates the number of samples in this batch. This represents the total number of categories for semantic tags. Indicates the first The sample at the th The actual labeled values for each category, The model represents the first The sample at the th Predicted values for each category For extremely small positive numbers, this embodiment preferably takes... This is used to avoid logarithmic divergence when the model's predicted values approach zero; this value is much smaller than the common order of magnitude of the predicted probability, so it does not affect the normal loss calculation. Cross-entropy loss is chosen because semantic label discrimination is essentially a multi-class classification problem, and cross-entropy can directly characterize the difference between the predicted distribution and the true distribution. Moreover, its gradient is larger when the prediction deviation is large and smaller when it is close to the true value, which is beneficial for rapid convergence in the early stage of training and stable refinement in the later stage.
[0053] During training, the semantic matching accuracy is evaluated using the validation set every 20 iterations. If the accuracy improvement between two consecutive evaluations is less than 0.1 percentage points, the learning rate is halved before training continues. The reason for using the accuracy improvement as the trigger for halving the learning rate is that a narrowing improvement indicates the model is approaching a local minimum at the current learning rate. Halving the learning rate allows it to approach a better solution with finer step sizes, thus avoiding oscillations near the minimum. The reduction in the learning rate is achieved by multiplying the current value by half. Training stops when the number of iterations reaches 100 or the validation set accuracy reaches 98% or higher. After training, the semantic matching accuracy, data integrity, and conflict resolution accuracy are validated using the test set. The values of these three metrics are obtained through backtracking and calibration of the acceptable error rate of government archive results. The conflict resolution accuracy requirement is the highest because conflict fields directly determine the credibility of the archive results; resolving errors can lead to misjudgments of the status of business matters. If any metric fails to meet the target, analyze the causes of semantic matching errors and conflict resolution failures, such as missing dictionary entries or extraction biases. Update the industry semantic dictionary and model parameters, and repeat the above training process until all metrics meet the target.
[0054] The coefficients, thresholds, and weights involved in the above embodiments can be set by default according to the present invention, or can be set by those skilled in the art.
[0055] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0056] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0057] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0058] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0059] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A closed-loop management method for the entire document lifecycle based on multi-source data fusion, characterized by: Includes the following steps: S1. Parse multi-source heterogeneous format documents to extract raw data, perform real-time verification and hierarchical warning based on industry-specific keyword library during optical character recognition, and clean and standardize the raw data to obtain raw text data; S2. Construct and dynamically update an industry semantic dictionary that includes core semantic entities, semantic association rules, and synonym mappings; S3. Based on the original text data and the industry semantic dictionary, perform multi-dimensional semantic feature extraction to obtain a set of core semantic feature vectors; S4. Based on the document source credibility and data integrity, assign differentiated weights to the core semantic feature vector set and perform weighted fusion. Correct semantic conflicts through conflict resolution rules and encapsulate the fusion result into standardized structured data output. S5. Assign a unique node identifier to each flow node and establish an association mapping with document data and operation logs. Verify the standardized structured data and locate the problem node, trigger linkage rectification, and form a closed loop through secondary verification. S6. Based on the node identifier state switching operation permission, perform hierarchical encryption according to the industry semantic tags in the standardized structured data.
2. The document lifecycle closed-loop management method based on multi-source data fusion according to claim 1, characterized in that: The real-time verification and graded early warning process includes: S11. In the process of optical character recognition, a string matching algorithm based on a partial matching table is used to scan the recognition results segment by segment and compare them with the industry-specific keyword library. The keywords in the keyword library are divided into three categories: risk keywords, sensitive keywords, and required keywords, and are stored separately. S12. When a risky keyword is identified, a Level 1 warning is triggered, the identification process is paused, and the information is sent to the reviewer for manual confirmation. When a sensitive keyword is identified, a Level 2 warning is triggered, the location of the sensitive content is marked, and the location mark is passed to the hierarchical encryption step. When no required keyword is identified, a Level 3 warning is triggered, and the information is sent to the data entry personnel for a reminder to supplement the information. All warning information is automatically stored in the warning log table and bound to the document identifier and the node identifier.
3. The document lifecycle closed-loop management method based on multi-source data fusion according to claim 1, characterized in that: The process of constructing and dynamically updating the industry semantic dictionary includes: S21. Abstractly define the core business elements of the industry as core semantic entities, establish semantic association rules between semantic entities to define the structural relationship between entities, and establish synonym mapping to unify and normalize semantically similar business terms. S22. Use web crawling technology to capture the latest industry business standards, connect to the interfaces of professional business systems, import professional organization standard documents, receive customized standards uploaded by users, and automatically update dictionary content and weights according to a preset cycle based on semantic deviations reported by users.
4. The document lifecycle closed-loop management method based on multi-source data fusion according to claim 1, characterized in that: The process of extracting multidimensional semantic features includes: The algorithm calculates word weights using the inverse document frequency (IVF) algorithm and transforms words into low-dimensional dense vectors using a word vector model. For core business fields in the industry, a dual extraction logic of format matching and semantic matching is employed. First, regular expressions are used to match the field format with the industry semantic dictionary to verify semantic legitimacy. Second, cosine similarity is used to verify the rationality of relationships between contextually related fields. Documents lacking core business fields or with semantic conflicts reaching a preset threshold are extracted for abnormal features and marked as requiring manual intervention. Redundant features are removed using variance analysis to obtain the core semantic feature vector set. The word vector model is trained using a labeled dataset formed by collecting multi-source document data from the target industry. Gradient descent is used to minimize semantic fusion error and dynamically adjust the learning rate. The final model is obtained after verifying semantic matching accuracy, fused data completeness, and conflict resolution accuracy using a test set.
5. The document lifecycle closed-loop management method based on multi-source data fusion according to claim 1, characterized in that: The process of assigning differentiated weights to perform weighted fusion includes: S41. Determine the source credibility based on the document acquisition method and assign a corresponding credibility weight; determine the data integrity based on the completeness of document fields and assign a corresponding integrity weight; the comprehensive weight of each source document is obtained by multiplying the credibility weight by the integrity weight. S42. Perform weighted fusion on the core semantic feature vector set of multiple source documents based on the comprehensive weight; the standardized structured data includes document basic information blocks, preprocessed original data blocks, semantic feature data blocks and fused structured data blocks, wherein the fused structured data blocks include business field values, semantic conflict records and fusion credibility scores.
6. The document lifecycle closed-loop management method based on multi-source data fusion according to claim 1, characterized in that: The process of assigning a unique node identifier and verifying and locating the problem node includes: Each circulation node is assigned a node identifier using a universally unique identification code algorithm. A mapping table is established in the underlying database to associate the node identifier with document data, operation logs, and permission information. During document circulation, a new node identifier is automatically generated while the parent node identifier is retained to form a node link. A dual verification mechanism of rule engine and semantic verification is used to verify the standardized structured data field by field and node by node. When a problem is found, a data fingerprint is generated using a hash algorithm and matched with the corresponding node identifier. The problem location is located by combining the content offset. A rectification notice is sent to the responsible personnel through a message queue. After rectification, a second verification is triggered. If the verification fails, the system locks the document and triggers a reminder again.
7. The document lifecycle closed-loop management method based on multi-source data fusion according to claim 1, characterized in that: The switching operation permissions and execution of hierarchical encryption include: Based on the changes in the node identifier status, operation permissions are automatically assigned: input nodes are assigned editing permissions, review nodes are assigned approval permissions, and archive nodes are assigned read-only permissions; industry semantic tags in the standardized structured data and the locations of sensitive content marked in the hierarchical warning are read, and high-strength symmetric encryption is enabled for core documents or documents with sensitive tags and operation logs, while basic-strength symmetric encryption is enabled for ordinary documents.
8. The document lifecycle closed-loop management method based on multi-source data fusion according to claim 5, characterized in that: The conflict resolution rules include a time-priority rule, a permission-priority rule, and an official document-priority rule, which are executed sequentially. The time-priority rule uses the latest timestamp data as the standard, the permission-priority rule overwrites low-permission role data with high-permission role data, and the official document-priority rule uses data obtained from official channels as the standard. When the number of conflicts reaches a preset threshold, forced fusion is stopped and the conflict is marked as requiring manual intervention. All resolution processes are recorded in the semantic conflict record field.
9. The document lifecycle closed-loop management method based on multi-source data fusion according to claim 6, characterized in that: In step S5, the closed-loop linkage is supported by four core data tables at the bottom layer: the node document association table, the operation log table, the problem record table, and the linkage rule table. The linkage rules are preset in the workflow engine's underlying configuration to establish the correspondence between nodes, responsible personnel, and reminder methods, and can be dynamically adjusted through the configuration interface. The underlying layer records the reminder message sending status, and automatically triggers a secondary reminder when the message is not read. The system lock is released only after the second verification is passed, and the entire feedback path is automatically recorded and bound to the node identifier.
10. A document lifecycle closed-loop management system based on multi-source data fusion, used to execute the document lifecycle closed-loop management method based on multi-source data fusion as described in any one of claims 1-9, characterized in that: include: The data preprocessing module is used to parse multi-source heterogeneous format documents to extract raw data, perform real-time verification and hierarchical warning based on an industry-specific keyword library during optical character recognition, and clean and standardize the raw data to obtain raw text data. The semantic dictionary module is used to build and dynamically update an industry semantic dictionary that includes core semantic entities, semantic association rules, and synonym mappings. The feature extraction module is used to perform multi-dimensional semantic feature extraction based on the original text data and the industry semantic dictionary to obtain a set of core semantic feature vectors; The data fusion module is used to assign differentiated weights to the core semantic feature vector set based on the credibility of the document source and the integrity of the data, perform weighted fusion, correct semantic conflicts through conflict resolution rules, and encapsulate the fusion result into standardized structured data output. The node management module is used to assign a unique node identifier to each flow node and establish an association mapping with document data and operation logs. It verifies the standardized structured data, locates problem nodes, and triggers linkage rectification to form a closed loop after secondary verification. The permission encryption module is used to switch operation permissions based on the node identifier status and to perform hierarchical encryption according to the industry semantic tags in the standardized structured data.