AI enterprise management document automatic filing system based on semantic analysis

CN122364167APending Publication Date: 2026-07-10SUZHOU GONGYING INTERNET TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZHOU GONGYING INTERNET TECH CO LTD
Filing Date
2026-06-05
Publication Date
2026-07-10

Smart Images

  • Figure CN122364167A_ABST
    Figure CN122364167A_ABST
Patent Text Reader

Abstract

This invention relates to the field of enterprise information management and intelligent document processing technology, specifically to an AI-based enterprise management document automatic archiving system based on semantic analysis. The system includes a cloud-based multi-source heterogeneous data acquisition module and a base semantic anchoring module, a spatiotemporal business graph construction module, a semantic context gravity coupling calculation module, and a fluid archiving execution module. By acquiring the location characteristics and flow range of business interface nodes and data source nodes, the system divides data into sub-regions, periodically collects document flow data and business pulse flow data, generates basic semantic vectors and business state weights, constructs a dynamic graph, calculates real-time correlation gravity, and executes virtual view archiving and permission mapping switching. This invention achieves the synchronous evolution of document archiving location, document permissions, and business visibility, avoiding the formation of knowledge silos due to fixed folder structures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of enterprise information management and intelligent document processing technology, specifically to an AI-based automatic archiving system for enterprise management documents based on semantic analysis. Background Technology

[0002] In enterprise management scenarios, procurement, R&D, legal affairs, auditing and other business operations will continuously generate a large number of documents. The documents in different business systems will also change constantly as projects progress, approval processes are processed and disputes are resolved. The efficiency of document archiving depends on the archiving system's ability to identify document content, business status and cross-system flow relationships.

[0003] In the traditional approach, enterprise documents are mostly stored separately by various business systems and archived according to fixed directories or preset rules. When the business context changes, documents are repeatedly referenced by multiple departments, or the circulation path is changed, it is still difficult to adjust the document archiving location and permission mapping in a timely manner, resulting in low document retrieval accessibility and low archiving stability across business scenarios. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention provides an AI-based automatic document archiving system for enterprise management based on semantic analysis. Specifically, the technical solution of this invention includes:

[0005] In the cloud, the cloud communication connection includes a multi-source heterogeneous data acquisition module, a base semantic anchoring module, a spatiotemporal service graph construction module, a semantic context gravity coupling calculation module, and a fluid archiving execution module;

[0006] The multi-source heterogeneous data acquisition module is used to acquire the location characteristics and flow range of several business interface nodes and data source nodes in the preset enterprise management system. Based on the location characteristics and flow range, the preset system is divided into several data sub-regions. According to the preset acquisition cycle, document flow data and business pulse flow data reflecting the real-time operating status of the business system are acquired from the business interface nodes of each data sub-region at regular intervals.

[0007] The base semantic anchoring module is used to perform digital signature verification and semantic parsing on document stream data to generate basic semantic vectors, and to calculate business state weights based on the basic semantic vectors.

[0008] The spatiotemporal business graph construction module is used to obtain business node prediction data for each data sub-region based on business pulse flow data. According to the business node prediction data of each data sub-region and the preset business state weight upper limit, the data sub-region is divided into high-quality data sub-regions or low-quality data sub-regions, and business graph collaborative scheduling is performed on the low-quality data sub-regions.

[0009] The semantic context gravity coupling calculation module is used to determine the gravity attenuation of business nodes in each data sub-region, calculate the real-time association gravity of business nodes based on the similarity between basic semantic vectors, and perform lightweight gravity compensation operation or adaptive collaborative gravity compensation operation based on the determination result.

[0010] The fluid archiving execution module is used to extract the business type and flow status of business nodes from document stream data and business pulse stream data, and to perform virtual view archiving and permission mapping switching operations based on the real-time association attraction, business type and flow status of business nodes.

[0011] Preferably, the process of the multi-source heterogeneous data acquisition module acquiring document stream data and business pulse stream data includes: deploying acquisition probes in each data sub-region, acquiring document stream data and business pulse stream data in the respective data sub-region in real time according to a preset acquisition cycle through the acquisition probes, and synchronizing the acquired document stream data and business pulse stream data to the cloud.

[0012] Preferably, the process by which the base semantic anchoring module performs digital signature verification and semantic parsing on document stream data to generate basic semantic vectors includes: presetting the public and private keys of the acquisition probes corresponding to each data sub-region through an asymmetric encryption algorithm; receiving document stream data sent by the multi-source heterogeneous data acquisition module after being signed by the acquisition probe's private key; digitally signing the document stream data in each acquisition cycle of the multi-source heterogeneous data acquisition module using the private key; verifying the validity of the digitally signed document stream data sent to the cloud by the multi-source heterogeneous data acquisition module using the public key; if the validity verification of the document stream data fails, the document stream data is removed and a parsing anomaly warning signal is generated; if the validity verification of the document stream data passes, the entity features and intent features of the document stream data are extracted, and basic semantic vectors are generated based on the entity features and intent features.

[0013] Basic semantic similarity is calculated based on the Euclidean distance between basic semantic vectors. The decay gravity value is calculated based on the time series change rate of the basic semantic similarity. The decay time range is determined based on the continuous interval of the decay gravity value being greater than zero. The duration of the business state is calculated based on the timestamp difference of the document stream data. The product of the basic semantic similarity and the duration is combined with a preset normalization upper bound for normalization processing to obtain the business state weight.

[0014] Preferably, the process of the spatiotemporal service graph construction module to obtain service node prediction data for each data sub-region includes: constructing a service node prediction model containing a long short-term memory network; obtaining service pulse flow data from several historical collection periods for each data sub-region as training data; training the service node prediction model using the training data; obtaining the completed training service node prediction model; inputting the service pulse flow data from the current collection period into the completed training service node prediction model; and outputting service node prediction data for each data sub-region containing a state weight time series from the completed training service node prediction model.

[0015] Preferably, the process of dividing a data sub-region into a high-quality data sub-region or a low-quality data sub-region includes: obtaining the preset upper limit of the business state weight for each data sub-region, extracting the time series sequence of the state weight in the business node prediction data of each data sub-region within the current collection period, comparing the time series sequence of the state weight of each data sub-region with the preset upper limit of the business state weight, and obtaining the cumulative time during which the state weight is continuously greater than or equal to the preset upper limit of the business state weight.

[0016] A preset cumulative time error threshold is set. If the cumulative time of a data sub-region is greater than or equal to the cumulative time error threshold, the data sub-region is marked as a low-quality data sub-region. If the cumulative time of a data sub-region is less than the error threshold, the data sub-region is marked as a high-quality data sub-region.

[0017] Preferably, the process of performing business graph collaborative scheduling on low-priority data sub-regions includes: Step 1: Obtain the time period in which the state weights of the low-priority data sub-regions exceed the preset upper limit of the business state weights and the maximum value of the excess weights within the time period; obtain the preset network topology containing each data sub-region; obtain the high-priority data sub-region that is closest to the low-priority data sub-region in the preset network topology; obtain the minimum idle weight of the high-priority data sub-region within the time period; determine whether the minimum idle weight is greater than or equal to the maximum value of the excess weights; if it is greater than or equal to, allocate the graph node corresponding to the minimum idle weight of the high-priority data sub-region to the low-priority data sub-region and end the business graph collaborative scheduling; if it is less than, allocate the graph node corresponding to the minimum idle weight of the high-priority data sub-region to the low-priority data sub-region, remove the high-priority data sub-region, and proceed to Step 2;

[0018] Step 2: Obtain the maximum value of the weight exceeding the low-quality data sub-region within the time period after the graph node allocation, re-obtain the high-quality data sub-region that is closest to the low-quality data sub-region in the preset network topology, obtain the minimum value of the idle weight of the re-obtained high-quality data sub-region within the time period, and then execute Step 3.

[0019] Step 3: Determine whether the minimum idle weight of the newly acquired high-priority data sub-region is greater than or equal to the maximum excess weight of the low-priority data sub-region within the time period. If it is greater than or equal to, allocate the graph node corresponding to the minimum idle weight of the newly acquired high-priority data sub-region to the low-priority data sub-region and end the business graph collaborative scheduling. If it is less than, allocate the graph node corresponding to the minimum idle weight of the newly acquired high-priority data sub-region to the low-priority data sub-region, remove the newly acquired high-priority data sub-region, and execute Step 2.

[0020] Preferably, the semantic context gravity coupling calculation module performs gravity attenuation determination on business nodes in each data sub-region, and performs lightweight gravity compensation operation or adaptive collaborative gravity compensation operation based on the determination result. This process includes: extracting the attenuation gravity value, attenuation time range, basic semantic similarity, duration of business state and weight of each basic semantic vector in the business nodes generated in real time in the data sub-region.

[0021] Determine whether the decay time range of the basic semantic vector is adjacent to or overlaps with the duration of the business state. If the decay time range is adjacent to or overlaps with the duration of the business state, the ratio of the decay gravity value to the basic semantic similarity is used as the comprehensive gravity ratio of the basic semantic vector. A preset comprehensive gravity ratio threshold is set. If the comprehensive gravity ratio is less than the comprehensive gravity ratio threshold, the business node to which the basic semantic vector belongs is marked as a gravity decay node. If the comprehensive gravity ratio is greater than or equal to the comprehensive gravity ratio threshold, the business node to which the basic semantic vector belongs is marked as a normal node.

[0022] Obtain the basic semantic vector and the number and location features of gravity decay nodes within the data sub-region, divide the data sub-region into several virtual view units of the same size, and obtain the gravity decay density of each virtual view unit based on the ratio of the number of gravity decay nodes to the total capacity of the topological nodes of the virtual view unit.

[0023] If the gravity decay density of the virtual view unit is less than or equal to the density threshold and there is a gravity decay node in the virtual view unit, then a lightweight gravity compensation operation is performed on the gravity decay node in the virtual view unit.

[0024] If the gravitational attenuation density of the virtual view cell is less than or equal to the density threshold and there are no gravitational attenuation nodes within the virtual view cell, no compensation operation is performed; if the gravitational attenuation density of the virtual view cell is greater than the density threshold, an adaptive cooperative gravitational compensation operation is performed on the area covered by the virtual view cell.

[0025] Preferably, the process of performing lightweight gravity compensation operation includes: obtaining several virtual directory paths in the preset archive link corresponding to the gravity decay node, extracting virtual directory paths that are not adjacent to or overlap with the decay time range, and marking the virtual directory paths that are not adjacent to or overlap with the decay time range as virtual directory paths to be allocated.

[0026] The search matching rate of each virtual directory path to be assigned is calculated based on the ratio of the historical search hits to the total search count. The virtual directory path with the lowest search matching rate is selected, the business status data is migrated to the selected virtual directory path, and the corresponding business status weight parameters are updated according to the attributes of the virtual directory path to be assigned.

[0027] Preferably, the process of performing adaptive cooperative gravity compensation includes: extracting the flow path corresponding to the gravity attenuation node in the virtual view unit coverage area, obtaining the data sub-region emitting the gravity attenuation node according to the flow path, and marking the data sub-region as the attenuation source; constructing an adaptive gravity compensation model based on a graph convolutional neural network, inputting the attenuation gravity value and attenuation time range of the attenuation source at the current moment and the topology graph corresponding to the virtual view unit coverage area into the adaptive gravity compensation model, extracting node features and graph structure features through the adaptive gravity compensation model, and outputting the optimal gravity adjustment parameters of the attenuation source.

[0028] Preferably, the process by which the fluid archiving execution module performs virtual view archiving and permission mapping switching operations based on the real-time correlation attraction, business type, and flow status of business nodes includes: extracting the real-time correlation attraction, business type, and flow status of each business node from the real-time collected data within the data sub-region;

[0029] The configuration is as follows: When the business node's flow status is far from the data sub-region and the real-time correlation attraction is less than the preset correlation attraction threshold corresponding to the business type, extract the preset usage data volume and real-time correlation attraction corresponding to the business type, obtain the idle data volume valley value and real-time correlation attraction valley value of other data sub-regions in the current collection period, obtain the target data sub-region where the idle data volume valley value is greater than the preset usage data volume and the real-time correlation attraction valley value is greater than the real-time correlation attraction, calculate the Euclidean distance between the target data sub-region and the business node, and select the target data sub-region with the shortest Euclidean distance to perform virtual view archiving and permission mapping switching operations.

[0030] When the business node's flow status is close to the data sub-region or the real-time association attraction is greater than or equal to the preset association attraction threshold corresponding to the business type, the current virtual view archiving and permission mapping status remains unchanged.

[0031] Compared with the prior art, the present invention has the following beneficial effects:

[0032] 1. By adopting the above technical solution, the present invention first divides the business interface nodes and data source nodes according to the data sub-regions by the multi-source heterogeneous data acquisition module, and then synchronously acquires document stream data and business pulse stream data according to the preset acquisition cycle through the regional acquisition probe. This effectively improves the continuous acquisition capability of document data and real-time business status in heterogeneous business scenarios such as procurement, production, and legal affairs, and reduces the distortion of archived input caused by inconsistent interfaces, permission isolation, or short-term disconnection.

[0033] 2. This invention verifies and semantically parses document stream data through a base semantic anchoring module and generates basic semantic vectors. Then, it calculates business state weights by combining basic semantic similarity, decay gravity value, and business state duration. This enables the system to not only identify document content but also quantify the effective association strength of documents in continuous business processes, thereby improving the ability to perceive dynamic business contexts.

[0034] 3. This invention uses a spatiotemporal business graph construction module to combine business pulse flow data and business node prediction model to output a state weight time series. Based on the cumulative time of continuous over-limit, the data sub-regions are divided into high-priority data sub-regions or low-priority data sub-regions. Furthermore, in conjunction with the business graph collaborative scheduling mechanism, the low-priority data sub-regions are released in a neighbor-first, step-by-step manner. This enables the graph maintenance and view mounting capability to be reconfigured before the business pressure increases, thereby improving the stability of cross-regional archiving.

[0035] 4. This invention performs gravity attenuation determination on business nodes through a semantic context gravity coupling calculation module, and selects lightweight gravity compensation operation or adaptive collaborative gravity compensation operation according to the gravity attenuation density in the virtual view unit. This can make targeted corrections to local archiving offsets and regional context migrations, and avoid fixed archiving paths from forming a long-term low-accessibility state after business shifts.

[0036] 5. This invention, through a fluid archiving execution module, combines real-time correlation gravity, business type, and flow status to perform virtual view archiving and permission mapping switching operations. While maintaining the original storage location of documents without physical migration or only logical migration, it achieves the synchronous evolution of document views, access permissions, and business flow relationships. This effectively solves the problem that traditional fixed directory archiving methods are difficult to adjust the archiving location and permission mapping in a timely manner when business context changes, frequent cross-departmental references, and flow path transfers. Ultimately, it can improve the retrieval accessibility, archiving stability, and compliance controllability of enterprise management documents in cross-business scenarios. Attached Figure Description

[0037] The present invention will be further explained below with reference to the accompanying drawings and embodiments:

[0038] Figure 1This is a schematic diagram of the modules of the AI-based enterprise management document automatic archiving system based on semantic analysis provided in the embodiments of this application. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.

[0040] The AI-based enterprise management document automatic archiving system based on semantic analysis includes a cloud platform. The cloud communication connection includes a multi-source heterogeneous data acquisition module, a base semantic anchoring module, a spatiotemporal business graph construction module, a semantic context gravity coupling calculation module, and a fluid archiving execution module.

[0041] The multi-source heterogeneous data acquisition module is used to acquire the location characteristics and flow range of several business interface nodes and data source nodes in the preset enterprise management system. Based on the location characteristics and flow range, the preset system is divided into several data sub-regions. According to the preset acquisition cycle, document flow data and business pulse flow data reflecting the real-time operating status of the business system are acquired from the business interface nodes of each data sub-region at regular intervals.

[0042] The base semantic anchoring module is used to perform digital signature verification and semantic parsing on document stream data to generate basic semantic vectors, and to calculate business state weights based on the basic semantic vectors.

[0043] The spatiotemporal business graph construction module is used to obtain business node prediction data for each data sub-region based on business pulse flow data. According to the business node prediction data of each data sub-region and the preset business state weight upper limit, the data sub-region is divided into high-quality data sub-regions or low-quality data sub-regions, and business graph collaborative scheduling is performed on the low-quality data sub-regions.

[0044] The semantic context gravity coupling calculation module is used to determine the gravity attenuation of business nodes in each data sub-region, calculate the real-time association gravity of business nodes based on the similarity between basic semantic vectors, and perform lightweight gravity compensation operation or adaptive collaborative gravity compensation operation based on the determination result.

[0045] The fluid archiving execution module is used to extract the business type and flow status of business nodes from document stream data and business pulse stream data, and to perform virtual view archiving and permission mapping switching operations based on the real-time association attraction, business type and flow status of business nodes.

[0046] This embodiment provides an AI-based automatic archiving mechanism for enterprise management documents based on semantic analysis, such as... Figure 1As shown; specifically, taking a large equipment manufacturing group as an example scenario, the group simultaneously has four types of concurrent businesses: new energy production line construction, cross-departmental procurement, external litigation response, and annual audit; the group internally connects to at least an enterprise resource planning system, a product lifecycle management system, an office automation system, a human resources system, and a legal management system. After all systems are uniformly connected to the cloud, documents are no longer fixed in a single tree-like directory, but documents, business nodes, and dynamic contexts are jointly mapped to a changeable virtual archive view;

[0047] The specific process is as follows: the multi-source heterogeneous data acquisition module obtains the location characteristics and flow range of the business interface nodes and data source nodes in the preset system; the location characteristics can be the network topology location, the system domain to which it belongs, the department domain to which it belongs, or the geographical data center location, and the flow range can be the system link range that the document is allowed to pass through;

[0048] For ease of explanation, assume that the group has three data sub-regions: production sub-region Z1, procurement sub-region Z2, and legal sub-region Z3. If the product lifecycle management system and manufacturing execution system interface are located in the production domain, the enterprise resource planning procurement interface is located in the procurement domain, and the litigation evidence database interface is located in the legal domain, then the system can, based on interface proximity and historical data flow, classify documents involving R&D drawings and test reports into Z1, procurement contracts and supplier letters into Z2, and lawyer's letters and infringement defense materials into Z3.

[0049] The system collects document stream data and business pulse stream data at the interface nodes of each sub-region at a preset collection cycle, such as once every 5 minutes. The former includes the document body, metadata, version number, attachment fingerprint, etc., while the latter includes real-time status data such as process approval status, project milestone changes, personnel reassignment, and external regulatory events.

[0050] The base semantic anchoring module performs encryption verification and semantic parsing on the document stream data to generate a basic semantic vector. For ease of micro-deduction, it is assumed that a procurement contract document D1, after parsing, forms a four-dimensional basic semantic vector. Their meanings can respectively correspond to procurement intensity, legal relevance, project dependence, and sensitivity to personnel changes; a certain R&D test report D2 is formed. ;

[0051] A lawyer's letter D3 was formed The system does not require the vector dimension to be four-dimensional; the aforementioned four-dimensional vector is only an example. The business status weight is calculated based on the similarity of the basic semantic vectors and the business status continuously reflected by the document within a certain time window. For example, if D1 is cited in three consecutive collection cycles of procurement approval, supplier claims, and project delays, its business status weight can be higher than that of a regular invoice attachment that has not been accessed after being uploaded once.

[0052] The spatiotemporal business graph construction module builds a dynamic graph based on business pulse flow data. Nodes in the graph may include supplier A, battery production line project P, legal case L, project manager U, etc., and edges may include relationships such as responsibility for citation, approval, and dispute association. To further illustrate with an illustrative deduction, if the risk of project P's delay increases during the current data collection period, the human resources system shows a change in project manager, and the legal system adds a new infringement objection, then the weight of the nodes connected to project P in the graph will increase accordingly.

[0053] Based on this, the system generates business node prediction data for each sub-region and compares it with the preset business status weight upper limit. When a sub-region continuously approaches or exceeds the upper limit in the next few cycles, it indicates that the pressure of reorganizing the document and business relationship in that region is increasing, and it is suitable to be marked as a low-priority data sub-region and start collaborative scheduling. Otherwise, it is retained as a high-priority data sub-region and undertakes the diversion tasks of other regions.

[0054] Based on this, the semantic context gravity coupling calculation module performs gravity attenuation determination on business nodes in each data sub-region; here gravity is not an abstract metaphor, but a calculation result used to measure the coupling strength between the current semantics of the document and the business context; schematically, D1 can be used to calculate three real-time associated gravity values ​​of 0.72, 0.68, and 0.81 with project P, supplier A, and case L respectively.

[0055] When a procurement contract is originally mainly related to procurement matters, but is frequently cited by legal departments due to infringement disputes, its attraction to case L may be higher than that of the procurement process nodes. Based on this, the system determines that the document is more suitable to be projected into the legal view in the current context. If the number of nodes with local attenuation in a certain area is less than the first preset threshold, lightweight attraction compensation is performed. If the area or number of nodes in the semantically and business-related regions exceeds the second preset threshold, adaptive collaborative attraction compensation is performed.

[0056] The fluid archiving execution module combines the real-time correlation of business nodes, business types, and flow status to switch the virtual view archiving location and permission mapping. Here, the virtual view is not a copy of the document itself, but rather a dynamically visible directory generated for different roles while keeping the original storage location unchanged or only logically migrated. For example, the original storage of procurement contract D1 may still be in the procurement database, but during litigation, it can appear simultaneously in the case L-evidence directory and the project P-key risk view, with the legal role having enhanced read permissions and the procurement role only retaining historical version viewing permissions. In this way, even if the business context changes, the document will not remain in a low retrieval reach state for a long time due to the fixed initial archiving path.

[0057] As an exception handling mechanism, if business pulse stream data cannot be retrieved from a certain business interface within a certain collection period, the valid state of the previous period is retained and a missing flag is added to the interface to avoid global misjudgment due to short-term disconnection; if a document cannot complete semantic parsing, a temporary low-confidence node is first established according to the metadata and retried in the next period; if the real-time association attraction of multiple candidate views is the same, the existing archived view is retained first, and then secondary sorting is performed according to the compliance priority of the business type to avoid frequent jitter switching.

[0058] In the new energy project of this equipment manufacturing group, a "Battery Cell Module Durability Test Report" was initially a research and development test document, which was then projected to the production sub-region Z1. Subsequently, due to a dispute over supplier materials that led to a claim, and the requirement to verify the consistency of the production line due to an overseas listing audit, the document was frequently cited by both the legal and audit systems. The system identified the superposition of three states—project delay, supplier dispute, and audit sampling—through the business pulse flow, enhancing the real-time correlation between the document and the legal case nodes and audit nodes, and automatically adding it to the claim evidence view and audit penetration view. At the same time, it restricted the editing permissions of irrelevant R&D personnel, retaining only the viewing permissions.

[0059] To further clarify, the preceding description of continuously approaching or exceeding the upper limit is a module-level general description used to illustrate the overall trend of regional load pressure. In actual hierarchical execution, the cumulative time during which the state weight is continuously greater than or equal to the upper limit of the business state weight, along with the error threshold, is still used as the specific criterion for whether to mark a data sub-region as low-priority. That is, approaching the upper limit itself can serve as an early warning signal to enter the map observation queue, but it does not directly replace the subsequent formal high / low priority division conditions, thereby ensuring that this embodiment and subsequent regional hierarchical implementation methods maintain consistency in triggering logic.

[0060] The purpose of this step is to integrate static document semantics and dynamic business status into a computable archiving framework, thereby enabling the synchronous evolution of document archiving location, document permissions, and business visibility, and avoiding data access isolation and data interoperability difficulties caused by fixed folder structures when business scenarios change drastically.

[0061] Furthermore, the process of the multi-source heterogeneous data acquisition module acquiring document stream data and business pulse stream data includes: deploying acquisition probes in each data sub-region, acquiring document stream data and business pulse stream data in the respective data sub-region in real time according to a preset acquisition cycle through the acquisition probes, and synchronizing the acquired document stream data and business pulse stream data to the cloud.

[0062] This embodiment provides a probe-based data acquisition mechanism for multiple data sub-regions. Specifically, in the continuous business scenario of the aforementioned equipment manufacturing group, relying solely on centralized data retrieval from the cloud can easily lead to data omissions due to inconsistent system interface protocols, network jitter, or isolated permission domains. In particular, different security levels often exist between the legal and production sub-regions, making it difficult for a single acquisition entry point to cover real-time changes. Therefore, independent acquisition probes are deployed in each data sub-region to collect data locally, preprocess it locally, and then synchronize it to the cloud.

[0063] The specific process is as follows: A first data collection probe P1 is deployed in the production sub-region Z1, connecting to the product lifecycle management system, manufacturing execution system, and quality inspection database; a second data collection probe P2 is deployed in the procurement sub-region Z2, connecting to the enterprise resource planning procurement unit, supplier portal, and contract ledger; a third data collection probe P3 is deployed in the legal sub-region Z3, connecting to the litigation case database, lawyer's letter sending and receiving interface, and evidence preservation document database; each probe operates according to the same or different data collection cycles; for example, the production sub-region collects quality inspection anomalies and test report updates every 2 minutes, the procurement sub-region collects contract status and payment approvals every 5 minutes, and the legal sub-region collects case progress and external correspondence access every 1 minute.

[0064] For ease of explanation, assume that at a certain moment, P1 collects the document stream data set as {D2 Test Report, D4 Abnormal Work Order}, and the business pulse stream data as {Project P Delay Probability 0.6, Production Line Downtime Flag 1}; P2 collects {D1 Purchase Contract, D5 Supplier Claim Letter}, and the business pulse stream data as {Payment Process Blockage, Supplier Rating Downgrade}; P3 collects {D3 Lawyer's Letter, D6 Defense Outline}, and the business pulse stream data as {Case L Enters Evidence Exchange Stage, External Audit Concerns Added}.

[0065] Each probe first adds a local timestamp and source region identifier to the collected results, and then synchronizes them to a unified cache queue in the cloud through an encrypted channel. During synchronization, a dual message structure of document stream + pulse stream can be used, where the former carries a document summary and version fingerprint, and the latter carries a status value and its effective time window, which facilitates the subsequent alignment of the map module by time.

[0066] As an anomaly tolerance mechanism, if a probe fails to collect any new documents within a preset collection period, it will still report an empty document identifier and heartbeat information to prove that there are currently no new additions in the sub-region, rather than the probe failing; if a probe loses connection with the cloud for a short time, it will cache the data packets of the most recent several periods locally and resend them in chronological order after the connection is restored; if the resent data is duplicated with existing data in the cloud, it will be deduplicated based on the document fingerprint and timestamp, retaining the latest valid version; if a sub-region cannot directly upload the document body due to permission isolation, it can also upload the de-identified features and business pulse items first, and then complete the document entity after subsequent authorization.

[0067] On the eve of the delivery of the new energy production line project, P1 in the production sub-region detected a version update of the "Battery Cell Module Durability Test Report", and at the same time, MES generated three consecutive process anomalies; P2 in the procurement sub-region collected a new supplier claim letter within the same hour; P3 in the legal sub-region synchronized with the evidence exchange notification for case L; the three probes uploaded these data to the cloud according to the sub-region dimension, enabling the system to identify the linkage changes in the R&D, procurement and legal contexts within the same business time window, rather than relying on a single system for manual summary afterward;

[0068] The purpose of this mechanism is to achieve continuous acquisition and reliable uploading of heterogeneous interfaces through regional probes, thereby realizing time alignment, regional alignment and permission alignment of document stream and business pulse stream, and providing stable input for subsequent semantic anchoring and graph prediction.

[0069] Furthermore, the process by which the base semantic anchoring module performs digital signature verification and semantic parsing on document stream data to generate basic semantic vectors includes: presetting the public and private keys of the acquisition probes corresponding to each data sub-region through an asymmetric encryption algorithm; receiving document stream data sent by the multi-source heterogeneous data acquisition module after being signed by the acquisition probe's private key; digitally signing the document stream data in each acquisition cycle of the multi-source heterogeneous data acquisition module using the private key; verifying the validity of the digitally signed document stream data sent to the cloud by the multi-source heterogeneous data acquisition module using the public key; if the validity verification of the document stream data fails, the document stream data is removed and a parsing anomaly warning signal is generated; if the validity verification of the document stream data passes, the entity features and intent features of the document stream data are extracted, and basic semantic vectors are generated based on the entity features and intent features.

[0070] Basic semantic similarity is calculated based on the Euclidean distance between basic semantic vectors. The decay gravity value is calculated based on the time series change rate of the basic semantic similarity. The decay time range is determined based on the continuous interval of the decay gravity value being greater than zero. The duration of the business state is calculated based on the timestamp difference of the document stream data. The product of the basic semantic similarity and the duration is combined with a preset normalization upper bound for normalization processing to obtain the business state weight.

[0071] This embodiment provides an anchoring mechanism that includes signature verification, semantic vector generation, and weight calculation. Specifically, while regional probe collection alone can improve coverage, it still has two drawbacks: First, data from different sub-regions may be tampered with or uploaded repeatedly, causing subsequent maps to be built on distorted data. Second, even if the document itself is successfully uploaded, if the continuous coupling relationship between the document's semantics and business status is not further quantified, the system can only identify that a document has arrived, but cannot quantify whether the document is still strongly related to the core business. Therefore, this embodiment first verifies the validity of the digital signature after the collection results enter the cloud, and then generates basic semantic vectors and calculates business status weights.

[0072] The specific process is as follows: The system pre-configures a pair of asymmetric keys for each data sub-region; taking Z1, Z2, and Z3 as examples, they correspond to key pairs K1, K2, and K3 respectively; after each collection cycle, each probe digitally signs the document stream data digest of this cycle using its own private key; after receiving the data packet, the cloud calls the corresponding public key to verify the signature; if the purchase contract D1 digest uploaded by P2 matches the signature, the semantic parsing process begins; if the lawyer's letter D3 signature uploaded by P3 does not match, it indicates that the data packet may be corrupted or have an abnormal source, and it is directly rejected, and a parsing anomaly warning signal is generated and sent to the operation and maintenance and security audit modules.

[0073] After the verification is approved, the system extracts entity features and intent features. Entity features may include project name, supplier name, amount, date, equipment model, case number, etc.; intent features may include procurement performance, quality objection, evidence preservation claim, etc.

[0074] For ease of micro-level deduction, assume that entity features identified in D1 are {supplier A, project P, amount 3 million, delivery batch B7}, and intent features are {performance, delay liability, breach of contract claim}; these are encoded into a basic semantic vector containing entity features and intent features. For D3, the identified entity features are {Case L, Patent No. X, Law Firm M}, and the intent features are {Infringement Warning, Evidence Exchange, Cease Sales}, coded as follows: ;

[0075] The system calculates basic semantic similarity based on the Euclidean distance between vectors; for example, if Vector of a certain historical claim document The Euclidean distance between them is approximately 0.076, so their basic semantic similarity can be mapped to 0.924; if the D2 test report vector and If the Euclidean distance is large, for example 0.87, the similarity is only about 0.13; the system examines the rate of change of similarity in continuous acquisition cycles;

[0076] It should be noted that, to maintain consistency with the zero-valued attenuation gravity in the embodiments, the system does not directly use the similarity change rate itself as the attenuation gravity value. Instead, it maps the magnitude of the similarity decrease to a positive attenuation gravity value. That is, when similarity increases, it is considered as semantic coupling enhancement, and the corresponding attenuation gravity value is recorded as 0; when similarity decreases, the absolute magnitude of the decrease is recorded as a positive attenuation gravity value. Schematic, this calculation method can be understood as follows:

[0077] Decaying gravity value = max(0, previous cycle basic semantic similarity - current cycle basic semantic similarity)

[0078] in The function represents taking the maximum value among the values ​​listed in parentheses; in this way, the decay gravity value will only be greater than zero when semantic alienation actually occurs, and the interval of continuous greater than zero can be directly used as the decay time range;

[0079] For example, assuming the similarity between D1 and the set of documents related to case L in the three periods t1, t2, and t3 are 0.40, 0.55, and 0.70 respectively, it indicates that the semantic coupling is continuously enhanced, and the gravity value at this stage is recorded as 0. If it then becomes 0.62 and 0.51 in t4 and t5 respectively, the decrease relative to the previous period is 0.08 and 0.11, corresponding to gravity values ​​of 0.08 and 0.11. Therefore, t4 to t5 can form a continuous decay time range. Through this approach, the upward trend in the first stage and the decay trend in the second stage can be uniformly distinguished in calculation, which avoids misjudging the enhancement process as decay and also provides a stable input for subsequent gravity decay determination.

[0080] The duration of a business status is determined by the difference in timestamps of the document stream data. Assuming that D1 spans 6 hours from its first reference in the procurement process to its reference in a legal case, and that the continuous effective association window is 4 hours, then the duration can be recorded as 4. The system multiplies the basic semantic similarity by the duration and then normalizes the result to obtain the business status weight.

[0081] To illustrate, if the similarity of 0.70 multiplied by the duration of 4 equals 2.8, and the business state weight is approximately 0.56 given the current normalization upper bound of 5; if another document has a similarity of 0.90 but is only briefly referenced within 10 minutes, its final weight may be lower than D1; this can prevent the system from over-prioritizing the archiving process due to a single instantaneous hit.

[0082] Furthermore, to make the normalization process more reproducible, the upper bound of normalization is preferably a preset reference upper limit of the basic semantic similarity × duration sample value within the current statistical time window. This reference upper limit can be a fixed empirical value, a historical high quantile value, or a preset upper bound value according to the business type. In the aforementioned example, the upper bound of 5 is used only for demonstration purposes. In actual operation, it can be configured separately according to business types such as purchase contracts, test reports, and lawyer's letters, so as to avoid different types of documents being improperly compressed by the same upper bound due to different time scales.

[0083] As an exception handling mechanism, if the signature verification fails but there is a valid version of the same document in the previous period, the valid version of the previous period will be retained for subsequent calculations, and the exception package in this period will only be used for audit evidence storage; if the entity features are missing but the intent features are complete, it is allowed to generate a dimension-reduced semantic vector first, and the missing components will be filled with default neutral values.

[0084] If the vector dimensions are found to be inconsistent when calculating the Euclidean distance, they are unified to the same dimension by using a preset projection matrix before the calculation is performed. If the normalized denominator is zero or there are no valid samples in the current time window, the business state weight is set to a preset lower limit instead of being set to zero to prevent important but sparse documents from being completely ignored by the system.

[0085] As the dispute over the new energy production line continued to escalate, the "Supplementary Procurement Agreement" uploaded by the procurement sub-region was verified and signed by the public key corresponding to Z2; the system identified features such as the redistribution of the breach of contract liability of supplier A project P, and calculated that the basic semantic similarity between the agreement and the historical claim documents increased from 0.48 to 0.74 in the past three collection cycles, lasting for 5 hours. Therefore, its business status weight was increased to a level greater than the preset high priority weight threshold.

[0086] Conversely, although a typical office supplies application also contains the word "contract" in its text, its entity features and intent features do not match the claim scenario. The distance between its vector and the core case document in the vector space is greater than the preset spatial distance threshold, and its weight remains low.

[0087] The purpose of this step is to ensure the credibility of input data through signature verification and to model the business validity of documents through semantic vectors and duration, thereby achieving the transition from parsable document content to quantifiable document status.

[0088] Furthermore, the process of the spatiotemporal service graph construction module to obtain service node prediction data for each data sub-region includes: constructing a service node prediction model containing a long short-term memory network; obtaining service pulse flow data from several historical acquisition periods for each data sub-region as training data; training the service node prediction model using the training data; obtaining the completed training service node prediction model; inputting the service pulse flow data from the current acquisition period into the completed training service node prediction model; and outputting the service node prediction data for each data sub-region containing a state weight time series from the completed training service node prediction model.

[0089] This embodiment provides a business node prediction mechanism based on long short-term memory networks. Specifically, if the graph is constructed solely based on the business pulse flow that has already occurred in the current collection period, the system can reflect the current situation, but it cannot predict in advance the congestion, migration, or permission reorganization pressure that will occur in a certain sub-region. For example, in the legal sub-region, before case L formally enters the evidence exchange stage, there are often precursors such as an increase in the frequency of lawyer's letters, an increase in the access volume of project nodes, and multiple searches of related procurement contracts. If this cannot be predicted in advance, the archiving switch may still be delayed. Therefore, this embodiment introduces a prediction model that includes long short-term memory networks to perform time-series prediction of the business node status of each sub-region.

[0090] The specific process is as follows: The system extracts business pulse flow data from several historical collection periods from each data sub-region as training data; taking Z1, Z2, and Z3 as examples, the state sequence of the most recent 90 days, with one sampling point every 5 minutes, can be selected; the input features of each time point include four items: project delay index, approval congestion, number of personnel changes, and case progress indicators.

[0091] For ease of explanation, assume that in a certain sub-region, the state vector values ​​of four consecutive sampling points are simplified as follows: At time 1, it is... The second moment is The third time is The fourth moment is This sequence indicates that project risks, approval bottlenecks, and external concerns are all accumulating and increasing; Long Short-Term Memory (LSTM) networks retain long-term trends through gating units while also taking into account short-term sudden changes, thereby extracting temporal precursor features that will lead to an increase in the weight of a certain type of node in the future.

[0092] After training, the system inputs the business pulse flow of the current collection period into the trained model and outputs the predicted business node data for each data sub-region, which includes the time sequence of state weights. For example, for the Z2 procurement sub-region, the model may output the predicted weights of the contract dispute node in the next three collection periods as 0.62, 0.77, and 0.83; for the Z3 legal sub-region, it outputs the predicted weights of the evidence catalog node as 0.71, 0.85, and 0.89. In this way, subsequent modules can pre-determine whether a region needs to be downgraded to a low-priority region and trigger collaborative scheduling before actual congestion occurs.

[0093] To enhance the readability of the spatiotemporal graph, the system can map the model output to the color, size, or heat value of the graph nodes; for example, if the heat value of node project P gradually increases over the next three periods, and the connection strength between procurement contract D1 and case L increases at the same time, it indicates that the contract has gradually evolved from a procurement transaction into key legal evidence.

[0094] As an anomaly tolerance mechanism, if the historical samples of a certain data sub-region are insufficient, for example, if a new legal division has only been running for a week, the system can use the group-level general model for migration initialization, and then fine-tune it with a small number of samples in this region; if the number of missing sampling points at a certain moment exceeds the preset missing number ratio threshold, then interpolation of the nearest time or the previous valid value is used to fill the gaps before it is fed into the prediction model; if the model output oscillates abnormally, for example, if it jumps from 0.2 to 0.95 and then falls back to 0.1 for two consecutive periods, then a smoothing constraint can be introduced, and only when the change exceeds the threshold and lasts for more than two periods is it considered a valid trend, so as to avoid the display status of the graph nodes changing frequently in adjacent periods;

[0095] A week before the equipment manufacturing group was scheduled to undergo an overseas audit, the system extracted a temporal characteristic pattern from historical data streams: whenever the project delay index, cross-departmental search frequency, and number of legal letters all increased simultaneously within 24 hours, the evidence node weight of the legal sub-region would significantly increase in the following 48 hours. In the current cycle, procurement and legal-related indicators again met this pattern, so the model predicted that the evidence node weight of Z3 would continue to rise in the next three cycles. Based on this, the system reserved processing capacity in advance for subsequent archiving and reorganization, rather than passively responding after the audit was officially launched.

[0096] The purpose of this mechanism is to extend the static current state into a predictable future state, thereby enabling the business graph to perceive in advance the upcoming trend of increasing state weights or migration pressure.

[0097] Furthermore, the process of dividing the data sub-region into high-quality data sub-regions or low-quality data sub-regions includes: obtaining the preset upper limit of the business state weight for each data sub-region, extracting the time series sequence of the state weight in the business node prediction data of each data sub-region within the current collection period, comparing the time series sequence of the state weight of each data sub-region with the preset upper limit of the business state weight, and obtaining the cumulative time during which the state weight is continuously greater than or equal to the preset upper limit of the business state weight.

[0098] A preset cumulative time error threshold is set. If the cumulative time of a data sub-region is greater than or equal to the cumulative time error threshold, the data sub-region is marked as a low-quality data sub-region. If the cumulative time of a data sub-region is less than the error threshold, the data sub-region is marked as a high-quality data sub-region.

[0099] This embodiment provides a regional classification mechanism based on the duration of continuous over-limit of predicted weights. Specifically, classifying sub-regions solely based on the level of a single prediction result is easily affected by instantaneous peaks. For example, the legal sub-region may experience a temporary increase due to a batch import, but this does not mean that the region is under continuous pressure. Conversely, if the procurement sub-region remains at a critically high level for an extended period, it is more likely to cause delays in document archiving and view refresh. Therefore, this embodiment does not classify sub-regions as high-priority or low-priority based on the absolute value at a certain moment, but rather on the cumulative time during which the state weight continuously reaches or exceeds the upper limit.

[0100] The specific process is as follows: The system presets a maximum upper limit for the business status weight of each data sub-region; schematically, Z1 is set to 0.85, Z2 to 0.80, and Z3 to 0.88; the system extracts the time series sequence of the status weight corresponding to the current collection period of each sub-region; for example, the prediction sequence of the next five sampling points for Z2 is... Since its upper limit is 0.80, the time consecutively greater than or equal to the upper limit covers the 2nd to 4th sampling points; if each sampling point represents 5 minutes, the cumulative time is 15 minutes; assuming the error threshold is set to 10 minutes, Z2 is marked as the low-quality data sub-region; for example, the sequence of Z1 is... If the error threshold is not reached even if the sampling time exceeds the limit by 5 minutes at the second sampling point, it will still be marked as a high-quality data sub-region.

[0101] Here, "high priority" and "low priority" do not indicate the importance of the business, but rather whether the area is currently capable of handling collaboration from other areas or whether it is nearing its capacity limit. Generally, a high priority area means there is still processing capacity, while a low priority area means that it needs to be diverted, reorganized, or compensated.

[0102] Using cumulative time instead of single-point values ​​has another advantage: it can tolerate prediction errors within a reasonable range. For example, if a model outputs fluctuations between 0.79 and 0.80, the region label will frequently flip if judged directly by a single-point value. By setting an error threshold, as long as the error exceeds the limit for a set duration, the region-level switching will not be triggered.

[0103] As an anomaly-tolerant fault-handling mechanism, if a sub-region lacks an effective state weight time series, such as if the prediction fails in this period, the region label of the previous period is temporarily used and recorded as a low-confidence flag; if multiple sub-regions simultaneously reach the low-optimal condition, the subsequent collaborative scheduling phase can prioritize comparing their over-limit magnitude and duration to determine which one gets the resources first; if the cumulative time is exactly equal to the error threshold, it is treated as reaching the low-optimal condition to avoid repeated jumps in the critical point.

[0104] On the afternoon of the third day of the audit preparation week, the predicted weight of contract dispute nodes in the procurement sub-region Z2 was above 0.80 for 15 consecutive minutes, while the weight of the legal sub-region Z3, although reaching 0.89 at one point, only lasted for 5 minutes. As a result, the system marked Z2 as a low-priority data sub-region and kept Z3 as a high-priority data sub-region. This means that the virtual archiving and view management on the procurement side will soon be under pressure, while the legal side will still be able to handle some of the graph nodes related to the case.

[0105] The purpose of this step is to replace instantaneous over-limit judgment with continuous over-limit time, thereby achieving a more stable regional carrying capacity assessment and establishing clear triggering conditions for subsequent map-based collaborative scheduling.

[0106] Furthermore, the process of performing business graph collaborative scheduling on low-quality data sub-regions includes:

[0107] Step 1: Obtain the time period in which the state weight of the low-quality data sub-region exceeds the preset upper limit of the business state weight, and the maximum value of the excess weight within the time period. Obtain the preset network topology containing each data sub-region. Obtain the high-quality data sub-region that is closest to the low-quality data sub-region in the preset network topology. Obtain the minimum idle weight of the high-quality data sub-region within the time period. Determine whether the minimum idle weight is greater than or equal to the maximum value of the excess weight. If it is greater than or equal to, allocate the graph node corresponding to the minimum idle weight of the high-quality data sub-region to the low-quality data sub-region and end the business graph collaborative scheduling. If it is less than, allocate the graph node corresponding to the minimum idle weight of the high-quality data sub-region to the low-quality data sub-region, remove the high-quality data sub-region, and proceed to Step 2.

[0108] Step 2: Obtain the maximum value of the weight exceeding the low-quality data sub-region within the time period after the graph node allocation, re-obtain the high-quality data sub-region that is closest to the low-quality data sub-region in the preset network topology, obtain the minimum value of the idle weight of the re-obtained high-quality data sub-region within the time period, and then execute Step 3.

[0109] Step 3: Determine whether the minimum idle weight of the newly acquired high-priority data sub-region is greater than or equal to the maximum excess weight of the low-priority data sub-region within the time period. If it is greater than or equal to, allocate the graph node corresponding to the minimum idle weight of the newly acquired high-priority data sub-region to the low-priority data sub-region and end the business graph collaborative scheduling. If it is less than, allocate the graph node corresponding to the minimum idle weight of the newly acquired high-priority data sub-region to the low-priority data sub-region, remove the newly acquired high-priority data sub-region, and execute Step 2.

[0110] This embodiment provides a graph collaborative scheduling mechanism for low-priority data sub-regions. Specifically, after completing the regional classification, if only low-priority regions are labeled without resource reallocation, the classification itself cannot solve the actual congestion problem, and the procurement sub-regions may still cause the relevant document view to be delayed due to continuous over-limit. Therefore, this embodiment further performs stepwise allocation of graph nodes based on network topology proximity and idle weight of high-priority regions.

[0111] The specific process is as follows: Suppose the current procurement sub-region Z2 is marked as low priority, with a preset upper limit of 0.80. If the predicted weight peak reaches 0.93 within a certain time period T, then the maximum weight exceeds the limit of 0.13. In the preset network topology, Z2 is 1 hop away from Z1 and 2 hops away from Z3. Therefore, the nearest high priority region Z1 is selected first. Then, the minimum idle weight of Z1 within the time period T is obtained. Assuming that the minimum idle weight of Z1 within T is 0.08, it is insufficient to completely cover the 0.13 excess load of Z2. At this time, instead of not allocating, the corresponding graph node with 0.08 is first collaboratively mapped from Z1 to Z2, Z1 is removed, and the search continues for the next nearest high priority region.

[0112] It needs to be further explained here that, in the embodiment, the graph node corresponding to the minimum idle weight of the high-priority data sub-region is assigned to the low-priority data sub-region. In this embodiment, it is preferred to understand that the graph computing capabilities, index maintenance capabilities, or view mounting capabilities that were originally available to the high-priority data sub-region are assigned to the low-priority data sub-region for use in the form of collaborative service mapping, rather than reassigning the core business entities that are already busy in the high-priority region to the low-priority region.

[0113] In other words, the assigned graph node processing share, calculation quota, or maintenance task entry is matched with the idle weight. The result is that the low-priority area can directly obtain external support. Therefore, although it is described as being allocated to the low-priority data sub-region, it actually reflects the output carrying capacity of the high-priority area to the low-priority area.

[0114] The allocation of graph nodes can be understood as transferring some document association calculations, view mounting, or index maintenance tasks to idle areas. For ease of explanation, assume that the three busiest graph nodes in Z2 are N21 supplier dispute, N22 supplementary agreement, and N23 payment freeze. If Z1 can only handle 0.08, the lighter parts of N22 and N23 can be transferred first. After allocation, the remaining excess weight of Z2 becomes 0.05. The nearest high-priority area is re-acquired, and at this time, only Z3 can be selected. If the minimum idle weight of Z3 in T is 0.07, which is greater than the remaining 0.05, then the remaining corresponding nodes, such as N21 and its adjacent evidence relationships, are handed over to Z3 for collaborative handling, and the scheduling ends.

[0115] Furthermore, the corresponding graph nodes are preferably sorted from low to high according to their maintenance load and then used sequentially until the cumulative load is close to but does not exceed the minimum idle weight. If a node cannot be split, the smallest complete relation cluster to which the node belongs is used as an allocation unit. This not only corresponds to the quantified result of the idle weight, but also avoids the situation where there is an idle weight in the scope, but no transferable unit can be matched in reality.

[0116] If the available weight of the high-priority area obtained the second time is still insufficient, for example, only 0.03, then it is also partially allocated first, then the area is removed and the steps are repeated until the remaining excess is covered or there are no available high-priority areas. This constitutes a scheduling path of gradual stripping and gradual takeover, rather than a direct allocation of the whole in a single time.

[0117] The minimum idle weight is used here instead of the average idle weight to ensure continuous availability throughout the target time period, rather than having slack only at local moments; otherwise, secondary congestion may occur just after the migration is completed due to the instantaneous busyness of the target area.

[0118] As an anomaly tolerance mechanism, if no high-priority data sub-region exists, the system will not perform cross-region scheduling. Instead, tasks in low-priority regions will be arranged in descending order of priority, with real-time updates reserved only for nodes with high compliance risk and high retrieval frequency, while the remaining nodes will be updated with delays. If multiple high-priority regions have the same network distance, the region with the larger minimum idle weight will be selected first. If a graph node cannot be divided, for example, if a node involves a complete chain of evidence in a case, it will be migrated according to the smallest complete unit, and it is prohibited to split it into multiple discontinuous segments to avoid destroying the context of the evidence chain.

[0119] When a collective claim by suppliers occurred in the group's new energy project, the procurement sub-region Z2 continuously exceeded the limit for two hours, with a peak exceeding the limit by 0.13. The system found that the production sub-region Z1 was the closest to it in terms of network distance and could provide an idle weight of 0.08. Therefore, the graph maintenance task of payment freeze and supplementary agreement version index was first transferred to Z1. The remaining 0.05 was then taken over by the legal sub-region Z3, which was responsible for the associated maintenance of the claim claim - case L. In this way, the graph update pressure that was originally concentrated in the procurement domain was dispersed and distributed, and the core contract approval view of Z2 was able to return to stability.

[0120] The purpose of this mechanism is to release the carrying pressure of low-priority areas through a cooperative scheduling method that prioritizes proximity and allocates resources gradually, while reducing the latency and semantic breaks caused by cross-domain migration.

[0121] Furthermore, the semantic context gravity coupling calculation module performs gravity attenuation determination on business nodes in each data sub-region, and performs lightweight gravity compensation operation or adaptive collaborative gravity compensation operation based on the determination result. This process includes: extracting the attenuation gravity value, attenuation time range, basic semantic similarity, duration of business state, and business state weight of each basic semantic vector in the business nodes generated in real time within the data sub-region.

[0122] Determine whether the decay time range of the basic semantic vector is adjacent to or overlaps with the duration of the business state. If the decay time range is adjacent to or overlaps with the duration of the business state, the ratio of the decay gravity value to the basic semantic similarity is used as the comprehensive gravity ratio of the basic semantic vector. A preset comprehensive gravity ratio threshold is set. If the comprehensive gravity ratio is less than the comprehensive gravity ratio threshold, the basic semantic vector is marked as a gravity decay node. If the comprehensive gravity ratio is greater than or equal to the comprehensive gravity ratio threshold, the basic semantic vector is marked as a normal node.

[0123] Obtain the basic semantic vector and the number and location features of gravity decay nodes within the data sub-region, divide the data sub-region into several virtual view units of the same size, and obtain the gravity decay density of each virtual view unit based on the ratio of the number of gravity decay nodes to the total capacity of the topological nodes of the virtual view unit.

[0124] If the gravitational attenuation density of the virtual view unit is less than or equal to the density threshold and there are gravitational attenuation nodes within the virtual view unit, then a lightweight gravitational compensation operation is performed on the gravitational attenuation nodes within the virtual view unit; if the gravitational attenuation density of the virtual view unit is less than or equal to the density threshold and there are no gravitational attenuation nodes within the virtual view unit, then no compensation operation is performed; if the gravitational attenuation density of the virtual view unit is greater than the density threshold, then an adaptive cooperative gravitational compensation operation is performed on the area covered by the virtual view unit.

[0125] This embodiment provides a gravitational coupling mechanism that provides hierarchical compensation based on attenuation intensity and spatial density. Specifically, while completing map-based collaborative scheduling can alleviate regional congestion, it cannot automatically identify which documents have begun to deviate from their original context. For example, some procurement agreements may have gradually been transformed into legal evidence, while others still belong to the normal procurement process. If individual attenuation and regional attenuation are not distinguished, over-migration or insufficient compensation may easily occur. Therefore, this embodiment introduces a gravitational attenuation determination and selects lightweight compensation or adaptive collaborative compensation based on the attenuation density in the virtual view unit.

[0126] The specific process is as follows: the system extracts multiple quantities from the real-time generated business nodes: the decay gravity value of the basic semantic vector, the decay time range, the basic semantic similarity, the duration of the business state, and the business state weight; taking document D1 as an example, assuming its basic semantic similarity is 0.72, the decay gravity value is 0.18, the decay time range is from 14:00 to 14:20, and the duration of the business state is from 13:50 to 14:30;

[0127] Since the two are adjacent and overlap, the combined gravity ratio can be calculated as 0.18 / 0.72=0.25. If the combined gravity ratio threshold is set to 0.30, then D1 is marked as a gravity decay node. For example, D5 is a supplier claim letter with a similarity of 0.80, a gravity decay value of 0.30, and a ratio of 0.375, which is greater than the threshold, so it is retained as a normal node.

[0128] It should be further clarified here that the aforementioned gravity decay nodes are used to identify objects that need to enter the compensation link, and are not equivalent to all objects that have already undergone context migration; since the comprehensive gravity ratio is based on the premise that the decay time range and the duration of the business state are adjacent or overlapped, and it is subsequently combined with the real-time associated gravity and fluid archiving execution module for remounting, the comprehensive gravity ratio threshold in this embodiment is preferably used to screen out nodes that are still in the early compensable window:

[0129] When the ratio is below the threshold, it indicates that although the attenuation of the node has not yet increased, it has already shown anchoring loosening within the effective business continuity range, making it suitable as a gravity attenuation node for compensation. When the ratio is greater than or equal to the threshold, it is preferred to consider that the node has formed a relatively clear new contextual trend or remains stable, and is no longer the priority object for compensation in this step, but is handed over to the subsequent real-time correlation gravity calculation and fluid archiving execution steps for further processing. Thus, the judgment criteria of D1 and D5 in the aforementioned numerical example are consistent with the threshold rules in the embodiment, and do not conflict with the subsequent archiving switching process.

[0130] After completing the node-level determination, the system divides the data sub-region into several virtual view units of the same size based on the location characteristics; the location characteristics here can be understood as the coordinates of the document in the virtual archive space, such as the cells after three-dimensional projection according to the business domain—project domain—compliance domain;

[0131] For ease of explanation, assume Z2 is divided into four view units G1, G2, G3, and G4, with the total capacity of topological nodes in each unit uniformly set to a preset baseline value, such as 100 nodes. If G1 has 5 basic semantic vectors, and 2 of them are identified as gravity decay nodes, then its gravity decay density is 2 / 100 = 0.02. If the density threshold is set to 0.01, then G1 triggers adaptive collaborative compensation. Conversely, if G2 has 4 nodes, only 1 of which is a decay node, and the density is 1, then lightweight compensation is performed. If G3 has no decay nodes, then no compensation is performed.

[0132] This design, which first identifies individual cases and then stratifies regions, can distinguish between scattered anomalies and group anomalies. Scattered anomalies are suitable for local migration path correction, while group anomalies indicate an overall change in the business context of the region, requiring stronger collaborative adjustments.

[0133] As an anomaly tolerance mechanism, if the decay time range of a node is neither adjacent to nor overlaps with the duration of the business state, then the node will not participate in the comprehensive gravity ratio calculation and can be directly regarded as a normal node to be observed; if the number of nodes in a virtual view unit is less than the preset minimum sample number threshold, for example, only 1 node, then even if the density value is greater than the density threshold, the minimum sample condition can be increased to prevent a single sample from triggering large-scale compensation; if multiple units exceed the density threshold at the same time, then units that are close to the center of the low-quality sub-region and involve compliant business types will be processed first.

[0134] In the afternoon following the escalation of the dispute over the new energy project, contracts, supplementary agreements, and payment applications related to supplier A within the procurement sub-region were mapped to multiple virtual view units. Among them, G1, covering the procurement performance-dispute handling area, showed two consecutive gravity decay nodes, indicating that the document was rapidly shifting from a procurement context to a legal context. G2, on the other hand, only had one ordinary supplementary agreement with slight decay, which was still suitable for local correction. Therefore, the system performed stronger adaptive collaborative compensation on G1 and only lightweight compensation on G2.

[0135] The purpose of this step is to calculably stratify the degree to which documents are disconnected from the business context, thereby enabling the precise allocation of compensation resources and avoiding unnecessary disturbances to normal nodes.

[0136] Furthermore, the process of performing lightweight gravity compensation operation includes: obtaining several virtual directory paths in the preset archive link corresponding to the gravity decay node, extracting virtual directory paths that are not adjacent to or overlap with the decay time range, and marking the virtual directory paths that are not adjacent to or overlap with the decay time range as virtual directory paths to be allocated.

[0137] The search matching rate of each virtual directory path to be assigned is calculated based on the ratio of the historical search hits to the total search count. The virtual directory path with the lowest search matching rate is selected, the business status data is migrated to the selected virtual directory path, and the corresponding business status weight parameters are updated according to the attributes of the virtual directory path to be assigned.

[0138] This embodiment provides a lightweight gravity compensation mechanism for locally decayed nodes. Specifically, in the previous layer judgment, if a virtual view unit has only a few decayed nodes, directly starting a large-scale collaborative compensation would cause unnecessary resource consumption and may also cause the originally stable view structure to be over-rearranged. Therefore, this embodiment only reselects a more suitable virtual directory path in its preset archive link for a single or a few decayed nodes.

[0139] The specific process is as follows: The system first obtains the existing preset archive links of the gravity decay node; taking D1 "Supplementary Procurement Agreement" as an example, its original virtual directory path may include: Path R1: Procurement Center / Project P / Framework Contract; Path R2: Supplier A / Performance Tracking / Supplementary Agreement; Path R3: Legal Warning / Dispute Materials / Pending Confirmation; Assuming that the decay time range of D1 is 14:00 to 14:20, the system will remove paths that are adjacent to or overlap with this time range; For example, R3 was just created by the legal department at 14:05, which is a sensitive path that is adjacent in time and is not considered as a candidate for lightweight migration; R1 and R2 are not adjacent to or overlap with this time range, and are marked as virtual directory paths to be allocated;

[0140] The system calculates the search matching rate based on the ratio of historical search hits to total search counts for each path to be assigned. For example, if R1 was searched 100 times in the past 30 days and related documents were opened 80 times, its matching rate is 0.80; if R2 was searched 50 times and hit 15 times, its matching rate is 0.30. The system selects the path with the lowest matching rate, R2, as the migration target. The selection of the lowest rate is not due to its low importance, but rather because it indicates that the path currently does not closely match user search needs, making it suitable as an adjustment point for re-attaching business status data. Improving its semantic orientation by migrating in new status data is crucial.

[0141] After the business status data is migrated, the system updates the corresponding business status weight parameters according to the attributes of the path. For example, if R2 has the attribute of supplier performance dispute, the business status weight of D1 under this path will be increased from 0.56 to 0.68, while its regular procurement weight under R1 will be weakened. In this way, when users search from the supplier dimension, it is easier to hit the document that has undergone context shift, without having to search through layers of the traditional procurement catalog.

[0142] As an anomaly tolerance mechanism, if the virtual directory path to be assigned is empty, it means that the original links are highly coupled with the current decay time range. In this case, lightweight migration is not performed, but the node is upgraded and submitted to the collaborative compensation candidate set. If the retrieval matching rate of multiple paths to be assigned is the same, the path that is closer to the current business type attribute is selected first. If the total number of historical retrievals of a certain path is zero, its matching rate can be regarded as a preset low value, but migration is only allowed when the path attribute passes the compliance verification, so as to avoid mounting critical documents to experimental directories that have never been officially used.

[0143] When the group was processing a claim from supplier A, D1 was originally listed under both the framework contract and supplier performance tracking views for project P. The system detected that its attractiveness in the procurement performance chain was beginning to decline, but had not yet formed a regional decline. Therefore, instead of a large-scale reorganization, the system compared the historical search performance of the two directories. Since the supplier performance tracking path had a long-term low hit rate, the system moved the business status data related to the dispute into this path and increased its dispute attribute weight, so that procurement managers and legal liaisons could obtain the agreement more quickly when searching from the supplier's perspective.

[0144] The purpose of this step is to correct the archive offset of local nodes with minimal adjustment cost, thereby achieving lightweight, continuous and low-disturbance virtual directory compensation.

[0145] Furthermore, the process of performing adaptive cooperative gravity compensation operation includes: extracting the flow path corresponding to the gravity attenuation node in the virtual view unit coverage area, obtaining the data sub-region that emits the gravity attenuation node according to the flow path, and marking the data sub-region as the attenuation source;

[0146] An adaptive gravity compensation model is constructed based on a graph convolutional neural network. The decay gravity value and decay time range of the decay source at the current moment, as well as the topological structure graph corresponding to the coverage area of ​​the virtual view unit, are input into the adaptive gravity compensation model. The model extracts node features and graph structure features and outputs the optimal gravity adjustment parameters of the decay source.

[0147] This embodiment provides an adaptive collaborative gravity compensation mechanism for high-density decay regions. Specifically, when a large number of decay nodes appear within a virtual view unit, simply migrating individual directory paths using lightweight methods is insufficient to restore the coupling relationship between documents and business context. At this point, the problem is no longer the deviation of the archiving position of a single document node, but rather the rearrangement of the entire business relationship network in the region. For example, if procurement, legal, and project management parties interact frequently around the same batch of contracts and test reports, it indicates that the original gravity parameters have been distorted. Therefore, this embodiment introduces an adaptive compensation model based on graph convolutional neural networks.

[0148] The specific process is as follows: The system first extracts the flow paths corresponding to each gravity attenuation node in the covered area; taking the high-density unit G1 as an example, assuming that the three attenuation nodes correspond to the following flow paths: L1: Z2 procurement sub-region → Z1 production sub-region → Z3 legal sub-region; L2: Z2 procurement sub-region → Z3 legal sub-region; L3: Z1 production sub-region → Z2 procurement sub-region → Z3 legal sub-region; if most attenuation nodes originate from Z2, then Z2 is marked as the attenuation source; the system obtains the attenuation gravity value and attenuation time range of Z2 at the current time, for example, the current attenuation gravity value set is {0.18, 0.21, 0.25}, and the attenuation time range covers 13:40 to 14:30;

[0149] The system constructs a topology diagram of the coverage area. For easier microscopic explanation, nodes can be simplified to N1 (purchase contract), N2 (supplementary agreement), N3 (test report), N4 (claim letter), and N5 (case node), with edges representing relationships such as citation, approval, and evidence association. If N1 is strongly correlated with N4 and N5, N2 is correlated with N1 and N5, and N3 is moderately correlated with N5, an adjacency graph can be formed. The graph convolutional neural network receives two types of inputs: one is the decay features of the node layer, such as the decay gravity value, decay time range encoding, and business state weight of each node; the other is the connection relationship of the graph structure layer. The model aggregates information from adjacent nodes through graph convolution, thereby outputting the optimal gravity adjustment parameters for the decay source.

[0150] The optimal gravity adjustment parameter here can be understood as a suggestion for increasing or decreasing gravity for nodes related to the attenuation source; for example, the model may output: increase the gravity coefficient of the contract node issued by Z2 to the legal case node by 0.12, decrease its gravity coefficient to the ordinary procurement approval node by 0.08, and add a temporary reinforcement edge between the test report and case nodes; the system recalculates the real-time correlation gravity and subsequent archiving projection based on this parameter;

[0151] As an anomaly-tolerant mechanism, if the topology graph is too sparse, making it difficult for graph convolution to form effective neighborhood aggregation, it can degenerate into a first-order compensation mode based on direct adjacency. If the attenuation source is not unique, for example, Z1 and Z2 simultaneously constitute the main source, the adjustment parameters of the two attenuation sources are calculated separately, and then weighted and synthesized according to the proportion of nodes in the coverage area. If the model output parameters exceed the preset safety range, for example, the single increase exceeds the preset increment limit, the system should truncate according to the upper limit to prevent the view structure adjustment range from exceeding the preset safety ratio limit, thereby affecting the continuity of user experience.

[0152] After the group's supplier claim escalated into formal litigation, multiple contracts, supplementary agreements, and correspondence in the G1 procurement area simultaneously experienced high-density attenuation. The system tracked the flow path of these documents, and the tracking results showed that most of them originated from Z2 and eventually converged at the legal case node. Therefore, Z2 was marked as the attenuation source. After combining the topological relationship of the contract-claim letter-case node with the graph convolution model, a set of adjustment parameters were output to enhance the overall attractiveness of these documents to the legal view and moderately reduce their attractiveness to the ordinary procurement view. The visibility of these documents in the evidence view was significantly improved, while their weight in the daily procurement catalog gradually decreased.

[0153] The purpose of this step is to provide structural compensation for regional contextual shifts, thereby achieving an integrated gravity reassessment across nodes and paths, rather than just a partial directory reconstruction.

[0154] Furthermore, the process by which the fluid archiving execution module performs virtual view archiving and permission mapping switching operations based on the real-time correlation attraction, business type, and flow status of business nodes includes: extracting the real-time correlation attraction, business type, and flow status of each business node from the real-time collected data within the data sub-region;

[0155] The configuration is as follows: When the business node's flow status is far from the data sub-region and the real-time correlation attraction is less than the preset correlation attraction threshold corresponding to the business type, extract the preset usage data volume and real-time correlation attraction corresponding to the business type, obtain the idle data volume valley value and real-time correlation attraction valley value of other data sub-regions in the current collection period, obtain the target data sub-region where the idle data volume valley value is greater than the preset usage data volume and the real-time correlation attraction valley value is greater than the real-time correlation attraction, calculate the Euclidean distance between the target data sub-region and the business node, and select the target data sub-region with the shortest Euclidean distance to perform virtual view archiving and permission mapping switching operations.

[0156] When the business node's flow status is close to the data sub-region or the real-time association attraction is greater than or equal to the preset association attraction threshold corresponding to the business type, the current virtual view archiving and permission mapping status remains unchanged.

[0157] This embodiment provides a fluid archiving execution mechanism based on real-time correlation gravity, business type, and flow status jointly triggered. Specifically, the aforementioned implementations have been able to identify attenuation, predict trends, and provide gravity adjustment parameters. However, if they are not implemented in the final virtual view archiving and permission mapping switching, the system remains at the analysis layer and cannot truly change the user's search entry and access boundaries. Therefore, this embodiment converts the analysis results into specific archiving location migration and permission status switching actions.

[0158] The specific process is as follows: The system extracts three core quantities for each business node from the real-time data collected in each data sub-region: real-time correlation attraction, business type, and flow status. The business type can include purchase contracts, test reports, lawyer's letters, payment applications, audit working papers, etc. The flow status describes whether the node is close to or far from the current data sub-region. When the frequency of a business node being called by other data sub-regions in the current collection period exceeds the preset cross-domain reference threshold, and the frequency of access within its original data sub-region shows a month-on-month decreasing trend, the flow status is determined to be far away. Conversely, if the frequency of access within the original data sub-region is dominant or does not reach the cross-domain reference threshold, it is determined to be close.

[0159] Taking procurement contract D1 as an example, assuming it is currently located in Z2, its business type is procurement contract, and the preset association attraction threshold for this type is 0.65; after compensation, the real-time association attraction of D1 in Z2 drops to 0.48, and its main source of reference in the last three collection cycles has gradually shifted from procurement approval to legal case view, so it can be determined that its flow status is far away from Z2; at this time, the system extracts the preset usage data volume corresponding to the procurement contract type, such as the required archive cache space of 20 units, and checks the valley value of idle data volume and real-time association attraction value of other data sub-regions in the current collection cycle;

[0160] Here, the valley value of idle data volume is preferably understood as the minimum remaining capacity that the target data sub-region can stably provide within the statistical window of the current collection cycle; the valley value of real-time correlation gravity is preferably understood as the minimum value among a set of correlation gravitations related to the candidate mounting view of the service node within the target data sub-region; therefore, only when a region's minimum remaining capacity is greater than the preset data usage volume, and its minimum candidate correlation gravity is higher than the current gravity of the service node, can it be said that the region has stable acceptance conditions throughout the entire cycle, rather than only temporarily meeting the requirements at individual moments; assuming that the valley value of idle data volume of Z1 is 18, which does not meet 20; the valley value of idle data volume of Z3 is 25, and its corresponding correlation gravity valley value is 0.55, which is higher than the current 0.48 of D1, so Z3 enters the candidate list;

[0161] The system calculates the Euclidean distance between the target data sub-region and the business node. This distance can be based on business semantic coordinates or a combination of network topology and graph location. To avoid the problem of the target region being filtered but the meaning of the distance being unclear, this embodiment preferably establishes the Euclidean distance on a unified business coordinate space, which includes at least three dimensions: business type features, current associated context features, and flow path features.

[0162] Specifically, business type features are mapped to fixed-dimensional one-hot encoded vectors through a pre-defined business enumeration dictionary; current context features directly reuse the dimensionality-reduced basic semantic vectors; flow path features are transformed into path penalty weight vectors through the number of hops across systems at nodes; each candidate target region is represented in this space by a concatenated vector formed by normalizing the above three types of vectors corresponding to its candidate view center point; and business nodes are represented by a concatenated vector formed by normalizing the three types of vectors corresponding to their real-time semantic positions.

[0163] For demonstration purposes, if the distance between D1 and Z3 in the multi-dimensional business coordinates is 0.22, and the distance to another candidate area Z4 is 0.35, where Z4 is a newly added audit data sub-area connected to the cloud in addition to Z1, Z2, and Z3, then Z3 will be selected first. The system performs virtual view archiving and permission mapping switching: the main display level of D1 in the Procurement Center / Project P / Framework Contract is reduced, while the main display level in Case L / Evidence Catalog / Supplier A is increased; at the same time, the legal role's permissions are upgraded from read-only to annotateable and indexable, the procurement role is changed from editable to read-only historical version, and the project management role retains viewing and referencing permissions.

[0164] Furthermore, the preferred order for permission mapping switching is to switch views first and then verify permissions. That is, new virtual view mounting candidates are first generated, and then compared with the enterprise's existing job permission matrix, confidentiality level matrix, and project participation relationship matrix item by item. Only roles that pass the verification are granted the corresponding operation level. If a role should have access to the document due to business flow, but its static confidentiality level does not meet the requirements, the system is only allowed to display the existence of the index within its visible range, and the ability to access the main text or annotate is not granted. Through this order, the fluid archiving action can be compatible with compliance control, and there will be no unauthorized access due to view migration.

[0165] To further clarify, the Euclidean distance used for calculating the similarity of basic semantic vectors mentioned above, and the Euclidean distance used for filtering target data sub-regions in this embodiment, both refer to the same type of Euclidean distance metric. They only correspond to vector space and business coordinate space respectively due to the different calculation objects. To ensure consistent terminology throughout the document, both are treated as the same distance paradigm in the interpretation of the specification: the former is used for distance calculation between basic semantic vectors, and the latter is used for distance calculation between business nodes and candidate data sub-regions in a unified business coordinate space, and do not represent two different distance algorithms.

[0166] Conversely, if a test report D2 is temporarily referenced by the legal department, but its circulation status is still close to Z1, or its real-time association attraction in Z1 is still higher than the corresponding threshold, such as 0.78 ≥ 0.70, then the current virtual view archive and permission mapping status remain unchanged, and only auxiliary links are added at the association recommendation level without changing the main archive position; this can avoid frequent movement of document views due to short-term cross-department access.

[0167] As an exception handling mechanism, if the migration conditions are met but no other data sub-regions simultaneously meet the conditions of free data volume and gravitational valley value, the system will not perform a switch, but will retain the original view and add a temporary recommended entry; if multiple target regions have the same Euclidean distance, the region with the higher compliance level will be selected first; if the permission switch would conflict with the company's existing confidentiality policy, such as ordinary procurement personnel not being able to see the legal evidence chain, the stricter permission rules will prevail, and only the view mount will be switched without expanding the accessible scope;

[0168] During the overlapping phase of the group's preparations for dealing with overseas audits and supplier lawsuits, the Supplementary Procurement Agreement was originally stably classified under the procurement sub-region. As the legal case progressed, its real-time correlation gravity in the procurement domain gradually decreased to 0.48, and the main flow path was significantly away from Z2. The system detected that the legal sub-region Z3 not only had sufficient free capacity, but also had the shortest semantic distance from the case of the agreement. Therefore, it automatically switched its main archive view to the case L evidence directory and adjusted the permissions accordingly, allowing the legal team to directly mark evidence points, while the procurement team retained read-only traceability capabilities.

[0169] Meanwhile, a "Battery Cell Module Durability Test Report" still mainly serves production rectification. Although it is called by the auditing party, its attraction in Z1 is still higher than the threshold. Therefore, it maintains its original archiving and permission status and only forms a mirror entry in the audit view.

[0170] The purpose of this step is to translate the results of the preceding calculations into perceptible archiving and permission actions, thereby enabling the document view to switch automatically as business processes change, while maintaining compliance boundaries and access stability.

[0171] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. An AI-based automatic document archiving system for enterprise management based on semantic analysis, characterized in that: Including the cloud, the cloud communication connection includes a multi-source heterogeneous data acquisition module, a base semantic anchoring module, a spatiotemporal service graph construction module, a semantic context gravity coupling calculation module, and a fluid archiving execution module; The multi-source heterogeneous data acquisition module is used to acquire the location characteristics and flow range of several business interface nodes and data source nodes in the preset enterprise management system. Based on the location characteristics and flow range, the preset system is divided into several data sub-regions. According to the preset acquisition cycle, document flow data and business pulse flow data reflecting the real-time operating status of the business system are acquired from the business interface nodes of each data sub-region at regular intervals. The base semantic anchoring module is used to perform digital signature verification and semantic parsing on document stream data to generate basic semantic vectors, and to calculate business state weights based on the basic semantic vectors. The spatiotemporal business graph construction module is used to obtain business node prediction data for each data sub-region based on business pulse flow data. According to the business node prediction data of each data sub-region and the preset business state weight upper limit, the data sub-region is divided into high-quality data sub-regions or low-quality data sub-regions, and business graph collaborative scheduling is performed on the low-quality data sub-regions. The semantic context gravity coupling calculation module is used to determine the gravity attenuation of business nodes in each data sub-region, calculate the real-time association gravity of business nodes based on the similarity between basic semantic vectors, and perform lightweight gravity compensation operation or adaptive collaborative gravity compensation operation based on the determination result. The fluid archiving execution module is used to extract the business type and flow status of business nodes from document stream data and business pulse stream data, and to perform virtual view archiving and permission mapping switching operations based on the real-time association attraction, business type and flow status of business nodes.

2. The AI-based enterprise management document automatic archiving system based on semantic analysis according to claim 1, characterized in that, The process by which the multi-source heterogeneous data acquisition module acquires document stream data and business pulse stream data includes: Within each data sub-region, a collection probe is deployed. The collection probes collect document stream data and business pulse stream data within their respective data sub-regions in real time according to a preset collection cycle, and then synchronize the collected document stream data and business pulse stream data to the cloud.

3. The AI-based enterprise management document automatic archiving system based on semantic analysis according to claim 2, characterized in that, The process by which the base semantic anchoring module performs digital signature verification and semantic parsing on document stream data to generate basic semantic vectors includes: The system uses an asymmetric encryption algorithm to pre-define the public and private keys of the acquisition probes corresponding to each data sub-region. It receives document stream data sent by multi-source heterogeneous data acquisition modules and has been signed by the private key of the acquisition probe. It digitally signs the document stream data in each acquisition cycle of the acquisition probe using the private key and verifies the validity of the digitally signed document stream data sent to the cloud by the acquisition probe using the public key. If the validity verification of the document stream data fails, the document stream data is removed and a parsing anomaly warning signal is generated. If the validity verification of the document stream data passes, the entity features and intent features of the document stream data are extracted, and a basic semantic vector is generated based on the entity features and intent features. Basic semantic similarity is calculated based on the Euclidean distance between basic semantic vectors. The decay gravity value is calculated based on the time series change rate of the basic semantic similarity. The decay time range is determined based on the continuous interval of the decay gravity value being greater than zero. The duration of the business state is calculated based on the timestamp difference of the document stream data. The product of the basic semantic similarity and the duration is combined with a preset normalization upper bound for normalization processing to obtain the business state weight.

4. The AI-based enterprise management document automatic archiving system based on semantic analysis according to claim 3, characterized in that, The process by which the spatiotemporal business graph construction module obtains the business node prediction data for each data sub-region includes: A business node prediction model incorporating a long short-term memory network is constructed. Business pulse flow data from several historical acquisition periods in each data sub-region is obtained as training data. The business node prediction model is trained using the training data to obtain a fully trained business node prediction model. Business pulse flow data from the current acquisition period is input into the fully trained business node prediction model. Based on the fully trained business node prediction model, business node prediction data for each data sub-region containing a time-series sequence of state weights within the current acquisition period is output.

5. The AI-based enterprise management document automatic archiving system based on semantic analysis according to claim 4, characterized in that, The process of dividing a data sub-region into high-priority or low-priority data sub-regions includes: Obtain the preset upper limit of business state weight for each data sub-region, extract the time series sequence of state weight in the business node prediction data of each data sub-region within the current collection period, compare the time series sequence of state weight of each data sub-region with the preset upper limit of business state weight, and obtain the cumulative time when the state weight is continuously greater than or equal to the preset upper limit of business state weight. A preset cumulative time error threshold is set. If the cumulative time of a data sub-region is greater than or equal to the cumulative time error threshold, the data sub-region is marked as a low-quality data sub-region. If the cumulative time of a data sub-region is less than the error threshold, the data sub-region is marked as a high-quality data sub-region.

6. The AI-based enterprise management document automatic archiving system based on semantic analysis according to claim 5, characterized in that, The process of performing business graph collaborative scheduling on low-quality data sub-regions includes: Step 1: Obtain the time period in which the state weight of the low-quality data sub-region exceeds the preset upper limit of the business state weight, and the maximum value of the excess weight within the time period. Obtain the preset network topology containing each data sub-region. Obtain the high-quality data sub-region that is closest to the low-quality data sub-region in the preset network topology. Obtain the minimum idle weight of the high-quality data sub-region within the time period. Determine whether the minimum idle weight is greater than or equal to the maximum value of the excess weight. If it is greater than or equal to, allocate the graph node corresponding to the minimum idle weight of the high-quality data sub-region to the low-quality data sub-region and end the business graph collaborative scheduling. If it is less than, allocate the graph node corresponding to the minimum idle weight of the high-quality data sub-region to the low-quality data sub-region, remove the high-quality data sub-region, and proceed to Step 2. Step 2: Obtain the maximum value of the weight exceeding the low-quality data sub-region within the time period after the graph node allocation, re-obtain the high-quality data sub-region that is closest to the low-quality data sub-region in the preset network topology, obtain the minimum value of the idle weight of the re-obtained high-quality data sub-region within the time period, and then execute Step 3. Step 3: Determine whether the minimum idle weight of the newly acquired high-priority data sub-region is greater than or equal to the maximum excess weight of the low-priority data sub-region within the time period. If it is greater than or equal to, allocate the graph node corresponding to the minimum idle weight of the newly acquired high-priority data sub-region to the low-priority data sub-region and end the business graph collaborative scheduling. If it is less than, allocate the graph node corresponding to the minimum idle weight of the newly acquired high-priority data sub-region to the low-priority data sub-region, remove the newly acquired high-priority data sub-region, and execute Step 2.

7. The AI-based enterprise management document automatic archiving system based on semantic analysis according to claim 6, characterized in that, The semantic context gravity coupling calculation module performs gravity attenuation determination on business nodes within each data sub-region, and executes either lightweight gravity compensation or adaptive cooperative gravity compensation based on the determination results. Extract the decay gravity value, decay time range, basic semantic similarity, duration of business state, and business state weight of each basic semantic vector in the real-time generated business nodes within the data sub-region. Determine whether the decay time range of the basic semantic vector is adjacent to or overlaps with the duration of the business state. If the decay time range is adjacent to or overlaps with the duration of the business state, the ratio of the decay gravity value to the basic semantic similarity is used as the comprehensive gravity ratio of the basic semantic vector. A preset comprehensive gravity ratio threshold is set. If the comprehensive gravity ratio is less than the comprehensive gravity ratio threshold, the business node to which the basic semantic vector belongs is marked as a gravity decay node. If the comprehensive gravity ratio is greater than or equal to the comprehensive gravity ratio threshold, the basic semantic vector is marked as a normal node. Obtain the basic semantic vector and the number and location features of gravity decay nodes within the data sub-region, divide the data sub-region into several virtual view units of the same size, and obtain the gravity decay density of each virtual view unit based on the ratio of the number of gravity decay nodes to the total capacity of the topological nodes of the virtual view unit. If the gravitational attenuation density of the virtual view unit is less than or equal to the density threshold and there are gravitational attenuation nodes within the virtual view unit, then a lightweight gravitational compensation operation is performed on the gravitational attenuation nodes within the virtual view unit; if the gravitational attenuation density of the virtual view unit is less than or equal to the density threshold and there are no gravitational attenuation nodes within the virtual view unit, then no compensation operation is performed; if the gravitational attenuation density of the virtual view unit is greater than the density threshold, then an adaptive cooperative gravitational compensation operation is performed on the area covered by the virtual view unit.

8. The AI-based enterprise management document automatic archiving system based on semantic analysis according to claim 7, characterized in that, The process of performing lightweight gravity compensation operations includes: Obtain several virtual directory paths in the preset archive link corresponding to the gravity decay node, extract the virtual directory paths that are not adjacent to or overlap with the decay time range, and mark the virtual directory paths that are not adjacent to or overlap with the decay time range as virtual directory paths to be assigned. The search matching rate of each virtual directory path to be assigned is calculated based on the ratio of the historical search hits to the total search count. The virtual directory path with the lowest search matching rate is selected, the business status data is migrated to the selected virtual directory path, and the corresponding business status weight parameters are updated according to the attributes of the virtual directory path to be assigned.

9. The AI-based enterprise management document automatic archiving system based on semantic analysis according to claim 8, characterized in that, The process of performing adaptive cooperative gravity compensation includes: Extract the flow path corresponding to the gravity attenuation node in the virtual view unit coverage area, obtain the data sub-region that emits the gravity attenuation node based on the flow path, and mark the data sub-region as the attenuation source; An adaptive gravity compensation model is constructed based on a graph convolutional neural network. The decay gravity value and decay time range of the decay source at the current moment, as well as the topological structure graph corresponding to the coverage area of ​​the virtual view unit, are input into the adaptive gravity compensation model. The model extracts node features and graph structure features and outputs the optimal gravity adjustment parameters of the decay source.

10. The AI-based enterprise management document automatic archiving system based on semantic analysis according to claim 9, characterized in that, The process by which the fluid archiving execution module performs virtual view archiving and permission mapping switching operations based on the real-time correlation gravity, business type, and flow status of business nodes includes: Extract the real-time correlation, business type, and flow status of each business node in the real-time data collected within the data sub-region; The configuration is as follows: When the business node's flow status is far from the data sub-region and the real-time correlation attraction is less than the preset correlation attraction threshold corresponding to the business type, extract the preset usage data volume and real-time correlation attraction corresponding to the business type, obtain the idle data volume valley value and real-time correlation attraction valley value of other data sub-regions in the current collection period, obtain the target data sub-region where the idle data volume valley value is greater than the preset usage data volume and the real-time correlation attraction valley value is greater than the real-time correlation attraction, calculate the Euclidean distance between the target data sub-region and the business node, and select the target data sub-region with the shortest Euclidean distance to perform virtual view archiving and permission mapping switching operations. When the business node's flow status is close to the data sub-region or the real-time association attraction is greater than or equal to the preset association attraction threshold corresponding to the business type, the current virtual view archiving and permission mapping status remains unchanged.