Online document processing method and system based on document processing model
By generating temporary lock identification and distributed transaction processing technology, combined with hierarchical cache and incremental synchronization, the data conflict and loading speed problems in multi-user collaboration in online document processing are solved, and the security and efficiency of the system are improved.
Patent Information
- Application Number
- CN202510798937.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing online document processing methods are prone to data conflicts in multi-user collaboration, imperfect permission control leads to inefficient editing processes, limited document loading speed, insufficient intelligent auditing capabilities, and difficult to ensure data consistency and content quality.
By generating temporary lock identification restriction modification operations, distributed transaction processing technology is used to record concurrent operations in sequence, hierarchical cache structure and incremental synchronization technology are designed, and cross-terminal synchronization process is optimized in combination with compression algorithms.
It realizes data consistency guarantee in multi-user collaboration, improves system response speed and content quality, reduces manual intervention costs, adapts to high-concurrency scenarios, and reduces network bandwidth consumption.
Smart Images

Figure CN120337879A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of document processing, and in particular discloses an online document processing method and system based on a document processing model. Background Art
[0002] As a core technical field in the information age, online document processing is of irreplaceable importance in aspects such as enterprise collaborative office, educational resource management, and cloud data storage, directly affecting the information flow efficiency and data security guarantee. However, current mainstream solutions often expose significant limitations when dealing with complex scenarios. For example, data conflicts are prone to occur in multi-user collaboration, the document loading speed is limited by the system architecture design, and the accuracy of intelligent review is also difficult to meet high-standard requirements. These problems limit the wide application of the technology.
[0003] By deeply analyzing the challenges faced in this field, it can be found that several core technical factors are interrelated and form obstacles. First, in the scenario of multi-user collaborative editing, the imperfect permission control mechanism leads to conflicts easily occurring in concurrent operations of chapter content, and it is difficult to guarantee data consistency. This conflict problem further causes inefficiency in the editing process. Because of the lack of a dynamic locking permission management strategy, user operations often need to be coordinated repeatedly, increasing the time cost. And the inefficient process directly affects the document processing speed. Especially when dealing with large-scale documents or cross-terminal synchronization, the system lacks an optimized data loading and caching mechanism, resulting in a significant response delay. More critically, the existing intelligent review technology lacks the ability in logical error identification and format optimization, making it difficult to guarantee the document quality through automated means, and the manual intervention cost remains high.
[0004] Therefore, how to design an online document processing method integrating dynamic permission locking, efficient data loading strategy, and intelligent review mechanism to ensure data consistency, improve the system response speed, and enhance the content quality in multi-user collaboration has become a key problem that needs to be solved urgently in this research. Summary of the Invention
[0005] The present invention provides an online document processing method and system based on a document processing model, aiming to solve at least one of the above-mentioned defects existing in the existing online document processing methods.
[0006] One aspect of the present invention relates to an online document processing method based on a document processing model, including the following steps: Obtain an editing request of a user for a target data unit, generate a temporary locking identifier according to the user permission identifier and the request timestamp, restrict the modification operation of other users on the target data unit, and obtain a conflict isolation result; Construct a data consistency module based on the conflict isolation result, extract the edit status information of the current data unit from the temporary lock identifier, and use distributed transaction processing technology to sequentially record concurrent operations; If a version conflict is detected in the sequential record, trigger a rollback mechanism and reassign the operation priority to determine the data version after consistency verification; Design a hierarchical cache structure based on the data version after consistency verification, store frequently accessed data in the memory cache area, and store infrequently accessed data in the disk cache area to obtain a hierarchical storage structure; Optimize the cross-terminal synchronization process based on the hierarchical storage structure, use incremental synchronization technology to extract the changed data set, and generate a compressed transmission packet in combination with a compression algorithm.
[0007] Further, the steps of obtaining the edit request of the user for the target data unit, generating a temporary lock identifier based on the user permission identifier and the request timestamp, and restricting the modification operation of other users on the target data unit to obtain the conflict isolation result include: According to the edit request submitted by the user, obtain the user permission identifier and the request timestamp from the database; Verify the user permission identifier through a pre-established permission verification rule. If the user permission identifier meets the preset access conditions, generate a temporary lock identifier, which is used to restrict the modification of the target data unit by other users; For the temporary lock identifier, obtain the access record of the target data unit from the server log, determine whether there is a concurrent edit request from other users. If there is a concurrent edit request, determine the priority through the order of the request timestamps to obtain the final lock ownership result; According to the final lock ownership result, use the distributed lock mechanism to update the status of the target data unit. When updating, obtain the current version number of the target data unit and determine whether the current version number is consistent with the locked version. If they are consistent, allow the edit operation to obtain the lock protection status; Through the lock protection status, obtain the specific content of the edit operation from the operation log, compare the difference between the specific content and the original content of the target data unit, use a version control tool to record the difference information, and determine the final result of conflict isolation.
[0008] Further, the steps of constructing a data consistency module based on the conflict isolation result, extracting the edit status information of the current data unit from the temporary lock identifier, and using distributed transaction processing technology to sequentially record concurrent operations include: According to the conflict isolation result, extract the edit status information of the target data unit from the temporary lock identifier, compare the edit status information with the operation log to determine whether there are unfinished concurrent operations. If there are unfinished concurrent operations, sort them according to the priority order to obtain a preliminary sequence record table; Through the preliminary sequence record table, use a distributed transaction processing tool to divide the concurrent operations into batches, obtain the status update information for the operation log in each batch, and determine the final edit status of each current data unit in the current batch; According to the final edit status, perform version verification for each current data unit, obtain the historical version information from the pre-established version record. If the historical version information does not match the current status, suspend the corresponding operation to obtain the operation queue after version verification; Through the operation queue after version verification, use a transaction processing tool to update the status of the current data unit, and determine whether the status update meets the data consistency requirements to obtain the final sequenced record result.
[0009] Furthermore, if a version conflict is detected in the sequenced record, trigger the rollback mechanism and reassign the operation priorities. The steps to determine the data version after consistency verification include: Obtain the version conflict information of the target operation queue from the operation record, compare the version conflict information with the pre-established historical data. If it is detected that the conflicting data in the version conflict information is inconsistent with the historical data, trigger the rollback mechanism to obtain the operation queue after initial rollback; According to the operation queue after initial rollback, use a distributed transaction tool to reassign the operation priorities, sort the status update information of the operation record, and determine the reallocated priority sequence; Through the reallocated priority sequence, obtain the consistency verification information of each operation record. If the consistency verification information does not match the preset threshold, adjust the status of the relevant operation record to obtain the adjusted operation queue; For the adjusted operation queue, use a log recording tool to finally confirm the data version, compare the verification result with the historical data to determine whether the final data version meets the consistency requirements, and obtain the confirmed data version.
[0010] Furthermore, design a hierarchical cache structure according to the data version after consistency verification, store the frequently accessed data in the memory cache area, and store the infrequently accessed data in the disk cache area. The steps to obtain the hierarchical storage structure include: Retrieve access log data from the database. By statistically processing the access log data, determine the access frequency value for each piece of access log data, and compare the access frequency value with a preset frequency threshold. If the access frequency value is higher than the frequency threshold, classify it as high-frequency access data; if the access frequency value is lower than the frequency threshold, classify it as low-frequency access data to obtain a classified data set. For the classified data set, use a memory caching tool to store and process the high-frequency access data. Through the fast read and write interface of the memory caching tool, load the high-frequency access data into the memory cache area to obtain the high-frequency data storage result in the memory cache area, and at the same time obtain the storage path of the low-frequency access data. According to the storage path of the low-frequency access data, use a disk storage tool to store and process the low-frequency access data. Allocate space in the disk cache area and write the low-frequency access data into the disk cache area to obtain the low-frequency access data storage result in the disk cache area. By performing consistency verification on the high-frequency data storage result in the memory cache area and the low-frequency access data storage result in the disk cache area, obtain the storage status of the two areas of the memory cache area and the disk cache area. If the storage status is abnormal, trigger a data reallocation mechanism to obtain the final hierarchical storage structure.
[0011] Furthermore, based on the hierarchical storage structure, optimizing the cross-terminal synchronization process, the steps of using the incremental synchronization technology to extract the changed data set and generating a compressed transmission package in combination with a compression algorithm include: Retrieve the changed data set from the hierarchical storage structure, classify the data in the memory cache area and the disk cache area. By comparing the timestamps and modification records, determine the specific range of the changed data to obtain a classified changed data set. Use an incremental extraction tool to process the classified changed data set. When extracting, if the timestamp of the changed data is later than the last synchronization time, classify the changed data set into the data group to be synchronized. By comparing item by item, obtain the complete content of the data group to be synchronized and determine the data range that needs to be transmitted. For the data group to be synchronized, use a compression tool to process it. Pack the data group to be synchronized into a compressed transmission package through a preset compression algorithm, and dynamically adjust according to the data size and the bandwidth of the terminal device to obtain the final compressed transmission package.
[0012] Distribute the final compressed transmission package to the target terminal device through a network transmission tool. When transmitting, if a network interruption or delay is detected, trigger a retransmission mechanism, obtain the transmission status feedback, and determine whether the final compressed transmission package is completely delivered.
[0013] Another aspect of the present invention relates to an online document processing system based on a document processing model for implementing the above-mentioned online document processing method based on a document processing model, including: A first acquisition module, configured to acquire an editing request of a user for a target data unit, generate a temporary lock identifier according to the user permission identifier and the request timestamp, restrict the modification operation of other users on the target data unit, and obtain a conflict isolation result; A recording module, configured to construct a data consistency module according to the conflict isolation result, extract the editing status information of the current data unit from the temporary lock identifier, and sequentially record concurrent operations by using distributed transaction processing technology; A determination module, configured to trigger a rollback mechanism and reassign operation priorities if a version conflict is detected in the sequential record, and determine the data version after consistency verification; A second acquisition module, configured to design a hierarchical cache structure according to the data version after consistency verification, store frequently accessed data in a memory cache area, and store infrequently accessed data in a disk cache area to obtain a hierarchical storage structure; A generation module, configured to optimize the cross-terminal synchronization process based on the hierarchical storage structure, extract a set of changed data by using incremental synchronization technology, and generate a compressed transmission packet in combination with a compression algorithm.
[0014] Further, the first acquisition module includes: A first acquisition unit, configured to acquire the user permission identifier and the request timestamp from a database according to the editing request submitted by the user; A generation unit, configured to verify the user permission identifier through a pre-established permission verification rule. If the user permission identifier meets the preset access condition, generate a temporary lock identifier, and the temporary lock identifier is used to restrict the modification of other users on the target data unit; A second acquisition unit, configured to obtain the access record of the target data unit from the server log for the temporary lock identifier, determine whether there is a concurrent editing request of other users. If there is a concurrent editing request, determine the priority according to the sequence of the request timestamps to obtain a final lock ownership result; A third acquisition unit, configured to update the status of the target data unit by using a distributed lock mechanism according to the final lock ownership result. When updating, obtain the current version number of the target data unit, and determine whether the current version number is consistent with the version at the time of locking. If they are consistent, allow the execution of the editing operation to obtain a lock protection status; A first determination unit, configured to obtain the specific content of the editing operation from the operation log through the lock protection status, compare the specific content with the original content of the target data unit for differences, record the difference information by using a version control tool, and determine the final result of conflict isolation.
[0015] Furthermore, the recording module includes: A fourth acquisition unit, configured to extract the edit status information of the target data unit from the temporary lock identifier according to the conflict isolation result, compare the edit status information with the operation log to determine whether there are unfinished concurrent operations, and if there are unfinished concurrent operations, sort them according to the priority order to obtain a preliminary sequence record table; A second determination unit, configured to divide the concurrent operations into batches by using a distributed transaction processing tool through the preliminary sequence record table, obtain status update information for the operation logs in each batch, and determine the final edit status of each current data unit in the current batch; A fifth acquisition unit, configured to perform version verification on each current data unit according to the final edit status, obtain historical version information from the pre-established version record, and if the historical version information does not match the current status, suspend the corresponding operation to obtain an operation queue after version verification; A sixth acquisition unit, configured to update the status of the current data unit by using a transaction processing tool through the operation queue after version verification, and determine whether the status update meets the data consistency requirement to obtain a final sequenced record result.
[0016] Furthermore, the determination module includes: A seventh acquisition unit, configured to obtain version conflict information of the target operation queue from the operation record, compare the version conflict information with the pre-established historical data, and if it is detected that the conflict data in the version conflict information is inconsistent with the historical data, trigger a rollback mechanism to obtain an operation queue after initial rollback; A third determination unit, configured to reassign the operation priorities by using a distributed transaction tool according to the operation queue after initial rollback, sort the status update information of the operation record, and determine the reallocated priority sequence; An eighth acquisition unit, configured to obtain consistency verification information of each operation record through the reallocated priority sequence, and if the consistency verification information does not match the preset threshold, adjust the status of the relevant operation record to obtain an adjusted operation queue; A ninth acquisition unit, configured to perform a final confirmation on the data version by using a log recording tool for the adjusted operation queue, compare the verification result with the historical data, and determine whether the final data version meets the consistency requirement to obtain a confirmed data version.
[0017] The beneficial effects achieved by the present invention are: The present invention provides an online document processing method and system based on a document processing model. By obtaining an editing request from a user for a target data unit, a temporary locking identifier is generated according to the user permission identifier and the request timestamp to restrict the modification operations of other users on the target data unit, and a conflict isolation result is obtained; a data consistency module is constructed according to the conflict isolation result, the editing status information of the current data unit is extracted from the temporary locking identifier, and a distributed transaction processing technology is used to sequentially record concurrent operations; if a version conflict is detected in the sequential record, a rollback mechanism is triggered and the operation priority is reallocated to determine the data version after consistency verification; a hierarchical cache structure is designed according to the data version after consistency verification, high-frequency access data is stored in the memory cache area, and low-frequency access data is stored in the disk cache area to obtain a hierarchical storage structure; based on the hierarchical storage structure, the cross-terminal synchronization process is optimized, and an incremental synchronization technology is used to extract a set of changed data and generate a compressed transmission packet in combination with a compression algorithm. The beneficial effects obtained by the online document processing method and system based on the document processing model provided by the present invention are as follows: I. Concurrent Conflict Isolation and Security Control 1. By restricting concurrent modification by multiple users through a temporary locking identifier and combining permission identifier and timestamp verification, the legality of operations is ensured, and problems such as data overwrite or unauthorized access are avoided.
[0018] 2. By combining a lock mechanism with a distributed transaction, fine-grained access control is realized, and the security of the system is enhanced.
[0019] II. Data Consistency Guarantee 1. The distributed transaction processing technology sequentially records concurrent operations, and combines version conflict detection and rollback mechanism to ensure the atomicity and eventual consistency of operations, and prevent problems such as dirty reads / phantom reads.
[0020] 2. By reallocating the version priority, the transaction processing order is optimized, and the retry of invalid operations caused by concurrency is reduced.
[0021] III. Improvement of Storage and Access Efficiency 1. The hierarchical cache structure dynamically divides the high-frequency / low-frequency data storage areas, reduces memory occupancy and accelerates the access to hot data, comprehensively improving the response speed and resource utilization rate.
[0022] 2. By combining the incremental synchronization technology with the compression algorithm, the cross-terminal data transmission volume is reduced, the network bandwidth consumption is reduced, and lightweight and efficient synchronization is realized.
[0023] IV. Enhancement of System Reliability and Scalability 1. The rollback mechanism and distributed transaction support the fault recovery ability and improve the fault tolerance of the system.
[0024] 2. The hierarchical storage and incremental synchronization design adapts to multi - terminal and high - concurrency scenarios, facilitating horizontal expansion and dynamic resource scheduling.
[0025] V. Cost and Maintenance Optimization 1. Automated version management and conflict resolution reduce the need for manual intervention and lower the complexity of operation and maintenance.
[0026] 2. Elastic resource allocation (such as on - demand caching and compressed transmission) avoids excessive hardware investment, conforming to the low - cost concept of the Serverless architecture. Brief Description of the Drawings
[0027] Figure 1 It is a schematic flowchart of an embodiment of an online document processing method based on a document processing model of the present invention. Detailed Embodiment
[0028] To better understand the above - mentioned technical solution, the following will describe the above - mentioned technical solution in detail in conjunction with the drawings in the specification and specific embodiments.
[0029] As Figure 1 shown, a first embodiment of the present invention proposes an online document processing method based on a document processing model, including the following steps: Step S100: Obtain an edit request from the user for a target data unit, generate a temporary lock identifier according to the user permission identifier and the request timestamp, restrict the modification operation of other users on the target data unit, and obtain a conflict isolation result.
[0030] The target data unit (TDU) is the smallest operable data entity defined in a data management system or processing flow, used to carry specific business logics, processing rules, or storage requirements. The design of the target data unit needs to meet the requirements of standardization, structurization, and traceability to adapt to the input specifications of the target system (such as a database, API interface, or document model). Example: The "clause unit" in an online contract management system may include fields such as clause text, version number, effective time, and signing status.
[0031] The user permission identifier is the core control mechanism in a database or information system used to uniquely identify the user identity and its operation permission scope. It restricts the access and operation permissions of users to data objects (such as tables, views) and functional modules through predefined rules to ensure system security and data integrity.
[0032] The timestamp is a numerical value or string in a computer system used to uniquely identify the occurrence time of an event. It calculates the time interval through a specific reference time point to provide a standardized time recording method.
[0033] The conflict isolation result is a key objective of transaction concurrency control in a database system. It refers to eliminating or restricting data inconsistency problems (such as dirty reads, non-repeatable reads, and phantom reads) generated during concurrent operations of multiple transactions through isolation level settings and lock mechanisms, and ultimately ensuring the data consistency state after transaction execution. Its essence is the implementation effect of transaction isolation (one of the ACID properties), and conflict avoidance is achieved by controlling the visibility rules between transactions.
[0034] Step S200: Construct a data consistency module based on the conflict isolation result, extract the edit status information of the current data unit from the temporary lock flag, and use distributed transaction processing technology to sequentially record concurrent operations.
[0035] The data consistency module is the core component of a database or distributed system, responsible for ensuring that data always meets the preset business logic and integrity requirements after operations through rule constraints, transaction control, and synchronization mechanisms.
[0036] A data unit is the smallest logical data management unit in a database or distributed system, referring to an independent data set with a clear business meaning and that cannot be further divided (such as a single record in a database table, a message in a message queue, or the smallest atomic object of a transaction operation).
[0037] Distributed transaction processing technology is a transaction management mechanism used for transactions across multiple independent nodes or services (such as databases, message queues, microservices, etc.), ensuring that multiple operations in a distributed system have atomicity, consistency, isolation, and durability (ACID properties). Its core goal is to solve the problem of data state inconsistency caused by network partitions, node failures, or concurrent conflicts, and to ensure the global effectiveness of business operations.
[0038] Concurrent operations refer to multiple independent operations that are executed simultaneously within an overlapping time period, which may involve accessing or modifying shared resources (such as data, devices, services). Its core goal is to improve system throughput and response speed, but coordination mechanisms (such as locks, transaction isolation, version control) are required to solve data competition (Race Condition) and state inconsistency problems caused by parallel execution.
[0039] Sequential recording refers to irreversible data units that are generated, stored, or operated in strict time or logical order during the data processing process. Its core goal is to ensure the timeliness and state consistency of operations. For example, database transaction logs, operation records of file systems, or atomic append operations in distributed systems (such as the atomic record append of GFS) are typical scenarios of sequential recording.
[0040] Step S300: If a version conflict is detected in the serialized record, trigger the rollback mechanism, reassign the operation priorities, and determine the data version after consistency verification.
[0041] The rollback mechanism is a technical means to restore data, status, or business processes to a previous consistent version when system operations fail or transaction executions are abnormal. Its core goal is to ensure the atomicity and consistency of the system. For example, a database transaction is rolled back to the state before it was committed, or in a distributed system, completed subtasks are revoked through compensating operations.
[0042] Operation priorities refer to the rules set to coordinate resource access or operation execution order in concurrent or distributed systems, used to determine the priority of execution of critical operations or resource requests to ensure system efficiency, data consistency, and the correctness of business logic. For example, the nested order of distributed locks and transactions, and the scheduling strategies for transaction isolation levels in high-concurrency scenarios all rely on priority rules.
[0043] Consistency verification refers to the verification mechanism in a distributed system or data processing process to ensure that the data states and operation results of multiple nodes, multiple replicas, or cross-services conform to the preset consistency rules. The core goal is to eliminate data differences caused by network latency, concurrent conflicts, or node failures, and ensure the reliability of the system and the correctness of business logic. For example, in a distributed transaction, verify whether all participants have submitted successfully, or verify the consistency of data replicas through version comparison.
[0044] Step S400: Design a hierarchical cache structure based on the data version after consistency verification, store frequently accessed data in the in-memory cache area, and store infrequently accessed data in the disk cache area to obtain a hierarchical storage structure.
[0045] A hierarchical cache structure refers to the collaborative design of multiple cache levels (such as local cache, distributed cache, persistent storage). According to data access frequency, cost, and performance requirements, requests are filtered layer by layer to maximize system throughput and response speed while reducing resource consumption. The core goal of a hierarchical cache structure is to balance performance and cost, and achieve efficient cache management through mechanisms such as data eviction, preheating, and backhaul between levels.
[0046] High-frequency access data refers to a dataset that is repeatedly requested or updated within a specific time window, usually showing a significantly higher access frequency than other data. Such data needs to be preferentially stored in high-speed storage media (such as local memory, L1 / L2 cache) through technologies such as cache layering and preloading to reduce access latency and improve system throughput. Its core characteristics include short-term intensive access and locality rules (temporal locality and spatial locality).
[0047] Low-frequency access data refers to inactive data with a significantly lower access frequency than the system average within a specific time window, usually characterized by being not read or updated for a long time and having a low requirement for real-time performance. For example, historical orders, archived logs, expired inventory records, etc., and their access intervals may be in hours, days, or even months.
[0048] The hierarchical storage structure refers to classifying and storing data in different media levels with different performance or costs according to dimensions such as data access frequency, business value, and storage cost, to achieve the optimal balance between resource utilization and performance. The core goal of the hierarchical storage structure is to meet differentiated data access requirements at a reasonable cost. Through mechanisms such as hot-cold separation and dynamic migration, it ensures that high-frequency access data preferentially uses high-performance storage (such as SSD, memory), and low-frequency access data sinks to low-cost media (such as HDD, object storage).
[0049] Step S500: Optimize the cross-terminal synchronization process based on the hierarchical storage structure, adopt the incremental synchronization technology to extract the changed data set, and generate a compressed transmission package in combination with the compression algorithm.
[0050] The cross-terminal synchronization process refers to achieving real-time or near-real-time consistent updates of data among multiple terminal devices through a distributed architecture and data management technology, ensuring that when users access and modify the same data on different devices, the data status and content remain unified. The core goal of the cross-terminal synchronization process is to eliminate data islands between devices and provide a seamless user experience, such as the automatic synchronization of contacts, calendars, files, etc. among devices such as mobile phones, tablets, and smart watches.
[0051] The incremental synchronization technology is a data synchronization strategy that refers to only transmitting the data that has changed in the source system since the last synchronization, rather than replicating the entire data set in full. The core goal of the incremental synchronization technology is to improve the efficiency of large-scale data synchronization by reducing the data transmission volume and shortening the synchronization time, especially suitable for high-frequency update scenarios (such as financial transactions, Internet of Things device status synchronization).
[0052] A changed data set refers to a subset of data that has changed (added, deleted, or modified) in the source system since the last synchronization during the data synchronization or update process, rather than the entire data set. Its core objective is to reduce redundant transmission and storage overhead and improve synchronization efficiency by accurately identifying and transmitting differential data, especially applicable to business scenarios with high-frequency updates (such as real-time trading systems and IoT device status synchronization).
[0053] A compression algorithm is a computational method that reduces the data volume by eliminating data redundancy or optimizing the data representation form. Its goal is to reduce storage or transmission costs while ensuring data recoverability (lossless compression) or acceptable information loss (lossy compression). The core value of the compression algorithm is reflected in resource efficiency improvement (such as saving storage space and bandwidth) and performance optimization (such as accelerating data transmission and reducing I / O latency).
[0054] A compressed transmission package refers to an encapsulated form that integrates multiple files or data into a single file through a specific compression algorithm, aiming to reduce the amount of data transmitted, improve transmission efficiency, and reduce storage and network resource consumption. The core function of the compressed transmission package is to significantly reduce the data volume by eliminating redundant data (such as duplicate bytes and invalid information), and at the same time support security mechanisms such as encryption protection.
[0055] Furthermore, the online document processing method based on the document processing model provided in this embodiment, step S100 includes: Step S110: Obtain the user permission identifier and the request timestamp from the database according to the edit request submitted by the user.
[0056] Exemplarily, in an enterprise internal document management system, assume that a user Zhang submitted an edit request. The system first obtains Zhang's permission identifier as "Editor-003" from the database, and at the same time records the request timestamp as "2023-10-15 14:30:00".
[0057] Step S120: Verify the user permission identifier through the pre-established permission verification rules. If the user permission identifier meets the preset access conditions, generate a temporary lock identifier, which is used to restrict other users from modifying the target data unit.
[0058] Through the preset permission verification rules, the system verifies that "Editor-003" meets the access conditions, so a temporary lock identifier "Lock-20231015-001" is generated to restrict other users from modifying the target document "Quarterly Report.docx", and it enters the preliminary locked state. This mechanism effectively avoids data chaos caused by multiple people modifying at the same time.
[0059] Step S130: For the temporary lock identifier, obtain the access record of the target data unit from the server log, and determine whether there are concurrent editing requests from other users. If there are concurrent editing requests, determine the priority according to the chronological order of the request timestamps to obtain the final lock ownership result.
[0060] In a possible implementation, for the temporary lock identifier, the system extracts the access record of "Quarterly Report.docx" from the server log and finds that another user, Xiao Li, also submitted an editing request at "2023-10-15 14:29:50".
[0061] Since Xiao Li's timestamp is earlier than Zhang's, the system assigns the final lock ownership to Xiao Li according to the time priority. This priority judgment based on timestamps ensures the fairness and orderliness of request processing and reduces conflicts.
[0062] Step S140: According to the final lock ownership result, use the distributed lock mechanism to update the status of the target data unit. When updating, obtain the current version number of the target data unit and determine whether the current version number is the same as that at the time of locking. If they are the same, allow the execution of the editing operation to obtain the lock protection status.
[0063] Specifically, for the final lock ownership result, the system uses the distributed lock mechanism to update the status of "Quarterly Report.docx".
[0064] When updating, the system obtains the current version number of the document as "V1.2" and compares it with the version number "V1.2" at the time of locking. After finding that they are the same, it allows Xiao Li to execute the editing operation and enters the lock protection state. This version number verification mechanism effectively prevents data overwrite problems and ensures the security of the editing operation.
[0065] Step S150: Through the lock protection status, obtain the specific content of the editing operation from the operation log, perform a difference comparison between the specific content and the original content of the target data unit, use a version control tool to record the difference information, and determine the final result of conflict isolation.
[0066] It should be noted that through the lock protection status, the system obtains Xiao Li's editing content from the operation log. For example, he added a paragraph of "Sales Data Analysis" text to the document. The system performs a difference comparison between this content and the original document content and finds that there is no conflict in the added part.
[0067] With the help of the version control tool, the system records the difference information, forms version "V1.3", and completes conflict isolation. This difference comparison and version recording method not only ensures data integrity but also facilitates subsequent traceability and rollback.
[0068] Preferably, the above permission verification rules can be configured as multi-level permissions in the system. For example, an "editor" can only modify, while an "auditor" can view and approve, ensuring refined management of permission allocation. Such a design enhances the security of the system.
[0069] In one embodiment, the distributed lock mechanism can be achieved through node coordination in the cluster. For example, when Xiao Li obtains the lock, other nodes synchronously update the status, avoiding multi-node conflicts. This approach significantly improves the stability of the system in a high-concurrency environment.
[0070] It can be understood that the above solution forms a complete document editing protection process from permission verification to final conflict isolation, ensuring data consistency and operation fairness. At the same time, through version control and difference comparison, the conflict risk is reduced and the collaboration efficiency is improved.
[0071] Furthermore, the online document processing method based on the document processing model provided in this embodiment, step S200 includes: Step S210: According to the conflict isolation result, extract the editing status information of the target data unit from the temporary lock identifier, compare the editing status information with the operation log to determine whether there are unfinished concurrent operations. If there are unfinished concurrent operations, sort them according to the priority order to obtain a preliminary order record table.
[0072] Exemplarily, in an enterprise internal document collaboration platform, for the processing of the conflict isolation result, the system needs to extract the editing status information of the target data unit from the temporary lock identifier.
[0073] Suppose the temporary lock identifier of a document "annual plan.docx" is "Lock-20231016-002", and the system parses the current editing status as "to be submitted".
[0074] By comparing with the operation log, it is found that the editing operations of two users, Xiao Wang and Xiao Li, are unfinished. Xiao Wang's request time is "2023-10-16 09:00:00", and Xiao Li's is "2023-10-16 09:01:00". Sorted according to the time priority, Xiao Wang ranks first, forming a preliminary order record table. This time-based sorting method ensures the fairness of the operations.
[0075] Step S220: Through the preliminary order record table, use a distributed transaction processing tool to divide the concurrent operations into batches, obtain the status update information for the operation log in each batch, and determine the final editing status of each current data unit in the current batch.
[0076] In a possible implementation, for the preliminary sequence record table, the system uses a distributed transaction processing tool to divide concurrent operations into batches.
[0077] Suppose Xiao Wang's editing operation is divided into the first batch and Xiao Li's into the second batch. The system obtains Xiao Wang's operation log for the first batch and finds that the content he edited is the addition of a "project schedule", and the status update information shows "modified". Through analysis, the system determines that the final editing status of this document in the current batch is "pending review", laying a foundation for subsequent operations.
[0078] Step S230: According to the final editing status, perform version verification for each current data unit, obtain historical version information from the pre-established version record. If the historical version information does not match the current status, suspend the corresponding operation to obtain an operation queue after version verification.
[0079] Specifically, for the version verification of the final editing status, the system extracts the historical version information of "annual plan.docx" from the version record as "V2.0", while the current status shows the version as "V2.1".
[0080] Due to the mismatch, the system suspends Xiao Wang's editing operation and adds it to the operation queue after version verification. This verification mechanism avoids the risk of data overwrite and ensures version consistency.
[0081] Step S240: Through the operation queue after version verification, use the transaction processing tool to update the status of the current data unit, and determine whether the status update meets the data consistency requirements to obtain the final sequential record result.
[0082] It should be noted that through the operation queue after version verification, the system uses the transaction processing tool to update the status of the data unit.
[0083] Suppose after Xiao Wang's editing operation passes the verification, the system attempts to update the document status to "reviewed" and checks whether it meets the data consistency requirements. It is found that there is no contradiction in the data logic before and after the update, and finally a sequential record result is formed. This method ensures the reliability of the operation.
[0084] Preferably, for sorting uncompleted concurrent operations, the system can also introduce the user role priority as an auxiliary rule. For example, Xiao Wang is a "project manager" with a higher priority than Xiao Li, an "ordinary editor". Even if the time is slightly later, his operation may be processed first. This multi-dimensional sorting improves the flexibility of the operation.
[0085] In one embodiment, the batch division can also be adjusted according to the operation type. For example, "adding content" and "deleting content" are divided into different batches to reduce the probability of conflicts. This refined management optimizes the resource allocation efficiency.
[0086] It is understandable that the above method ensures the operation sequence and data consistency in document collaboration by the complete process from status extraction to final record, providing stable support for multi-person collaboration.
[0087] Furthermore, the online document processing method based on the document processing model provided in this embodiment, step S300 includes: Step S310, obtain the version conflict information of the target operation queue from the operation record, compare the version conflict information with the pre-established historical data. If it is detected that the conflict data in the version conflict information is inconsistent with the historical data, trigger the rollback mechanism to obtain the operation queue after the initial rollback.
[0088] Exemplarily, in a document collaboration platform within an enterprise, regarding the theme of obtaining the version conflict information of the target operation queue from the operation record, it can be understood in principle that the system needs to identify the conflict points in the operation queue that may cause data inconsistency.
[0089] Suppose there is a version conflict in the operation queue for a certain document "Quarterly Report.docx". The system detects that the current version is "V3.2", while the version shown in the operation record is "V3.1".
[0090] By comparing with the historical data, after detecting the inconsistency, trigger the rollback mechanism to restore the operation queue to the initial state, that is, roll back to the operation queue of the "V3.1" version. This rollback mechanism ensures that data will not be incorrect due to conflicts.
[0091] Step S320, according to the operation queue after the initial rollback, use a distributed transaction tool to reassign the operation priorities, sort the status update information of the operation record, and determine the reallocated priority sequence.
[0092] In a possible implementation manner, for the operation queue after the initial rollback, the system uses a distributed transaction tool to reassign the operation priorities.
[0093] Suppose three users, Zhang, Li, and Wang, edit "Quarterly Report.docx" simultaneously. Their operation times are "2023-10-18 10:00:00", "2023-10-18 10:01:00", and "2023-10-18 10:02:00" respectively. The system reassigns the priorities according to the time sequence. Zhang has the highest priority, followed by Li, and finally Wang. This time-based sorting method ensures the orderliness of the operations.
[0094] Step S330: Obtain the consistency verification information of each operation record according to the reallocated priority sequence. If the consistency verification information does not match the preset threshold, adjust the status of the relevant operation records to obtain an adjusted operation queue.
[0095] Specifically, for the reallocated priority sequence, the system obtains the consistency verification information of each operation record.
[0096] Suppose Zhang's operation record shows that his edited content is "update data table", and the verification information indicates that the current status matches the preset threshold. While Li's operation record shows "modify title", but the verification information does not match. The system will adjust Li's operation status to "pending confirmation", thus forming an adjusted operation queue. This verification mechanism avoids the execution of invalid operations.
[0097] Step S340: For the adjusted operation queue, use a log recording tool to perform a final confirmation of the data version. Compare the verification result with the historical data to determine whether the final data version meets the consistency requirements, and obtain the confirmed data version.
[0098] It should be noted that for the adjusted operation queue, the system uses a log recording tool to perform a final confirmation of the data version. Suppose Zhang's operation passes the verification, the system updates the document version to "V3.3", and combines with the historical data comparison to confirm that the current version is consistent with the operation record and meets the consistency requirements. While Li's operation needs to wait for subsequent confirmation due to status adjustment. This final confirmation method ensures the reliability of the data.
[0099] Preferably, from multiple perspectives to handle version conflict information, user roles can be introduced as auxiliary rules. For example, although Wang is the latest in time, his role is "supervisor", and the system may elevate his priority under certain circumstances. This multi-dimensional consideration improves the flexibility of operation allocation.
[0100] In one embodiment, for the acquisition of consistency verification information, it can be refined and managed according to the operation type. For example, "content modification" and "format adjustment" are verified separately to avoid interference between different types of operations. This classification method optimizes the verification efficiency.
[0101] It can be understood that the above method ensures data consistency in document collaboration through a complete process from conflict identification to final version confirmation, providing stable support for multi-person collaboration.
[0102] Furthermore, the online document processing method based on the document processing model provided in this embodiment, step S400 includes: Step S410: Obtain access log data from the database. By statistically processing the access log data, determine the access frequency value for each piece of access log data, and compare the access frequency value with a preset frequency threshold. If the access frequency value is higher than the frequency threshold, it is classified as high-frequency access data; if the access frequency value is lower than the frequency threshold, it is classified as low-frequency access data, resulting in a classified data set.
[0103] Exemplarily, in a document management platform within an enterprise, regarding the topic of obtaining access log data from the database and performing statistical processing, it can be understood in principle that the system needs to analyze the access behavior of users to documents in order to optimize the storage strategy.
[0104] Suppose the platform processes thousands of access logs every day. The system calculates the access frequency value by counting the number of accesses to each piece of data. For example, the access frequency of a certain document "Annual Plan.docx" is 50 times per day, while that of another "Meeting Minutes.docx" is 5 times per day.
[0105] Set the frequency threshold to 20 times per day. The former is classified as high-frequency access data, and the latter as low-frequency access data, forming a classified data set.
[0106] Step S420: For the classified data set, use a memory caching tool to store and process the high-frequency access data. Through the fast read-write interface of the memory caching tool, load the high-frequency access data into the memory cache area, obtaining the high-frequency data storage result in the memory cache area, and at the same time obtain the storage path of the low-frequency access data.
[0107] In a possible implementation, for the storage and processing of high-frequency access data, the system uses a memory caching tool to improve the access speed. Suppose "Annual Plan.docx" is loaded into the memory cache area. Through the fast read-write interface, the response time for users to access this document is significantly shortened. For low-frequency access data such as "Meeting Minutes.docx", the system only records its storage path, pointing to the disk storage area, to avoid occupying memory resources.
[0108] Step S430: According to the storage path of the low-frequency access data, use a disk storage tool to store and process the low-frequency access data. Allocate space in the disk cache area and write the low-frequency access data into the disk cache area, obtaining the low-frequency access data storage result in the disk cache area.
[0109] Specifically, for the disk storage processing of low-frequency access data, space can be allocated in the disk cache area through the disk storage tool. Suppose "Meeting Minutes.docx" is written into the disk cache area. The system allocates a fixed storage space for it and records the access path for quick positioning when needed. This hierarchical storage method ensures reasonable resource allocation.
[0110] Step S440: By performing consistency verification on the storage results of high-frequency data in the memory buffer and the storage results of low-frequency access data in the disk buffer, obtain the storage status of the memory buffer and the disk buffer. If the storage status is abnormal, trigger the data redistribution mechanism to obtain the final hierarchical storage structure.
[0111] It should be noted that for the consistency verification of the data storage results in the memory buffer and the disk buffer, the system will regularly check the storage status of the two areas.
[0112] Suppose that in a verification, it is found that part of the data of "Annual Plan.docx" is missing in the memory buffer. The system determines that the storage status is abnormal and triggers the data redistribution mechanism to reload the document data from the disk to the memory to ensure data integrity.
[0113] Preferably, from the perspective of the data redistribution mechanism, priority rules can be introduced to assist in processing. For example, if multiple high-frequency access data simultaneously show storage abnormalities, the system preferentially restores the document with the highest access frequency according to the access frequency. This method ensures the availability of core data.
[0114] In one embodiment, for the implementation of consistency verification, the verification frequency and scope can be refined. For example, the high-frequency data in the memory buffer is verified once a day, while the low-frequency data in the disk buffer is verified once a week to avoid resource waste. This differential management improves the system efficiency.
[0115] It can be understood that the above hierarchical storage and verification mechanism, through the accurate analysis of access frequency and the reasonable allocation of storage resources, ensures the efficient access and stability of data in the document management platform, provides a smooth user experience, and reduces the storage cost at the same time.
[0116] Furthermore, in the online document processing method based on the document processing model provided in this embodiment, step S500 includes: Step S510: Obtain the changed data set from the hierarchical storage structure, classify the data in the memory buffer and the disk buffer, and determine the specific scope of the changed data by comparing the timestamps and modification records to obtain the classified changed data set.
[0117] Exemplarily, in a document management platform within an enterprise, for the acquisition and processing of changed data in the hierarchical storage structure, it can be understood from the principle that the system needs to identify the data content that has changed in the memory buffer and the disk buffer for subsequent synchronization and transmission.
[0118] Assume that the platform stores a large amount of document data, and the system determines the scope of changes by comparing timestamps and modification records. For example, the last modification time of a document "Project Report.docx" in the memory buffer is 2023-10-15 14:00, while the backup time in the disk buffer is 2023-10-15 10:00. The system determines that the document has changes and classifies it into the change data set.
[0119] In a possible implementation, for the classification processing of change data, the system will distinguish according to the storage location and modification records.
[0120] Assume that "Project Report.docx" belongs to high-frequency access data and is stored in the memory buffer, while another "Historical Archive.docx" belongs to low-frequency access data and is stored in the disk buffer. The system confirms the scope of changes for the two documents through timestamp comparison and forms a classified change data set. This classification method facilitates subsequent differential processing strategies for different storage areas.
[0121] Step S520: Use an incremental extraction tool to process the classified change data set. When extracting, if the timestamp of the change data is later than the last synchronization time, classify the classified change data set into the data group to be synchronized. By comparing item by item, obtain the complete content of the data group to be synchronized and determine the data range that needs to be transmitted.
[0122] Specifically, for the application of the incremental extraction tool, the system will filter out the data with a timestamp later than the last synchronization time. Assume that the last synchronization time is 2023-10-15 12:00. Then, "Project Report.docx" is classified into the data group to be synchronized because its modification time is later than this time point. While the modification time of "Historical Archive.docx" is 2023-10-15 11:00, which does not exceed the synchronization time, so it is excluded. The system compares the content of the data group to be synchronized item by item to confirm the specific data range that needs to be transmitted, ensuring that only necessary data is processed.
[0123] Step S530: For the data group to be synchronized, use a compression tool for processing. Pack the data group to be synchronized into a compressed transmission package through a preset compression algorithm, and dynamically adjust it according to the data size and the bandwidth of the terminal device to obtain the final compressed transmission package.
[0124] It should be noted that for the compression processing of the data group to be synchronized, the system will use a preset compression algorithm to generate a compressed transmission package. Assume that the original size of "Project Report.docx" is 5MB. The system compresses it to 2MB through the compression tool and dynamically adjusts the transmission parameters according to the bandwidth situation of the target terminal device. For example, if the terminal bandwidth is low, the system will further fragment the transmission to ensure smooth transmission. This method effectively reduces the transmission burden.
[0125] Step S540: Distribute the final compressed transmission package to the target terminal device through a network transmission tool. When detecting network interruption or latency during transmission, trigger the retransmission mechanism, obtain the transmission status feedback, and determine whether the final compressed transmission package is completely delivered.
[0126] Preferably, during the process of the network transmission tool distributing the compressed transmission package, the system will monitor the network status in real time. Assume that network interruption is detected when transmitting the compressed package of "Project Report.docx". The system will trigger the retransmission mechanism and confirm whether it is completely delivered through the transmission status feedback.
[0127] For example, if the feedback shows that only 80% of the transmission package is delivered, the system will resend the remaining part until the target terminal device confirms complete reception. This mechanism guarantees the reliability of data transmission.
[0128] It can be understood that the above-mentioned links form a complete data synchronization chain from the acquisition of changed data to the final transmission.
[0129] Whether it is classification processing, incremental extraction, or compression and transmission, each step closely focuses on the data synchronization requirements of the document management platform, ensuring the efficient flow of data in the hierarchical storage structure, while taking into account the balance between resource occupation and transmission efficiency.
[0130] The present invention relates to an online document processing system based on a document processing model for implementing the above-mentioned online document processing method based on a document processing model, including a first acquisition module, a recording module, a determination module, a second acquisition module, and a generation module. Among them, the first acquisition module is used to obtain the editing request of the user for the target data unit, generate a temporary lock identifier according to the user permission identifier and the request timestamp, restrict the modification operation of other users on the target data unit, and obtain the conflict isolation result; the recording module is used to construct a data consistency module according to the conflict isolation result, extract the editing status information of the current data unit from the temporary lock identifier, and use the distributed transaction processing technology to sequentially record the concurrent operations; the determination module is used to trigger the rollback mechanism and reassign the operation priority if a version conflict is detected in the sequentially recorded data, and determine the data version after consistency verification; the second acquisition module is used to design a hierarchical cache structure according to the data version after consistency verification, store the frequently accessed data in the memory cache area, and store the infrequently accessed data in the disk cache area to obtain a hierarchical storage structure; the generation module is used to optimize the cross-terminal synchronization process based on the hierarchical storage structure, extract the changed data set using the incremental synchronization technology, and generate a compressed transmission package in combination with a compression algorithm.
[0131] Furthermore, for the online document processing system based on the document processing model provided in this embodiment, the first acquisition module includes a first acquisition unit, a generation unit, a second acquisition unit, a third acquisition unit, and a first determination unit. Among them, the first acquisition unit is used to obtain the user permission identifier and the request timestamp from the database according to the editing request submitted by the user; the generation unit is used to verify the user permission identifier through a pre-established permission verification rule. If the user permission identifier meets the preset access conditions, a temporary lock identifier is generated, and the temporary lock identifier is used to restrict other users from modifying the target data unit; the second acquisition unit is used to obtain the access record of the target data unit from the server log for the temporary lock identifier, and determine whether there is a concurrent editing request from other users. If there is a concurrent editing request, the priority is determined according to the order of the request timestamps to obtain the final lock ownership result; the third acquisition unit is used to update the status of the target data unit by using the distributed lock mechanism according to the final lock ownership result. When updating, the current version number of the target data unit is obtained, and it is determined whether the current version number is consistent with the locked version. If they are consistent, the editing operation is allowed to obtain the lock protection status; the first determination unit is used to obtain the specific content of the editing operation from the operation log through the lock protection status, compare the specific content with the original content of the target data unit, record the difference information by using the version control tool, and determine the final result of conflict isolation.
[0132] Preferably, for the online document processing system based on the document processing model provided in this embodiment, the recording module includes a fourth acquisition unit, a second determination unit, a fifth acquisition unit, and a sixth acquisition unit. Among them, the fourth acquisition unit is used to extract the editing status information of the target data unit from the temporary lock identifier according to the conflict isolation result, compare the editing status information with the operation log, and determine whether there are unfinished concurrent operations. If there are unfinished concurrent operations, they are sorted according to the priority order to obtain a preliminary order record table; the second determination unit is used to divide the concurrent operations into batches by using the distributed transaction processing tool through the preliminary order record table, obtain the status update information for the operation log in each batch, and determine the final editing status of each current data unit in the current batch; the fifth acquisition unit is used to perform version verification for each current data unit according to the final editing status, obtain the historical version information from the pre-established version record. If the historical version information does not match the current status, the corresponding operation is suspended to obtain the operation queue after version verification; the sixth acquisition unit is used to update the status of the current data unit by using the transaction processing tool through the operation queue after version verification, and determine whether the status update meets the data consistency requirements to obtain the final sequential record result.
[0133] Furthermore, for the online document processing system based on the document processing model provided in this embodiment, the determination module includes a seventh acquisition unit, a third determination unit, an eighth acquisition unit, and a ninth acquisition unit. Among them, the seventh acquisition unit is used to obtain the version conflict information of the target operation queue from the operation record, compare the version conflict information with the pre-established historical data, and if it is detected that the conflict data in the version conflict information is inconsistent with the historical data, trigger the rollback mechanism to obtain the initially rolled-back operation queue; the third determination unit is used to reassign the operation priorities according to the initially rolled-back operation queue, using a distributed transaction tool to sort the status update information of the operation record to determine the reallocated priority sequence; the eighth acquisition unit is used to obtain the consistency verification information of each operation record through the reallocated priority sequence, and if the consistency verification information does not match the preset threshold, adjust the status of the relevant operation record to obtain the adjusted operation queue; the ninth acquisition unit is used to finally confirm the data version for the adjusted operation queue, using a log recording tool to compare the verification result with the historical data to determine whether the final data version meets the consistency requirements to obtain the confirmed data version.
[0134] For the online document processing method and system based on the document processing model provided in this embodiment, compared with the prior art, by obtaining the edit request of the user for the target data unit, generating a temporary lock identifier according to the user permission identifier and the request timestamp, restricting the modification operation of other users on the target data unit to obtain the conflict isolation result; constructing a data consistency module according to the conflict isolation result, extracting the edit status information of the current data unit from the temporary lock identifier, and using the distributed transaction processing technology to sequentially record the concurrent operations; if a version conflict is detected in the sequential record, triggering the rollback mechanism and reassigning the operation priorities to determine the data version after consistency verification; designing a hierarchical cache structure according to the data version after consistency verification, storing the frequently accessed data in the memory cache area and the infrequently accessed data in the disk cache area to obtain the hierarchical storage structure; optimizing the cross-terminal synchronization process based on the hierarchical storage structure, using the incremental synchronization technology to extract the changed data set, and generating a compressed transmission packet in combination with the compression algorithm. The beneficial effects obtained by the online document processing method and system based on the document processing model provided in this embodiment are as follows: I. Concurrent Conflict Isolation and Security Control 1. Restrict multi-user concurrent modification through the temporary lock identifier, and combine permission identifier and timestamp verification to ensure the legality of operations and avoid data overwrite or unauthorized access problems.
[0135] 2. Combine the lock mechanism with distributed transactions to achieve fine-grained access control and enhance the security of the system.
[0136] II. Data Consistency Guarantee 1. The distributed transaction processing technology records concurrent operations in sequence, combines version conflict detection and rollback mechanisms to ensure operation atomicity and eventual consistency, and prevent problems such as dirty reads / phanton reads.
[0137] 2. By reallocating version priorities, optimize the transaction processing order and reduce the retry of invalid operations caused by concurrency.
[0138] III. Improvement of Storage and Access Efficiency 1. The hierarchical cache structure dynamically divides the storage areas for high-frequency / low-frequency data, reduces memory occupancy and accelerates the access to hot data, comprehensively improving the response speed and resource utilization rate.
[0139] 2. The combination of incremental synchronization technology and compression algorithms reduces the amount of cross-terminal data transmission, consumes less network bandwidth, and realizes lightweight and efficient synchronization.
[0140] IV. Enhancement of System Reliability and Scalability 1. The rollback mechanism and distributed transactions support the fault recovery ability and improve the system fault tolerance.
[0141] 2. The hierarchical storage and incremental synchronization design adapt to multi-terminal and high-concurrency scenarios, facilitating horizontal expansion and dynamic resource scheduling.
[0142] V. Cost and Maintenance Optimization 1. Automated version management and conflict resolution reduce the need for manual intervention and lower the operation and maintenance complexity.
[0143] 2. Elastic resource allocation (such as on-demand caching and compressed transmission) avoids excessive hardware investment and conforms to the low-cost concept of the Serverless architecture.
[0144] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments as well as all changes and modifications falling within the scope of the present invention. Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. An online document processing method based on a document processing model, characterized in that, It includes the following steps: Obtain the user's edit request for the target data unit, generate a temporary lock identifier based on the user permission identifier and the request timestamp, restrict the modification operations of other users on the target data unit, and obtain the conflict isolation result; Construct a data consistency module according to the conflict isolation result, extract the edit status information of the current data unit from the temporary lock identifier, and use the distributed transaction processing technology to record the concurrent operations in sequence; If a version conflict is detected in the sequenced record, trigger a rollback mechanism and reassign the operation priority to determine the data version after consistency verification; Design a hierarchical cache structure according to the data version after consistency verification, store the frequently accessed data in the memory cache area, and store the infrequently accessed data in the disk cache area to obtain a hierarchical storage structure; Optimize the cross-terminal synchronization process based on the hierarchical storage structure, use the incremental synchronization technology to extract the change data set, and generate a compressed transmission packet in combination with the compression algorithm.
2. The online document processing method based on a document processing model according to claim 1, wherein, The step of obtaining the user's edit request for the target data unit, generating a temporary lock identifier based on the user permission identifier and the request timestamp, restricting the modification operations of other users on the target data unit, and obtaining the conflict isolation result includes: According to the edit request submitted by the user, obtain the user permission identifier and the request timestamp from the database; Verify the user permission identifier through the pre-established permission verification rules. If the user permission identifier meets the preset access conditions, generate a temporary lock identifier, and the temporary lock identifier is used to restrict the modification of the target data unit by other users; For the temporary lock identifier, obtain the access record of the target data unit from the server log, determine whether there is a concurrent edit request from other users. If there is the concurrent edit request, determine the priority according to the sequence of the request timestamps to obtain the final lock ownership result; According to the final lock ownership result, use the distributed lock mechanism to update the status of the target data unit. When updating, obtain the current version number of the target data unit, and determine whether the current version number is consistent with the locked version. If they are consistent, allow the execution of the edit operation to obtain the lock protection status; Through the lock protection status, obtain the specific content of the edit operation from the operation log, compare the specific content with the original content of the target data unit, use the version control tool to record the difference information, and determine the final result of conflict isolation.
3. The online document processing method based on a document processing model according to claim 1, characterized in that, The step of constructing a data consistency module according to the conflict isolation result, extracting the edit status information of the current data unit from the temporary lock identifier, and using the distributed transaction processing technology to record the concurrent operations in sequence includes: According to the conflict isolation result, extract the edit status information of the target data unit from the temporary lock identifier, compare the edit status information with the operation log, determine whether there are unfinished concurrent operations. If there are unfinished concurrent operations, sort them according to the priority order to obtain a preliminary sequence record table; Through the preliminary sequence record table, a distributed transaction processing tool is used to divide the concurrent operations into batches, obtain status update information for the operation logs in each batch, and determine the final editing status of each current data unit in the current batch; According to the final editing status, version verification is performed for each current data unit, historical version information is obtained from the pre-established version record. If the historical version information does not match the current status, the corresponding operation is suspended to obtain an operation queue after version verification; Through the operation queue after version verification, a transaction processing tool is used to update the status of the current data unit, and it is judged whether the status update meets the data consistency requirements to obtain the final sequenced record result.
4. The online document processing method based on a document processing model according to claim 1, wherein If a version conflict is detected in the sequenced record, a rollback mechanism is triggered and the operation priorities are reallocated. The steps to determine the data version after consistency verification include: Obtain the version conflict information of the target operation queue from the operation record, compare the version conflict information with the pre-established historical data. If it is detected that the conflicting data in the version conflict information is inconsistent with the historical data, a rollback mechanism is triggered to obtain an operation queue after initial rollback; According to the operation queue after initial rollback, a distributed transaction tool is used to reallocate the operation priorities, sort the status update information of the operation record, and determine the reallocated priority sequence; Through the reallocated priority sequence, obtain the consistency verification information of each operation record. If the consistency verification information does not match the preset threshold, the status of the relevant operation record is adjusted to obtain an adjusted operation queue; For the adjusted operation queue, a log recording tool is used to finally confirm the data version, compare the verification result with the historical data, and judge whether the final data version meets the consistency requirements to obtain the confirmed data version.
5. The online document processing method based on a document processing model according to claim 1, wherein The steps of designing a hierarchical cache structure according to the data version after consistency verification, storing high-frequency access data in the memory cache area, and storing low-frequency access data in the disk cache area to obtain a hierarchical storage structure include: Obtain access log data from the database, perform statistical processing on the access log data to determine the access frequency value of each access log data, and compare the access frequency value with the preset frequency threshold. If the access frequency value is higher than the frequency threshold, it is classified as high-frequency access data; if the access frequency value is lower than the frequency threshold, it is classified as low-frequency access data to obtain a classified data set; For the classified data set, a memory cache tool is used to store the high-frequency access data. Through the fast read-write interface of the memory cache tool, the high-frequency access data is loaded into the memory cache area to obtain the high-frequency data storage result in the memory cache area, and at the same time obtain the storage path of the low-frequency access data; According to the storage path of the low-frequency access data, a disk storage tool is used to store the low-frequency access data, allocate space in the disk cache area, and write the low-frequency access data into the disk cache area to obtain the low-frequency access data storage result in the disk cache area; By performing consistency verification on the high-frequency data storage results in the memory buffer and the low-frequency access data storage results in the disk buffer, the storage status of the memory buffer and the disk buffer is obtained. If the storage status is abnormal, a data reallocation mechanism is triggered to obtain the final hierarchical storage structure.
6. The online document processing method based on a document processing model according to claim 1, wherein The steps of optimizing the cross-terminal synchronization process based on the hierarchical storage structure, extracting the changed data set using the incremental synchronization technology, and generating a compressed transmission packet in combination with the compression algorithm include: Obtain the changed data set from the hierarchical storage structure, classify the data in the memory buffer and the disk buffer, and determine the specific range of the changed data by comparing the timestamps and modification records to obtain the classified changed data set; Use an incremental extraction tool to process the classified changed data set. When extracting, if the timestamp of the changed data is later than the last synchronization time, the classified changed data set is included in the data group to be synchronized. By comparing item by item, the complete content of the data group to be synchronized is obtained, and the data range to be transmitted is determined; For the data group to be synchronized, use a compression tool to process it, pack the data group to be synchronized into a compressed transmission packet through a preset compression algorithm, and dynamically adjust according to the data size and the bandwidth of the terminal device to obtain the final compressed transmission packet; Distribute the final compressed transmission packet to the target terminal device through a network transmission tool. When transmitting, if a network interruption or delay is detected, a retransmission mechanism is triggered to obtain the transmission status feedback and determine whether the final compressed transmission packet is completely delivered.
7. An online document processing system based on a document processing model, which is used to implement the online document processing method based on the document processing model as described in any one of claims 1 to 6, characterized in that, Including: The first acquisition module is used to obtain the editing request of the user for the target data unit, generate a temporary lock identifier according to the user permission identifier and the request timestamp, and restrict the modification operation of other users on the target data unit to obtain a conflict isolation result; The recording module is used to construct a data consistency module according to the conflict isolation result, extract the editing status information of the current data unit from the temporary lock identifier, and use the distributed transaction processing technology to sequentially record the concurrent operations; The determination module is used to trigger a rollback mechanism and reallocate the operation priority if a version conflict is detected in the sequential record, and determine the data version after consistency verification; The second acquisition module is used to design a hierarchical cache structure according to the data version after consistency verification, store the high-frequency access data in the memory buffer, and store the low-frequency access data in the disk buffer to obtain a hierarchical storage structure; The generation module is used to optimize the cross-terminal synchronization process based on the hierarchical storage structure, extract the changed data set using the incremental synchronization technology, and generate a compressed transmission packet in combination with the compression algorithm.
8. The online document processing system based on a document processing model according to claim 7, wherein The first acquisition module includes: The first acquisition unit is used to obtain the user permission identifier and the request timestamp from the database according to the editing request submitted by the user; The generation unit is used to verify the user permission identifier through a pre-established permission verification rule. If the user permission identifier meets the preset access conditions, a temporary lock identifier is generated, and the temporary lock identifier is used to restrict the modification of the target data unit by other users; A second acquisition unit, configured to obtain, from the server log, the access record of the target data unit for the temporary lock identifier, determine whether there is a concurrent editing request from another user, and if there is such a concurrent editing request, determine the priority through the order of the request timestamps to obtain the final lock ownership result; A third acquisition unit, configured to update the status of the target data unit by using a distributed lock mechanism according to the final lock ownership result, obtain the current version number of the target data unit during the update, and determine whether the current version number is consistent with that at the time of locking. If they are consistent, allow the execution of the editing operation to obtain the lock protection status; A first determination unit, configured to obtain the specific content of the editing operation from the operation log through the lock protection status, perform a difference comparison between the specific content and the original content of the target data unit, record the difference information by using a version control tool, and determine the final result of conflict isolation.
9. The online document processing system based on a document processing model according to claim 7, characterized in that, The recording module includes: A fourth acquisition unit, configured to extract the editing status information of the target data unit from the temporary lock identifier according to the conflict isolation result, compare the editing status information with the operation log, determine whether there are unfinished concurrent operations, and if there are unfinished concurrent operations, sort them according to the priority order to obtain a preliminary order record table; A second determination unit, configured to divide the concurrent operations into batches by using a distributed transaction processing tool through the preliminary order record table, obtain the status update information for the operation log in each batch, and determine the final editing status of each current data unit in the current batch; A fifth acquisition unit, configured to perform a version check on each current data unit according to the final editing status, obtain the historical version information from the pre-established version record, and if the historical version information does not match the current status, suspend the corresponding operation to obtain an operation queue after version check; A sixth acquisition unit, configured to update the status of the current data unit by using a transaction processing tool through the operation queue after version check, and determine whether the status update meets the data consistency requirement to obtain the final sequential record result.
10. The online document processing system based on a document processing model according to claim 7, characterized in that The determination module includes: A seventh acquisition unit, configured to obtain the version conflict information of the target operation queue from the operation record, compare the version conflict information with the pre-established historical data, and if it is detected that the conflict data in the version conflict information is inconsistent with the historical data, trigger a rollback mechanism to obtain an operation queue after initial rollback; A third determination unit, configured to reassign the operation priorities by using a distributed transaction tool according to the operation queue after initial rollback, sort the status update information of the operation record, and determine the reallocated priority sequence; An eighth acquisition unit, configured to obtain the consistency check information of each operation record through the reallocated priority sequence, and if the consistency check information does not match the preset threshold, adjust the status of the relevant operation record to obtain an adjusted operation queue; A ninth acquisition unit is configured to, for the adjusted operation queue, use a logging tool to perform a final confirmation on the data version, compare the verification result with historical data, determine whether the final data version meets the consistency requirements, and obtain the confirmed data version.
Citation Information
Patent Citations
Word document online collaborative editing method and system
CN108269063A
Collaborative document management service method and system
CN114064568A
Data processing method and device, storage medium and program product
CN118069665A
Extremely cold object storage system and access and management method
CN119311699A
WebSocket and JSON-PATCH-based conflict resolution method and system for realizing multi-person collaborative editing
CN119902754A
Cited By
Data security sharing and management method and system
CN120729639A
Dynamic workflow collaboration method and related equipment
CN121193754A
Page editing locking method, system and equipment and computer readable storage medium
CN121255487A
Page editing locking method, system, device and computer readable storage medium
CN121255487B
Data synchronization method and device for cloud document management system
CN121478738A