Document data processing method and device, computer equipment and storage medium

By splitting the document data into attribute pools and index data, and generating read replicas in parallel processing, the problem that online documents cannot be executed concurrently in editing and typography in large-scale user collaboration scenarios is solved, and processing performance and response speed are improved.

CN120257945APending Publication Date: 2025-07-04TENCENT TECH WUHAN
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410007844.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the scenario of large-scale user collaboration, online documents have frequent editing changes, resulting in the inability to execute editing and typesetting concurrently, resulting in high memory usage and slow response speed.

Method used

By splitting the document data into attribute pools and index data, a read replica is generated in parallel processing, and parallel execution of editing and typesetting is realized.

Benefits of technology

In the large-scale user collaboration scenario, document data processing performance and editing response speed are improved, and memory usage and time overhead are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257945A_ABST
    Figure CN120257945A_ABST
Patent Text Reader

Abstract

The invention relates to a document data processing method and device, computer equipment and a storage medium. The method can be applied to an application scene for processing document data in a vehicle-mounted terminal, a cloud server, a block chain node or other equipment, and comprises the following steps: in response to an editing operation executed on a target document by a collaborator, obtaining an attribute tree corresponding to each document element in the target document and index data corresponding to the attribute tree; modifying at least one part of data in the index data to obtain modified index data; generating a copy used for typesetting the target document based on the modified index data; and typesetting the target document based on the copy, and editing the target document in parallel. By adopting the method, the processing performance of the document data can be effectively improved, and meanwhile, the response speed of the document editing operation is also improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technologies, and in particular, to a method, apparatus, computer device, and storage medium for processing document data. Background Art

[0002] With the development of computer technologies and Internet technologies, collaborative office work has become increasingly common. In the field of collaborative office work, users have a strong demand for completing document compilation by using online documents. An online document is a technology that supports multi-user collaborative processing of documents that can be created and edited anytime and anywhere. With the support of online document technologies, users can edit and share the same document on multiple platforms, and users can also edit and share online documents on various device terminals such as mobile phones and tablet computers. Online documents can also support multiple people to view and edit simultaneously. An online document is a document processing method based on network technologies. The emergence of online documents has greatly expanded the communication and collaboration capabilities among users.

[0003] However, in the current methods for processing document data, since online documents can support multiple users to edit and process them, in the scenario of large-scale user collaboration, the editing of documents changes frequently, and multiple document copies will cause high memory occupancy and additional time overhead, which are likely to cause performance problems such as document editing lags, and thus lead to poor processing performance of document data. Summary of the Invention

[0004] Based on this, to solve the above technical problems, it is necessary to provide a method, apparatus, computer device, and storage medium for processing document data, which can effectively improve the processing performance of document data and also improve the response speed of document editing operations.

[0005] In a first aspect, the present application provides a method for processing document data. The method includes: in response to an editing operation performed by a collaborating party on a target document, obtaining an attribute tree corresponding to each document element in the target document and index data corresponding to the attribute tree; modifying at least a part of the data in the index data to obtain modified index data; generating a copy for typesetting the target document based on the modified index data; and typesetting the target document based on the copy and simultaneously performing document editing on the target document.

[0006] Second aspect, the present application also provides an apparatus for processing document data. The apparatus includes: an acquisition module, configured to acquire an attribute tree corresponding to each document element in the target document and index data corresponding to the attribute tree in response to an editing operation performed by a collaborating party on the target document; a modification module, configured to modify at least a part of the data in the index data to obtain modified index data; a generation module, configured to generate a copy for typesetting the target document based on the modified index data; and a processing module, configured to perform typesetting on the target document based on the copy and perform document editing on the target document in parallel.

[0007] Third aspect, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented: acquiring an attribute tree corresponding to each document element in the target document and index data corresponding to the attribute tree in response to an editing operation performed by a collaborating party on the target document; modifying at least a part of the data in the index data to obtain modified index data; generating a copy for typesetting the target document based on the modified index data; and performing typesetting on the target document based on the copy and performing document editing on the target document in parallel.

[0008] Fourth aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented: acquiring an attribute tree corresponding to each document element in the target document and index data corresponding to the attribute tree in response to an editing operation performed by a collaborating party on the target document; modifying at least a part of the data in the index data to obtain modified index data; generating a copy for typesetting the target document based on the modified index data; and performing typesetting on the target document based on the copy and performing document editing on the target document in parallel.

[0009] Fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented: acquiring an attribute tree corresponding to each document element in the target document and index data corresponding to the attribute tree in response to an editing operation performed by a collaborating party on the target document; modifying at least a part of the data in the index data to obtain modified index data; generating a copy for typesetting the target document based on the modified index data; and performing typesetting on the target document based on the copy and performing document editing on the target document in parallel.

[0010] The method, apparatus, computer device, storage medium, and computer program product for processing the above-mentioned document data obtain the attribute tree corresponding to each document element in the target document and the index data corresponding to the attribute tree by responding to the editing operation performed by the collaborating party on the target document; modify at least a part of the data in the index data to obtain the modified index data; generate a copy for typesetting the target document based on the modified index data; typeset the target document based on the copy, and concurrently perform document editing on the target document. Since each document element in the target document can be expressed as a corresponding attribute tree, in response to the editing operation performed by the collaborating party on the target document, the attribute tree corresponding to each document element in the target document and the index data corresponding to the attribute tree can be directly obtained quickly and accurately, and at least a part of the data in the index data can be modified to obtain the modified index data. Furthermore, a copy for typesetting the target document can be generated more quickly and accurately based on the modified index data, enabling a read-only copy to be generated more efficiently with extremely low memory occupancy, so that the editing process and the typesetting process can be executed concurrently, and high document editing performance can still be achieved in the scenario of large-scale user collaboration with a large amount of document data. That is, while effectively improving the processing performance of document data, the response speed of document editing operations is also effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 FIG. is an application environment diagram of the method for processing document data in an embodiment;

[0012] Figure 2 FIG. is a flowchart of the method for processing document data in an embodiment;

[0013] Figure 3 FIG. is a document diagram of large-scale collaborative editing of a document in an embodiment;

[0014] Figure 4 FIG. is a flowchart of the process when an actual modification is made to a certain attribute tree in an embodiment;

[0015] Figure 5 FIG. is a flowchart of the process during online document collaborative editing in an embodiment;

[0016] Figure 6 FIG. is a design diagram of the data model of a document in an embodiment;

[0017] Figure 7 FIG. is a flowchart of the steps for constructing each attribute tree based on the document elements in a sample document in an embodiment;

[0018] Figure 8 FIG. is a flowchart of the steps for responding to the editing operation performed by other collaborating parties on the target document in an embodiment;

[0019] Figure 9 It is a structural block diagram of a processing device for document data in an embodiment;

[0020] Figure 10 It is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0021] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0022] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing.

[0023] Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used as needed, and is flexible and convenient. Cloud computing technology will become an important support. The services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the highly developed application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the system for logical processing. Data at different levels will be processed separately, and various industry data requires a powerful system back-end support, which can only be achieved through cloud computing.

[0024] Cloud storage is a new concept extended and developed from the cloud computing concept. A distributed cloud storage system (hereinafter referred to as the storage system) refers to a storage system that combines a large number of different types of storage devices (storage devices are also called storage nodes) in the network through cluster applications, grid technology, and distributed file systems, and collaborates through application software or application interfaces to jointly provide data storage and business access functions to the outside world.

[0025] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0026] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0027] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, vehicle networking, autonomous driving, intelligent transportation, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0028] Model parallel computing: It refers to distributing the computing tasks of the model to multiple computing devices (such as CPUs, GPUs, TPUs, etc.) for simultaneous computing, thereby accelerating the training and inference of the model. Model parallel computing can effectively utilize computing resources and improve the computing efficiency and training speed of the model.

[0029] It should be noted that in the following description, the terms "first, second, and third" only distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first, second, and third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0030] The method for processing document data provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set separately, integrated on the server 104, or placed on the cloud or other network servers. The device used by the collaborator can be the terminal 102. In response to the editing operation performed by the collaborator on the target document, the terminal 102 can obtain the attribute tree corresponding to each document element in the target document and the index data corresponding to the attribute tree from the server 104. Or, in response to the editing operation performed by the collaborator on the target document, the terminal 102 can also obtain the attribute tree corresponding to each document element in the target document and the index data corresponding to the attribute tree from the local, and modify at least a part of the data in the index data to obtain the modified index data; further, the terminal 102 can generate a copy for typesetting the target document based on the modified index data; the terminal 102 can typeset the target document based on the copy and perform document editing on the target document in parallel.

[0031] Among them, the terminal 102 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart TV, a smart watch, an Internet of Things device, and a portable wearable device. The Internet of Things device can be a smart vehicle device, etc. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc.

[0032] The server 104 can be an independent physical server or a service node in a blockchain system. A peer-to-peer (Peer To Peer) network is formed among the service nodes in the blockchain system. The Peer To Peer protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP).

[0033] In addition, the server 104 can also be a server cluster composed of multiple physical servers, and can be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0034] The terminal 102 and the server 104 can be connected through communication connection methods such as Bluetooth, USB (Universal Serial Bus), or the network. This application does not make any restrictions here.

[0035] In one embodiment, as Figure 2As shown, a method for processing document data is provided. This method can be executed independently by a server or a terminal, or jointly by a server and a terminal. Taking the terminal in Figure 1 as an example for illustration, the method includes the following steps:

[0036] Step 202: In response to an editing operation performed by a collaborating party on a target document, obtain the attribute tree corresponding to each document element in the target document and the index data corresponding to the attribute tree.

[0037] The collaborating party refers to an object that can perform an editing operation on the target document. In this application, the collaborating party can include different collaborating users. For example, the collaborating party in this application can include collaborating user 1, collaborating user 2, and collaborating user 3.

[0038] The target document refers to a specific document. For example, the target document in this application can be an online document. An online document refers to a cloud document that can be simultaneously edited and collaborated on by multiple users, and the edited content can be displayed and saved in real time.

[0039] The editing operation refers to an editing operation triggered by the collaborating party on the content of the target document. For example, the editing operations in this application include but are not limited to deletion operations, modification operations, update operations, etc.

[0040] The document element refers to the basic element included in the target document. For example, if the target document is Document 1, the document elements included in the document structure of Document 1 include but are not limited to: table elements, image elements, text elements, etc. It can be understood that the document element here can refer to the identifier of the document element. In this application, the document element can also be called an element, and the element is used to describe the specific content in the document. For example, a certain element is used to describe a text paragraph, that is, the text data within a certain range belongs to the content corresponding to this element.

[0041] The attribute tree refers to a tree structure constructed with document elements as nodes. It can be understood as expressing each document element in the target document as an attribute tree in tree structure form. A document can contain multiple attribute trees. By accessing the nodes in the attribute tree, the elements described by each node and the attribute data corresponding to the elements can be obtained.

[0042] The index data refers to the index data in each attribute tree included in the target document. In this application, the index data is used to reflect the mapping relationship between the nodes in the attribute tree and the attribute data in the attribute data pool. For example, the index data in this application can be: obtained by combining the node index data stored in each node in the attribute tree.

[0043] Specifically, when a user wants to edit a certain document, the user can open the document application (Application, APP) on the terminal through a triggering operation, and enter the target document page of the document application through a selection operation, that is, the user can log in to the document application through the triggering operation. Further, in the main interface displayed by the document application, the user can enter the target document page through a triggering operation. In the target document page displayed on the terminal, each collaboration party (collaborative user) can view the specific content and relevant information of the target document, and each collaboration party can also trigger an editing operation on the content of the displayed target document. Then, the terminal responds to the editing operation performed by the collaboration party on the target document, performs a document editing operation on the target document, and can obtain the modified target document, and uses the modified target document as the first version of the document. That is, in response to the editing operation performed by the collaboration party on the target document, the terminal can obtain the attribute tree corresponding to each document element in the target document, and obtain the index data corresponding to each attribute tree, so that the terminal can modify the index data corresponding to the editing operation subsequently, quickly and accurately obtain the modified index data, and use the modified index data as the first version of the document.

[0044] For example, take the target document as an online document for illustration. As Figure 3 shown, it is a document schematic diagram of large-scale collaborative document editing. Assume the target document is the Figure 3 online document A for collaborative editing shown in. When collaborative user 1 wants to edit the online document A, collaborative user 1 can open the document application on the terminal through a triggering operation, and enter the page of the online document A in the document application through a selection operation. In the page of the online document A shown in Figure 3 on the terminal, the collaboration party, that is, collaborative user 1, can trigger an editing operation on the content of the displayed online document A. Then, in response to the editing operation performed by collaborative user 1 on the online document A, the terminal can obtain the attribute tree corresponding to each document element in the online document A, and obtain the node index data corresponding to the nodes in each attribute tree. For example, as Figure 4 shown, it is a process schematic diagram when a certain attribute tree is actually modified. That is, in response to the editing operation performed by collaborative user 1 on the online document A, the terminal can obtain one of the attribute trees corresponding to each element in the online document A, which is the original attribute tree shown on the left in Figure 4 , and the index data corresponding to the original attribute tree is the index data stored in segments in the actual index memory (i.e., the dotted box) shown below the original attribute tree in Figure 4 .

[0045] Step 204, modify at least a part of the data in the index data to obtain modified index data.

[0046] Among them, the modified index data refers to the index data obtained after modifying the index data corresponding to the attribute tree. For example, the modified index data in this application can be Figure 4 the respective node index data stored in segments in the actual index memory (i.e., the dashed box) shown below the edited attribute tree shown in

[0047] Specifically, after the terminal obtains the attribute tree corresponding to each document element in the target document and the index data corresponding to each attribute tree in response to the editing operation performed by the collaborating party on the target document, the terminal can sequentially copy the index data corresponding to each attribute tree according to the storage order, and modify at least a part of the data in the copied index data corresponding to each attribute tree, thereby obtaining the modified index data. For example, as Figure 4 shown, it is a schematic flow diagram of actual modification of a certain attribute tree. Assume that the terminal responds to the editing operation performed by the collaborating user 1 on the online document A, and this editing operation can be an operation of inserting characters. One of the attribute trees obtained by the terminal for each element in the online document A is Figure 4 the original attribute tree shown on the left in Figure 4 and the index data corresponding to this original attribute tree is the index data stored in segments in the actual index memory (i.e., the dashed box) shown below the original attribute tree in Figure 4 After that, the terminal can sequentially copy the index data corresponding to the original attribute tree according to the storage order, and obtain the index data stored in segments shown on the right in Figure 4 ; Further, the terminal can modify the index data corresponding to the copied attribute tree (the data of two nodes in the original attribute tree has changed), that is, the terminal modifies the node index data stored in the first node and the last node of the first segment in the index data corresponding to the copied attribute tree shown in Figure 4 That is, modify the index value stored in the first node of the first segment: from 18 to 21, and at the same time modify the index value stored in the last node of the first segment: from 9 to 12, to obtain the modified index data as:

[0048] Step 206, generate a copy for typesetting the target document based on the modified index data.

[0049] Among them, a copy refers to a read-only copy of the target document for typesetting access. For example, the generation method of the copy in this application can be: copying the modified index data to obtain the copied index data, and using the copied index data as the copy for typesetting the target document.

[0050] Specifically, after the terminal modifies at least a part of the data in the index data corresponding to each attribute tree to obtain the modified index data, the terminal can generate a copy for typesetting the target document based on the modified index data. For example, the terminal can copy the modified index data to obtain the copied index data, and use the copied index data as the copy for typesetting the target document.

[0051] For example, assume that the terminal responds to the editing operation performed by the collaborative user 1 on the online document A. The editing operation can be an operation of inserting characters. The terminal obtains the attribute tree corresponding to each document element in the online document A and the index data corresponding to the attribute tree as Figure 4 shown. The terminal modifies at least a part of the data in the index data corresponding to the original attribute tree shown on the left in Figure 4 to obtain the modified index data shown on the right in Figure 4 , that is, the modified index data is: Figure 4 the index data stored in segments in the actual index memory (i.e., within the dotted box) below the edited attribute tree shown in Figure 4 on the right; further, the terminal can copy the modified index data within the dotted box shown on the right in

[0052] Step 208, typesetting the target document based on the copy and performing document editing on the target document in parallel.

[0053] Among them, typesetting refers to the operation of performing typesetting processing and rendering processing on the target document.

[0054] Specifically, after the terminal generates a copy for typesetting the target document based on the modified index data, the terminal can typeset the target document based on the generated copy and perform document editing on the target document in parallel. For example, the terminal can typeset the target document based on the generated copy through the second thread to obtain the first visual document of the target document. At the same time, the terminal can also obtain the index data corresponding to each document element in the target document through the first thread and perform document editing on the index data to obtain the first version document of the target document. That is, the first thread in this application can be an independent thread for performing document editing, and the second thread can be an independent thread for performing document typesetting.

[0055] For example, as Figure 5 shown, it is a schematic diagram of the process during online document collaborative editing. As Figure 5 shown in, if three collaborative users perform editing operations on the same document, the terminal can, based on a preset policy, sequentially respond to the editing operations of each collaborative user. For example, the terminal's sequential response to the editing operations of each collaborative user can be in the order shown in Figure 5 : Collaborative User 1 → Collaborative User 2 → Collaborative User 3. Suppose the terminal responds to the editing operation performed by Collaborative User 1 on online document A. This editing operation can be an operation of inserting characters. The terminal obtains the attribute tree corresponding to each document element in online document A and the index data corresponding to this attribute tree through the first thread, as Figure 4 shown. The terminal modifies at least a part of the data in the index data corresponding to the original attribute tree shown on the left in Figure 4 through the first thread, and after obtaining the modified index data shown on the right in Figure 4 , the terminal can use the obtained modified index data as document data version 1 of online document A.

[0056] Furthermore, the terminal can copy, through the first thread, Figure 4 the modified index data within the dashed box shown on the right in, that is, copy document data version 1, to obtain a read-only copy 1 of the document, and use the read-only copy 1 of the document as the copy for typesetting online document A. The terminal can perform typesetting and rendering processing on online document A based on this read-only copy 1 of the document through the second thread, and thus obtain visual version 1 of online document A. In addition, during the process of the terminal performing typesetting and rendering processing on online document A based on this read-only copy 1 of the document through the second thread, the terminal can also continue to perform document editing on online document A through the first thread. That is, during the process of the terminal performing typesetting and rendering processing on online document A based on this read-only copy 1 of the document through the second thread, the terminal can respond to the editing operation performed by Collaborative User 2 on online document A, and obtain the attribute tree corresponding to each document element in document data version 1 of online document A and the index data corresponding to this attribute tree through the first thread. The terminal modifies at least a part of the data in the index data corresponding to the attribute tree, and after obtaining the modified index data, the terminal can use the obtained modified index data as document data version 2 of online document A. The terminal can copy document data version 2 through the first thread to obtain a read-only copy 2 of the document, and perform typesetting and rendering processing on online document A based on this read-only copy 2 of the document through the second thread or the third thread (another new thread), and thus obtain visual version 2 of online document A.

[0057] In this embodiment, by responding to the editing operation performed by a collaborating party on a target document, the attribute tree corresponding to each document element in the target document and the index data corresponding to the attribute tree are obtained; at least a part of the data in the index data is modified to obtain modified index data; a copy for typesetting the target document is generated based on the modified index data; the target document is typeset based on the copy, and document editing of the target document is performed in parallel. Since each document element in the target document can be expressed as a corresponding attribute tree, in response to the editing operation performed by the collaborating party on the target document, the attribute tree corresponding to each document element in the target document and the index data corresponding to the attribute tree can be directly obtained quickly and accurately, and at least a part of the data in the index data is modified to obtain modified index data. Furthermore, a copy for typesetting the target document can be generated more quickly and accurately based on the modified index data, enabling a read-only copy to be generated more efficiently with extremely low memory occupancy, such that the editing process and the typesetting process can be executed concurrently, and high document editing performance is still maintained in the scenario of large-scale user collaboration with a large amount of document data. That is, while effectively improving the processing performance of document data, the response speed of document editing operations is also effectively increased.

[0058] In one embodiment, the step of obtaining the attribute tree corresponding to each document element in the target document and the index data corresponding to the attribute tree includes:

[0059] Query the attribute tree corresponding to each document element in the target document;

[0060] Based on the tree identifier of each attribute tree, obtain the node index data stored in the nodes of each attribute tree;

[0061] Combine the node index data stored in the nodes of each attribute tree into the index data corresponding to the attribute tree.

[0062] Among them, the tree identifier is used to identify a unique attribute tree, and each node in the attribute tree stores node index data. For example, the node index data in this application may further include position index data (such as a position index value) and attribute index data (such as an attribute index value). The position index data is used to reflect the index position of the node index data corresponding to each node in the attribute tree in the segmented storage index memory; the attribute index data is used to reflect the index position of the attribute data corresponding to each node in the attribute tree in the attribute data pool.

[0063] Specifically, taking the target document as an online document as an example for illustration. Assume the target document is Figure 3 the collaborative editing online document A shown in Figure 3In the page of the online document A shown, the collaborating party, i.e., the collaborating user 1, can trigger an editing operation on the content of the displayed online document A. Then, in response to the editing operation performed by the collaborating user 1 on the online document A, the terminal obtains the attribute tree corresponding to each document element in the online document A, and obtains the node index data corresponding to the nodes in each attribute tree. For example, as Figure 6 shown, it is a design schematic diagram of the data model of the document, that is, in response to the editing operation performed by the collaborating user 1 on the online document A, the terminal can obtain that the attribute trees corresponding to each element in the online document A are as Figure 6 shown in three attribute trees, and the index data corresponding to the three attribute trees respectively are as Figure 6 the node index data stored in the form of a continuous memory segment shown below each attribute tree. That is, in response to the editing operation performed by the collaborating user 1 on the online document A, the terminal can query that the attribute trees included in the online document A include Figure 6 the attribute tree 1, attribute tree 2, and attribute tree 3 shown, and based on the tree identifiers of each attribute tree (attribute tree 1, attribute tree 2, and attribute tree 3), obtain the node index data stored in the nodes of each attribute tree as Figure 6 shown, and combine the node index data stored in the nodes of each attribute tree into the index data corresponding to the attribute tree, that is Figure 6 the three segments of node index data corresponding to the three attribute trees below are combined into the index data corresponding to the online document A, Figure 6 and the attribute data corresponding to each element in the online document A is stored in the attribute data pool below the node index data of each attribute tree.

[0064] In this embodiment, by proposing a brand-new method for organizing the document data model, the document data based on this method can achieve fast copying of read-only copies, and the memory occupancy is extremely small, enabling editing and typesetting to be executed concurrently, and still having high document editing performance in the scenario of large-scale user collaboration with a large amount of document data.

[0065] In one of the embodiments, the attribute tree includes a table attribute tree, an image attribute tree, and a text attribute tree; the step of obtaining the node index data stored in the nodes of each attribute tree based on the tree identifier of each attribute tree includes:

[0066] Based on the tree identifier of the table attribute tree, obtain the table node index data stored in the nodes of the table attribute tree;

[0067] Based on the tree identifier of the image attribute tree, obtain the image node index data stored in the nodes of the image attribute tree;

[0068] Based on the tree identifier of the text attribute tree, obtain the text node index data stored in the nodes of the text attribute tree;

[0069] The step of combining the node index data stored in the nodes of each attribute tree into the index data corresponding to the attribute tree includes:

[0070] Combining the table node index data, the image node index data, and the text node index data into the index data corresponding to the attribute tree.

[0071] Among them, the table attribute tree, the image attribute tree, and the text attribute tree in this application are only used to distinguish different types of attribute trees; the table node index data, the image node index data, and the text node index data are only used to distinguish different types of node index data. It can be understood that the attribute trees included in the document in this application include but are not limited to: the table attribute tree, the image attribute tree, and the text attribute tree. For example, it can also include the table row attribute tree, the table cell attribute tree, the paragraph attribute tree, the chapter attribute tree, the annotation attribute tree, etc., which are not specifically limited in this application.

[0072] Specifically, as Figure 6 shown, it is a design schematic diagram of the data model of the document. In response to the editing operation performed by the collaborative user 1 on the online document A, the terminal can obtain that the attribute trees corresponding to the elements in the online document A are the three attribute trees as Figure 6 shown, and the index data corresponding to the three attribute trees respectively are the node index data stored in the form of continuous memory segments shown below each attribute tree as Figure 6 shown. That is, in response to the editing operation performed by the collaborative user 1 on the online document A, the terminal can query that the attribute trees included in the online document A include Figure 6 the attribute tree 1 (i.e., the table attribute tree), the attribute tree 2 (i.e., the image attribute tree), and the attribute tree 3 (i.e., the text attribute tree) shown as Figure 6 shown, and based on the tree identifiers of each attribute tree (attribute tree 1, attribute tree 2, and attribute tree 3), obtain the node index data stored in the nodes of each attribute tree as Figure 6 shown, and combine the node index data stored in the nodes of each attribute tree into the index data corresponding to the attribute tree, that is, Figure 6 combine the three segments of node index data respectively corresponding to the three attribute trees below into the index data corresponding to the online document A. Figure 6 The attribute data corresponding to each element in the online document A is stored in the attribute data pool below the node index data of each attribute tree. Thus, by splitting the document data into an attribute pool and index data, the attribute data in the attribute pool describes the attributes such as text, paragraphs, and pictures at a specific location in a document, and the index data records the index of the attributes at a specific location in the document in the attribute pool and is stored continuously in segments, that is, a method for efficiently expressing document data is provided. This method can be applied to online document applications, enabling document editing and document layout to be processed concurrently, and thus effectively improving the editing performance of online document applications.

[0073] In one embodiment, the index data includes first position index data and second position index data; the first position index data and the second position index data are stored in consecutive segments;

[0074] The step of modifying at least a part of the index data to obtain the modified index data includes:

[0075] Copy the first position index data and the second position index data in sequence according to the storage order to obtain the copied first index data and the copied second index data;

[0076] In the attribute tree corresponding to each document element in the target document, query the target node corresponding to the editing operation;

[0077] When the position index value corresponding to the target node belongs to the first position index data, modify the copied first index data to obtain the modified first index data;

[0078] Use the modified first index data and the copied second index data as the modified index data.

[0079] Among them, the first position index data and the second position index data are only used to distinguish index data located in different storage positions. For example, the first position index data in this application can be Figure 4 the first segment of index data in the segmented storage index memory as shown in Figure 4 and the second position index data can be the second to third segments of index data in the segmented storage index memory as shown in

[0080] Specifically, as shown in Figure 4 is the flow diagram of the actual modification of a certain attribute tree. Assume that the terminal responds to the editing operation performed by the collaborative user 1 on the online document A, and this editing operation can be an operation of inserting characters. The terminal obtains one of the attribute trees corresponding to each element in the online document A as Figure 4 the original attribute tree shown on the left in Figure 4 and obtains that the index data corresponding to this original attribute tree is the index data stored in segments in the index memory (i.e., within the dotted box) shown below the original attribute tree in Figure 4 After that, the terminal can copy the index data (including the first position index data and the second position index data) corresponding to the original attribute tree in sequence according to the storage order, and thus obtain the copied index data stored in segments (including the copied first index data and the copied second index data) shown on the right in Figure 4In the original attribute tree shown in the figure, the target nodes corresponding to the edit operations triggered by the collaborative user 1 are node 1 and node 5. Since the position index values corresponding to node 1 and node 5 are located in Figure 4 the first segment position in the index memory stored in segments as shown in the figure, that is, the position index values corresponding to node 1 and node 5 belong to the first position index data. Therefore, the terminal can, for example, Figure 4 modify the node index data stored in the first node (node 1) and the last node (node 5) in the first segment of the copied index data as shown in the figure. That is, modify the index value stored in the first node (node 1) of the first segment: from 18 to 21, and at the same time modify the index value stored in the last node (node 5) of the first segment: from 9 to 12. Then the modified index data can be obtained as: Figure 4 the modified index data stored in segments in the index memory (i.e., within the dashed box) shown below the edited attribute tree in the figure. That is, the modified index data in the embodiment of the present application includes three segments of node index data stored in segments. The terminal copies the three segments of index data corresponding to the original attribute tree as a whole, and only modifies the first segment of index data (i.e., the first position index data) among the three segments of index data corresponding to the original attribute tree, and combines the modified first segment of index data and the copied second and third segments of index data into the modified index data. Thus, when copying document data once, only the continuous memory segment where the node is located needs to be copied, the copied content is small and the speed is fast, and the memory actually occupied by the copied data only increases the memory occupied by the node index. The actual memory occupation is small. That is, after the document editing application, a copy can be quickly copied, effectively reducing the time overhead, enabling the data editing and typesetting processing to be fully parallel, and the response speed of the document editing operation is also faster.

[0081] In one embodiment, the steps of generating a copy for typesetting the target document based on the modified index data include:

[0082] Copy the modified index data to obtain the copied index data;

[0083] Use the copied index data as the copy for typesetting the target document.

[0084] Specifically, assume that the terminal responds to the edit operation performed by the collaborative user 1 on the online document A. This edit operation can be an operation of inserting characters. The terminal obtains the attribute tree corresponding to each document element in the online document A and the index data corresponding to this attribute tree as Figure 4 shown in the figure. The terminal modifies at least a part of the data in the index data corresponding to the original attribute tree shown on the left in Figure 4 the figure to obtain the modified index data as shown on the right in Figure 4 the figure. That is, the modified index data is: Figure 4The index data stored in segments in the actual index memory (i.e., within the dashed box) shown below the edited property tree in the Chinese version; further, the terminal can copy Figure 4 The modified index data within the dashed box shown on the right in the Chinese version to obtain the copied index data, and use the copied index data as a copy for typesetting the target document. That is, the data in the copy generated in this application is the modified index data. This enables, after the document editing is completed, only by using the method of whole-segment copying for all index segments, without copying any real attribute data, a read-only copy of the document for typesetting access can be generated. Based on the generation of this read-only copy, editing and typesetting can be executed in parallel, thereby effectively improving the processing performance of document data, and still having relatively high document editing performance even in the scenario of large-scale user collaboration with a large amount of document data.

[0085] In one embodiment, the index data is used to reflect the mapping relationship between each node in the property tree and the attribute data in the attribute data pool; the steps of typesetting the target document based on the copy include:

[0086] Based on the index data in the copy, obtain the attribute data corresponding to each document element from the attribute data pool;

[0087] Typeset the target document based on the attribute data.

[0088] Among them, the attribute data pool is used to store the attribute data corresponding to each document element. For example, the attribute data pool in this application can be as Figure 4 shown in the Chinese version.

[0089] Specifically, assume that the terminal responds to the editing operation performed by collaborative user 1 on the online document A. This editing operation can be an operation of inserting characters. The terminal obtains the property tree corresponding to each document element in the online document A and the index data corresponding to this property tree through the first thread as Figure 4 shown in the Chinese version. The terminal modifies at least a part of the data in the index data corresponding to the original property tree shown on the left in the Chinese version through the first thread, and uses the obtained modified index data as the document data version 1 of the online document A. Figure 4

[0090] Figure 4 Further, the terminal can copy through the first thread Figure 4 the modified index data within the dashed box shown on the right in the Chinese version, that is, copy the document data version 1, to obtain the document read-only copy 1, and use the document read-only copy 1 as a copy for typesetting the online document A. That is, the terminal can, through the second thread, based on the index data in the document read-only copy 1, from the Chinese version as Figure 4Retrieve the specific attribute data corresponding to each document element from the attribute data pool shown, and perform typesetting and rendering processing on the online document A based on the retrieved attribute data, then the visual version 1 of the online document A can be obtained. Thus, by proposing a brand-new method for organizing document data models, the document data based on this method can achieve fast copying of read-only copies with extremely small memory occupancy, enabling editing and typesetting to be executed concurrently, and still having high document editing performance in scenarios of large-scale user collaboration with a large amount of document data.

[0091] In one embodiment, the index data is used to reflect the mapping relationship between each node in the attribute tree and the attribute data in the attribute data pool; the step of typesetting the target document based on the copy includes:

[0092] Based on the index data and change description information in the copy, retrieve the attribute data corresponding to each document element from the attribute data pool, and retrieve the document data corresponding to each document element from the text pool;

[0093] Typeset the target document based on the attribute data and the document data.

[0094] Among them, the change description information is information used to describe which node index data in the index data has changed. For example, the change description information is: the node index data stored in node 1 and node 5 has changed.

[0095] Specifically, assume that the terminal responds to the editing operation performed by collaborative user 1 on the online document A. This editing operation can be an operation of inserting characters. The terminal obtains the attribute tree corresponding to each document element in the online document A and the index data corresponding to this attribute tree through the first thread as Figure 4 shown, and the terminal modifies at least a part of the data in the index data corresponding to the original attribute tree shown on the left in Figure 4 through the first thread, and uses the obtained modified index data as the document data version 1 of the online document A.

[0096] Furthermore, the terminal can copy Figure 4 the modified index data within the dashed box shown on the right in Figure 4 through the first thread, that is, copy the document data version 1 to obtain the document read-only copy 1, and use the document read-only copy 1 as the copy for typesetting the online document A. That is, the terminal can, through the second thread, based on the index data and change description information in the document read-only copy 1, retrieve the specific attribute data corresponding to each document element from the attribute data pool shown in

[0097] In addition, in some cases, the terminal does not need to perform global typesetting processing. Instead, the terminal can perform typesetting processing on the changed parts. That is, the terminal in the embodiments of the present application can also, through the second thread, based on the index data and change description information in the read-only copy 1 of the document, obtain the specific attribute data corresponding to the changed nodes (i.e., changed elements) from the attribute data pool as shown in Figure 4 and obtain the document data corresponding to the changed nodes (i.e., changed elements) from the text pool; the terminal performs typesetting and rendering processing on the online document A based on the obtained attribute data and document data corresponding to the changed elements, and thus can obtain the visual version 1 of the online document A. This enables concurrent processing of document editing and document typesetting, effectively improving the editing performance of the online document application.

[0098] In one embodiment, before obtaining the attribute tree corresponding to each document element in the target document and the index data corresponding to the attribute tree in response to the editing operation performed by the collaborating party on the target document, the method further includes:

[0099] Extracting the document elements in the sample document;

[0100] Constructing each attribute tree based on the document elements in the sample document;

[0101] Obtaining the attribute data corresponding to each document element according to the node index data of each node in each attribute tree;

[0102] Storing the attribute data corresponding to each document element in the attribute data pool.

[0103] Wherein, the sample document refers to a document pre-imported for constructing the attribute tree structure.

[0104] Specifically, as shown in Figure 6 , for the design schematic diagram of the data model of the document, the user can use the online document B as the sample document and import the sample document into the document application program. Then, in response to the import operation performed by the user on the online document B, the terminal can extract all the document elements included in the online document B and construct different types of attribute trees based on the extracted document elements. For example, the terminal constructs 3 different types of attribute trees as shown in Figure 6 . Further, the terminal can obtain the attribute data corresponding to each document element according to the node index data of each node in each attribute tree and store the attribute data corresponding to each document element in the attribute data pool. That is, the terminal can respectively obtain the attribute data corresponding to each document element according to the node index data of each node in the 3 attribute trees as shown in Figure 6 and store the obtained attribute data corresponding to each document element in Figure 6In the property data pool shown. Thus, by proposing a completely new method for organizing the document data model, the document data based on this method can achieve a quick copy of a read-only copy, with extremely low memory occupancy, enabling editing and typesetting to be executed concurrently, and still having high document editing performance in the scenario of large-scale user collaboration with a large amount of document data.

[0105] In one embodiment, as Figure 7 shown, the steps of constructing each attribute tree based on the document elements in the sample document include:

[0106] Step 702, classify the document elements to obtain groups of document elements of the same category;

[0107] Step 704, respectively use each element in each group of document elements as a node to construct attribute trees of different categories.

[0108] Specifically, in response to the import operation performed by the user on the online document B, after the terminal extracts all the document elements included in the online document B, the terminal can classify the document elements to obtain groups of document elements of the same category, and respectively use each element in each group of document elements as a node to construct attribute trees of different categories. For example, the terminal classifies all the document elements included in the online document B to obtain 3 groups of document elements of the same category. For example, the first group of document elements is table elements, the second group of document elements is image elements, and the third group of document elements is text elements; further, the terminal can respectively use each element in each group of document elements as a node, and then can construct 3 attribute trees of different categories as shown in Figure 6 That is, the attribute tree 1 shown in Figure 6 can be a table attribute tree, the attribute tree 2 can be an image attribute tree, and the attribute tree 3 can be a text attribute tree. Thus, by proposing a completely new method for organizing the document data model, the document data based on this method can achieve a quick copy of a read-only copy, with extremely low memory occupancy, enabling editing and typesetting to be executed concurrently, and still having high document editing performance in the scenario of large-scale user collaboration with a large amount of document data.

[0109] In one embodiment, the steps of typesetting the target document based on the copy and concurrently performing document editing on the target document include:

[0110] Obtain the index data corresponding to each document element in the target document through the first thread, and perform document editing on the index data to obtain the first version of the target document;

[0111] Through the second thread, typeset the target document based on the copy to obtain the first visual document of the target document.

[0112] Among them, the first thread, the second thread, and the third thread are only used to distinguish different threads. For example, the first thread in this application can be an independent thread for document editing, and the second thread and the third thread can be independent threads for document typesetting.

[0113] Specifically, as Figure 5 shown, it is a schematic flowchart for online document collaborative editing. Assume that the terminal responds to Figure 5 the editing operation performed by the collaborative user 1 on the online document A shown in Figure 4 . The editing operation can be an operation of inserting characters. The terminal obtains the attribute tree corresponding to each document element in the online document A and the index data corresponding to the attribute tree through the first thread, as shown in Figure 4 . The terminal modifies at least a part of the data in the index data corresponding to the original attribute tree shown on the left in Figure 4 through the first thread to obtain the modified index data shown on the right in

[0114] . After that, the terminal can use the obtained modified index data as the first version document of the online document A. Figure 4 Furthermore, the terminal can copy the modified index data within the dashed box shown on the right in

[0115] through the first thread, that is, copy the first version document, to obtain a read-only copy 1 of the document, and use the read-only copy 1 of the document as the copy for typesetting the online document A. That is, the terminal can perform typesetting and rendering processing on the online document A based on the read-only copy 1 of the document through the second thread, and thus obtain the first visual document of the online document A. Therefore, by proposing a brand-new method for organizing the document data model, the document data based on this method can achieve fast copying of read-only copies, and the memory occupancy is extremely small, enabling editing and typesetting to be executed concurrently, and still having high document editing performance in the scenario of large-scale user collaboration with a large amount of document data. Figure 8 In one embodiment, as

[0116] shown, after performing document editing on the index data to obtain the first version document of the target document, the method includes:

[0117] Step 802, in response to the editing operation performed by other collaborating parties on the target document, perform document editing on the first version document through the first thread to obtain the second version document of the target document;

[0118] Step 804, during the process of performing document editing on the first version document, concurrently execute the step of typesetting the target document based on the copy through the second thread. Figure 5As shown, if three collaborative users perform editing operations on the same document, the terminal can sequentially respond to the editing operations of each collaborative user. When the terminal responds to Figure 5 the editing operation of collaborative user 1 shown, collaborative users 2 and 3 are other collaborating parties.

[0119] Specifically, as Figure 5 shown, it is a schematic diagram of the process during online document collaborative editing. As Figure 5 shown, if three collaborative users perform editing operations on the same document, the terminal can, based on a preset policy, sequentially respond to the editing operations of each collaborative user. For example, the terminal's sequential response to the editing operations of each collaborative user can be in the order shown in Figure 5 : collaborative user 1 → collaborative user 2 → collaborative user 3. Assume that the terminal responds to the editing operation performed by collaborative user 1 on online document A. This editing operation can be an operation of inserting characters. The terminal obtains the attribute tree corresponding to each document element in online document A and the index data corresponding to this attribute tree through the first thread as Figure 4 shown. The terminal modifies at least a part of the data in the index data corresponding to the original attribute tree shown on the left in Figure 4 through the first thread to obtain the modified index data shown on the right in Figure 4 . After that, the terminal can use the obtained modified index data as document data version 1 of online document A.

[0120] Furthermore, the terminal can copy Figure 4 the modified index data within the dashed box shown on the right, that is, copy document data version 1, to obtain a read-only copy 1 of the document, and use the read-only copy 1 of the document as the copy for typesetting the online document A. That is, the terminal can perform typesetting and rendering processing on the online document A based on the read-only copy 1 of the document through the second thread to obtain the visual version 1.

[0121] In addition, during the process of the terminal performing typesetting and rendering processing on the online document A based on the read-only copy 1 of the document through the second thread, the terminal can also continue to perform document editing on the online document A through the first thread. That is, during the process of the terminal performing typesetting and rendering processing on the online document A based on the read-only copy 1 of the document through the second thread, the terminal can respond to the editing operations performed on the online document A by other collaborating parties, namely the collaborating user 2, and obtain the attribute tree corresponding to each document element in the document data version 1 of the online document A and the index data corresponding to the attribute tree through the first thread. After the terminal modifies at least a part of the data in the index data corresponding to the attribute tree through the first thread to obtain the modified index data, the terminal can use the obtained modified index data as the document data version 2 of the online document A. The terminal can copy the document data version 2 through the first thread to obtain the read-only copy 2 of the document, and perform typesetting and rendering processing on the online document A based on the read-only copy 2 through the second thread or the third thread, thereby obtaining the visual version 2 of the online document A. That is, in the embodiments of the present application, the threads for editing processing must be the same, and the threads for typesetting processing can be the same or different independent threads. This enables data editing and typesetting to be fully parallel, and the response speed during multi-user collaborative document editing is faster. Especially for the scenario of large-scale document collaboration, it has obvious performance advantages.

[0122] In one embodiment, the present application further provides an application scenario that applies the above method for processing document data. Specifically, the application of the method for processing document data in this application scenario is as follows:

[0123] In the scenario of large-scale document collaboration, when developers want to achieve full parallelism between data editing and typesetting while quickly and effectively improving the processing performance of document data, they can adopt the above method for processing document data. The method provided by the present application can be applied to online document applications, enabling concurrent processing of document editing and document typesetting, and improving the editing performance of online document applications. Different users can open the online document application on the terminal through a trigger operation and enter the page of the target document in the document application through a selection operation. As shown on the terminal Figure 3In the page of the target document shown, the collaborating party, i.e., the collaborating user, can trigger an editing operation on the content of the displayed target document. Then, in response to the editing operation performed by the collaborating party on the target document, the terminal obtains the attribute tree corresponding to each document element in the target document and the index data corresponding to the attribute tree, and modifies at least a part of the data in the index data to obtain modified index data. Further, the terminal can generate a copy for typesetting the target document based on the modified index data, and the terminal can typeset the target document based on this copy and concurrently perform document editing on the target document. Compared with the traditional method, the present application proposes a brand-new method for organizing a document data model. Based on this method, document data can achieve fast copying of read-only copies, with extremely low memory occupancy, and can realize concurrent execution of editing and typesetting, and still has high document editing performance in the scenario of large-scale user collaboration with a large amount of document data.

[0124] The method provided in the embodiments of the present application can be applied to various scenarios of document data processing. Taking the scenario of an online document application as an example, the method for processing document data provided in the embodiments of the present application will be described below.

[0125] Online document: A cloud document that can be edited and collaborated on by multiple users simultaneously, and the edited content can be displayed and saved in real time.

[0126] In the traditional method, in the processing flow of online documents, editing and typesetting mostly adopt a serial model. That is, after the user triggers a document editing operation, the terminal first modifies the document data. After the document data modification is completed, the terminal triggers a typesetting and rendering again. After the typesetting is completed, the next editing operation is processed.

[0127] That is, in the traditional method, editing and typesetting are serial: Document data editing and corresponding typesetting are serial, that is, after the terminal finishes editing the document, it then performs typesetting, and editing and typesetting cannot be executed concurrently at the same time. Copying document data: After the editing operation of the document is completed, the terminal needs to copy the current document data into a read-only copy and pass it to typesetting so that subsequent document editing can continue, enabling editing and typesetting to be executed concurrently. However, this processing method will have high memory occupancy and time overhead for copying data when the document data scale is large.

[0128] The disadvantages of the above traditional technologies are as follows:

[0129] 1. Editing and typesetting are serial: Since editing and typesetting cannot be executed concurrently, in the scenario of large-scale user collaboration for online documents, document editing changes will be more frequent. When using the serial method of typesetting and editing, the response speed of the document will be reduced due to mutual waiting, resulting in performance problems such as document editing jams.

[0130] 2. Copying document data: Since it is necessary to additionally copy the document data into a separate read-only copy after each editing, when the scale of the document data is large, multiple document copies will cause high memory occupancy and additional time overhead, thus aggravating the lag problem of document editing.

[0131] Therefore, to solve the above problems, this application provides a solution, that is, this application proposes a method for efficiently expressing document data, which can be used in online document applications, enabling concurrent processing of document editing and document layout, and improving the editing performance of online document applications. Specifically, it includes:

[0132] 1. Limited by the fact that the scale of the document data may be large, first, the terminal splits the document data into an attribute pool and index data. The attribute data in the attribute pool is used to describe the attributes of text, paragraphs, pictures, etc. at a specific position in a document. The index data records the indexes of the attributes at specific positions in the document in the attribute pool and is stored continuously in segments.

[0133] 2. After document editing, the terminal generates newly modified attributes and puts them into the index pool, and modifies the attribute pool indexes at the corresponding positions. This operation can be completed in an independent thread of document editing.

[0134] 3. After editing is completed, only by using whole-segment copying of all the index segments, without copying any real attribute data, a read-only copy of the document for layout access can be generated.

[0135] 4. Based on the generated read-only copy, editing and layout can be executed in parallel.

[0136] On the product side, the method provided by this application can be applied to online document applications. As Figure 3 shown in the schematic diagram of large-scale collaborative document editing, multiple users can simultaneously edit large-scale documents and have a smooth editing experience.

[0137] On the technical side, 1.1 Document data model design

[0138] The core data describing the structure of a document is the basic elements contained in the document and the attributes corresponding to the elements. Since the data volume of these basic elements and attributes is large, the data model of the technical solution designed in this application is as Figure 6 shown, including:

[0139] 1. Express the document data elements as corresponding attribute trees. A document usually contains multiple attribute trees. By accessing the nodes of the attribute trees, the descriptions of the corresponding elements and their attributes can be obtained.

[0140] 2. The tree nodes only store lightweight data such as data indexes and attribute indexes, and the actual memory of each node is stored in multiple continuous memory segments.

[0141] 3. All the attribute data is stored in a shared attribute data pool, where the reference count of the node corresponding to the attribute data can also be recorded, and nodes with no references are periodically eliminated.

[0142] 1.2 Copying and Editing Scheme for Document Data Model

[0143] The document data copying scheme based on this data model is as follows:

[0144] 1. The terminal finds all the attribute trees describing the document data.

[0145] 2. The terminal finds the continuous memory occupied by the attribute tree according to the attribute tree.

[0146] 3. The terminal performs a memory copy on the occupied continuous memory to obtain a new attribute tree.

[0147] 4. The terminal forms the copied attribute trees into a new document data, thus completing a full copy of the document data.

[0148] Advantages of the scheme:

[0149] 1. For a copy of the document data, only the continuous memory segment where the node is located needs to be copied, with less content to copy and fast speed.

[0150] 2. The memory actually occupied by the copied document data only increases by the memory occupied by the node index.

[0151] As Figure 4 shown in Figure 4 illustrates the copy-on-write process during the actual modification of an attribute tree when inserting text once.

[0152] 1.3 Text Editing Collaboration Process

[0153] As Figure 5 shown in

[0154] the collaboration editing process of the online document based on this data model is as follows:

[0155] 1. The terminal applies the editing operations of each collaborator in sequence.

[0156] 2. After each editing operation is applied by the terminal, the terminal can quickly copy a read-only copy of the document based on the above data model.

[0157] 3. The terminal can perform parallel typesetting based on this read-only copy while processing the next editing operation.

[0158] Advantages of the scheme:

[0159] 2. Data editing and typesetting are fully parallel, resulting in a faster response speed.

[0160] 3. For scenarios of large-scale document collaboration, it has obvious performance advantages.

[0161] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0162] Based on the same inventive concept, an embodiment of the present application further provides a script program upgrade device for implementing the above-mentioned script program upgrade method. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the following script program upgrade device can refer to the limitations on the script program upgrade method in the above text, and will not be repeated here.

[0163] In one embodiment, as Figure 9 shown, a document data processing device is provided, including: an acquisition module 902, a modification module 904, a generation module 906, and a processing module 908, where:

[0164] The acquisition module 902 is configured to obtain the attribute tree corresponding to each document element in the target document and the index data corresponding to the attribute tree in response to an editing operation performed by a collaborator on the target document.

[0165] The modification module 904 is configured to modify at least a part of the data in the index data to obtain modified index data.

[0166] The generation module 906 is configured to generate a copy for typesetting the target document based on the modified index data.

[0167] The processing module 908 is configured to perform typesetting on the target document based on the copy and perform document editing on the target document in parallel.

[0168] In one embodiment, the device further includes: a query module and a combination module. The query module is configured to query the attribute tree corresponding to each document element in the target document; the obtaining module is further configured to obtain the node index data stored in the nodes of each attribute tree based on the tree identifier of each attribute tree; the combination module is configured to combine the node index data stored in the nodes of each attribute tree into the index data corresponding to the attribute tree.

[0169] In one embodiment, the attribute tree includes a table attribute tree, an image attribute tree, and a text attribute tree; the obtaining module is further configured to obtain the table node index data stored in the nodes of the table attribute tree based on the tree identifier of the table attribute tree; obtain the image node index data stored in the nodes of the image attribute tree based on the tree identifier of the image attribute tree; obtain the text node index data stored in the nodes of the text attribute tree based on the tree identifier of the text attribute tree; the combination module is further configured to combine the table node index data, the image node index data, and the text node index data into the index data corresponding to the attribute tree.

[0170] In one embodiment, the index data includes first position index data and second position index data; the first position index data and the second position index data are stored in continuous segments; the device further includes: a copying module and a query module. The copying module is configured to sequentially copy the first position index data and the second position index data according to the storage order to obtain copied first index data and copied second index data; the query module is configured to query the target node corresponding to the editing operation in the attribute tree corresponding to each document element in the target document; the modification module is further configured to, when the position index value corresponding to the target node belongs to the first position index data, modify the copied first index data to obtain modified first index data; use the modified first index data and the copied second index data as the modified index data.

[0171] In one embodiment, the device further includes: a copying module, configured to copy the modified index data to obtain copied index data; use the copied index data as a copy for typesetting the target document.

[0172] In one embodiment, the index data is used to reflect the mapping relationship between each node in the attribute tree and the attribute data in the attribute data pool; the obtaining module is further configured to obtain the attribute data corresponding to each document element from the attribute data pool based on the index data in the copy; the processing module is further configured to perform typesetting on the target document based on the attribute data.

[0173] In one embodiment, the index data is used to reflect the mapping relationship between each node in the attribute tree and the attribute data in the attribute data pool; the obtaining module is further configured to obtain the attribute data corresponding to each of the document elements from the attribute data pool based on the index data and the change description information in the copy, and obtain the document data corresponding to each of the document elements from the text pool; the processing module is further configured to typeset the target document based on the attribute data and the document data.

[0174] In one embodiment, the apparatus further includes: an extraction module, a construction module, and a storage module. The extraction module is configured to extract document elements from a sample document; the construction module is configured to construct each of the attribute trees based on the document elements in the sample document; the obtaining module is further configured to obtain the attribute data corresponding to each of the document elements according to the node index data of each node in each of the attribute trees; the storage module is configured to store the attribute data corresponding to each of the document elements in the attribute data pool.

[0175] In one embodiment, the apparatus further includes: a classification module, configured to classify the document elements to obtain each group of the document elements of the same category; the construction module is further configured to construct the attribute trees of different categories by using each element in each group of the document elements as a node respectively.

[0176] In one embodiment, the processing module is further configured to obtain the index data corresponding to each of the document elements in the target document through a first thread, and perform document editing on the index data to obtain a first version document of the target document; perform typesetting on the target document based on the copy through a second thread to obtain a first visual document of the target document.

[0177] In one embodiment, the processing module is further configured to, in response to an editing operation performed by other collaborating parties on the target document, perform document editing on the first version document through the first thread to obtain a second version document of the target document; during the process of performing document editing on the first version document, concurrently execute the step of performing typesetting on the target document based on the copy through the second thread.

[0178] Each module in the above document data processing apparatus can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above respective modules.

[0179] In one embodiment, a computer device is provided. The computer device can be a terminal or a server. In this embodiment, taking the computer device being a terminal as an example for illustration, its internal structure diagram can be as Figure 10As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a method for processing document data. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0180] Those skilled in the art can understand that Figure 10 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0181] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0182] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0183] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0184] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0185] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.

[0186] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.

[0187] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for processing document data, characterized in that, The method includes: In response to an editing operation performed by a collaborating party on a target document, obtaining an attribute tree corresponding to each document element in the target document and index data corresponding to the attribute tree; Modifying at least a part of the data in the index data to obtain modified index data; Generating a copy for typesetting the target document based on the modified index data; Typesetting the target document based on the copy, and concurrently performing document editing on the target document.

2. The method according to claim 1, characterized in that, The obtaining of the attribute tree corresponding to each document element in the target document and the index data corresponding to the attribute tree includes: Querying the attribute tree corresponding to each document element in the target document; Based on the tree identifier of each attribute tree, obtaining the node index data stored in the nodes of each attribute tree; Combining the node index data stored in the nodes of each attribute tree into the index data corresponding to the attribute tree.

3. The method according to claim 2, wherein The attribute tree includes a table attribute tree, an image attribute tree, and a text attribute tree; the obtaining of the node index data stored in the nodes of each attribute tree based on the tree identifier of each attribute tree includes: Based on the tree identifier of the table attribute tree, obtaining the table node index data stored in the nodes of the table attribute tree; Based on the tree identifier of the image attribute tree, obtaining the image node index data stored in the nodes of the image attribute tree; Based on the tree identifier of the text attribute tree, obtaining the text node index data stored in the nodes of the text attribute tree; The combining of the node index data stored in the nodes of each attribute tree into the index data corresponding to the attribute tree includes: Combining the table node index data, the image node index data, and the text node index data into the index data corresponding to the attribute tree.

4. The method according to claim 1, characterized in that, The index data includes first position index data and second position index data; the first position index data and the second position index data are stored continuously in segments; The modifying of at least a part of the data in the index data to obtain modified index data includes: Sequentially copying the first position index data and the second position index data according to the storage order to obtain copied first index data and copied second index data; In the attribute tree corresponding to each document element in the target document, querying a target node corresponding to the editing operation; When the position index value corresponding to the target node belongs to the first position index data, modifying the copied first index data to obtain modified first index data; Using the modified first index data and the copied second index data as the modified index data.

5. The method according to claim 1, characterized in that, The generating of a copy for typesetting the target document based on the modified index data includes: Copying the modified index data to obtain copied index data; Using the copied index data as a copy for typesetting the target document.

6. The method according to claim 1, characterized in that, The index data is used to reflect the mapping relationship between each node in the attribute tree and the attribute data in the attribute data pool; the typesetting of the target document based on the copy includes: Based on the index data in the copy, obtain the attribute data corresponding to each of the document elements from the attribute data pool; Typeset the target document based on the attribute data.

7. The method according to claim 1, characterized in that, The index data is used to reflect the mapping relationship between each node in the attribute tree and the attribute data in the attribute data pool; the typesetting the target document based on the copy includes: Based on the index data and the change description information in the copy, obtain the attribute data corresponding to each of the document elements from the attribute data pool, and obtain the document data corresponding to each of the document elements from the text pool; Typeset the target document based on the attribute data and the document data.

8. The method according to claim 7, wherein Before obtaining the attribute tree corresponding to each document element in the target document and the index data corresponding to the attribute tree in response to the editing operation performed by the collaborator on the target document, the method further includes: Extract the document elements in the sample document; Construct each of the attribute trees based on the document elements in the sample document; Obtain the attribute data corresponding to each of the document elements according to the node index data of each node in each of the attribute trees; Store the attribute data corresponding to each of the document elements in the attribute data pool.

9. The method according to claim 8, characterized in that The constructing each of the attribute trees based on the document elements in the sample document includes: Classify the document elements to obtain each group of the document elements of the same category; Construct the attribute trees of different categories respectively with each element in each group of the document elements as a node.

10. The method according to any one of claims 1 to 9, characterized in that The typesetting the target document based on the copy and concurrently performing document editing on the target document includes: Obtain the index data corresponding to each of the document elements in the target document through the first thread, and perform document editing on the index data to obtain the first version document of the target document; Perform typesetting on the target document based on the copy through the second thread to obtain the first visual document of the target document.

11. The method according to claim 10, characterized in that, After performing document editing on the index data to obtain the first version document of the target document, the method includes: In response to the editing operation performed by other collaborators on the target document, perform document editing on the first version document through the first thread to obtain the second version document of the target document; During the process of performing document editing on the first version document, concurrently execute the step of performing typesetting on the target document based on the copy through the second thread.

12. A processing device for document data, characterized in that, The apparatus includes: An obtaining module, configured to obtain the attribute tree corresponding to each document element in the target document and the index data corresponding to the attribute tree in response to the editing operation performed by the collaborator on the target document; A modifying module, configured to modify at least a part of the data in the index data to obtain modified index data; A generating module, configured to generate a copy for typesetting the target document based on the modified index data; A processing module, configured to typeset the target document based on the copy and concurrently perform document editing on the target document.

13. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 11.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.

15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.