Document data processing method and device, storage medium and computer equipment

By using attribute balance binary tree and character balance binary tree to store document data in document applications, the problem of inefficient document editing caused by DOM tree storage is solved, and more efficient document editing is achieved.

CN120257944APending Publication Date: 2025-07-04TENCENT TECH WUHAN
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410007802.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In existing document applications, the use of DOM trees to store document data, resulting in low document editing efficiency, especially when editing the middle position of the document, which requires full node update, which is inefficient.

Method used

The combination of attribute balance binary tree and character balance binary tree is used to store document data. By receiving document editing operations, the target tree is determined in the attribute balance binary tree and character balance binary tree respectively, and updated to reduce the time complexity of node search and update.

Benefits of technology

It improves the efficiency of document editing, reduces the rate of node search and update, avoids the problem of full node update, and improves document data processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257944A_ABST
    Figure CN120257944A_ABST
Patent Text Reader

Abstract

The invention provides a document data processing method and device, a storage medium and computer equipment. The method comprises the steps of receiving a document editing operation for a target document; when the document editing operation is an attribute editing operation, determining a target attribute balanced binary tree in attribute balanced binary trees corresponding to the plurality of document attributes; updating the target attribute balanced binary tree according to the attribute editing operation; when the document editing operation is a character editing operation, updating a first character balanced binary tree and a corresponding attribute balanced binary tree of the target document based on the character editing operation; and updating the second character balanced binary tree and the node mapping relation of the target document according to the character editing operation. According to the method, the time complexity of node searching and node updating can be reduced, and then the efficiency of editing the document data can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and particularly to a method, an apparatus, a storage medium, and a computer device for processing document data. Background Art

[0002] Document applications are very common office applications and are used very frequently in daily office work. Based on various text editing functions provided by document applications, users can efficiently process text, thereby realizing the addition, deletion, modification, and query of text information.

[0003] In related technologies, the in-memory data of each document in a document application is stored in the form of a Document Object Model (DOM) tree, which results in low efficiency in editing documents currently. Summary of the Invention

[0004] Embodiments of the present disclosure provide a method, an apparatus, a storage medium, and a computer device for processing document data, which can improve the efficiency of processing document data.

[0005] According to one aspect of the present disclosure, a method for processing document data is provided. The method includes:

[0006] Receiving a document editing operation for a target document;

[0007] When the document editing operation is an attribute editing operation, determining a target attribute balanced binary tree in a plurality of attribute balanced binary trees corresponding to document attributes, where nodes in the attribute balanced binary tree store attribute information of a section of text in the target document;

[0008] Updating the target attribute balanced binary tree according to the attribute editing operation;

[0009] When the document editing operation is a character editing operation, updating a first character balanced binary tree and a corresponding attribute balanced binary tree of the target document based on the character editing operation, where nodes in the first character balanced binary tree correspond to a section in the document display interface of the target document; and

[0010] Updating a second character balanced binary tree and a node mapping relationship of the target document according to the character editing operation, where nodes in the second character balanced binary tree store a section in the character array corresponding to the target document, and the node mapping relationship is a mapping relationship between nodes of the first character balanced binary tree and the second character balanced binary tree.

[0011] According to one aspect of the present disclosure, a device for processing document data is provided. The device includes:

[0012] A receiving unit, configured to receive a document editing operation for a target document;

[0013] A determining unit, configured to, when the document editing operation is an attribute editing operation, determine a target attribute balanced binary tree in a plurality of attribute balanced binary trees corresponding to document attributes, where nodes in the attribute balanced binary tree store attribute information of a section of text of the target document;

[0014] A first updating unit, configured to update the target attribute balanced binary tree according to the attribute editing operation;

[0015] A second updating unit, configured to, when the document editing operation is a character editing operation, update a first character balanced binary tree of the target document and a corresponding attribute balanced binary tree based on the character editing operation, where nodes in the first character balanced binary tree correspond to a section in a document display interface of the target document; and

[0016] A third updating unit, configured to update a second character balanced binary tree of the target document and a node mapping relationship according to the character editing operation, where nodes in the second character balanced binary tree store a section in a character array corresponding to the target document, and the node mapping relationship is a mapping relationship between nodes of the first character balanced binary tree and the second character balanced binary tree.

[0017] Optionally, in some embodiments, the first updating unit includes:

[0018] An obtaining subunit, configured to obtain an attribute editing range and attribute editing content corresponding to the attribute editing operation;

[0019] A searching subunit, configured to search for a target attribute node in the target attribute balanced binary tree according to the attribute editing range;

[0020] A first updating subunit, configured to update the target attribute balanced binary tree according to an intersection relationship between the attribute editing range and a target node range corresponding to the target attribute node, and the attribute editing content.

[0021] Optionally, in some embodiments, the first updating subunit includes:

[0022] A first updating module, configured to, when an intersection of the attribute editing range and a target node range corresponding to the target attribute node is the target node range, update attribute information corresponding to the target attribute node according to the attribute editing content;

[0023] A splitting module, configured to split the target attribute node into at least two new attribute nodes according to the sub-interval when the intersection of the attribute editing interval and the target node interval corresponding to the target attribute node is a sub-interval of the target node interval;

[0024] A processing module, configured to update the attribute information of the corresponding new attribute node according to the attribute editing content, and perform a balancing process on the target attribute AVL tree based on the new attribute node.

[0025] Optionally, in some embodiments, the second update unit includes:

[0026] An identification subunit, configured to identify the type of the character editing operation when the document editing operation is a character editing operation;

[0027] A second update subunit, configured to determine a character insertion interval and insertion character content when the type of the character editing operation is an insert character type, and update the first character AVL tree, the second character AVL tree, the mapping relationship between the nodes of the first character AVL tree and the second character AVL tree, and the corresponding attribute AVL tree of the target document according to the character insertion interval and the insertion character content;

[0028] A third update subunit, configured to determine a character editing interval corresponding to the character editing operation when the type of the character editing operation is not an insert character type, and update the first character AVL tree, the second character AVL tree, the mapping relationship between the nodes of the first character AVL tree and the second character AVL tree, and the corresponding attribute AVL tree of the target document according to the character editing interval.

[0029] Optionally, in some embodiments, the second update subunit includes:

[0030] A second update module, configured to update the character array of the target document according to the insertion character content, and determine an updated character interval corresponding to the insertion character content;

[0031] A third update module, configured to update the first character AVL tree, the second character AVL tree, and the mapping relationship between the nodes of the first character AVL tree and the second character AVL tree of the target document based on the updated character interval and the character insertion interval;

[0032] A fourth update module, configured to update the corresponding attribute AVL tree according to the character insertion interval.

[0033] Optionally, in some embodiments, the second update module includes:

[0034] Add a sub-module to add the characters corresponding to the inserted character content at the end of the character array of the target document to obtain an updated character array;

[0035] Determine a sub-module to obtain the number information of the characters in the updated character array and determine the updated character range corresponding to the inserted character content according to the number information.

[0036] Optionally, in some embodiments, the third update module includes:

[0037] An expansion sub-module, when it is detected that there are first character nodes and second character nodes with expandable ranges in the first character balanced binary tree and the second character balanced binary tree of the target document, expand the first character node based on the character insertion range and expand the second character node based on the updated character range;

[0038] The first character node is a node in the first character balanced binary tree, the second character node is a node in the second character balanced binary tree, there is a mapping relationship between the first character node and the second character node, the range corresponding to the first character node can be expanded based on the character insertion range, and the range corresponding to the second character node can be expanded based on the updated character range;

[0039] A first update sub-module to update the character range corresponding to each node in the first character balanced binary tree according to the range length of the character insertion range;

[0040] A second update sub-module, when it is detected that there are no first character nodes and second character nodes with expandable ranges in the first character balanced binary tree and the second character balanced binary tree of the target document, create a third character node and a fourth character node respectively based on the character insertion range and the updated character range, and update the first character balanced binary tree according to the third character node and update the second character balanced binary tree according to the fourth character node.

[0041] Optionally, in some embodiments, the root node of the first character balanced binary tree stores character position data, and other nodes in the first character balanced binary tree store first relative position distance data relative to a first prior node, where the first prior node is the corresponding left child node or parent node in the first character balanced binary tree, and the first update sub-module can also be used to:

[0042] Search for a target character node that needs to update stored data in the first character balanced binary tree according to the character position data and the character insertion position corresponding to the character insertion range;

[0043] Update the character position data or the first relative position distance data stored in the target character node based on the interval length of the character insertion interval.

[0044] Optionally, in some embodiments, attribute position data is stored in the root node of the attribute balanced binary tree, and second relative position distance data relative to a second prior node is stored in other nodes in the attribute balanced binary tree, where the second prior node is the corresponding left child node or parent node in the attribute balanced binary tree. The fourth update module includes:

[0045] A third update sub-module, configured to find a target attribute node whose position data needs to be updated in a corresponding attribute balanced binary tree according to the attribute position data and the attribute change position determined according to the character insertion interval;

[0046] A fourth update sub-module, configured to update the attribute position data or the second relative position distance data of the target attribute node based on the interval length of the character insertion interval.

[0047] Optionally, in some embodiments, the third update sub-unit includes:

[0048] A fifth update module, configured to determine a character editing interval corresponding to the character editing operation when the type of the character editing operation is inserting a space or deleting a space, and update the first character balanced binary tree of the target document, the mapping relationship between the nodes of the first character balanced binary tree and the second character balanced binary tree, and the corresponding attribute balanced binary tree according to the character editing interval;

[0049] A sixth update module, configured to determine a character deletion interval corresponding to the character editing operation when the type of the character editing operation is deleting a character type, and update the first character balanced binary tree based on the intersection relationship between the character deletion interval and the node interval corresponding to the node in the first character balanced binary tree, and update the second character balanced binary tree and the mapping relationship between the first character balanced binary tree and the second character balanced binary tree according to the updated first character balanced binary tree.

[0050] Optionally, in some embodiments, the sixth update module is further configured to:

[0051] Find a target deletion character node corresponding to the character deletion interval in the first character balanced binary tree;

[0052] When the character deletion interval is a sub - interval of the character interval corresponding to the target deletion character node, split the target deletion character node into two fifth - character nodes, and update the node intervals corresponding to the fifth - character nodes;

[0053] When the character interval corresponding to the target deletion character node is a sub - interval of the character deletion interval, delete the target character node;

[0054] When the character interval corresponding to the target deletion character node and the character deletion interval are cross - intervals, delete the overlapping part between the character interval corresponding to the target deletion character node and the character deletion interval, generate a sixth - character node according to the remaining part, and update the node interval corresponding to the sixth - character node.

[0055] According to one aspect of the present disclosure, there is provided a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the document data processing method described above is implemented.

[0056] According to one aspect of the present disclosure, there is provided a storage medium that stores a computer program, and when the computer program is executed by a processor, the document data processing method described above is implemented.

[0057] According to one aspect of the present disclosure, there is provided a computer program product. The computer program product includes a computer program, and the computer program is read and executed by a processor of a computer device, so that the computer device executes the document data processing method described above.

[0058] The document data processing method provided by the embodiments of the present disclosure includes receiving a document editing operation for a target document; when the document editing operation is an attribute editing operation, determining a target attribute balanced binary tree in a plurality of attribute balanced binary trees corresponding to document attributes, where the nodes in the attribute balanced binary tree store attribute information of a section of interval text of the target document; updating the target attribute balanced binary tree according to the attribute editing operation; when the document editing operation is a character editing operation, updating the first character balanced binary tree and the corresponding attribute balanced binary tree of the target document based on the character editing operation, where the nodes in the first character balanced binary tree correspond to a section of interval in the document display interface of the target document; and updating the second character balanced binary tree and the node mapping relationship of the target document according to the character editing operation, where the nodes in the second character balanced binary tree store a section of interval in the character array corresponding to the target document, and the node mapping relationship is the mapping relationship between the nodes of the first character balanced binary tree and the second character balanced binary tree.

[0059] In the embodiments of the present disclosure, an attribute balanced binary tree is adopted to store the attribute information of the text in each interval of the document, and a first character balanced binary tree and a second character balanced binary tree with a mapping relationship between nodes are adopted to store the text in each interval of the document display interface and the text in each interval of the actual corresponding character array respectively; relying on the time complexity advantages of the balanced binary tree in node search and node update, that is, the time complexity of node search and node update can be reduced to the logarithmic level, so that the rate of node search and update can be greatly reduced, and further the problem of full node update required during document editing caused by storing document data in the form of a DOM tree in the related art can be avoided, thereby greatly improving the processing efficiency of document data. That is, the editing efficiency of adding, deleting, modifying, and querying the document can be improved.

[0060] Other features and advantages of the present disclosure will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present disclosure. The objectives and other advantages of the present disclosure can be achieved and obtained through the structures specifically pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] The drawings are used to provide a further understanding of the technical solutions of the present disclosure, and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure, and do not constitute a limitation to the technical solutions of the present disclosure.

[0062] Figure 1 It is a system architecture diagram applied to the document data processing method of the embodiments of the present disclosure;

[0063] Figure 2 It is a schematic diagram of the present disclosure applied to the editing scenario of an offline document application;

[0064] Figure 3 It is a flowchart of the document data processing method provided by the present disclosure;

[0065] Figure 4 It is a schematic diagram of the document data model in the present disclosure;

[0066] Figure 5 It is a schematic diagram of constructing a mapping between the characters displayed in the document and the characters in the character array based on the mapping tree;

[0067] Figure 6 It is a schematic diagram of inserting a character into the document;

[0068] Figure 7 It is another schematic diagram of inserting a character into the document;

[0069] Figure 8aSchematic diagram of a first character balanced binary tree for storing absolute position information in a node;

[0070] Figure 8b For Figure 8a Schematic diagram for updating the described balanced binary tree;

[0071] Figure 9a Schematic diagram of a first character balanced binary tree constructed by the method using a floating - point tree provided by the present disclosure;

[0072] Figure 9b For Figure 9a Schematic diagram for updating the shown floating - point tree;

[0073] Figure 10 Schematic diagram for updating a mapping tree based on a character deletion operation;

[0074] Figure 11 Another flowchart schematic diagram of the document data processing method provided by the present disclosure;

[0075] Figure 12 Schematic diagram of a piece table processing method adopted in the present disclosure;

[0076] Figure 13 Schematic diagram of the structure of a document data processing device provided by an embodiment of the present disclosure;

[0077] Figure 14 Structure diagram of a terminal for implementing the methods according to an embodiment of the present disclosure;

[0078] Figure 15 Structure diagram of a server for implementing the methods according to an embodiment of the present disclosure. Detailed implementation manners

[0079] In order to make the objectives, technical solutions and advantages of the present disclosure more clear and understandable, the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.

[0080] Before further elaborating on the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure are explained. The nouns and terms involved in the embodiments of the present disclosure are applicable to the following explanations:

[0081] Binary tree: A binary tree is an important type of tree structure. The data structures abstracted from many practical problems are often in the form of binary trees. Even general trees can be simply converted into binary trees. Moreover, the storage structure and algorithms of binary trees are relatively simple. Therefore, binary trees are particularly important. The characteristic of a binary tree is that each node can have at most two subtrees, and there is a distinction between left and right subtrees.

[0082] Balanced binary tree: A balanced binary tree is a special type of binary tree. The height difference between its left subtree and right subtree will never be greater than 1. Compared with ordinary binary trees, a balanced binary tree has an additional balance factor to ensure that the height difference between the left and right subtrees is always less than or equal to 1. The emergence of balanced binary trees is used to optimize the query efficiency of binary trees. If we give an ordered sequence and use it to generate a binary tree, a skew tree will be generated, making the binary tree become a single-chain structure and losing the advantage of fast query speed. Therefore, balanced binary trees appear, making the height difference between the left subtree and right subtree of any node not exceed 1, so that all nodes are more evenly distributed on both sides of the binary tree, improving the query efficiency.

[0083] Segment tree: A segment tree is a binary search tree, similar to an interval tree. It divides an interval into some unit intervals, and each unit interval corresponds to a leaf node in the segment tree.

[0084] Red-black tree: A red-black tree is a self-balancing binary search tree and an efficient search tree. A red-black tree has good efficiency and can complete operations such as searching, adding, and deleting in O(logN) time.

[0085] In related technologies, a word document in the format of extensible markup language (XML) (a text document, hereinafter referred to as an XML document) describes the in-memory data of the document in the form of a complete DOM tree. Specifically, first obtain the XML data of the text document, then parse the XML data and generate a DOM tree according to the parsing results. In the DOM processing standard, all documents, elements, attributes, and texts in the XML document are described in the form of nodes, and then the data content is set for each node, and then the node relationships are set. The DOM tree is a multi-way tree with an uncertain number of levels. When editing operations need to be performed on the text in the document, full-node searching is required, resulting in low node searching efficiency. Moreover, when the position modified by the editing operation is in the middle of the document, the node data of the subsequent document needs to be modified, which leads to very low document editing efficiency. To solve the above problem of low document editing efficiency to a certain extent, the present disclosure provides a method for processing document data. The method will be introduced in detail below.

[0086] System architecture and scenario description applied in the embodiments of the present disclosure

[0087] Figure 1 It is a system architecture diagram applied to the document data processing method according to an embodiment of the present disclosure. It includes a terminal 140, the Internet 130, a gateway 120, a server 110, etc.

[0088] The terminal 140 includes various forms such as a desktop computer, a laptop computer, a PDA (Personal Digital Assistant), a mobile phone, a vehicle-mounted terminal, a home theater terminal, a dedicated terminal, a smart voice interaction device, a smart home appliance, an aircraft, etc. Additionally, it can be a single device or a collection of multiple devices. The terminal 140 can communicate with the Internet 130 in a wired or wireless manner to exchange data. A document application can be installed in the terminal 140, and this document application can specifically be an online document application, and the online document application can specifically be a document application that supports multiple people to edit a document online simultaneously.

[0089] The server 110 refers to a computer system that can provide certain services to the terminal 140. Compared with an ordinary terminal 140, the server 110 has higher requirements in terms of stability, security, performance, etc. The server 110 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part (such as a virtual machine) allocated from a high-performance computer, a combination of parts (such as virtual machines) allocated from multiple high-performance computers, etc. The server 110 can specifically be the server corresponding to the aforementioned online document application, and multiple terminals 140 can simultaneously access the document data of the same online document in the server through the installed online document application client.

[0090] The gateway 120 is also called an internetwork connector and a protocol converter. The gateway realizes network interconnection at the transport layer and is a computer system or device that acts as a conversion role. Between two systems using different communication protocols, data formats, or languages, and even with completely different architectures, the gateway is a translator. At the same time, the gateway can also provide filtering and security functions. The message sent by the terminal 140 to the server 110 needs to be sent to the corresponding server 110 through the gateway 120. The message sent by the server 110 to the terminal 140 also needs to be sent to the corresponding terminal 140 through the gateway 120.

[0091] The document data processing method provided by the embodiment of the present disclosure can be implemented in the aforementioned terminal 140, or can be partially implemented in the terminal 140 and partially implemented in the server 110.

[0092] When the document data processing method provided by the embodiments of the present disclosure is implemented in the terminal 140, the terminal 140 may receive a document editing operation for a target document; when the terminal 140 detects that the document editing operation is an attribute editing operation, the terminal 140 determines a target attribute balanced binary tree in the attribute balanced binary trees corresponding to multiple document attributes, and the nodes in the attribute balanced binary tree store the attribute information of a section of text of the target document; the terminal 140 updates the target attribute balanced binary tree according to the attribute editing operation; when the terminal 140 detects that the document editing operation is a character editing operation, the terminal 140 updates the first character balanced binary tree of the target document and the corresponding attribute balanced binary tree based on the character editing operation, and the nodes in the first character balanced binary tree correspond to a section in the document display interface of the target document; and updates the second character balanced binary tree of the target document and the node mapping relationship according to the character editing operation, the nodes in the second character balanced binary tree store a section in the character array corresponding to the target document, and the node mapping relationship is the mapping relationship between the nodes of the first character balanced binary tree and the second character balanced binary tree.

[0093] When a part of the document data processing method provided by the embodiments of the present disclosure is implemented in the server 110 and another part is implemented in the terminal 140, the terminal 140 may receive a document editing operation for a target document, and then send the document editing operation to the server 110; when the server 110 detects that the document editing operation is an attribute editing operation, the server 110 determines a target attribute balanced binary tree in the attribute balanced binary trees corresponding to multiple document attributes, and the nodes in the attribute balanced binary tree store the attribute information of a section of text of the target document, and updates the target attribute balanced binary tree according to the attribute editing operation; when the server 110 detects that the document editing operation is a character editing operation, the server 110 updates the first character balanced binary tree of the target document and the corresponding attribute balanced binary tree based on the character editing operation, and the nodes in the first character balanced binary tree correspond to a section in the document display interface of the target document; and updates the second character balanced binary tree of the target document and the node mapping relationship according to the character editing operation, the nodes in the second character balanced binary tree store a section in the character array corresponding to the target document, and the node mapping relationship is the mapping relationship between the nodes of the first character balanced binary tree and the second character balanced binary tree.

[0094] The above system architecture diagram is only the system architecture diagram in a scenario where the embodiments of the present disclosure can be applied, and specifically can be the system architecture diagram when the embodiments of the present disclosure are applied to an online document application. In addition, the embodiments of the present disclosure can also be applied to the scenario of an offline document application. In this scenario, the system architecture only needs to be implemented by the terminal, and the data of the document application can be stored in the memory or disk of the terminal.

[0095] Such asFigure 2 As shown in the figure, it is a schematic diagram of the application of the embodiments of the present disclosure in the editing scenario of an offline document application. When the embodiments of the present disclosure are applied to the editing scenario of an offline document application, an XML-formatted text document can be displayed using the offline document application on the display 210 of the terminal, and document editing operations on the displayed text document by an object can be received. Then, the display 210 sends the received document editing operations to the central processing unit 220 of the terminal for processing, specifically, updating the property balanced binary tree, the first character balanced binary tree, the second character balanced binary tree, and the mapping relationship between the nodes of the first character balanced binary tree and the second character balanced binary tree based on the document editing operations. The character array corresponding to the text document, as well as the mapping relationship data between the foregoing property balanced binary tree, the first character balanced binary tree, the second character balanced binary tree, and the two character balanced binary trees, are stored in the data memory 230 of the terminal. When an object performs a text insertion operation on the text document, the central processing unit 220 also updates the character array.

[0096] When the embodiments of the present disclosure are applied to the editing scenario of an online document application, reference can specifically be made to Figure 1 the system architecture diagram shown in the figure. First, a display interface of an online text document can be simultaneously displayed in the online document application clients installed in multiple terminals 140. The interfaces of the same online text document displayed in different terminals can be different, that is, different parts of the online text document can be displayed in different terminals. Then, multiple terminals 140 can simultaneously receive document editing operations of an object and send the received document editing operations on the online text document to the server 110, so that the server 110 updates the property balanced binary tree, the first character balanced binary tree, the second character balanced binary tree, the mapping relationship between the two character balanced binary trees, and the character array of the stored online text document. When a large number of objects simultaneously perform simultaneous editing on the same online text document, the improvement effect of this solution on document editing efficiency is more obvious.

[0097] General description of the embodiments of the present disclosure

[0098] According to an embodiment of the present disclosure, a document data processing method is provided. This method can be used in the scenario of editing the aforementioned offline text document or the scenario of editing the aforementioned online text document.

[0099] As Figure 3 shown in the figure, it is a flowchart of a document data processing method provided by the present disclosure. This method can be applied to a document editing device, and the document editing device can be integrated in a computer device. The computer device can specifically be a terminal or a server. The document data processing method can include:

[0100] Step 310, receive a document editing operation for a target document.

[0101] In an embodiment of the present disclosure, a method for processing document (specifically a text document in this embodiment) data is provided. This method is applied to improve the efficiency of document editing during the process of editing a document. This method specifically depends on a design for storing document memory data. Therefore, before introducing the document data processing method provided by the present disclosure, the design scheme for storing document memory data in this method is first introduced.

[0102] In this embodiment, the document memory data is divided into a text stream and attribute streams, where the attribute streams are specifically multiple attribute streams corresponding to multiple attributes. Specifically, the attribute streams may include: text attributes, which are used to describe the character attributes of text, such as bold, red, and font; paragraph attributes, such as line height and item numbering; graphic and image attributes, such as shape attributes and picture attributes; comment attributes, such as the comment content itself; comment range attributes, such as comment range and commenter; in addition, it may also include object attributes, table attributes, table row attributes, table cell attributes, revision attributes, footnote attributes, and endnote attributes, etc., which are not listed one by one here. As Figure 4 shown, it is a schematic diagram of the document data model in the present disclosure. As shown in the figure, in this embodiment, the document data is divided into a text stream and multiple attribute streams. There is a certain mapping relationship between the text stream and the attribute streams. This mapping relationship can be specifically represented by a certain interval. This interval can be specifically an interval determined by character numbers. For example, characters 1 to 99 determine a section of text in the interval [1, 100). The character attributes corresponding to the interval [1, 100) are red font, and the paragraph attributes corresponding to the interval [1, 100) are paragraph attribute 1.

[0103] For each attribute stream, the attribute information of this type of document can be described by a balanced binary tree. This balanced binary tree can be called an attribute balanced binary tree. Each node in the attribute balanced binary tree stores the attribute information of a section of the text document. For the text stream, in this embodiment, two balanced binary trees are used for description. These two balanced binary trees are specifically the first character balanced binary tree and the second character balanced binary tree in the embodiment of the present disclosure. These two balanced binary trees can also be called a pair of mirror trees because these two balanced binary trees have the following properties:

[0104] 1. The nodes of the two trees are the same, that is, the two trees have the same number of nodes and the same node structure;

[0105] 2. Each node of the two trees corresponds to an interval;

[0106] 3. The intervals corresponding to the nodes within the same tree do not have overlapping regions;

[0107] 4. The nodes in the two trees are in one-to-one correspondence, and the sizes of the intervals corresponding to the corresponding nodes are the same.

[0108] Specifically, the first character balanced binary tree corresponds to the characters displayed in the document interface. Each node in the first character balanced binary tree corresponds to an interval within the interval [0, document content length), that is, a sub-interval of this interval, and the intervals corresponding to all nodes in the first character balanced binary tree are continuous. All the intervals added together will be a complete [0, document content length). The second character balanced binary tree corresponds to the character array (textpool) that actually stores the character content of the document, that is, the character array is data stored in the memory of the terminal (such as a disk or cache). Therefore, the intervals corresponding to the nodes in the second character balanced binary tree are within [0, character array length). During the editing process of the object on the document, some content may be deleted, and the characters corresponding to this content will not be deleted from the character array. Therefore, the intervals corresponding to the nodes in the second character balanced binary tree can be discontinuous. When the object inserts new characters into the document, no matter where the object inserts the characters in the document, the inserted characters will be appended to the end of the character array. Therefore, all the characters displayed in the document interface can find their corresponding characters in the character array of the document; while the characters in the character array of the document do not necessarily all appear in the document interface. It may be that the character is deprecated due to a deletion operation during the document editing process. After deprecating the character, the character no longer appears in the document interface but still exists in the character array of the document and will not be deleted.

[0109] As Figure 5 shown, it is a schematic diagram of constructing a mapping between the characters displayed in the document and the characters in the character array based on the mapping tree. As shown in the figure, the characters "XX Document is an online collaboration" are displayed on the display interface of the document. Each character corresponds to a numbered information. According to the numbered information of these characters, it can be determined that the intervals corresponding to the nodes in the first character balanced binary tree corresponding to these characters are [0, 11). In addition, the characters "XX Document is an online collaboration" are also stored in the character array corresponding to this document. Each character also corresponds to a numbered information in the character array; according to the numbered information of these characters in the character array, it can be determined that the intervals corresponding to the nodes in the second character balanced binary tree corresponding to these characters are [2, 13). Then, a mapping relationship can be constructed between the nodes corresponding to the interval [0, 11) in the first character balanced binary tree and the nodes corresponding to the interval [2, 13) in the second character balanced binary tree. When the document is opened and displayed next time, the characters displayed in the interval [0, 11) in the display interface can be determined according to the characters stored in the interval [2, 13) in the character array.

[0110] On this basis, when the document file is opened, the mapping relationship between each node interval in the first character balanced binary tree and each node interval in the second character balanced binary tree can be obtained, so as to determine the correspondence between each interval displayed in the document interface and the intervals in the second character balanced binary tree. Then, the character data corresponding to each interval in the second character balanced binary tree in the character array is obtained correspondingly, so as to determine the character data that should be displayed in each interval in the document interface. Further, the attribute data corresponding to the characters in each interval in the document interface is obtained, and the character data is displayed in the display interface of the document in combination with the attribute data.

[0111] Based on the above design, the attribute data of the text document can be described by the combination of multiple attribute balanced binary trees, and the text character data of the text document can be described by the mapping tree combined with the character array. Compared with the case of the DOM tree with multiple branches and uncertain number of layers, traversal is required to find nodes, and its time complexity is O(n), where n is the number of nodes. The balanced binary tree is a strict binary branch, and the node search process starts from the root node and searches layer by layer to one side of the child nodes. Therefore, the time complexity of node search is O(logn). Based on the unique advantage of high node query efficiency of the balanced binary tree, the time complexity of finding the nodes to be modified during the document editing process can be reduced to the logarithmic level, so that the efficiency of finding the nodes to be modified during the document editing process can be improved, and further the document editing efficiency can be improved. The document editing process corresponds to the maintenance process of multiple attribute balanced binary trees, mapping trees and character arrays. The following details the maintenance process of attribute balanced binary trees, mapping trees and character data during the document editing process based on the above design.

[0112] After the document interface of the text document is displayed, the document editing operations from the object can be received. The document editing operations include attribute editing operations and character editing operations. Among them, the attribute editing operation can specifically be an operation to modify the attributes of the characters already displayed in the text document, such as changing the font, color, font size, etc. of the characters in a certain interval; or changing the paragraph attributes of the characters in a certain interval, splitting a paragraph of text into multiple paragraphs of text, etc.; of course, other attributes of the document can also be edited, and no further examples are given here. The attribute editing operation needs to update the target attribute balanced binary tree corresponding to the target attribute to be edited. The character editing operation can specifically be adding characters or deleting characters, or text editing operations such as adding or deleting spaces. When adding characters, the character array corresponding to the text document needs to be updated, and in other cases, the character array corresponding to the text document does not need to be updated. The character editing operation not only needs to update the mapping tree, that is, update the mapping relationship between the two character balanced binary trees and the nodes of the two character balanced binary trees, but also needs to update the corresponding attribute balanced binary tree.

[0113] In this embodiment, the process of maintaining multiple attribute balanced binary trees, mapping trees, and character arrays corresponding to the target document during the document editing process is described in detail by taking the process of document editing of the target document as an example. As described above, the target document can specifically be a text document, which can be an offline-edited text document or an online text document edited simultaneously by multiple people.

[0114] First, in response to the trigger operation of opening the target document, the document data corresponding to the target document is obtained. The document data can specifically be data stored in Extensible Markup Language. Then, the document data is parsed to obtain multiple attribute balanced binary trees and mapping trees corresponding to the document data. The aforementioned character array can also be included in the document data, and the character array can be stored on the disk of the terminal or on the server corresponding to the online document application. Then, the document interface of the target document is displayed according to the multiple attribute balanced binary trees, mapping trees, and character arrays. In the document interface, the editing operation of the target document by the document editing object can be received.

[0115] Step 320: When the document editing operation is an attribute editing operation, determine the target attribute balanced binary tree among the attribute balanced binary trees corresponding to multiple document attributes.

[0116] Since the maintenance processes of the relevant data of the target document corresponding to different editing operations of the target document are different, when the document editing operation of the target document is received, the specific editing operation content of the document editing operation can be identified first to determine the corresponding document editing type, and then the corresponding data maintenance can be performed according to different document editing types.

[0117] When the result of identifying the document editing operation determines that the document editing operation is an attribute editing operation, the content of the attribute editing can be further obtained, and the target attribute balanced binary tree corresponding to the content of the attribute editing is determined among the multiple attribute balanced binary trees according to the content of the attribute editing.

[0118] Generally, the operations of modifying the attribute information of the document are independent. Therefore, for one document editing operation, the determined target attribute balanced binary tree can be one attribute balanced binary tree. After the target attribute balanced binary tree is determined, the target attribute balanced binary tree can be updated according to the content of the attribute editing corresponding to the attribute editing operation.

[0119] Step 330: Update the target attribute balanced binary tree according to the attribute editing operation.

[0120] Among them, as described above, each node in the attribute balanced binary tree stores the attribute information of the text in a section, that is, each node stores two pieces of information. On the one hand, it is a section, and on the other hand, it is the attribute information corresponding to this section. Therefore, updating the target attribute balanced binary tree can specifically include updating the attribute data stored in the node, updating the section stored in the node, and updating the tree structure itself.

[0121] Specifically, in the embodiments of the present disclosure, updating the target attribute balanced binary tree according to the attribute editing operation includes:

[0122] Obtaining the attribute editing section and the attribute editing content corresponding to the attribute editing operation;

[0123] Searching for the target attribute node in the target attribute balanced binary tree according to the attribute editing section;

[0124] Updating the target attribute balanced binary tree according to the intersection relationship between the attribute editing section and the target node section corresponding to the target attribute node, and the attribute editing content.

[0125] That is, in the embodiments of the present disclosure, when it is determined that the document editing operation is an attribute editing operation, the attribute editing section and the attribute editing content corresponding to the attribute editing operation can be obtained first. Among them, the attribute editing section can specifically also be the section determined according to the character numbers in the target document. For example, if the attribute editing operation is to change the font color of the 55th character to the 80th character to red, then the attribute editing section corresponding to this attribute editing operation can be determined as [55, 81), and the attribute editing content can be determined as changing the font color to red.

[0126] After determining the attribute editing section, the target attribute node that needs to be updated can be further searched in the target attribute balanced binary tree according to the attribute editing section. Among them, the target attribute node can specifically be the section whose corresponding section and the attribute editing section have an intersection. For example, there are several intervals corresponding to attribute nodes as [40, 45), [45, 66), [66, 77), and [77, 88). Since [40, 45) has no intersection with [55, 81), and the three intervals [45, 66), [66, 77), and [66, 88) have intersections with [55, 81), the attribute nodes corresponding to the latter three intervals can be determined as the target attribute nodes. That is, in the embodiments of the present disclosure, the determined target attribute node can be one or more.

[0127] After determining the target attribute node, the target attribute AVL tree can be updated according to the attribute editing content and the determined target attribute node. Specifically, since only some sub-intervals in the target node may be within the attribute editing interval, and the sub-intervals not within the attribute editing interval do not need to have their attributes changed. In this case, the target node needs to be split, and then the target attribute AVL tree is updated according to the split nodes. Therefore, after finding the target attribute node in the target attribute AVL tree according to the attribute editing interval, the intersection relationship between the attribute editing interval and the target node interval corresponding to the target attribute node can be further obtained, and then the target attribute AVL tree is updated according to the intersection relationship between the attribute editing interval and the target node interval corresponding to the target attribute node and the attribute editing content.

[0128] In some embodiments, updating the target attribute AVL tree according to the intersection relationship between the attribute editing interval and the target node interval corresponding to the target attribute node, and the attribute editing content, includes:

[0129] When the intersection of the attribute editing interval and the target node interval corresponding to the target attribute node is the target node interval, update the attribute information corresponding to the target attribute node according to the attribute editing content;

[0130] When the intersection of the attribute editing interval and the target node interval corresponding to the target attribute node is a sub-interval of the target node interval, split the target attribute node into at least two new attribute nodes according to the sub-interval;

[0131] Update the attribute information of the corresponding new attribute nodes according to the attribute editing content, and perform a balancing process on the target attribute AVL tree based on the new attribute nodes.

[0132] Among them, after determining the attribute editing interval and the target node interval corresponding to each target attribute node, the intersection relationship between the attribute editing interval and each target node interval can be analyzed one by one. When the intersection of the attribute editing interval and the target node interval is the target node interval, that is, the target node interval is a sub-interval of the attribute editing interval, which means that the attribute information of all texts in the target node interval needs to be updated. At this time, the attribute information corresponding to the target attribute node can be updated according to the attribute editing content.

[0133] When the intersection of the attribute editing interval and the target node interval is a sub-interval of the target node interval, that is, the attribute editing interval only contains a part of the target node interval, that is, only the attribute information of part of the text in the target node interval needs to be updated. However, in this case, there will be a problem that the attribute information of different intervals in a node is different, resulting in the inability to clarify the attribute information stored in the target node. In this case, the target node needs to be split, specifically, it can be split into two new attribute nodes or three new attribute nodes. When the attribute editing interval and the target node interval are in an intersection relationship, the target attribute node corresponding to the target node interval can be split into two new attribute nodes; when the attribute editing interval is a sub-interval in the target node interval, the target attribute node corresponding to the target node interval can be split into three new attribute nodes.

[0134] After splitting the target attribute node into multiple new attribute nodes, the attribute information of the new attribute nodes that need to update the attribute information can be updated according to the attribute editing content. Then, since new attribute nodes are generated in the target attribute AVL tree, it may cause the target attribute AVL tree to become unbalanced. At this time, it is necessary to further perform a balance detection on the target attribute AVL tree after generating the new attribute nodes to determine whether the target attribute AVL tree still maintains balance. If the target attribute AVL tree is no longer balanced, at this time, the target attribute AVL tree can be balanced to achieve the update of the target attribute AVL tree.

[0135] Step 340, when the document editing operation is a character editing operation, update the first character AVL tree of the target document and the corresponding attribute AVL tree based on the character editing operation.

[0136] When it is detected that the document editing operation is a character editing operation, the mapping tree, character array, and attribute AVL tree of the target document can be updated based on the character editing operation. That is, update the first character AVL tree, the second character AVL tree, the mapping relationship between the first character AVL tree and the second character AVL tree, the character array, and the corresponding attribute AVL tree of the target document based on the character editing operation. That is, the operations between step 340 and step 350 here can be parallel operations, and the step numbers do not limit the execution order of these two steps.

[0137] In some embodiments, when it is detected that the document editing operation is a character editing operation, the type of the character editing operation can be further identified according to the content of the character editing operation. In this embodiment, the type of the character editing operation can be divided into an inserted character type and other types (which can also be referred to as non-inserted character types). The main difference in classification in this way is that for a character editing operation of the inserted character type, the character array of the target document needs to be updated, while for other types of character editing operations, the character array of the target document does not need to be updated.

[0138] When it is recognized that the type of the character editing operation is the inserted character type, the character insertion interval and the inserted character content can be determined, and then the mapping tree, the character array, and the corresponding attribute balanced binary tree of the target document are updated according to the character insertion interval and the inserted character content. When it is recognized that the type of the character editing operation is not the inserted character type, the character editing interval corresponding to the character editing operation can be determined. Here, the character editing interval can specifically be any one of a character deletion interval, a space insertion interval, or a space deletion interval. Then, without based on the character content, the mapping tree and the corresponding attribute balanced binary tree of the target document can be directly updated based on the character editing interval.

[0139] In some embodiments, updating the first character balanced binary tree, the second character balanced binary tree, the mapping relationship between the nodes of the first character balanced binary tree and the second character balanced binary tree, and the corresponding attribute balanced binary tree of the target document according to the character insertion interval and the inserted character content includes:

[0140] Updating the character array of the target document according to the inserted character content, and determining the updated character interval corresponding to the inserted character content;

[0141] Updating the first character balanced binary tree, the second character balanced binary tree, and the mapping relationship between the nodes of the first character balanced binary tree and the second character balanced binary tree of the target document based on the updated character interval and the character insertion interval;

[0142] Updating the corresponding attribute balanced binary tree according to the character insertion interval.

[0143] Among them, when it is recognized that the type of character editing operation is the character insertion type, the purpose of updating the mapping tree of the target document is to establish a mapping relationship of the inserted character content in the mapping tree. Specifically, the mapping relationship is the mapping relationship between the node intervals corresponding to the nodes of the inserted character content in the first character balanced binary tree and the nodes of the inserted character content in the second character balanced binary tree. The node interval corresponding to the inserted character content in the first character balanced binary tree can specifically be the aforementioned character insertion interval. The node interval corresponding to the inserted character content in the second character balanced binary tree needs to be determined according to the position of the newly added characters in the character array.

[0144] Therefore, in the embodiments of the present disclosure, when it is recognized that the type of character editing operation is the character insertion type, the character array of the target document can be updated according to the inserted character content, and the updated character interval corresponding to the inserted character content can be determined. Then, the mapping tree of the target document can be updated according to the updated character interval and the character insertion interval. In addition, the corresponding attribute balanced binary tree also needs to be updated according to the character insertion interval.

[0145] In some embodiments, updating the character array of the target document according to the inserted character content and determining the updated character interval corresponding to the inserted character content includes:

[0146] Adding the characters corresponding to the inserted character content to the end of the character array of the target document to obtain an updated character array;

[0147] Obtaining the number information of the characters in the updated character array and determining the updated character interval corresponding to the inserted character content according to the number information.

[0148] In the embodiments of the present disclosure, when a document editing object inserts characters into a document, the corresponding inserted character content can be determined, and then the characters corresponding to the inserted character content are added to the end of the character array of the target document to obtain an updated character array. For example, as Figure 6 shown, it is a schematic diagram of inserting characters into a document. For example, adding the two characters "duo ren" after the 6th character "kuan" in the display interface of the document. At this time, regardless of what the character insertion interval corresponding to the added characters in the display interface is, the newly added two characters will be added to the end of the character array, that is, added to the positions corresponding to the 13th and 14th, to implement the update of the character array. Then, the number information of the characters in the updated character array can be obtained, and the updated character interval corresponding to the inserted character content can be determined according to the number information. As shown in the figure, the corresponding updated character interval is [13, 15).

[0149] In some embodiments, updating the first character balanced binary tree, the second character balanced binary tree of the target document, and the mapping relationship between the nodes of the first character balanced binary tree and the second character balanced binary tree based on the updated character interval and the character insertion interval includes:

[0150] When it is detected that there are first character nodes and second character nodes with expandable intervals in the first character balanced binary tree and the second character balanced binary tree of the target document, expanding the first character nodes based on the character insertion interval and expanding the second character nodes based on the updated character interval;

[0151] The first character node is a node in the first character balanced binary tree, the second character node is a node in the second character balanced binary tree, there is a mapping relationship between the first character node and the second character node, the interval corresponding to the first character node can be expanded based on the character insertion interval, and the interval corresponding to the second character node can be expanded based on the updated character interval;

[0152] Updating the character interval corresponding to each node in the first character balanced binary tree according to the interval length of the character insertion interval;

[0153] When it is detected that there are no first character nodes and second character nodes with expandable intervals in the first character balanced binary tree and the second character balanced binary tree of the target document, creating a third character node and a fourth character node respectively based on the character insertion interval and the updated character interval, and updating the first character balanced binary tree according to the third character node and updating the second character balanced binary tree according to the fourth character node.

[0154] As introduced above, after inserting a string of length len into the target document, the update of the mapping tree is to construct the mapping relationship from the interval [pos1, pos1 + len) in the first character balanced binary tree to the interval [pos2, pos2 + len) in the second character balanced binary tree. Here, pos1 is the position where the string is inserted in the display interface of the target document, pos2 is the position of the newly added characters in the character array, and len is the length of the inserted string. After updating the character array of the target document based on the inserted string, it is possible to first detect whether there are first character nodes and second character nodes with expandable intervals in the first character balanced binary tree and the second character balanced binary tree of the target document. Among them, the first character node is a node in the first character balanced binary tree, the second character node is a node in the second character balanced binary tree, there is a mapping relationship between the first character node and the second character node, the interval corresponding to the first character node can be expanded based on the character insertion interval, and the interval corresponding to the second character node can be expanded based on the updated character interval. If there are, the first character node can be expanded based on the character insertion interval and the second character node can be expanded based on the updated character interval. For example, if there is a mapping from [s1, pos1) (the interval corresponding to the first character node) to [s2, pos2) (the interval corresponding to the second character node), instead of creating a new node for mapping, the interval corresponding to the first character node can be directly expanded to [s1, pos1 + len), and the interval corresponding to the second character node can be expanded to [s2, pos2 + len).

[0155] As Figure 7 shown, it is another schematic diagram of inserting characters into a document. As shown in the figure, when inserting the string "application" into the character insertion interval [11, 13) in the display interface of the target document, at this time, the string "application" needs to be newly added in the updated character interval [13, 15) of the character array, and the mapping relationship from the character insertion interval [11, 13) to the updated character interval [13, 15) needs to be constructed. And there is an interval [0, 11) corresponding to an expandable first character node in the first character balanced binary tree and an interval [2, 13) corresponding to an expandable second character node in the second character balanced binary tree. At this time, the interval corresponding to the first character node can be directly expanded to [0, 13), and the interval corresponding to the second character node can be expanded to [2, 15), and the mapping relationship between the first character node and the second character node remains unchanged.

[0156] In some other embodiments, when there is a mapping from [pos1+len, end1) to [pos2+len, end2), it is also possible to directly expand it to a mapping from [pos1, end1) to [pos2, end2) without creating new nodes, that is, only the intervals corresponding to the two nodes with the original mapping relationship need to be expanded.

[0157] When it is detected that there are no first and second character nodes that can expand the interval corresponding to the inserted string, new nodes need to be recreated and the mapping relationship between the new nodes needs to be constructed. Please continue to refer to Figure 6 , Figure 6 The example of Figure 6 is a case where there are no expandable nodes. It is necessary to recreate the third character node corresponding to the interval [7, 9) in the first character balanced binary tree and the fourth character node corresponding to the interval [13, 15) in the second character balanced binary tree, and then construct the mapping relationship between these two newly created nodes. In addition, in the

[0158] example, it is also necessary to split the nodes with the original mapping relationship into two new nodes and recreate the mapping relationship. Moreover, in this case, since new nodes are created in the first character balanced binary tree and the second character balanced binary tree, it is necessary to further perform rebalancing detection on these two character balanced binary trees. If these two character balanced binary trees are no longer balanced after adding new nodes, it is necessary to perform rebalancing processing on these two character balanced binary trees.

[0158] In the embodiments of the present disclosure, after creating a node with a length of len, the intervals corresponding to the nodes after this node need to be shifted accordingly, that is, it is necessary to further update the character intervals corresponding to each node in the first character balanced binary tree according to the interval length of the character insertion interval.

[0159] Step 350, update the second character balanced binary tree and the node mapping relationship of the target document according to the character editing operation.

[0160] Among them, using an AVL tree to describe the attribute stream and the text stream can improve the node search speed when searching for nodes during document editing. However, in the scenario of inserting a string in the middle of a document, when updating the first-character AVL tree, not only new nodes need to be created or nodes need to be expanded, but also the node intervals corresponding to all nodes in the subsequent interval need to be updated. This makes the time complexity of node update unable to reach O(logn). Therefore, the present disclosure further provides a first-character AVL tree constructed based on the relative position relationship, which can be referred to as a floating-point tree in the present disclosure. Specifically, in the embodiments of the present disclosure, the root node of the first-character AVL tree stores character position data, and other nodes in the first-character AVL tree store first relative position distance data relative to the first prior node, where the first prior node is the corresponding left child node or parent node in the first-character AVL tree. Updating the character interval corresponding to each node in the first-character AVL tree according to the interval length of the character insertion interval includes:

[0161] Search for the target character node whose stored data needs to be updated in the first-character AVL tree according to the character position data and the character insertion position corresponding to the character insertion interval;

[0162] Update the character position data or the first relative position distance data stored in the target character node based on the interval length of the character insertion interval.

[0163] That is, in the embodiments of the present disclosure, the root node of the first-character AVL tree stores character position data, and the character position data can specifically be absolute character position data. Then, in other nodes of the first-character AVL tree, first relative position distance data relative to the first prior node is stored, where the first prior node can specifically be the left child node or the parent node of the current node. In this way, when a new node with an interval length of len is created, only the character position data or the first relative position distance data of the corresponding target character node needs to be updated, and there is no need to update the node intervals corresponding to all nodes after this node. Thus, the time complexity of node update can be greatly reduced, and further the efficiency of document editing can be greatly improved.

[0164] Such as Figure 8aAs shown, it is a schematic diagram of the first character balanced binary tree for storing absolute position information in a node. As shown in the figure, this first character balanced binary tree contains 6 nodes, and each node stores an absolute position information, which can specifically be the number information of the characters displayed in the document interface. As introduced above, each node corresponds to an interval. The starting point of the interval corresponding to each node in the figure is the value stored in the previous node, and the ending point is the value stored in the current node. For example, the interval corresponding to the node storing the value 1 is [0, 1), the interval corresponding to the node storing the value 6 is [1, 6), the interval corresponding to the node storing the value 15 is [6, 15), and the interval corresponding to the root node storing the value 18 is [15, 18). In this case, if 3 characters are inserted after the 7th character, and if no new node is established in this case, but the existing nodes are extended, the new first character balanced binary tree as shown in Figure 8b can be obtained. In this case, the node update process needs to update 4 nodes, and the time complexity of node update is greater than O(logn).

[0165] As Figure 9a shown, it is a schematic diagram of the first character balanced binary tree constructed by the method using a floating-point tree provided by the present disclosure. The interval corresponding to each node in the first character balanced binary tree schematically shown in this figure is the same as the interval corresponding to each node in the first character balanced binary tree schematically shown in Figure 8a , except that the value stored in each right child node in this figure is the distance value relative to its parent node, or the relative position value. In this case, if 3 characters are inserted after the 7th character, and if no new node is established in this case, but the existing nodes are extended, the new first character balanced binary tree as shown in Figure 9b can be obtained. In this case, the node update process only needs to update 2 nodes, and the time complexity of node update is O(logn). In this way, the efficiency of node update can be greatly improved, and thus the efficiency of document editing can be greatly improved.

[0166] Similarly, the attribute balanced binary tree, like the first character balanced binary tree, is also a balanced binary tree with continuous node intervals, and the attribute balanced binary tree corresponding to each attribute can also be constructed in the above floating-point tree manner.

[0167] Specifically, in some embodiments, the root node of the attribute balanced binary tree stores attribute position data, and the other nodes in the attribute balanced binary tree store the second relative position distance data relative to the second previous node. The second previous node is the corresponding left child node or parent node in the attribute balanced binary tree. Updating the corresponding attribute balanced binary tree according to the character insertion interval includes:

[0168] Based on the attribute change position determined by the attribute position data and the character insertion interval, search for the target attribute node whose position data needs to be updated in the corresponding attribute balanced binary tree;

[0169] Update the attribute position data or the second relative position distance data of the target attribute node based on the interval length of the character insertion interval.

[0170] In the embodiments of the present disclosure, in the attribute balanced binary tree corresponding to each attribute data, the root node may store the attribute position data, and the attribute position data is absolute position data, which can be specifically determined according to the character number or the gap position between characters. In other nodes, the relative position distance data relative to the corresponding preceding node is stored, where the preceding node here is the parent node or the left child node of the current node. When the attribute balanced binary tree corresponding to each attribute data is constructed in the above floating-point tree manner, when the corresponding attribute balanced binary tree is updated based on the character editing operation, it is only necessary to determine the attribute change position according to the attribute position data and the character insertion interval, then search for the target attribute node whose position data needs to be updated in the corresponding attribute balanced binary tree, and update the attribute position data or the second relative position distance data of the target attribute node based on the interval length of the character insertion interval. This method can avoid updating the node intervals of all subsequent nodes, thereby improving the update efficiency of the attribute balanced binary tree, and further improving the document editing efficiency.

[0171] Among them, in some embodiments, when the type of the character editing operation is not the character insertion type, determine the character editing interval corresponding to the character editing operation, and update the first character balanced binary tree, the second character balanced binary tree, the mapping relationship between the nodes of the first character balanced binary tree and the second character balanced binary tree, and the corresponding attribute balanced binary tree of the target document according to the character editing interval, including:

[0172] When the type of the character editing operation is inserting a space or deleting a space, determine the character editing interval corresponding to the character editing operation, and update the first character balanced binary tree, the mapping relationship between the nodes of the first character balanced binary tree and the second character balanced binary tree, and the corresponding attribute balanced binary tree of the target document according to the character editing interval;

[0173] When the type of the character editing operation is the character deletion type, determine the character deletion interval corresponding to the character editing operation, and update the first character balanced binary tree based on the intersection relationship between the character deletion interval and the node interval corresponding to the node in the first character balanced binary tree, and update the second character balanced binary tree and the mapping relationship between the first character balanced binary tree and the second character balanced binary tree according to the updated first character balanced binary tree.

[0174] Among them, when the type of the character editing operation is other than the character insertion operation, for example, when it is the space insertion or deletion type, the character editing range corresponding to the character editing operation can be determined. Specifically, the character editing range can be the space insertion range or the space deletion range. In the embodiments of the present disclosure, the space insertion range or the space deletion range may not require updating the second character balanced binary tree, but only requires shifting the node range corresponding to the nodes after the insertion or deletion position in the first character balanced binary tree forward or backward by a length that is the same as the length of the space. For the corresponding attribute balanced binary tree, it is also necessary to shift the node range corresponding to the nodes after the insertion or deletion position forward or backward accordingly.

[0175] When the type of the character editing operation is the character deletion type, the character deletion range corresponding to the character editing operation can be determined, and then the intersection relationship between the character deletion range and the node range corresponding to the nodes in the first character balanced binary tree can be determined. Further, the mapping tree and the corresponding attribute balanced binary tree can be updated based on this intersection relationship.

[0176] In some embodiments, updating the first character balanced binary tree based on the intersection relationship between the character deletion range and the node range corresponding to the nodes in the first character balanced binary tree includes:

[0177] Searching for the target deletion character node corresponding to the character deletion range in the first character balanced binary tree;

[0178] When the character deletion range is a sub-range of the character range corresponding to the target deletion character node, splitting the target deletion character node into two fifth character nodes and updating the node ranges corresponding to the fifth character nodes;

[0179] When the character range corresponding to the target deletion character node is a sub-range of the character deletion range, deleting the target character node;

[0180] When the character range corresponding to the target deletion character node and the character deletion range are overlapping ranges, deleting the overlapping part between the character range corresponding to the target deletion character node and the character deletion range, generating a sixth character node according to the remaining part, and updating the node range corresponding to the sixth character node.

[0181] In the embodiments of the present disclosure, the intersection relationship between the character deletion range corresponding to the character deletion operation and the node range corresponding to the corresponding target deletion character node can be divided into three types: In one case, the character deletion range is a sub-range of the node range corresponding to the target deletion character node; in one case, the node range corresponding to the target deletion character node is a sub-range of the character deletion range; and in another case, the node range corresponding to the target deletion character node and the character deletion range partially overlap.

[0182] When the character deletion range is a sub-range of the character range corresponding to the target deletion character node, the target deletion character node can be split into two fifth character nodes, and the node ranges corresponding to the fifth character nodes are updated. For example, if the character range corresponding to the target deletion character node is [s1, e1), and the character deletion range is [pos1, pos1 + len), where s1 < pos1 and e1 > pos1 + len. Then the character deletion operation splits the target deletion character node into two fifth character nodes with character ranges [s1, pos1) and [pos1, e1 - len) respectively. And the node ranges of all nodes after e1 need to be shifted forward by len. Correspondingly, the node in the second character balanced binary tree corresponding to the target deletion character node also needs to be split accordingly, but the change in the ranges of the nodes obtained after splitting is different from the way the node ranges in the first character balanced binary tree are updated.

[0183] As Figure 10 shown, it is a schematic diagram of the mapping tree update based on the character deletion operation. As shown in the figure, the range corresponding to the target deletion character node is [0, 11), and the character deletion range is [7, 9). Then the ranges of the two fifth character nodes obtained after splitting by the character deletion operation are [0, 7) and [7, 9) respectively. And the node in the second character balanced binary tree corresponding to the target deletion character node is split into two new nodes with ranges [2, 9) and [11, 13) respectively, so the mapping relationship between the new nodes needs to be updated.

[0184] When the node range corresponding to the target deletion character node partially overlaps with the character deletion range, for example, the node range corresponding to the target character node is [s1, e1), and the character deletion range is [pos1, pos1 + len), where s1 < pos1 and e1 <= pos1 + len. At this time, the right half of the node range corresponding to the target deletion character node needs to be deleted, that is, the part [pos1, e1), and at the same time, the range of the node mapped in the second character balanced binary tree is reduced. In addition, for the ranges after pos1 in the first character balanced binary tree, they also need to be shifted forward by len.

[0185] When the node range corresponding to the target deletion character node is a sub-range of the character deletion range, the target deletion character node can be directly deleted, and the rest can be processed with reference to the situation where the node range corresponding to the target deletion character node partially overlaps with the character deletion range.

[0186] Therefore, the document data processing method provided by the embodiments of the present disclosure includes receiving a document editing operation for a target document; when the document editing operation is an attribute editing operation, determining a target attribute balanced binary tree among a plurality of attribute balanced binary trees corresponding to document attributes, where the nodes in the attribute balanced binary tree store the attribute information of a section of text in the target document; updating the target attribute balanced binary tree according to the attribute editing operation; when the document editing operation is a character editing operation, updating the first character balanced binary tree of the target document and the corresponding attribute balanced binary tree based on the character editing operation, where the nodes in the first character balanced binary tree correspond to a section in the document display interface of the target document; and updating the second character balanced binary tree and the node mapping relationship of the target document according to the character editing operation, where the nodes in the second character balanced binary tree store a section in the character array corresponding to the target document, and the node mapping relationship is the mapping relationship between the nodes of the first character balanced binary tree and the second character balanced binary tree.

[0187] In the embodiments of the present disclosure, by using an attribute balanced binary tree to store the attribute information of each section of text in the document, and using a first character balanced binary tree and a second character balanced binary tree with a mapping relationship between nodes to store the text of each section in the document display interface and the text of each section in the actual corresponding character array respectively; relying on the time complexity advantages of the balanced binary tree in terms of node search and node update, that is, the time complexity of node search and node update can be reduced to the logarithmic level, thus greatly reducing the rate of node search and update, and further avoiding the problem of full node update required during document editing caused by storing document data in the form of a DOM tree in the related art, thereby greatly improving the processing efficiency of document data. That is, the editing efficiency of adding, deleting, modifying, and querying the document can be improved. In addition, compared with generating a DOM tree by loading all document-related data into memory when using the DOM tree form to carry document data, the present embodiment uses a mapping tree + character array form to carry text stream data, which does not require loading character data into memory data and only needs to read it according to the mapping relationship, so that the memory occupancy can be greatly reduced, thereby further improving the editing efficiency of the document.

[0188] Detailed description of the embodiments of the present disclosure in combination with specific application scenarios

[0189] As Figure 11 shown, it is another flowchart of the document data processing method provided by the present disclosure. In this embodiment, the document data processing method will be introduced in detail in combination with the execution subject of each step. The method specifically includes the following steps:

[0190] Step 1101, in response to the received document startup operation, the computer device obtains the XML data of the target document and parses the XML data of the target document to obtain the attribute position tree and the text mapping tree of the target document.

[0191] In the embodiments of the present disclosure, the in-memory data structure of the document data uses multiple attribute position trees (position trees) to describe the attribute data stream of the document, and uses a text mapping tree (or called a mapping tree, that is, a mirror tree) to describe the text data stream of the document. Among them, the attribute position tree is a balanced binary tree, which can specifically be the floating-point tree introduced in the foregoing embodiments. That is, in the attribute position tree, each node stores the attribute information of an interval, the root node stores an absolute position value, and the other nodes store the position distance values relative to the previous node. The text mapping tree includes two trees with a mapping relationship between nodes, which can specifically be a logical ranges tree and a physical ranges tree. Each node in the logical ranges tree corresponds to an interval in the document display interface, and the physical ranges tree corresponds to an interval in the character array of the document. The characters displayed in the interval in the document display interface are determined by the number of character data stored in an interval corresponding to the character array. In the embodiments of the present disclosure, the logical ranges tree can also be the foregoing floating-point tree. That is, the root node of the logical ranges tree stores the absolute position data of the characters in the document display interface, and the other nodes store the relative position distance data relative to the previous node. The mapping relationship between the logical ranges tree and the physical ranges tree can refer to the mapping relationship between the first character balanced binary tree and the second character balanced binary tree in the foregoing embodiments, which will not be elaborated here.

[0192] In the embodiments of the present disclosure, in response to the received document startup operation of the target document, after the computer device obtains the XML data of the target document, it does not automatically generate a DOM tree, but generates an attribute position tree and a text mapping tree corresponding to the target document.

[0193] Step 1102, the computer device obtains the character data in the character array according to the text mapping tree, and renders and displays the target document in combination with the attribute information in the attribute position tree.

[0194] After generating the attribute position tree and the text mapping tree corresponding to the target document by parsing the XML data of the target document, the computer device can obtain the forward mapping, that is, obtain the character data that should be displayed in each interval in the document interface from the character array of the target document according to the text mapping tree, and then render and display the target document in combination with the attribute information corresponding to each interval in the attribute position tree.

[0195] Step 1103, the computer device receives an attribute editing operation on the target document, and determines the target attribute location tree, the attribute editing range, and the attribute editing content corresponding to the attribute editing operation.

[0196] After the target document is displayed, the computer device can receive a document editing operation on the target document by the document editing object based on the display interface of the target document. The document editing operation can specifically include an attribute editing operation and a character editing operation. The character editing operation further includes a character insertion operation and a character deletion operation.

[0197] When the document editing operation on the target document is an attribute editing operation, the computer device can search for the target attribute location tree corresponding to the attribute editing operation in multiple attribute location trees of the target document. Then, the attribute editing range and the attribute editing content can be further determined according to the editing position in the attribute editing operation.

[0198] Step 1104, the computer device updates the target attribute location tree according to the intersection relationship between the attribute editing range and the range corresponding to the target attribute node in the target attribute location tree, and the attribute editing content.

[0199] After determining the attribute editing range, the computer device can further determine the nodes associated with the range of the attribute editing range in the target attribute location tree, and determine these nodes as target attribute nodes. There can be an inclusion or intersection relationship between the range corresponding to the target attribute node and the attribute editing range. For each target attribute node, when the range corresponding to it is within the attribute editing range, the attribute information stored in the node is directly modified. When there is a partial range overlap between the range corresponding to it and the attribute editing range, the target attribute node is split into two nodes, the attribute information stored in the node with the overlapping range and the attribute editing range is modified, and finally the target attribute location tree is rebalanced. When the attribute editing range is inside the range corresponding to the target attribute node, the target attribute node can be split into two or three new nodes, then the attribute information stored in the node corresponding to the attribute editing range is modified, and finally the target attribute location tree is rebalanced.

[0200] Step 1105, the computer device receives a character insertion operation on the target document, updates the character array according to the inserted character content corresponding to the character insertion operation, and detects whether there is an expandable node in the text mapping tree according to the character insertion range corresponding to the character insertion operation.

[0201] When the received document editing operation is a character insertion operation, the computer device first updates the character array according to the inserted character content corresponding to the character insertion operation, that is, adds the inserted characters to the end part of the character array of the target document, and determines the new character range corresponding to the inserted characters in the character array. Then, it detects whether there is an expandable node in the text mapping tree according to the character insertion range corresponding to the character insertion operation and the new character range. Among them, the expandable node can specifically be expressed as a node range that can be extended forward or backward by a length of len, where len is the length corresponding to the character insertion range and the new character range, and the position of the inserted character and the position of the new character array are both in front of or behind the character positions or character array positions corresponding to the two expandable nodes. Specific examples of expandable nodes have been Figure 3 introduced in detail in the corresponding embodiments and will not be elaborated here.

[0202] Step 1106, when there is an expandable node in the text mapping tree, the computer device expands the node range of the expandable node.

[0203] When it is detected that there is an expandable node in the text mapping tree, the computer device can expand the range of the expandable node, thereby realizing the update of the text mapping tree.

[0204] Step 1107, when there is no expandable node in the text mapping tree, the computer device establishes a pair of nodes with a mapping relationship according to the character insertion range and the new character range, and updates the text mapping tree according to the node ranges of the pair of nodes.

[0205] When there is no expandable node in the text mapping tree, the computer device establishes a pair of nodes with a mapping relationship according to the character insertion range and the new character range, that is, creates a new node in the logical interval tree and the physical interval tree respectively, and there is a corresponding mapping relationship between these two new nodes.

[0206] In addition, it is also necessary to detect whether the character insertion range interrupts the range corresponding to the existing nodes in the logical interval tree. If so, the existing nodes of the interrupted range need to be split into two new nodes, and correspondingly, the nodes in the physical interval tree that have a mapping relationship with the interrupted range also need to be split into two new nodes. Then, update the text mapping tree according to the newly generated multiple nodes. If it does not interrupt the range corresponding to the existing nodes in the logical interval, directly generate a new pair of nodes and update the text mapping tree according to the newly generated pair of nodes.

[0207] Step 1108, the computer device receives a character deletion operation on the target document, determines the character deletion range corresponding to the character deletion operation, and determines the target character node corresponding to the character deletion operation in the logical interval tree.

[0208] When the document editing operation on the target document is a character deletion operation, the character deletion range corresponding to the character deletion operation can be determined first. When deleting characters from the document, the content in the character array of the document will not be affected. It only needs to deprecate the characters in the character array corresponding to the deleted characters, that is, delete the range in the physical interval tree. Therefore, after determining the character deletion range corresponding to the character deletion operation, the logical interval tree can be updated according to the character deletion range first, and then the corresponding physical interval tree can be updated accordingly.

[0209] Specifically, the target character node corresponding to the character deletion operation can be determined in the logical interval tree according to the character deletion range first, and then the target character node can be updated.

[0210] Step 1109, the computer device updates the text mapping tree according to the intersection relationship between the character deletion range and the node range of the target character node.

[0211] After determining the target character node corresponding to the character deletion operation in the logical interval tree, the character deletion range can be compared with the node range corresponding to the target character node to determine the intersection relationship between the two. The intersection relationship specifically includes three types. These three intersection relationships and the methods for updating the text mapping tree corresponding to different intersection relationships have also been introduced in the foregoing embodiments and will not be elaborated here.

[0212] Among them, in this case, the character insertion operation and the character deletion operation specifically adopt the piece table method to replace the traditional string method. In this method, it only needs to distinguish which range is from the original data and which ranges are from the added data. As Figure 12 shown, it is a schematic diagram of the piece table processing method adopted in the present disclosure. Among them, in the new text sequence 1230, the content of some ranges comes from the original file 1210, and the content of some ranges comes from the added file 1220. By adopting this method, there is no need to modify the entire character array, which can greatly improve the efficiency of character insertion and deletion.

[0213] As shown in Table 1 and Table 2 below, it is a comparison table of the editing efficiency of the document editing method provided in this case and the document editing method in the related art.

[0214]

[0215] Table 1 Comparison table of character processing efficiency

[0216] Among them, Table 1 is a comparison table of the efficiency of character processing. As shown in Table 1, the solution provided by the present disclosure can achieve excellent character editing efficiency.

[0217]

[0218] Table 2 Efficiency Comparison Table for Interval Processing

[0219] Among them, Table 2 is the efficiency comparison table for interval processing. As shown in Table 2, the solution provided by the present disclosure can achieve higher interval processing efficiency.

[0220] Device and equipment description of the embodiments of the present disclosure

[0221] It can be understood that although each step in the above-mentioned various flowcharts is sequentially displayed according to the indication of the arrow, these steps do not necessarily need to be sequentially executed according to the order indicated by the arrow. Unless there is a clear description in this embodiment, the execution of these steps has no strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above-mentioned flowcharts may include multiple steps or multiple stages. These steps or stages do not necessarily need to be executed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily need to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0222] It should be noted that in each specific implementation manner of the present disclosure, when it comes to performing relevant processing according to data related to the characteristics of the target object, such as the target object attribute information or the attribute information set, etc., the permission or consent of the target object will be obtained first. Moreover, the collection, use, and processing of these data will comply with the relevant laws, regulations, and standards of the relevant region. In addition, when the embodiment of the present application needs to obtain the target object attribute information, it will obtain the separate permission or separate consent of the target object through methods such as popping up a window or jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the target object, the necessary data related to the target object for the normal operation of the embodiment of the present application will be obtained.

[0223] Figure 13 It is a schematic structural diagram of the document data processing device 1300 provided by the embodiment of the present disclosure. The device includes:

[0224] A receiving unit 1310, configured to receive a document editing operation for a target document;

[0225] A determining unit 1320, configured to determine a target attribute balanced binary tree in a plurality of attribute balanced binary trees corresponding to document attributes when the document editing operation is an attribute editing operation, and nodes in the attribute balanced binary tree store attribute information of a section of interval text of the target document;

[0226] A first updating unit 1330, configured to update the target attribute balanced binary tree according to the attribute editing operation;

[0227] A second update unit 1340, configured to update the first character AVL tree and the corresponding attribute AVL tree of the target document based on a character editing operation when the document editing operation is a character editing operation, where a node in the first character AVL tree corresponds to an interval in the document display interface of the target document; and

[0228] A third update unit 1350, configured to update the second character AVL tree and the node mapping relationship of the target document according to the character editing operation, where a node in the second character AVL tree stores an interval in the character array corresponding to the target document, and the node mapping relationship is a mapping relationship between nodes of the first character AVL tree and the second character AVL tree.

[0229] Optionally, in some embodiments, the first update unit includes:

[0230] An acquisition subunit, configured to acquire an attribute editing interval and attribute editing content corresponding to an attribute editing operation;

[0231] A search subunit, configured to search for a target attribute node in the target attribute AVL tree according to the attribute editing interval;

[0232] A first update subunit, configured to update the target attribute AVL tree according to an intersection relationship between the attribute editing interval and a target node interval corresponding to the target attribute node, and the attribute editing content.

[0233] Optionally, in some embodiments, the first update subunit includes:

[0234] A first update module, configured to update attribute information corresponding to the target attribute node according to the attribute editing content when an intersection of the attribute editing interval and a target node interval corresponding to the target attribute node is the target node interval;

[0235] A splitting module, configured to split the target attribute node into at least two new attribute nodes according to a sub-interval when an intersection of the attribute editing interval and a target node interval corresponding to the target attribute node is a sub-interval of the target node interval;

[0236] A processing module, configured to update attribute information of corresponding new attribute nodes according to the attribute editing content, and perform a balancing process on the target attribute AVL tree based on the new attribute nodes.

[0237] Optionally, in some embodiments, the second update unit includes:

[0238] An identification subunit, configured to identify a type of the character editing operation when the document editing operation is a character editing operation;

[0239] A second update subunit, configured to, when the type of the character editing operation is the character insertion type, determine a character insertion range and insertion character content, and update a first character balanced binary tree, a second character balanced binary tree, a mapping relationship between nodes of the first character balanced binary tree and the second character balanced binary tree, and a corresponding attribute balanced binary tree of the target document according to the character insertion range and the insertion character content;

[0240] A third update subunit, configured to, when the type of the character editing operation is not the character insertion type, determine a character editing range corresponding to the character editing operation, and update a first character balanced binary tree, a second character balanced binary tree, a mapping relationship between nodes of the first character balanced binary tree and the second character balanced binary tree, and a corresponding attribute balanced binary tree of the target document according to the character editing range.

[0241] Optionally, in some embodiments, the second update subunit includes:

[0242] A second update module, configured to update a character array of the target document according to the insertion character content, and determine an updated character range corresponding to the insertion character content;

[0243] A third update module, configured to update a first character balanced binary tree, a second character balanced binary tree, and a mapping relationship between nodes of the first character balanced binary tree and the second character balanced binary tree of the target document based on the updated character range and the character insertion range;

[0244] A fourth update module, configured to update a corresponding attribute balanced binary tree according to the character insertion range.

[0245] Optionally, in some embodiments, the second update module includes:

[0246] An adding submodule, configured to add characters corresponding to the insertion character content at the end of the character array of the target document to obtain an updated character array;

[0247] A determining submodule, configured to obtain number information of characters in the updated character array, and determine an updated character range corresponding to the insertion character content according to the number information.

[0248] Optionally, in some embodiments, the third update module includes:

[0249] An expanding submodule, configured to, when it is detected that there are first character nodes and second character nodes with expandable ranges in a first character balanced binary tree and a second character balanced binary tree of the target document, expand the first character nodes based on the character insertion range and expand the second character nodes based on the updated character range;

[0250] The first character node is a node in the first character balanced binary tree, the second character node is a node in the second character balanced binary tree, there is a mapping relationship between the first character node and the second character node, the interval corresponding to the first character node can be extended based on the character insertion interval, and the interval corresponding to the second character node can be extended based on the updated character interval;

[0251] The first update sub-module is used to update the character interval corresponding to each node in the first character balanced binary tree according to the interval length of the character insertion interval;

[0252] The second update sub-module is used to, when it is detected that there are no first character nodes and second character nodes in the first character balanced binary tree and the second character balanced binary tree of the target document whose intervals can be extended, create a third character node and a fourth character node respectively based on the character insertion interval and the updated character interval, and update the first character balanced binary tree according to the third character node and update the second character balanced binary tree according to the fourth character node.

[0253] Optionally, in some embodiments, the root node of the first character balanced binary tree stores character position data, and other nodes in the first character balanced binary tree store first relative position distance data relative to the first prior node, where the first prior node is the corresponding left child node or parent node in the first character balanced binary tree. The first update sub-module can also be used to:

[0254] Search for the target character node that needs to update the stored data in the first character balanced binary tree according to the character position data and the character insertion position corresponding to the character insertion interval;

[0255] Update the character position data or the first relative position distance data stored in the target character node based on the interval length of the character insertion interval.

[0256] Optionally, in some embodiments, the root node of the attribute balanced binary tree stores attribute position data, and other nodes in the attribute balanced binary tree store second relative position distance data relative to the second prior node, where the second prior node is the corresponding left child node or parent node in the attribute balanced binary tree. The fourth update module includes:

[0257] The third update sub-module is used to search for the target attribute node that needs to update the position data in the corresponding attribute balanced binary tree according to the attribute position data and the attribute change position determined according to the character insertion interval;

[0258] The fourth update sub-module is used to update the attribute position data or the second relative position distance data of the target attribute node based on the interval length of the character insertion interval.

[0259] Optionally, in some embodiments, the third update subunit includes:

[0260] A fifth update module, configured to, when the type of the character editing operation is inserting a space or deleting a space, determine a character editing range corresponding to the character editing operation, and update a first character balanced binary tree of the target document, a mapping relationship between nodes of the first character balanced binary tree and a second character balanced binary tree, and a corresponding attribute balanced binary tree according to the character editing range;

[0261] A sixth update module, configured to, when the type of the character editing operation is a character deletion type, determine a character deletion range corresponding to the character editing operation, and update the first character balanced binary tree based on an intersection relationship between the character deletion range and a node range corresponding to a node in the first character balanced binary tree, and update the second character balanced binary tree and a mapping relationship between the first character balanced binary tree and the second character balanced binary tree according to the updated first character balanced binary tree.

[0262] Optionally, in some embodiments, the sixth update module is further configured to:

[0263] Search for a target deletion character node corresponding to the character deletion range in the first character balanced binary tree;

[0264] When the character deletion range is a sub-range of the character range corresponding to the target deletion character node, split the target deletion character node into two fifth character nodes, and update the node ranges corresponding to the fifth character nodes;

[0265] When the character range corresponding to the target deletion character node is a sub-range of the character deletion range, delete the target character node;

[0266] When the character range corresponding to the target deletion character node and the character deletion range are intersecting ranges, delete the intersecting part of the character range corresponding to the target deletion character node and the character deletion range, generate a sixth character node according to the remaining part, and update the node range corresponding to the sixth character node.

[0267] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit including the function of the module or unit.

[0268] Refer to Figure 14 , Figure 14Block diagram of a part of the terminal 140 for implementing the document data processing method according to an embodiment of the present disclosure. The terminal 140 includes components such as a Radio Frequency (RF) circuit 1410, a memory 1415, an input unit 1430, a display unit 1440, a sensor 1450, an audio circuit 1460, a wireless fidelity (WiFi) module 1470, a processor 1480, and a power supply 1490. Those skilled in the art can understand that Figure 14 The shown structure of the terminal 140 does not limit a mobile phone or a computer, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0269] The RF circuit 1410 can be used for receiving and sending signals during information transceiver or call processes. Specifically, after receiving the downlink information from the base station, it is given to the processor 1480 for processing; in addition, the designed uplink data is sent to the base station.

[0270] The memory 1415 can be used to store software programs and modules. The processor 1480 executes various functional applications and document editing of the terminal by running the software programs and modules stored in the memory 1415.

[0271] The input unit 1430 can be used to receive input digital or character information, and generate key signal inputs related to the settings and function controls of the terminal. Specifically, the input unit 1430 may include a touch panel 1431 and other input devices 1432.

[0272] The display unit 1440 can be used to display input information or provided information and various menus of the terminal. The display unit 1440 may include a display panel 1441.

[0273] The audio circuit 1460, the speaker 1461, and the microphone 1462 can provide an audio interface.

[0274] In this embodiment, the processor 1480 included in the terminal 140 can execute the document data processing method of the previous embodiment.

[0275] The terminal 140 according to an embodiment of the present disclosure includes, but is not limited to, a mobile phone, a computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc.

[0276] Figure 15Structural block diagram of a part of server 110 for implementing the document data processing method of the embodiments of the present disclosure. Server 110 may vary significantly due to configuration or performance differences, and may include one or more central processing units (CPUs) 1522 (e.g., one or more processors) and a storage device 1532, and one or more storage media 1530 (e.g., one or more mass storage devices) for storing application programs 1542 or data 1544. Among them, the storage device 1532 and the storage medium 1530 may be transient storage or persistent storage. The program stored in the storage medium 1530 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server 110. Further, the central processing unit 1522 may be configured to communicate with the storage medium 1530 and execute a series of instruction operations in the storage medium 1530 on the server 110.

[0277] Server 110 may further include one or more power supplies 1526, one or more wired or wireless network interfaces 1550, one or more input / output interfaces 1558, and / or one or more operating systems 1541, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0278] The central processing unit 1522 in server 110 may be used to execute the document data processing method of the embodiments of the present disclosure.

[0279] The embodiments of the present disclosure further provide a storage medium for storing program code for executing the document data processing method of the foregoing respective embodiments.

[0280] The embodiments of the present disclosure further provide a computer program product, which includes a computer program. The processor of the computer device reads and executes the computer program, so that the computer device executes to implement the above-mentioned document data processing method.

[0281] The terms "first", "second", "third", "fourth", etc. (if any) in the description of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0282] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B may be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c may be single or multiple.

[0283] It should be understood that in the description of the embodiments of the present disclosure, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as greater than, less than, exceeding, etc. do not include the present number, and understandings such as above, below, within, etc. include the present number.

[0284] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0285] The unit described as a separation component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0286] In addition, each functional unit in various embodiments of the present disclosure may be integrated in a processing unit, may exist physically separately for each unit, or two or more units may be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0287] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present disclosure. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0288] It should also be understood that the various embodiments provided in the present disclosure can be combined arbitrarily to achieve different technical effects.

[0289] The above is a specific description of the embodiments of the present disclosure, but the present disclosure is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present disclosure, and these equivalent deformations or substitutions are all included in the scope defined by the claims of the present disclosure.

Claims

1. A method for processing document data, characterized in that, The method includes: Receiving a document editing operation for a target document; When the document editing operation is an attribute editing operation, determining a target attribute balanced binary tree among a plurality of attribute balanced binary trees corresponding to document attributes, where nodes in the attribute balanced binary tree store attribute information of a section of text of the target document; Updating the target attribute balanced binary tree according to the attribute editing operation; When the document editing operation is a character editing operation, updating the first character balanced binary tree of the target document and the corresponding attribute balanced binary tree based on the character editing operation, where nodes in the first character balanced binary tree correspond to a section in the document display interface of the target document; and Updating the second character balanced binary tree of the target document and the node mapping relationship according to the character editing operation, where nodes in the second character balanced binary tree store a section of the character array corresponding to the target document, and the node mapping relationship is the mapping relationship between nodes of the first character balanced binary tree and the second character balanced binary tree.

2. The method according to claim 1, characterized in that The updating the target attribute balanced binary tree according to the attribute editing operation includes: Obtaining an attribute editing range and attribute editing content corresponding to the attribute editing operation; Searching for a target attribute node in the target attribute balanced binary tree according to the attribute editing range; Updating the target attribute balanced binary tree according to the intersection relationship between the attribute editing range and the target node range corresponding to the target attribute node, and the attribute editing content.

3. The method according to claim 2, characterized in that, The updating the target attribute balanced binary tree according to the intersection relationship between the attribute editing range and the target node range corresponding to the target attribute node, and the attribute editing content includes: When the intersection of the attribute editing range and the target node range corresponding to the target attribute node is the target node range, updating the attribute information corresponding to the target attribute node according to the attribute editing content; When the intersection of the attribute editing range and the target node range corresponding to the target attribute node is a sub-range of the target node range, splitting the target attribute node into at least two new attribute nodes according to the sub-range; Updating the attribute information of the corresponding new attribute nodes according to the attribute editing content, and performing a balancing process on the target attribute balanced binary tree based on the new attribute nodes.

4. The method according to claim 1, characterized in that, The updating the first character balanced binary tree of the target document and the corresponding attribute balanced binary tree based on the character editing operation when the document editing operation is a character editing operation, and the updating the second character balanced binary tree of the target document and the node mapping relationship according to the character editing operation includes: When the document editing operation is a character editing operation, identifying the type of the character editing operation; When the type of the character editing operation is the character insertion type, determine the character insertion range and the inserted character content, and update the first character balanced binary tree, the second character balanced binary tree, the mapping relationship between the nodes of the first character balanced binary tree and the second character balanced binary tree, and the corresponding attribute balanced binary tree of the target document according to the character insertion range and the inserted character content; When the type of the character editing operation is not the character insertion type, determine the character editing range corresponding to the character editing operation, and update the first character balanced binary tree, the second character balanced binary tree, the mapping relationship between the nodes of the first character balanced binary tree and the second character balanced binary tree, and the corresponding attribute balanced binary tree of the target document according to the character editing range.

5. The method according to claim 4, characterized in that, The updating the first character balanced binary tree, the second character balanced binary tree, the mapping relationship between the nodes of the first character balanced binary tree and the second character balanced binary tree, and the corresponding attribute balanced binary tree of the target document according to the character insertion range and the inserted character content includes: Update the character array of the target document according to the inserted character content, and determine the updated character range corresponding to the inserted character content; Update the first character balanced binary tree, the second character balanced binary tree, and the mapping relationship between the nodes of the first character balanced binary tree and the second character balanced binary tree of the target document based on the updated character range and the character insertion range; Update the corresponding attribute balanced binary tree according to the character insertion range.

6. The method according to claim 5, characterized in that, The updating the character array of the target document according to the inserted character content, and determining the updated character range corresponding to the inserted character content includes: Add the characters corresponding to the inserted character content at the end of the character array of the target document to obtain an updated character array; Obtain the number information of the characters in the updated character array, and determine the updated character range corresponding to the inserted character content according to the number information.

7. The method according to claim 5, wherein The updating the first character balanced binary tree, the second character balanced binary tree, and the mapping relationship between the nodes of the first character balanced binary tree and the second character balanced binary tree of the target document based on the updated character range and the character insertion range includes: When it is detected that there are first character nodes and second character nodes with expandable ranges in the first character balanced binary tree and the second character balanced binary tree of the target document, expand the first character nodes based on the character insertion range and expand the second character nodes based on the updated character range; The first character node is a node in the first character AVL tree, the second character node is a node in the second character AVL tree, there is a mapping relationship between the first character node and the second character node, the interval corresponding to the first character node can be extended based on the character insertion interval, and the interval corresponding to the second character node can be extended based on the updated character interval; Update the character interval corresponding to each node in the first character AVL tree according to the interval length of the character insertion interval; When it is detected that there are no first character nodes and second character nodes with extensible intervals in the first character AVL tree and the second character AVL tree of the target document, create a third character node and a fourth character node respectively based on the character insertion interval and the updated character interval, and update the first character AVL tree according to the third character node and update the second character AVL tree according to the fourth character node.

8. The method according to claim 7, characterized in that, The root node of the first character AVL tree stores character position data, and other nodes in the first character AVL tree store first relative position distance data relative to a first prior node, where the first prior node is the corresponding left child node or parent node in the first character AVL tree. The updating of the character interval corresponding to each node in the first character AVL tree according to the interval length of the character insertion interval includes: Search for a target character node that needs to update the stored data in the first character AVL tree according to the character position data and the character insertion position corresponding to the character insertion interval; Update the character position data or the first relative position distance data stored in the target character node based on the interval length of the character insertion interval.

9. The method according to claim 5, wherein The root node of the attribute AVL tree stores attribute position data, and other nodes in the attribute AVL tree store second relative position distance data relative to a second prior node, where the second prior node is the corresponding left child node or parent node in the attribute AVL tree. The updating of the corresponding attribute AVL tree according to the character insertion interval includes: Search for a target attribute node that needs to update the position data in the corresponding attribute AVL tree according to the attribute position data and the attribute change position determined by the character insertion interval; Update the attribute position data or the second relative position distance data of the target attribute node based on the interval length of the character insertion interval.

10. The method according to claim 4, characterized in that When the type of the character editing operation is not the character insertion type, determine the character editing interval corresponding to the character editing operation, and update the first character AVL tree, the second character AVL tree, the mapping relationship between the nodes of the first character AVL tree and the second character AVL tree, and the corresponding attribute AVL tree according to the character editing interval, including: When the type of the character editing operation is inserting a space or deleting a space, determine the character editing range corresponding to the character editing operation, and update the first character AVL tree of the target document, the mapping relationship between the nodes of the first character AVL tree and the second character AVL tree, and the corresponding attribute AVL tree according to the character editing range; When the type of the character editing operation is deleting a character type, determine the character deletion range corresponding to the character editing operation, and update the first character AVL tree based on the intersection relationship between the character deletion range and the node range corresponding to the nodes in the first character AVL tree, and update the second character AVL tree and the mapping relationship between the first character AVL tree and the second character AVL tree according to the updated first character AVL tree.

11. The method according to claim 10, characterized in that, The updating the first character AVL tree based on the intersection relationship between the character deletion range and the node range corresponding to the nodes in the first character AVL tree includes: Search for a target deletion character node corresponding to the character deletion range in the first character AVL tree; When the character deletion range is a sub-range of the character range corresponding to the target deletion character node, split the target deletion character node into two fifth character nodes, and update the node ranges corresponding to the fifth character nodes; When the character range corresponding to the target deletion character node is a sub-range of the character deletion range, delete the target character node; When the character range corresponding to the target deletion character node and the character deletion range are overlapping ranges, delete the overlapping part of the character range corresponding to the target deletion character node and the character deletion range, generate a sixth character node according to the remaining part, and update the node range corresponding to the sixth character node.

12. A document data processing device, characterized in that, The apparatus includes: a receiving unit, configured to receive a document editing operation for a target document; a determining unit, configured to determine a target attribute AVL tree in the attribute AVL trees corresponding to multiple document attributes when the document editing operation is an attribute editing operation, where the nodes in the attribute AVL tree store the attribute information of a section of text in the target document; a first updating unit, configured to update the target attribute AVL tree according to the attribute editing operation; a second updating unit, configured to update the first character AVL tree of the target document and the corresponding attribute AVL tree based on the character editing operation when the document editing operation is a character editing operation, where the nodes in the first character AVL tree correspond to a section of the document display interface of the target document; and a third updating unit, configured to update the second character AVL tree and the node mapping relationship of the target document according to the character editing operation, where the nodes in the second character AVL tree store a section of the character array corresponding to the target document, and the node mapping relationship is the mapping relationship between the nodes of the first character AVL tree and the second character AVL tree.

13. A storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the document data processing method according to any one of claims 1 to 11.

14. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the document data processing method according to any one of claims 1 to 11.

15. A computer program product, which includes a computer program that is read and executed by a processor of a computer device, so that the computer device executes the document data processing method according to any one of claims 1 to 11.

Citation Information

Cited By

  • Document object model updating processing method, electronic equipment and storage medium

    CN120973804A