Merging tree for collaboration
Patent Information
- Application Number
- CN202080032868.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-03
- Filing Date
- 2020-04-24
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2040-04-24
Smart Images

Figure CN114127690B_ABST
Abstract
Description
[0001] Priority requirements
[0002] This patent application claims priority to U.S. Patent Application No. 16 / 402,730, filed May 3, 2019, entitled “MERGE TREES FOR COLLABORATION,” the entire contents of which are incorporated herein by reference. Background Technology
[0003] In fault-tolerant distributed systems, atomic broadcast or total-order broadcast can be used to ensure that multiple distributed processes receive operations in an equivalent sequence, regardless of which node in the distributed system initiates each operation. The operation can be propagated to every node in the distributed system, so that each operation is completed at each node, or the operation is rolled back at each node. Attached Figure Description
[0004] In accompanying drawings that are not necessarily drawn to scale, similar numbers may describe similar parts in different views. The same numbers with different letter suffixes may represent different instances of similar components. The accompanying drawings generally illustrate the various embodiments discussed in this document by way of example and not limitation.
[0005] Figure 1A This is an overview diagram of an exemplary collaborative messaging architecture.
[0006] Figure 1B This is an exemplary state transition diagram illustrating how an operation moves between three states.
[0007] Figure 2A An exemplary message format is shown for messages exchanged between a collaboration module instance and a synchronization service in one or more of the disclosed embodiments.
[0008] Figure 2B This is an overview diagram showing the relationship between messages and the set of operations ordered sequentially.
[0009] Figure 2C An exemplary message communication between two collaboration module instances 105a-b and synchronization service 106 is shown.
[0010] Figure 2D An exemplary message exchange between a collaboration module instance and a synchronization service is shown.
[0011] Figure 2E This is an overview diagram of the snapshot process.
[0012] Figure 3Two data structures that can be used to construct a tree are shown in some embodiments of the disclosed examples.
[0013] Figure 4 An example of a merged tree portion stored by a collaboration module instance is shown.
[0014] Figure 5 An exemplary merge tree section stored by a collaboration module instance is shown.
[0015] Figure 6 This shows an updated version of the merged tree portion stored by the collaboration module instance.
[0016] Figure 7 This shows an updated version of the merged tree portion stored by the collaboration module instance.
[0017] Figure 8 Showing from Figure 4 The updated version of the merged tree section.
[0018] Figure 9 This shows an updated version of the merged tree portion stored by the collaboration module instance.
[0019] Figure 10 The updated tree portion stored by the collaboration module instance is shown.
[0020] Figure 11 An exemplary merge tree section stored by a collaboration module instance is shown.
[0021] Figure 12 The merge tree section indicating confirmation of the insertion is shown.
[0022] Figure 13 The merged tree portion stored by the collaboration module instance is shown.
[0023] Figure 14 This is a flowchart illustrating an exemplary process for distributing operations on a distributed data structure across multiple instances of collaborating modules.
[0024] Figure 15 This is a flowchart illustrating an exemplary process for distributing operations on a distributed data structure across multiple instances of collaborating modules.
[0025] Figure 16 This is a flowchart illustrating an exemplary process for distributing operations on a distributed data structure across multiple instances of collaborating modules.
[0026] Figure 17 This is a flowchart illustrating an exemplary process for accessing a distributed data structure.
[0027] Figure 18This is a flowchart illustrating an exemplary process for accessing a distributed data structure.
[0028] Figure 19 This is a flowchart illustrating an exemplary process for accessing a distributed data structure.
[0029] Figure 20 This is a flowchart illustrating an exemplary process for accessing a distributed data structure.
[0030] Figure 21 This is a flowchart illustrating an exemplary process for accessing a distributed data structure.
[0031] Figure 22 This is a flowchart illustrating an exemplary process for accessing a distributed data structure.
[0032] Figure 23 A block diagram of an exemplary machine is illustrated, on which any one or more of the techniques (e.g., methods) discussed herein can be performed in one or more of the disclosed embodiments.
[0033] Figure 24 This is a flowchart of an exemplary process that can be implemented by a synchronization service. Detailed Implementation
[0034] The following description and accompanying drawings fully illustrate specific embodiments to enable those skilled in the art to practice thereon. Other embodiments may be incorporated with structural, logical, electrical, process, and other variations. Parts and features of some embodiments may be included in or replace those parts and features of other embodiments. The embodiments set forth in the claims cover all available equivalents of those claims.
[0035] As discussed above, total order broadcast is a technique for ensuring that operations are completed across every node in a distributed system or not performed by any of those nodes. The disclosed embodiments utilize a total order broadcast architecture to provide a collaborative environment for accessing distributed data structures (DDS). A centralized serialization process defines the order in which operations are initiated by multiple collaborating participants. In at least some aspects, this order can be represented by a sequence number assigned to each operation by the centralized serialization process. Once a sequence number is assigned to an operation, the information defining the operation and its assigned sequence number are distributed to all collaborating participants.
[0036] The disclosed embodiments also provide indications of which operations have been confirmed by all participants in the collaboration and which operations have not yet been confirmed. Operations confirmed by all participants form the lower bound (exclusively) of the "collaboration window," representing pending or otherwise incompletely confirmed operations across participants. The top of the collaboration window is indicated by the most recent or highest-ranked operations.
[0037] When the disclosed embodiments are operated, the "top" end advances through higher-order sequence numbers, and as time passes, the lower end also advances by receiving confirmation of the operation from each participant individually. Therefore, the collaboration window represents a "scrolling" window for serialization operations on a distributed data structure.
[0038] In the disclosed embodiments, each participant (e.g., a computing device as part of the collaboration, or an instance of a collaboration module, as discussed further below) is individually responsible for applying each serialization operation to its own copy of the distributed data structure. Thus, if an operation is initiated by a first participant on a distributed data structure accessed by a collaboration session comprising twenty participants, the operation will be physically executed twenty times on twenty local copies of the distributed data structure, once for each collaborating participant.
[0039] The disclosed embodiments are implemented in part by instructions that configure hardware processing circuitry to perform operations implementing the disclosed embodiments. For ease of discussion, these instructions are collectively referred to as cooperative modules. These instructions can be instantiated to provide a software program that can access instructions and data necessary to perform the instructions and functions described herein. This instantiation program is referred to throughout this disclosure as a cooperative module instance. Each cooperative module instance operates to configure appropriate hardware processing circuitry to perform one or more of the functions discussed herein and attributed to it. While a particular function may be attributed to a cooperative module instance, it is not required that the cooperative module instance explicitly or implicitly include instructions for implementing all the functions described herein. Rather, when a particular function is attributed to a cooperative module instance, the cooperative module instance only needs to include instructions necessary to implement that particular function.
[0040] Some embodiments disclosed provide a next-generation architecture for collaborative text editing. To achieve success, these improvements should offer fast load times, reduced latency in change propagation, a smooth user experience, and high scalability at a reasonable cost. This next-generation architecture should also provide a scalable context while offering a relatively simple implementation, for example, by providing a stateless operating mode. The next-generation architecture should combine intelligent services with online collaboration and support document branching.
[0041] The disclosed text editing process includes at least three operations: insertion, deletion, and annotation. Because multiple participants in the collaboration may simultaneously edit the same portion of the text, conflict rules are established to determine how to resolve conflicting edits across collaboration module instances when detected. The disclosed embodiments distribute changes occurring at each collaboration module instance to every other collaboration module instance via a single serialization service, which may be implemented via a centralized service (such as a synchronization service discussed below) or via a peer-to-peer protocol.
[0042] The serialization service assigns a sequence number to each operation and distributes each operation within the operation to all participants in the collaboration. When a particular operation has been successfully distributed to each collaboration module instance, the serialization service also notifies all collaboration module instances, enabling all instances to fully incorporate the operation into a permanent copy of their distributed data structure (e.g., moving the bottom of the collaboration window forward in operation order). Where appropriate, messages between participating collaboration module instances and the serialization service can be encrypted to ensure secure communication.
[0043] The disclosed embodiments can be implemented using a publish / subscribe data pattern. For example, Syphon is a highly available and reliable distributed publish / subscribe system built using Apache Kafka and can be used in some embodiments of the disclosed embodiments. The disclosed embodiments can provide a variety of different types of distributed data structures, including lists, maps, sequences, and streams. In some aspects, these data structures can be implemented via merge trees.
[0044] The disclosed embodiments provide a collaboration model that results in a consistent state across all collaborating module instances. Changes are optimistically replicated because when a change occurs, it is distributed to the collaborating module instances and then reversed in the rare event that the operation cannot be distributed to all collaborating module instances participating in the collaboration. The disclosed embodiments provide immediate local application of local changes, resulting in low latency for interactive environments. Furthermore, the operational primitives provided by the disclosed embodiments can be used to construct more complex distributed data structures, such as rich strings, tables, counters, graphs, and other types of distributed data structures.
[0045] Some implementations provide the use of a merge tree to represent a shared document structure. The disclosed merge tree implementation delegates the integration of changes from all participating collaborative module instances to each individual collaborative module instance. While this results in replication of each operation across each collaborative module instance, it also reduces the processing requirements for synchronization services. The merge tree supports hierarchical rich text constructions, such as nested tables, and enables multi-stream text with cross-references (such as footnotes). Some embodiments of the disclosed examples provide intelligent services capable of efficiently caching information in the merge tree and referencing the merge tree position in both space (position in the text) and time (position in the revision history).
[0046] The disclosed merge tree provides branching, merging, and continuous integration from another branch of the tree. The merge tree can efficiently utilize storage and provides retrieval of points in the history of the shared merge tree structure. Joining a collaborative session is facilitated by providing a copy of the recorded document and any pending changes to the joining collaboration module instance. The disclosed implementation can provide local garbage collection, resulting in stable memory usage and compact storage. The disclosed merge tree provides constant time and space requirements for operational processing in the synchronization service, leading to cost-effective synchronization service deployment.
[0047] The disclosed implementations further offer improved performance for local operations because local operations take into account the total length of the data structure under collaboration. These implementations can provide an equivalent representation between the local and collaborative models, in contrast to Google's Operation Transformation (OT) and Conflict Tree Copy Data Type (CRDT) implementations that provide a split representation. This simplifies the implementations relative to OT and CRDT. The disclosed embodiments also provide a simplified undo / redo architecture. By traversing the tree upwards to identify the correct location / range, information supporting undo / redo operations can be read directly from the segments of the merged tree. Change tracking is also inherent in the disclosed merged tree because the segments of the merged tree directly map to every change made to the underlying collaborative data structure. Furthermore, cloning the tree is relatively fast, assuming the merged tree has few internal nodes. A final observation is that the disclosed embodiments offer improved performance for common operations such as insertion, removal, and annotation, while maintaining performance no worse than other methods for more complex and less common operations.
[0048] Figure 1A This is an overview diagram of a collaborative messaging architecture. Collaborative system 100 shows two users 101a-b collaborating via individual computing devices 102a-b. Distributed data structure views 103a-b are displayed to the two users 101a-b on displays 104a-b, respectively. Figure 1AThe collaborative messaging architecture shown facilitates the propagation of edits made by any of the users 10la-b to a distributed data structure (at least a portion of which is shown in views 103a-b) to other participants (e.g., devices and / or collaborative module instances) within the collaboration. Although Figure 1A Only two users are shown, but the disclosed embodiments envision collaboration through any number of user / device / collaboration module instances.
[0049] Figure 1A Two collaboration module instances 105a-b are also shown. Each collaboration module instance in collaboration module instances 105a-b represents a set of instructions and data stored on a non-transitory computer-readable storage medium. The instructions in collaboration module instances 105a-b configure each client device in client devices 102a-b to perform one or more of the functions described herein. In some aspects, each collaboration module instance participating in a particular collaboration will run on a different physical device. In some cases, multiple collaboration module instances may run on a single physical device. In some aspects, one or more collaboration module instances in collaboration module instances 105a-b may run on devices other than client devices 102a-b (in... Figure 1A It can be executed on other devices (not shown), but can still receive input from and provide output to client devices 102a-b. In some respects, the collaboration module instance 105a can run, for example, on a computer that also runs the synchronization service 106.
[0050] The term "cooperative module" is not intended to limit the features disclosed herein in any way, but is merely used as a notational convenience for referring to those common features. For example, it is not intended to require that all instructions implementing the claimed features reside on a single storage device or a single computing device, or be physically contiguous, for example.
[0051] In the hypothetical example of the collaborative system 100, user 101a can edit a distributed data structure view 103a. The locally edited data is immediately displayed on user 101a's display screen 104a by a collaborative module instance 105a. Additionally, collaborative module instance 105a sends message 110a to synchronization service 106. Message 110a indicates the nature of the editing operation performed by user 101a. For example, message 110a may indicate whether the operation is an insertion, removal, or comment operation. Message 110a may also indicate the insertion point for an insertion operation; in other words, the location where data is inserted within the distributed data structure view 103a. If the operation is a removal operation, message 110a indicates the range of distributed data structure data removed. Message 110a also indicates a reference sequence number for the operation. The reference sequence number is the version of the distributed data structure that user 101a (and device 102a) operates on when performing a topic operation on the distributed data structure.
[0052] Upon receiving message 110a, synchronization service 106 generates a sequence number for the operation defined by message 110a. Synchronization service 106 then broadcasts a message identifying the operation performed by user 101a (and computing device 102a) to each of the collaboration module instances 105a-b participating in the collaboration. Figure 1A In the example, this includes both collaboration module instances 105a-b. In some embodiments, the broadcast may be an actual broadcast network message using a broadcast destination address. In other embodiments, the broadcast may include two or more unicast or multicast messages that co-address each collaboration module participating in the collaboration system 100. The broadcast messages are shown as messages 120a and 120b. In some cases, the broadcast message (e.g., in...) Figure 1A In the examples referred to collectively as 120a and 120b), it can also indicate that each collaborating module participating in the collaboration has been notified of the operation initiated by device 102a. Alternatively, as in Figure 1A As shown, individual broadcast messages (collectively shown as messages 130a and 130b) can indicate that an operation initiated by user 101a (and computing device 102a) has been successfully propagated to all collaborative modules participating in the collaboration.
[0053] exist Figure 1AIn the illustrated embodiments, the synchronization service 106 enforces a common operation order across all operations occurring on the distributed data structure, regardless of which collaborating module initiates the operation. Therefore, in some cases, the local operation order may differ from the operation order enforced by the synchronization service. To resolve the differences between the local operation order and the common operation order across all collaborating modules participating in the collaboration, the disclosed embodiments provide conflict rules to resolve these differences. For example, in some embodiments maintaining a string-oriented distributed data structure, an insertion operation with a larger sequence number (occurring later than the second insertion in the common order) is placed in the string earlier than the second insertion. In other cases, two deletions initiated by two different collaborating modules may overlap. Some embodiments resolve overlapping deletions by determining that the deletion operation with the earlier (e.g., smaller) sequence number is used to perform the deletion, while the overlapping portion of the later deletion has essentially no effect.
[0054] Although the collaborative system 100 has been discussed above with reference to synchronization service 106, in some other embodiments, a peer-to-peer protocol may be used to implement the order of operations. For example, some embodiments may assign sequence numbers to data based on an open-source library such as OrbitDB.
[0055] The collaborative system described above enables each edit transformation of the distributed data structure views 103a-b to be performed through up to three different states maintained by at least one of the collaborative module instances 105a-b.
[0056] Figure 1B This is a state transition diagram illustrating how an operation moves between three states. Figure 1B The first state 155, designated "Local Only," represents the operation state after a collaborative module (e.g., 102a) has performed a local edit, but the edit has not yet been confirmed by the synchronization service 106 (e.g., via message 120a). The operation exists only in the first state 155 on the collaborative module that initiated the operation. The operation will be initialized on a second collaborative module that is in the second state, not the first state.
[0057] Once the operation / edit has been confirmed by the synchronization service 106 (e.g., via messages 120a and / or 120b), the operation exists in a second state 160. This second state of the edit can exist at a collaboration module other than the collaboration module that performed or caused the edit. The third state 165 of the edit (in [the context is missing]) occurs when all collaboration modules have confirmed the edit and the instruction to do so has been propagated to the collaboration modules via the synchronization service 106 (e.g., via messages 130a or 130b). Figure 1BThe third state (named "synchronized") exists on the collaboration module. This third state can be considered "synchronized" because the operation has been synchronized across the collaboration module, and it may no longer be necessary to track the edited information, since the edited data is considered "recorded" within the distributed data structure.
[0058] The disclosed embodiments process edits or operations on a distributed data structure in a well-defined order through a collaboration module. Synchronization service 106 defines the order of operations by assigning a unique sequence number to each operation. Although each edit or operation is assigned a sequence number identifying itself, the edit operates on a specific version of the distributed data structure. This specific version can be identified by a second sequence number different from the sequence number assigned to the edit or operation. Throughout this disclosure, this second sequence number may be referred to as a reference sequence number for the edit / operation.
[0059] The following is an example of the distinction between an operation's sequence number and a reference sequence number. Specifically, after a first operation, identified by a first sequence number, is applied to a distributed data structure, the resulting version of the distributed data structure can be identified by the first sequence number. The first sequence number defines a first "version" of the distributed data structure. A second operation, defined by a second sequence number, modifies the first "version" of the distributed data structure, including modifications caused by the first operation (e.g., its result). The result of this second modification is a second version of the distributed data structure. Further subsequent operations, defined by additional sequence numbers, further modify the distributed data structure to form additional new "versions" defined by those sequence numbers.
[0060] Based on the sequence number of the operation and the version of the distributed data structure it operates on, each collaborating module can appropriately apply operations initiated at other collaborating modules to its own copy of the distributed data structure. This is possible even when multiple collaborating modules may modify the distributed data structure "simultaneously," and even when some collaborating modules may lag in synchronizing with the evolving version of the distributed data structure.
[0061] This is accomplished through two rules. First, if an operation is initiated by a local collaboration module, subsequent operations initiated by that local collaboration module operate on versions of the distributed data structure, including those modified by the first operation. In other words, collaboration module operations are executed sequentially without exception. Furthermore, local operations apply to all remote operations, regardless of their sequence numbers.
[0062] Regarding the application of a specific remotely initiated modification at a local device, the local device may only consider some of the operations for which it has already received notification. When applying a specific operation to a distributed data structure, operations with lower reference sequence numbers (initiated by both local and remote devices) are relevant compared to the operation sequence numbers of remotely initiated operations. Operations with higher operation sequence numbers are not visible at the initiating device when the specific operation is initiated, and are therefore irrelevant in determining how to apply the specific operation to the distributed data structure.
[0063] Figure 2A Exemplary message formats are shown for messages exchanged between a collaboration module instance (e.g., 105a-b) and a synchronization service 106 in one or more of the disclosed embodiments. In some aspects, any of the messages 110a, 120a-b, and 130a-b may include the following regarding... Figure 2A One or more of the fields being discussed.
[0064] Message 200 includes a collaboration identifier field 202, a reference sequence number or version field 204, an operation sequence number field 206, an operation type field 208, an operation scope field 210, an operation data field 212, and a maximum sequence number field 214. The collaboration identifier 202 uniquely identifies the collaboration module participating in the collaboration. The collaboration identifier 202 identifies the collaboration module that initiated the operation identified by message 200. In some aspects, when the collaboration module joins the collaboration, the synchronization service 106 assigns an identifier number to each collaboration module. The reference sequence number or version field 204 identifies the version of the distributed data structure to which the operation identified by message 200 (via field 206) is applied. Therefore, each version of the distributed data structure maintained by the collaboration module is identified via a different reference sequence number. The reference sequence number or version field 204 identifies the data segment synchronized by the collaboration module (via field 202) when a modification is made.
[0065] The operation sequence number field 206 identifies the sequence number of the operation identified by message 200. When the coordinating module initiating the operation sends a message including the operation sequence number field 206, the operation sequence number field 206 can be set to a predetermined value (such as -1) indicating that no sequence number has been assigned. This predetermined value indicates to the synchronization service that the operation identified by the message is new, and therefore the synchronization service assigns a sequence number to the operation. Before the sequence number is assigned by the synchronization service 106, such an operation can be considered to be in the state described above regarding... Figure 1B In the described "local only" state.
[0066] Synchronization service 106 may assign incrementing sequence numbers to new operations that are received in the order that the synchronization service identifies those new operations. After synchronization service 106 assigns a sequence number to a specific operation, messages related to that operation may include the assigned sequence number in the operation sequence number field 206.
[0067] The operation type field 208 indicates the type of operation indicated by message 200. In some embodiments, the operation type may indicate an operation type of insertion, removal, or annotation, but the operations contemplated in this disclosure are not limited to these operation types.
[0068] The operation range field 210 indicates the data range of the distributed data structure operated on by the operation. In some aspects, the range is a single value, such as the position of an insertion operation in a stream. In other aspects, for example, the data range can be indicated when a data range in the distributed data structure is deleted.
[0069] Operation data field 212 indicates the data to be applied as part of the operation. For example, if the operation type is insert, operation data field 212 indicates the data to be inserted.
[0070] Depending on the role of the sender of the field, the maximum sequence number field 214 can have two different meanings. When field 214 is transmitted by a collaborating module, it indicates the maximum sequence number received by the collaborating module from the synchronization service 106 or from a peer device using a peer-to-peer protocol for serialization. When field 214 is transmitted by the synchronization service 106 or from another collaborating module using a peer-to-peer protocol for serialization to a collaborating module, it indicates the maximum sequence number that has been acknowledged by all collaborating modules participating in the collaboration.
[0071] Figure 2B It is shown in Figure 2A The diagram 215 provides an overview of the relationship between message 200 and the sequentially ordered set of operations 220. Operations 220 are ordered according to the operation sequence order 219. Figure 2B Snapshot 218 of the distributed data structure is shown. Snapshot 218 represents the data values of the distributed data structure at a specific version. Figure 2B A set of operations 220 for sequential sorting performed on a version of the distributed data structure derived from snapshot 218 is also shown.
[0072] Message200 is also Figure 2B As shown above, including the information mentioned above. Figure 2A Each of the fields discussed. Figure 2BThe operation sequence number field 206 of message 200 is shown to identify the highest-ranking operation 222 in the sequentially ordered set of operations 220. Note that when message 200 is received by a collaborating module (e.g., 105a-b), field 206 identifies the highest-ranking operation 222 because message 200 indicates that a sequence number has been assigned to the operation. As discussed above, when a collaborating module initiates an operation, it sets the operation sequence number field 206 to a predetermined number, indicating that the operation needs to have a sequence number assigned to it. In this scenario, compared with... Figure 2B In contrast to the example shown, field 206 may not necessarily identify the highest sorting operation.
[0073] Message 200 also includes a reference sequence number or version field 204, which identifies operation 224 on the distributed data structure. Note that version field 204 identifies a version of the distributed data structure, including the result of operation 224 and all operations ordered after operation 224. In other words, if the value of version field 204 is 950, then this version of the distributed data structure includes the results of operations with sequence numbers 950, 949, 948, 947, etc. This example assumes that the operations are ordered such that operations with higher sequence numbers occur after operations with lower sequence numbers. Some embodiments may use alternative methods to order the operations (e.g., a lower sequence number indicates a later order than a higher sequence number).
[0074] When transmitted by synchronization service 106 (or received by a collaboration module from a peer-to-peer network), message 200 also identifies the maximum sequence number of the operation confirmed by all collaboration modules participating in the collaboration. This maximum sequence number is indicated by field 214. Some implementations may perform a garbage collection process 232 on the operation described below and including operation 226. Operations ordered after operation 226 (e.g., pending consecutive operations 230) have at least one pending confirmation. Figure 2B The pending continuous operation 230 shown may be referred to as a collaboration window throughout this disclosure. The collaboration window defines operations that have been assigned sequence numbers by the synchronization service 106 (or the peer-to-peer protocol for serialization) but have not yet been confirmed by all collaboration modules participating in the collaboration. When the disclosed embodiments operate, the collaboration window represents a scrolling window because it advances through sequential operations ordered by the synchronization service 106 (or the peer-to-peer protocol), wherein the top of the collaboration window is defined by the most recently assigned or highest-ranked sequence number, and the bottom of the collaboration window is defined (exclusively) by the largest sequence number confirmed by all collaboration modules participating in the collaboration.
[0075] Figure 2CExemplary message communication between two collaboration module instances 105a-b and synchronization service 106 is illustrated. Although synchronization service 106 is an example of a synchronization service, other embodiments may use peer-to-peer protocols to facilitate the serialization of operations between collaboration module instances.
[0076] Figure 2C The image shows message 234 transmitted from a synchronization service (e.g., 106) to a collaboration module instance 105a. Message 234 may include the information above regarding... Figure 2A One or more of the fields discussed. Specifically, message 234 is shown as transmitting an operation sequence number with a value of ten (10) (e.g., via version field 204), a version number of nine (9) (e.g., via field 206), and a maximum sequence number of eight (8) (e.g., via field 214). Since message 234 is transmitted by synchronization service 106 to collaboration module instance 105a, the maximum sequence number value eight (8) of message 234 indicates the maximum sequence number of the operation confirmed by all collaboration modules participating in the collaboration. Therefore, when message 234 is transmitted by synchronization service 106, the synchronization service has received confirmation from all collaboration modules participating in the collaboration up to and including the operation assigned sequence number eight (8).
[0077] Next, collaboration module instance 105a performs a new operation on the distributed data structure and transmits message 235. Message 235 may include the information above regarding message 200 and... Figure 2A One or more of the fields discussed. Specifically, message 235 is shown as instructing collaboration module instance 105a to initiate a first operation with an initial sequence number -1. "-1" is an example of a first predetermined (sequence) number defined by some embodiments in the disclosed embodiments to indicate that a sequence number has not yet been assigned to the first operation defined in message 235 (e.g., via one or more of fields 208, 210, and 212). Message 235 also indicates that the first operation is performed by collaboration module instance 105a on version 10 of the distributed data structure. This indicates that all operations up to and including those assigned sequence number ten (10) are applied to the distributed data structure before the operation defined by message 235 is performed. In other words, when the first operation defined by message 235 is performed, any results derived from operations up to and including operation ten (10) are taken into account. Therefore, if the first operation depends on a portion of the distributed data structure modified by any of those operations, the result of the first operation is based on those modifications.
[0078] Message 235 also indicates that the maximum operation sequence number received by the collaboration module instance 105a is ten (10) (as provided by message 234). In some aspects, the maximum operation sequence number illustrated in message 235 may be included in field 214. Message 235 serves as an acknowledgment by the collaboration module instance 105a of the synchronization service 106 for all operations up to sequence number ten (10).
[0079] Next, Figure 2C Message 236, transmitted from collaboration module instance 105b to synchronization service 106, is shown. Message 236 may include one or more of the fields discussed above with respect to message 200. Message 236 identifies collaboration module instance 105b (e.g., via field 202) and indicates that collaboration module instance 105b has initiated a new operation, which may be defined in the message (e.g., via fields 208, 210, 212, not shown). Message 236 also indicates a second operation performed on version eight (8) of the distributed data structure. In other words, when the second operation is performed on the distributed data structure, collaboration module instance 105b includes any operation result having a sequence number up to and including sequence number eight (8). Message 235 also indicates that collaboration module instance 105b has received the maximum operation sequence number eight (8). Therefore, collaboration module instance 105b is somewhat behind in notification of operations compared to collaboration module instance 105a. Compared to collaboration module instance 105b, collaboration module instance 105a has been notified of two additional operations (operations nine (9) and ten (10)).
[0080] Note that after message 236 is received by synchronization service 106, both operations need to have assigned sequence numbers. The first operation is initiated by collaboration module instance 105a, and the second operation is initiated by collaboration module instance 105b. Also note that, as discussed above, collaboration module instance 105b is somewhat behind because it is still unaware that the operations are sequenced as nine (9) and ten (10). In response, the synchronization service transmits messages 237 and 238 to collaboration module instance 105b.
[0081] One or more of messages 237 and 238 may include the information above regarding messages 200 and... Figure 2A One or more of the fields described. In some respects, messages 237 and 238 may be broadcast or multicast to more collaborative modules than just collaborative module instance 105b. Messages 237 and 238 notify collaborative module instance 105b at least, respectively, of the operations assigned sequence numbers nine (9) and ten (10). Messages 237 and 238 may provide additional information defining the operations identified by sequence numbers nine (9) and ten (10) (e.g., via fields 208, 210, and 212).
[0082] The transmission of messages 237 and 238 by synchronization service 106 demonstrates at least one design parameter of several embodiments in the disclosed embodiments, namely, the design parameter of enforcing a single order of operations across all collaborating modules, and the design parameter of ensuring that each collaborating module receives notification of ordered operations. Therefore, since collaborating module instance 105b indicates via message 236 that its maximum received sequence number is eight (8), synchronization service 106 responds by transmitting operations nine (9) and ten (10) to collaborating module instance 105b (via messages 236 and 237, respectively), enabling the distribution module to also transmit the subsequent operation assigned sequence number eleven (11) via message 239. As shown, message 239 indicates that the second operation initially indicated by message 236 has been assigned sequence number eleven (11) by synchronization service 106. A similar message 240 notifies collaborating module instance 105a of the second operation, its assignment of sequence number 11, and the version of the distributed data structure on which the second operation is performed (as indicated in message 236).
[0083] Figure 2C Message 241, transmitted from synchronization service 106 to collaboration module instance 105b, is also shown. Message 241 indicates that a first operation initiated by collaboration module instance 105a and indicated by message 234 has been assigned a sequence number of twelve (12) by the synchronization service. Message 241 may (e.g., via fields 208, 210, and 212) provide additional information defining the first operation. Note that the version indication in message 241 is equivalent to the indication provided in message 235, as both messages define the same operation. The maximum sequence number indication in message 241 is eight (8), indicating a lower limit of all collaboration modules participating in the collaboration (set by collaboration module instance 105b in this example). The synchronization service sends message 242 to collaboration module instance 105a. In some aspects, messages 241 and 242 may be the same message broadcast or multicast to collaboration module instances 105a and 105b.
[0084] Figure 2C An exemplary heartbeat message 243 is also shown. Heartbeat message 243 may include the information above regarding... Figure 2AOne or more fields discussed in message 200. Heartbeat message 243 may be transmitted by collaboration module instance 105b after a predetermined or configured period of inactivity. Inactivity may be defined by messages transmitted by collaboration module instance 105b to synchronization service 106. Messages received by collaboration module instance 105b may be disregarded in the inactivity determination. Since message 243 is a heartbeat message, the sequence number field (e.g., field 206) is set to a second predetermined value to distinguish it from a first predetermined value (e.g., in messages 235 and 236) for operations with unassigned sequence numbers. Heartbeat message 243 indicates (e.g., via field 214) that the maximum sequence number received by collaboration module instance 105b is twelve (12).
[0085] Message 244 indicates that collaboration module instance 105a has initiated a third operation on version 12 (12) of the distributed data structure (when the third operation is applied to the distributed data structure via collaboration module instance 105a, all operations ordered by number 12 and lower are considered). In response to receiving message 244, synchronization service 106 distributes a notification of the third operation to all collaboration modules participating in the collaboration. Figure 2C The synchronization service of sending messages 245 and 246 to collaboration module instances 105b and 105a, respectively, is illustrated. In some aspects, messages 245 and 246 may be the same physical messages simultaneously broadcast or multicast to at least two collaboration module instances 105a-b. Note that messages 245 and 246 indicate the same version information (12) and collaboration module identifier (105a) as initially indicated in message 244. Also note that heartbeat message 241 updates the maximum sequence number received by collaboration module instance 105b. Since collaboration module instance 105b previously represented the lower limit of the collaboration window, messages 245 and 246 indicate an update to the bottom of the collaboration window by indicating a maximum sequence number value of twelve (12) (which is consistent with heartbeat message 243).
[0086] Figure 2D An exemplary message exchange 250 is shown between collaboration module instances 105a-b and a synchronization service (in this case, synchronization service 106). The following section discusses... Figure 2D One or more of the messages under discussion may include the above regarding Figure 2A One or more of the fields discussed in message 200.
[0087] exist Figure 2DThe message exchange illustrated herein is intended to demonstrate how a synchronization service, such as synchronization service 106, adjusts the reference sequence number for an operation as it is distributed to participating collaboration module instances (or devices). These adjustments allow collaboration module instances to perform operations on distributed data structures without blocking or otherwise waiting for the synchronization service before proceeding. Specifically, when a collaboration module instance performs multiple operations, these adjustments may be appropriate before the synchronization service assigns sequence numbers to any of those operations.
[0088] Figure 2D A series of three messages 251a-c are shown, indicating a first, second, and third operation initiated respectively by the collaboration module instance 105a. Messages 251a-c all indicate that no sequence number was assigned to any of the first, second, or third operations (e.g., via an exemplary predetermined value of -1 for sequence numbers). Each message in 251a-c also indicates a reference sequence number + (10) for its respective operation. Note that the reference sequence number does not change as each of the three messages 251a-c is sent. This is a result of the collaboration module instance 105a not receiving any messages from the synchronization service (e.g., 106) between the initiation of the three operations.
[0089] Next, Figure 2D The synchronous service that sends a pair of messages 252a-b is illustrated. In some aspects, in Figure 2D The two messages 252a-b shown can be a single physical message broadcast or multicast to both cooperative module instances 105a-b. Message 252a-b indicates that sequence number 11 has been assigned to the first operation defined by message 251a. The version of the distributed data structure operated by the first operation is identified as version ten (10) by message(s) 252a-b.
[0090] Next, Figure 2DThe synchronization service 106 is shown transmitting messages 253a-b to collaboration module instances 105b and 105a, respectively. Similar to the case for message 252a-b, message 253a-b can be a single physical message broadcast or multicast to collaboration module instances 253b and 253a, respectively. Message 253a-b indicates that a sequence number has been assigned to a second operation initiated by collaboration module instance 105a (and is indicated by message 251b). Note that although collaboration module instance 105a indicates a reference sequence number of 10 for the second operation (see message 251b), the synchronization service indicates that the reference sequence number for the second operation (having sequence number twelve (12)) is eleven (11) when the second operation is assigned a sequence number. Note that this reference sequence number is equivalent to the sequence number assigned to the operation indicated in one or more messages 252a-b. Therefore, the synchronization service updates the reference sequence number for the second operation based on its knowledge of the sequence of multiple operations performed by collaboration module instance 105a.
[0091] Specifically, the synchronization service 106 is notified to the collaboration module instance 105a to perform a first operation, followed by a second operation, and then a third operation. This notification is provided by message sequences 251a-c. Thus, the synchronization service is provided with an indication that the second operation is performed on a version of a distributed data structure that includes the result of the first operation, which is assigned sequence number 11 via one or more messages 252a-b. As a result, the synchronization service instructs in one or more messages 253a-b that the second operation (sequence number 12 (12) is performed on version 11 (11) of the distributed data structure that includes the result of the first operation) is performed.
[0092] Similarly, synchronization service 106 updates the reference sequence number for the third operation in a similar manner. (As in...) Figure 2D As shown, synchronization service 106 transmits messages 254a-b (which may be the same as 252a-b and 253a-b) to notify the participating collaboration module instances that a third operation is initiated by collaboration module instance 105a. The third operation is initially indicated in message 251c. The third operation is assigned sequence number thirteen (13). The second operation is assigned sequence number twelve (12). Since the third operation is performed on a version of a distributed data structure that includes the result of the second operation, synchronization service 106 indicates the reference sequence number twelve (12) for the third operation in one or more messages 254a-b.
[0093] Figure 2E This is an overview diagram of the snapshot process. Snapshot process 260 shows... Figure 1AThe synchronization service 106 provides operation logs 265 to the persistence service 268. The persistence service 268 feeds the operation logs 265 to the operation log 270. In some aspects, the persistence service 268 is integrated with the synchronization service.
[0094] Operation log 270 includes individual records 272, each defining an operation. The individual records 272 included in operation log 270 can be included in the order defined by the operation sequence number assigned to each operation by synchronization service 106.
[0095] In some respects, each individual record of the operation log 270 may consist of one or more values from the fields contained in the message 200, as described above. The operation log 270 is read by the snapshot service 280. The snapshot service 280 generates snapshots, such as in... Figure 2E Snapshots 281a and 281b are shown. A snapshot represents the state of a distributed data structure at a specific point in time. For example, snapshot 281a could represent the state of the distributed data structure after all operations, up to and including those represented by individual record 272, have been applied to the data structure. Snapshot service 280 can receive a first snapshot (such as snapshot 281a) and an additional set of operation records as input, containing operations with sequence numbers greater than the largest sequence number included in the first snapshot. These operation records are in... Figure 2E The snapshot is represented as 285. The snapshot service 280 then applies additional operations 285 to the first snapshot 281a to generate a second snapshot 281b, which represents the second state of the distributed data structure up to and including the operation record 288.
[0096] Figure 2E Supply service 290 is also shown. Supply service 290 is responsible for bringing new collaboration modules online into existing collaborations. To do this, supply service 290 provides the new collaboration module ( Figure 2E 102c) provides the latest snapshot, Figure 2E 281b in the previous snapshot. Service 290 also provides operation records to the new collaboration module instance 105c for those operations with serial numbers greater than those included in the previous snapshot 281b. Figure 2E These operation records are shown as 292a, which are read from operation log 270 as 292a and provided to the new collaboration module instance 105c as record 292b. Collaboration module instance 105c then applies operation record 292b to snapshot 281b to obtain the "current" or latest version of the distributed data structure managed by the collaboration.
[0097] Figure 3View 300 illustrates two data structures that can be used to construct a tree in some embodiments of the disclosed embodiments. In some embodiments of these embodiments, the tree data structure is generated and maintained to include one or more blocks 301 and one or more elements 320. All levels of the tree, except for the leaf levels, include one or more fields described below with respect to the block data structure 301, while the leaf nodes of the tree may include one or more fields described below with respect to the leaf / element data structure. The exemplary merged tree data structures (leaf, node, block, tree, element, etc.) described below represent an exemplary format of data values provided in physical hardware memory. For example, each field described below may store one or more values in hardware memory. These values may be stored in memory by hardware processing circuitry (such as one or more hardware processors). These values may also be read from the hardware memory by one or more hardware processors as needed to perform one or more of the functions discussed herein. In some embodiments, the hardware processing circuitry may write to or read the values via a memory address that identifies each data value in the data values. In some aspects, the memory address may be a word-based address, and the data values may not necessarily be stored in a word-aligned physical location within the memory. In these cases, as is known in the art, the hardware processing circuitry can be configured to read a data word including a specific value and then perform additional processing on the word value within the hardware processing circuitry to isolate non-word-aligned values. The disclosed embodiments can utilize any existing method of accessing one or more data structures stored in hardware memory via hardware processing circuitry.
[0098] The following text is about Figure 3 The exemplary block 301 discussed has a variable length. In other words, block 301 can be stored in a variable number of portions of hardware memory (e.g., a variable number of bytes, words, etc.). The length varies based on both the number of collaborating modules participating in the collaboration and the size of the current collaboration window. The collaboration window can be considered as multiple edits or operations on a distributed data structure that have been initiated but not yet finalized (“synchronized”) across all participating collaborating modules. To track edits to the distributed data structure that have not yet been fully synchronized / confirmed, block 301 includes a minimum length field 302. The minimum length field of block 301 represents the length of the distributed data structure represented by the elements of the tree below the block that is common to all collaborating modules or synchronized. In other words, at least in some aspects, the minimum length field 302 represents the data length represented by the elements below the block and associated with a sequence number less than or equal to the maximum sequence number received in field 214, as discussed above.
[0099] Block 301 also includes a variable number of collaborative module identifier fields, such as field 304. Field 314 is shown as another example of a collaborative module identifier field, but the number of collaborative module identifier fields in merge tree block 301 can vary from zero to any possible upper limit, limited only by the number of collaborative modules participating in the collaboration.
[0100] For each collaborative module instance identifier included in merge tree block 301, merge tree block 301 also includes a variable number of data pairs. Each data pair is associated with a reference sequence number (e.g., 306). 1..n ) and length values (e.g., stored in field 308) 1..n (in Chinese). Stored in field 316. 1..n The serial number value is stored in field 318. 1..n The length values in the code are shown to be associated with different collaboration module instance identifiers 314, to show that the number of associations for each collaboration module instance identifier can vary.
[0101] The fields of block 301 introduced above track which part of the distributed data structure represented by the merge tree is visible to each instance of the collaborating module participating in the collaboration. The information provided by these fields is specific to the reference sequence number (e.g., 306). 1..n and 316 1..n Each version of the distributed data structure identified by the identifier is maintained. This information is used when determining how to apply operations originating from local or remote collaboration module instances to a specific merge tree. This information may be necessary because the version of the distributed data structure at the remote collaboration module when it initiates an operation may not be equivalent to a second version of the distributed data structure at the second collaboration module. In some embodiments of the disclosed embodiments, the second collaboration module needs to apply the operation to this second version of the distributed data structure in a manner that replicates the results obtained by the remote collaboration module. Partial length information supports these operations.
[0102] The partial length information for a specific block and for a specific collaborative module can be defined by Equation 1 below, wherein the summation is performed on all leaf node elements represented by the block and satisfying the defined conditions:
[0103] pLen(op client,op ref seq,block)=minLen(block)+
[0104] ∑len(leaf.op seq≤op ref seq)+len(leaf.client=
[0105] op client and leaf.op seq>op ref seq) (1)
[0106] in:
[0107] minLen() returns the minimum length of the data represented by the block (the length of the synchronization data represented by the block).
[0108] len() returns the length of data in a distributed data structure represented by elements that match one or more conditions identified.
[0109] The op client is the identifier of the collaborating module that initiated the operation.
[0110] op ref seq is the reference sequence number for the operation.
[0111] The leaf node.client identifier represents the collaborative module that initiates the operation represented by the leaf node.
[0112] The leaf node.op seq is the operation sequence number of the second operation represented by the leaf node.
[0113] Exemplary element 320 includes a data field 322, an operation sequence number field 324, a deletion sequence number field 325, a reference sequence number field 326, and a collaboration module identifier field 328. The data field 322 includes data representing a portion of the collaboration data (e.g., a distributed data structure) represented by a specific element. The operation sequence number field 324 identifies the sequence number assigned to the operation represented by the element. The deletion sequence number field 325 identifies the sequence number of the operation used to delete the data represented by element 320. In other words, the operation sequence number field 324 can indicate the sequence number of an operation that inserts or annotates data, and if that data is subsequently deleted, the deletion sequence number field 325 will indicate the sequence number of that (subsequent) operation.
[0114] The reference sequence number field 326 indicates the maximum sequence number of the synchronization operation received by the collaboration module when it performs the operation defined by element 320. Synchronization operations can be those that have been acknowledged by all participating collaboration modules. The collaboration module identifier field 328 identifies the collaboration module that initiated the operation.
[0115] the following Figure 4-13 The data structure representing collaboration between two collaborative modules is referred to as collaborative module instance 105a and collaborative module instance 105b for simplicity. Each of these two collaborative module instances 105a-b is editing a string structure, and the merge tree data structure supports the synchronization of these edits across the two collaborative module instances 105a-b.
[0116] Figure 4An example of a merge tree section 400 maintained by collaboration module instance 105a is shown. Merge tree section 400 represents the string “Cat on the mat”. This string is the result of a concatenation of two substrings. The first substring “on the mat” is represented by element 405a of merge tree section 400, and the second string “Cat” is inserted by collaboration module instance 105a at position zero (0) of the string “on the mat”. This insertion operation is represented by element 405b. Each element in elements 405a-b may utilize one or more fields from the fields discussed above regarding merge tree block 301.
[0117] The insertion operation represented by element 405b has not yet been confirmed by the server and is therefore assigned a sequence number of -1, as shown. The reference sequence number for the insertion of “Cat” represented by element 405a is zero (0) because the insertion occurs on a version of the distributed data structure (the string “on the mat”), where the highest-order operation is assigned a sequence number of zero (0).
[0118] Figure 4 Block 420 of tree portion 400 is also shown. Block 420 of the tree stores a minimum length value 422 (e.g., in some aspects, stored in field 302). The minimum length value 422 corresponds to the length of the data “on the mat”, represented by element 405a. The data represented by element 405a is acknowledgment data because all instances of the collaborating modules participating in the collaboration have an acknowledgment operation that results in the string “on the mat”. Because of the result of an operation that assigns a sequence number less than or equal to the current reference sequence number (e.g., in some aspects, as provided in field 214 from synchronization service 106), it is visible to all instances of the collaborating modules participating in the collaboration. Thus, the length of this data can be contained within the minimum length value 422.
[0119] Block 420 also includes partial length information for collaboration module instance 105a, as shown in 424. When collaboration module instance 105a accesses those segments with sequence numbers of zero or more, the partial length information 424 indicates that the segment below block 420 includes four (4) extra characters of data (exceeding the minimum length). This corresponds to the insertion of “cat” by collaboration module instance 105a. Block 420 does not include any partial length information for collaboration module instance 105b, indicating that no additional data is available for collaboration module instance 105b under any circumstances (because the insertion of “cat” has not yet been assigned a sequence number by synchronization service 106, and is therefore not visible to collaboration module instance 105b).
[0120] Figure 5An exemplary merge tree section 500 is shown on a collaboration module instance 105b. Figure 5 The exemplary merge tree section 500 represents the string "Big on the mat". The merge tree section 500 includes the string section "on the mat" represented by element 505a, which is synchronized with the "on the mat" string section represented by element 405a discussed above with respect to collaboration module instance 105a and merge tree section 400.
[0121] Tree portion 500 also includes a second element 505b, which represents the insertion operation of the string "Big" at position zero (0) of the string "on the mat". This insertion operation is initiated by the collaboration module instance 105b. The collaboration module instance 105b assigns a sequence number -1 to the "Big" insertion operation until the insertion is confirmed by the synchronization service. The reference sequence number / distributed data structure version of the insertion, represented by element 505b, is zero. This indicates that when performing the "Big" insertion operation on the distributed data structure, the result of an operation with a lower order or an equivalent sequence number of zero is considered.
[0122] Figure 5 Block 520 of tree section 500 is also shown. Block 520 indicates a minimum length value 522 for the leaf elements below block 520. The minimum length value 522 represents the length of the synchronization data represented by element 505a. In other words, the minimum length value 522 indicates the length of the data represented by block 520 that has been passed out of the bottom of the collaboration window (e.g., pending sequential operation 230).
[0123] Block 520 also includes partial length information 524 for collaboration module instance 105b. Partial length information 524 indicates that for reference sequence numbers of zero or higher, the leaf elements below block 520 include four additional characters of data in addition to the minimum length value 522. Block 520 does not include any partial length information for collaboration module instance 105a. This indicates that, apart from the data represented by the minimum length value 522, no additional data is visible to collaboration module instance 105a in the tree section 500.
[0124] Figure 6 An updated version of the merge tree section is shown on collaboration module instance 105b. Figure 6Merge tree section 600 reflects the additional insertion operation that occurred on collaboration module instance 105b, and is therefore a modification of merge tree section 500. The second insertion operation inserts the word “furry” at position four (4) in the string and is represented by element 505c. The word “furry” appears after the word “big” in the string and is represented by element 505b. The insertion operation represented by element 505c is assigned a sequence number -1 because the operation has not yet been assigned a sequence number by the server; and is assigned a reference sequence number zero because the insertion of the word “furry” occurred before any operation was received from the server.
[0125] The block of tree section 600 is shown as 620. Block 620 indicates that the minimum length value 622 is eleven, again representing the data represented by element 505a. The minimum length value 622 does not include the length of the data represented by elements 505b and 505c, because this data was not acknowledged by all the collaboration module instances participating in the collaboration. The block also includes partial length information 612. Since collaboration module instance 105b initiated the two insertion operations represented by elements 505b and 505c, and the reference sequence numbers of both operations are zero, the partial length information 612 indicates that collaboration module instance 105b accesses a larger reference of ten (10) characters of tree section 600 for data with a reference sequence number of zero or greater than the minimum length value 622. These ten characters represent the six characters represented by element 505c and the four (4) characters from element 505b (each string includes a space at the end).
[0126] Figure 7 This shows an updated version of the merged tree portion 600 stored at the collaboration module instance 105b. The updated portion is labeled 700. Tree portion 700 shows the data from... Figure 4 The insertion operation "cat" initiated by collaboration module instance 105a has been propagated (e.g., via synchronization service 106) to collaboration module instance 105b. This insertion operation is represented as element 505d in tree section 700.
[0127] The insertion of "cat" by collaboration module instance 105a conflicts with the insertions of "big" and "furry" by collaboration module instance 105b, because all these insertions would position the string at zero. The conflict rule determines the position of element 505d in tree section 700 based on the reference sequence number of the insertion operation. The reference sequence number for the insertion of "Cat" for collaboration module instance 105a is zero. Under the conflict rule discussed above, the "big" and "furry" insertion operations are shifted to the left because they will be assigned a sequence number greater than the sequence number of the "cat" insertion (i(1)).
[0128] Figure 7 Block 720 of tree section 700 is also shown. The block indicates a minimum length value 722. Block 720 also indicates partial length information 712a-b for collaboration module instances 105a and 102b, respectively. Partial length information 712a indicates a partial length of four (4) characters for a reference sequence number of one (1) or greater. This corresponds to the length of the data represented by element 505d. The data represented by elements 505b and 505c is not yet visible to collaboration module instance 105a, and therefore no partial length information for this data is provided for collaboration module instance 105a. Partial length information 712b for collaboration module instance 105b indicates that for a reference sequence number greater than or equal to zero (0), ten characters of data are included in the elements below block 720. The ten characters include the data represented by elements 505b and 505c. Partial length information 712b also includes an indication that for a reference sequence number of one (1) or greater, the elements below block 720 include 14 characters of data. Compared to reference sequence number zero (0) in partial length information 712b, the additional four (4) characters of the data for reference sequence number one (1) are data represented by element 505d. The data represented by element 505d is assigned sequence number one (1) and initiated by cooperative module instance 105a.
[0129] After receiving notification of the insertion of “Cat” from synchronization service 106, collaboration module instance 105b can receive confirmation of the insertion of “Big” from synchronization service 106. The synchronization service will assign sequence number two to the insertion operation of “big” represented by element 505b. This will be shown in subsequent examples below.
[0130] Figure 8 Showing from Figure 4 The updated version of the merged tree section 400 is tree section 800. Tree section 800 shows that the collaboration module instance 105a has been notified by the collaboration module instance 105b to insert “Big” into the string. This insertion operation is represented by element 405c. The insertion of “Big” uses the reference sequence number zero (0), which conflicts with the reference sequence number of the insertion of “Cat” represented by element 405b. Under the conflict rule, “Big” is placed before “Cat” because its sequence number (two (2)) is later than that of “Cat” (which is one (1)).
[0131] Tree section 800 includes block 820. Block 820 indicates a minimum length value of 822. Block 820 includes partial length information 812a-b for each of the collaboration modules 102a-b. Partial length information 812a indicates that, for collaboration module instance 105a, the element below block 820 represents four characters (represented by element 405b) of data for a reference sequence number of zero. Partial length information 812a also indicates that, for collaboration module instance 105a, when the reference sequence number is two or greater, the element below block 820 represents eight (8) characters of additional data (in addition to the minimum length value 822). These eight characters of data are represented by elements 405b and 405c. Regarding collaboration module instance 105b, when the reference sequence number is zero, partial length information 812b indicates that the element below block 820 represents four (4) additional characters of data (exceeding the minimum length value 822). These four (4) additional characters are represented by element 405c, which is initiated by the collaboration module instance 105b and has a reference sequence number of zero. When the reference sequence number of the collaboration module instance 105b is one (1), the partial length information 812b indicates that the element below block 820 represents a total of eight (8) characters of additional data, because the "Cat" insertion becomes visible to the collaboration module instance 105b when the reference sequence number is one.
[0132] Figure 9 An updated version of collaboration module instance 105a is shown, with the tree portion merged into tree portion 900. Tree portion 900 shows the additional insertion operation that has occurred on collaboration module instance 105a. The insertion operation inserts the string “top of” into position 11 of the previous string “big cat on the mat”. Since the insertion occurs in the middle of the string “on the mat” represented by element 405a in the merged tree of the previous collaboration module instance 105a, element 405a is split into two elements, labeled as elements 405d and 405e. The new character “top of” is represented by element 405f. Since both the character “on” represented by element 405d and the character “the mat” represented by element 405e originate from the original string “on the mat” with sequence number zero (0), sequence number zero is also assigned to each of elements 405d and 405e that represent “on” and “the mat” respectively. When collaboration module instance 105a inserts the string "top of", it also sends a message to the server indicating the insertion operation, the character to be inserted, and the insertion position. The sequence number for this insertion operation is assigned as -1 (as in...). Figure 9 (As shown in the figure), until the synchronization service 106 provides a confirmation sequence number for the insertion operation.
[0133] Block 920, contained within tree section 900, indicates minimum length information 922. For collaboration module instances 105a-b, partial length information for the merged tree section 900 is shown as 912a-b in block 920, respectively. Partial length information 912a indicates that, for collaboration module instance 105a, reference sequence number zero includes fourteen (14) additional characters of data in the elements below block 920. These fourteen (14) additional characters include data represented by elements 405b, 405d, and 405f. When the reference number on collaboration module instance 105a reaches two (2) or more, the insertion of “Big” by collaboration module instance 105b becomes visible to collaboration module instance 105a, and thus the partial length information increases by four (4) relative to sequence number zero, as shown by partial length information 912a.
[0134] Regarding collaboration module instance 105b, partial length information 912b indicates that the element below block 920 provides four (4) extra characters when the reference sequence number is zero. These four (4) extra characters are represented by element 405c. Partial length information 912b also indicates that the element below block 920 provides four (4) extra characters when the reference sequence number is one (1). These four (4) extra characters are represented by element 405b when the insertion of “Cat” in collaboration module instance 105a becomes visible to collaboration module instance 105b.
[0135] Figure 10 The updated tree portion 1000 on collaboration module instance 105b is shown. Tree portion 1000 reflects the state after collaboration module instance 105b receives confirmation of the insertion of "furry" from the server. As shown, the update assigns a sequence number to "furry".
[0136] Block 1020 includes minimum length information 1022. Block 1020 also indicates partial length information 1012a-b for collaboration module instances 105a and 102b, respectively. Partial length information 1012a indicates a partial length of four (4) characters for a reference sequence number of zero (0) or greater. This corresponds to the length of the data represented by element 505d initiated by collaboration module instance 105a. When the reference sequence number is zero, the data represented by elements 505a and 505c is not yet visible to collaboration module instance 105a, and therefore, when the reference sequence number is zero, no partial length information for this data is provided to collaboration module instance 105a.
[0137] Partial length information 1012a also indicates that when the reference sequence number for the collaboration module instance 105a is two, the four additional characters represented by element 505b become visible to the collaboration module instance 105a. When the reference sequence number is three, partial length information 1012a indicates that the elements below block 1020 include an additional fourteen (14) characters of information, including data represented by elements 505b and 505c (and 505d).
[0138] The partial length information 1012b for collaboration module instance 105b indicates that for reference sequence numbers greater than or equal to zero (0), ten characters of data are included in the elements below block 1020. These ten characters include data represented by elements 505b and 505c. The partial length information 1012b also includes an indication that for reference sequence numbers one (1) or greater, the elements below block 1020 include 14 characters of data. Compared to reference sequence number zero (0) in partial length information 1012b, the additional four (4) characters of data for reference sequence number one (1) are data represented by element 505d. The data represented by element 505d is assigned sequence number one (1) and initiated by collaboration module instance 105a.
[0139] Figure 11 An exemplary merge tree portion 1100 is shown on the collaboration module instance 105a after the synchronization service 106 has notified the collaboration module instance 105a of the insertion of “Furry”. The insertion is represented by element 405g. The notification from the server indicates to the collaboration module instance 105a that the Furry insertion is assigned sequence number three (3).
[0140] Block 1120, contained within tree section 1100, indicates minimum length information 1122. For collaboration module instances 105a-b, the partial length information for the merged tree section 1100 is shown as 1112a-b in block 1120, respectively. Partial length information 1112a indicates that, for collaboration module instance 105a, reference sequence number zero includes fourteen (14) additional characters of data in the elements below block 1120. These fourteen (14) additional characters include data represented by elements 405b, 405d, and 405f. When the reference number on collaboration module instance 105a reaches two (2) or more, the insertion of “Big” by collaboration module instance 105b becomes visible to collaboration module instance 105a, and thus the partial length information increases by four (4) relative to sequence number zero to a total of eighteen (18), as shown by partial length information 1112a. When the reference number on the collaboration module instance 105a reaches three (3) or more, the insertion of “furry” on the collaboration module instance 105b becomes visible, and therefore the partial length information 1112a indicates an additional six (6) characters for the data with reference number three (3), as shown.
[0141] Regarding collaboration module instance 105b, partial length information 1112b indicates that when the reference sequence number is zero, four (4) additional characters are provided by the element below block 1120. These four (4) additional characters are represented by element 405c. Partial length information 1112b also indicates that when the reference sequence number is one (1), the element below block 1120 provides four (4) additional characters. These four (4) additional characters are represented by element 405b when the insertion of “Cat” in collaboration module instance 105a becomes visible to collaboration module instance 105b.
[0142] Figure 12 The merge tree section 1200 is shown, indicating that the collaboration module instance 105a has received confirmation from the synchronization service 106 for the insertion of "top of". The synchronization service 106 assigns sequence number four (4) to the insertion operation of "top of". This is shown in element 405f.
[0143] Block 1220, contained within tree section 1200, indicates minimum length information 1222. For collaboration module instances 105a-b, partial length information for the merged tree section 1200 is shown as 1212a-b in block 1220, respectively. Partial length information 1212a indicates that, for collaboration module instance 105a, reference sequence number zero includes fourteen (14) additional characters of data in the elements below block 1220. These fourteen (14) additional characters include data represented by elements 405b, 405d, and 405f. When the reference number on collaboration module instance 105a reaches two (2) or more, the insertion of “Big” in collaboration module instance 105b becomes visible to collaboration module instance 105a, and thus the partial length information increases by four (4) relative to sequence number zero to a total of eighteen (18), as shown by partial length information 1212a. When the reference number on the collaboration module instance 105a reaches three (3) or more, the insertion of “furry” by the collaboration module instance 105b becomes visible, and therefore the partial length information 1212a indicates an additional six (6) characters of the data with reference number three (3), as shown.
[0144] Regarding collaboration module instance 105b, partial length information 1212b indicates that the element below block 1220 provides four (4) extra characters when the reference sequence number is zero. These four (4) extra characters are represented by element 405c. Partial length information 1212b also indicates that the element below block 1220 provides four (4) extra characters when the reference sequence number is one (1). These four (4) extra characters are represented by element 405b when the insertion of “Cat” in collaboration module instance 105a becomes visible to collaboration module instance 105b. This can then be indicated by updating partial length information 1212b by assigning the sequence number (4) to the insertion of “top of” represented by element 405f. In other words, when accessing tree section 1200, the collaboration module instance 105b has visibility for fifteen (15) characters of additional information (beyond what is indicated by minimum length information 1222) when its reference sequence number is four (4) or greater, as shown by partial length information 1212b. The data represented by element 405f provides an additional seven characters relative to sequence number one (1).
[0145] Figure 13The diagram shows a merged tree portion 1300 stored on collaboration module instance 105b after it receives a notification from collaboration module instance 105a regarding the insertion of the string "top of". The insertion is represented by element 505g. To complete the insertion, collaboration module instance 105b splits its previous "on the mat" element 505a into two elements 505e and 505f to represent "on" and "the mat" respectively. The "top of" insertion is then represented by element 505g, as shown.
[0146] Block 1320 includes minimum length values 1322 and partial length information 1312a-b for collaboration module instances 105a-b, respectively. Based on the "top of" insertion represented by element 505g, the indication of reference sequence number four (4) (equivalent to the sequence number assigned by synchronization service 106 to the "top of" insertion operation) provides an additional 7 characters of data relative to reference sequence number one, two, or three, as shown by partial length information 1312b.
[0147] Figure 14 This is a flowchart illustrating an exemplary process for distributing operations on a distributed data structure across multiple collaborating modules. In some aspects, the following section discusses... Figure 14 The process 1400 discussed can be performed by the synchronization service 106. In some aspects, instructions stored in memory (e.g., instruction 2324 below) can configure hardware processing circuitry (e.g., processor 2302 below) to perform one or more of the functions discussed below.
[0148] After starting operation 1402, process 1400 moves to operation 1405. In operation 1405, collaboration sessions are established with multiple collaboration module instances. For example, as mentioned above... Figure 1A As discussed, multiple collaborative module instances (e.g., collaborative module instances 105a and 105b) can each edit a distributed data structure, such as a text string. Establishing a collaborative session may include, for example, retrieving data defining a specific version of the distributed data structure, and in some aspects, obtaining data from one or more corresponding collaborative modules defining one or more in-process operations on the distributed data structure. In-process operations may not yet be synchronized with each collaborative module instance participating in the collaborative session.
[0149] In operation 1410, instructions for operations on the distributed data structure are received. For example, as mentioned above... Figure 1A The collaborative module instance discussed, which performs operations on the distributed data structure, can notify the synchronization service 106 of the operations (e.g., via...). Figure 2A(Message 200). Operation 1410 may also include receiving an indication of the version of the distributed data structure to which the operation is performed (e.g., via version field 204 of message 200). Further indications received in operation 1410 may include the type of operation (e.g., insert, delete, or comment), an indication of the cooperating module instance that initiated the operation (e.g., field 202), a value associated with the operation (e.g., a comment value or a string inserted as another example), and an identifier of the location in the distributed data structure where the operation is performed. For example, an offset in the distributed data structure may identify the insertion point used for insertion. As another example, an offset range may indicate a portion of the distributed data structure to be deleted by the operation.
[0150] In operation 1415, a sequence number is assigned to the operation. In some aspects, operation sequence numbers can be assigned to operations according to the order in which instructions for the operation are received by a public service (such as synchronization service 106) for the collaboration module instance. The sequence number assigned to each received operation can be a monotonically increasing number or a strictly increasing number. In some aspects, when a collaboration module instance initiates an operation and notifies the public service (e.g., synchronization service 106) of the operation, the collaboration module instance can initiate a sequence number for the new operation as a predefined number, such as -1. Upon receiving the notification, the public service can identify the need to assign a sequence number to the operation based at least in part on setting the sequence number for the operation to this predefined number.
[0151] In operation 1420, the information defining the operation is distributed to multiple collaborative module instances participating in the collaborative session. In some aspects, the information may be indicated as per the above regarding... Figure 2A One or more of the fields under discussion are consistent. Operation 1420 may include generating a message that includes the information and transmitting it to each collaborative module instance participating in the collaborative session. For example, Figure 1A The diagram illustrates messages 120a and 120b distributing messages to each of the collaboration module instances 105a and 105b, notifying each of those instances of an operation initiated by collaboration module instance 105a. In this particular example, collaboration module instance 105a, which has known about the operation since it was initiated by collaboration module instance 105a, identifies the sequence number assigned to the operation via message 120a.
[0152] Operation 1425 determines whether the information has been distributed to all collaborative module instances participating in the collaboration. If not, the process returns to operation 1420, where the information is distributed to one or more additional devices.
[0153] If the information is distributed to all collaboration module instances, process 1400 moves to operation 1430, which indicates to all devices participating in the collaboration that the sequence number assigned to the operation is synchronized. In other words, operation 1430 notifies all collaboration module instances that each collaboration module instance is now aware of the operation. In some aspects, operation 1430 is accomplished by updating field 214, as described above regarding... Figure 2A The discussed value indicates the equivalent of the operation sequence number assigned in operation 1415. Then, message 200 can be transmitted to all cooperating module instances. This is in Figure 1A Additional messages 130a and 130b are illustrated, which notify each of the collaboration module instances 105a and 105b that the operation has been fully synchronized across all collaboration module instances.
[0154] For example, as mentioned above... Figure 2A The messages discussed, transmitted from synchronization service 106 to collaboration module instances (e.g., 105a and / or 105b), define operations performed on the distributed data structure. In some aspects, operation 1410 also includes receiving (e.g., by synchronization service 106) an indication of a sequence number assigned to said operation. This indication may be included in a field of a message (e.g., 200) received by a collaboration field (e.g., field 206). Operation 1410 may also include receiving an indication of the version of the distributed data structure on which the operation is performed (e.g., a reference sequence number in version field 204).
[0155] Figure 15 This is a flowchart illustrating an exemplary process for distributing operations on a distributed data structure across multiple instances of collaborating modules. In some aspects, the following refers to... Figure 15 The process 1500 discussed can be executed by a cooperative module instance, such as any of the cooperative module instances 105a-b. In some aspects, instructions stored in memory (e.g., instruction 2324 hereinafter) can configure hardware processing circuitry (e.g., processor 2302 hereinafter) to perform one or more of the functions discussed below.
[0156] After initiating operation 1502, process 1500 moves to operation 1505. In operation 1505, the node initiating the operation, its location, reference sequence number, and the collaborating module instance are determined. For example, in some aspects, process 1500 may be executed to identify a portion of the tree to which the operation is to be applied. For example, insert, delete, or comment operations may be generated by a local operation or by a remote collaborating module instance. The operation may indicate a portion of the distributed data structure that the operation will affect. For example, an insert command may indicate an insertion point within the distributed data structure. A delete operation may indicate the range of data in the distributed data structure to be deleted. A comment operation may indicate the location in the distributed data structure for which a comment to be added is located.
[0157] Operation 1510 sets the offset variable to the position indicated in the search. As process 1500 continues, the offset variable can be changed.
[0158] Decision operation 1515 evaluates whether the node identified in operation 1505 is a block (a non-leaf node of the tree) or an element (a leaf node of the tree). If the node is a leaf, the position identified in operation 1505 is located or represented by the leaf node in operation 1518. Therefore, the identification of the leaf node can be returned as a result of the search. Process 1500 then moves from operation 1518 to end operation 1519. If the node is a non-leaf node, process 1500 moves from decision operation 1515 to operation 1520.
[0159] In operation 1520, the child nodes of the node are obtained. In some aspects, operation 1520 is configured to traverse the child nodes from a first position in the data structure (e.g., a node representing the beginning of the data structure) to a second position in the data structure (a node representing the end of the data structure). For example, when the distributed data structure represents a rich text string, operation 1520 can be configured to provide a node representing the beginning of the string, and subsequent calls to operation 1520 can provide nodes representing progressively later portions of the rich text string.
[0160] In operation 1525, a specific length of the collaborative module instance represented by child nodes is determined. This specific length is determined based on the collaborative module instance that initiated the operation and the reference sequence number at that collaborative module instance. For example, as described above for... Figure 4-13 As discussed in any of the exemplary tree structures, the disclosed embodiments may maintain partial length information that defines a specific length of a collaborative module instance of a distributed data structure represented by a specific portion of the tree structure. In some aspects, the partial length may be determined based on Equation 1 discussed above.
[0161] Operation 1525 can determine the specific length of the collaborative module instance based on the partial length and the minimum length value, wherein the minimum length value indicates the minimum size or length of the distributed data structure represented by a portion of the tree below the child node of operation 1520. The partial length value for the child node and the minimum length can be added together to determine the specific length of the collaborative module instance.
[0162] Decision operation 1530 determines whether the length determined in operation 1525 is greater than or equal to the offset value. If the length is greater than the offset, the position being searched (from operation 1505) is represented by a portion of the tree represented by the current node. In this case, process 1500 moves to operation 1535, which moves to a lower level in the tree by setting the node as a child of the previous node. If the length is not greater than the offset value, the current node does not represent a part of the distributed data structure including the position. Therefore, the offset value is adjusted in operation 1540 by the determined length. Then, process 1500 moves to operation 1545 and determines another child node of the node (the peer node of the current node).
[0163] Figure 16 This is a flowchart illustrating an exemplary process for distributing operations on a distributed data structure across multiple instances of collaborating modules. In some aspects, the following refers to... Figure 16 The process 1600 discussed can be executed by a cooperative module instance, such as any of the cooperative module instances 105a-b. In some aspects, instructions stored in memory (e.g., instruction 2324 hereinafter) can configure hardware processing circuitry (e.g., processor 2302 hereinafter) to perform one or more of the functions discussed below.
[0164] After initiating operation 1602, process 1600 moves to operation 1605. In operation 1605, information defining the operation is received. This information can be received from a public service for the collaboration session. This public service can communicate with collaboration module instances participating in the collaboration session. In some aspects, the public service is the synchronization service 106 discussed above. The information defining the operation may include the information above regarding message 200 and... Figure 2AOne or more items from the projects under discussion. For example, the information may include an operation sequence number assigned to the operation by the public service. The operation sequence number uniquely identifies the operation. The information may also include a reference sequence number for the operation. The reference sequence number identifies the version of the distributed data structure modified by the operation. The information may also include an indication of the operation type, such as whether the operation is an insertion, deletion, or annotation of the distributed data structure (such as a text string or rich text string). The information may include a value associated with the operation (e.g., data to be inserted into the data structure, or data used to annotate the data structure). The information may also include an indication of the maximum sequence number of the operation synchronized with each instance of the collaborative module participating in the collaborative session (e.g., the value of field 214). The information may also include the location within the distributed data structure for the operation to be applied (e.g., field 210 of message 200).
[0165] In operation 1610, elements are added or modified to the merged tree to represent the operation. In some aspects, operation 1610 may include updating collaboration module instance-specific partial length information for each block node between the added / modified element and the root of the tree. Thus, if collaboration module instance 105a initiates a specific operation, the partial length information specific to collaboration module instance 105a can be updated for each block node between the added node / element and the root.
[0166] For example, an operation that inserts data into a rich text string based on a specific reference sequence number may result in each node between the element representing the inserted data and the root node of the tree representing the rich text string being updated to correlate the specific reference sequence number with an increase in the partial length equivalent to the length of the inserted data. The partial length of the collaborative module instance that initiated the operation is also updated.
[0167] In operation 1615, receiving the operation associated with the sequence number is an indication that the operation is a synchronization operation. For example, as per... Figure 1A Each of the collaboration module instances 105a-h discussed can receive corresponding messages 130a-b indicating that the operation has been synchronized across all collaboration module instances participating in the collaboration. In some aspects, this can be achieved by setting field 214 of message 200 to a value equivalent to the sequence number of the operation assigned by a public service (e.g., synchronization service 106).
[0168] In operation 1620, the reference sequence number is updated based on the instruction received in operation 1615. In other words, the collaborative module implementation can track the maximum synchronization sequence number for the distributed data structure. When a collaborative module instance initiates an operation, it can use the tracked reference sequence number as the basis for the operation. In other words, the tracked reference sequence number is used to identify the version of the distributed data structure on which the collaborative module instance operates. This information is provided when these operations are shared with other collaborative module instances via a common service such as synchronization service 106. For example, the version of the distributed data structure used as the basis for a particular operation can be transmitted to other collaborative module instances via message 200, and specifically, in some embodiments, via version field 204, as discussed above. After operation 1620 is completed, process 1600 moves to end operation 1625.
[0169] Figure 17 This is a flowchart illustrating an exemplary process for accessing a distributed data structure. In some aspects, process 1700 may be executed by a device running a cooperative module instance. For example, process 1700 may be executed by one or more devices, namely devices 102a and / or 102b. In some aspects, the execution of the following... Figure 17 Instances of one or more of the collaborative modules discussed can run on the server side of the implementation, such as on a computer that also runs the synchronization service 106. In some aspects, the following section discusses… Figure 17 The process 1700 discussed can be executed by a cooperative module instance, such as any of the cooperative module instances 105a-b. In some aspects, instructions stored in memory (e.g., instruction 2324 hereinafter) can configure hardware processing circuitry (e.g., processor 2302 hereinafter) to perform one or more of the functions discussed below. Figure 17 In the discussion, the device that performs process 1700 can be referred to as the "execution device".
[0170] After starting operation 1702, process 1700 moves to operation 1705. In operation 1705, the least recently used node is identified. As discussed above, in embodiments representing distributed data structures via trees, a list of least recently used nodes on the tree can be maintained. The list of least recently used nodes can be used to facilitate garbage collection and / or other optimizations of the tree structure.
[0171] In some respects, in addition to the LRU list, nodes can be identified by operation 1705 in other ways. For example, as discussed above, as all collaborating module instances confirm the maximum sequence number of the operation, some maintenance on the data structure representing the distributed data structure can be performed, for example, as described below.
[0172] Decision operation 1715 evaluates whether the sequence number assigned to the operation represented by the node is less than the largest sequence number already confirmed by all participants in the collaboration (e.g., collaboration module instances). In other words, the node represents an operation within or below the collaboration window. In some embodiments, this can be achieved by comparing two sequence numbers identified by the node. For example, the node may indicate an insert or comment sequence number (e.g., via field 324) and a delete sequence number (via field 325). If multiple sequence numbers are indicated, all of them must pass through below the collaboration window before further operations (e.g., garbage collection) can be performed on the node.
[0173] If the node represents an operation that is still within the collaboration window, then procedure 1700 returns to operation 1705 and may select another node from the LRU list.
[0174] If the operation represented by the node has been passed out of the collaboration window, process 1700 moves from decision operation 1715 to operation 1720, which evaluates the operation as a deletion. To make this determination, operation 1720 may evaluate whether the node includes a deletion sequence number (e.g., field 325). If a deletion sequence number is indicated, process 1700 moves to operation 1722 and deletes the leaf node. Processing then returns to operation 1705. If the node does not represent a deletion operation, process 1700 moves to decision operation 1725, which determines whether the operation is an insertion operation. If the operation is an insertion, process 1700 moves from decision operation 1725 to operation 1730, which identifies adjacent parts of the distributed data structure (e.g., represented by sibling nodes). Operation 1740 may integrate adjacent parts of the distributed data structure into a single leaf node. This integration may depend on whether those other parts have also been passed down to the collaboration window. Process 1700 then returns to obtain another node from the LRU list.
[0175] If the operation is not an insertion, process 1700 moves from decision operation 1725 to decision operation 1750, which evaluates whether there are more tree nodes. If there are more tree nodes, process 1700 moves from decision operation 1750 to operation 1705, as discussed above. If there are no more tree nodes, process 1700 moves from decision operation 1750 to termination operation 1752.
[0176] Figure 18 This is a flowchart illustrating an exemplary process for accessing a distributed data structure. In some aspects, process 1800 may be executed by a device running a collaboration module instance. For example, process 1800 may be executed by one or more devices, namely devices 102a and / or 102b. In some aspects, the collaboration module instance performs the following... Figure 18 One or more of the functions discussed. The collaboration module instances (e.g., 105a-b) can run on the server side of the implementation, such as on a computer that also runs the synchronization service 106. In some aspects, the following refers to... Figure 18 The process 1800 discussed can be executed by a cooperative module instance, such as any of the cooperative module instances 105a-b. In some aspects, instructions stored in memory (e.g., instruction 2324 hereinafter) can configure hardware processing circuitry (e.g., processor 2302 hereinafter) to perform one or more of the functions discussed below. Figure 18 In the discussion, the device that performs process 1800 can be referred to as the "execution device".
[0177] After starting operation 1805, process 1800 transitions to operation 1810, which identifies the starting node. In some aspects, the starting node may be the root node of a tree representing a distributed data structure.
[0178] In operation 1815, the partial length information for the operation sequence number is less than or equal to the maximum sequence number confirmed by all collaborative module instances participating in the collaboration (e.g., the value of field 214). The reasoning supporting operation 1815 is that once a particular sequence number is less than or equal to the maximum sequence number confirmed by all collaborative module instances, the operation assigned to that particular sequence number has already been seen and / or applied by all collaborative module instances participating in the collaboration, and therefore, it is not necessary to maintain a partial length for that operation. Note that the minimum length value for the node can be adjusted based on the deleted partial length information.
[0179] Decision operation 1820 determines whether there are any additional nodes to check. If not, process 1800 completes at end operation 1840. Otherwise, process 1800 moves to operation 1830, which recursively traverses to the next node in the tree.
[0180] Figure 19This is a flowchart of an exemplary process for accessing a distributed data structure. In some aspects, process 1900 may be executed by a device running a collaboration module instance. For example, process 1900 may be executed by one or more devices, namely devices 102a and / or 102b. In some aspects, the collaboration module instance may run on the server side of an implementation, such as on a computer that also runs synchronization service 106. In some aspects, instructions stored in memory (e.g., instruction 2324 hereinafter) may configure hardware processing circuitry (e.g., processor 2302 hereinafter) to perform one or more of the functions discussed below. Figure 19 In the discussion, the device that performs process 1900 can be referred to as the "execution device".
[0181] After initiating operation 1905, process 1900 moves to operation 1910, which joins a collaborative session. The collaborative session provides access to a distributed data structure. Joining the collaborative session may include establishing a session with synchronization service 106. In some aspects, joining the collaborative session may include interfaced with other devices and / or collaborative module instances participating in the collaborative session via a peer-to-peer protocol.
[0182] In operation 1915, offline editing is performed on the distributed data structure. Performing offline editing may include generating operations on the distributed data structure. In embodiments utilizing a tree structure such as those discussed herein, performing the offline editing may include generating non-leaf nodes and / or leaf nodes, if necessary, to represent the offline operation locally. Because the editing can be generated when the execution device is unable to contact synchronization service 106 (or other peer devices when using a peer-to-peer protocol), the collaboration window maintained at the execution device may become relatively large, encompassing all operations for providing the offline editing.
[0183] In operation 1920, the execution device rejoins the cooperative session. For example, although the execution device may have lost network connectivity to synchronization service 106 during operation 1915, network connectivity between synchronization service 106 and the execution device is restored in 1920.
[0184] In operation 1925, a snapshot of the distributed data structure is received. As discussed above, a snapshot indicates the absolute value or state of the distributed data structure at a point in time. Thus, as an example, if one or more operations were applied to the distributed data structure prior to the snapshot, the snapshot represents the combined result of those one or more operations.
[0185] In operation 1930, any additional operations occurring after the snapshot (initiated by other cooperating module instances) are transmitted to the execution device (and the execution cooperating module instance). This allows the execution device to represent a version of the distributed data structure based on the snapshot and the additional operations. Note that only operations with sequence numbers greater than those of the snapshot version are included in the version.
[0186] In operation 1935, the offline edit is applied to the snapshot. In some aspects, operation 1935 includes transmitting message 200 for each offline edit, each message defining a particular offline edit as an operation on the distributed data structure. Each message will be constructed by process 1900, consistent with the description of message 200 discussed above. Operation 1935 may also include receiving a corresponding number of responses / messages from the synchronization service, wherein the synchronization service 106 assigns a sequence number to each offline edit / operation. The execution device may then represent an additional version of the distributed data structure, which includes both the operations of 1930 and 1935. This process will generate a version of the distributed data structure maintained by the execution device, consistent with the distributed data structure maintained by other instances of the collaborating modules participating in the collaboration.
[0187] Decision operation 1940 determines whether there is any conflict between the offline editing and the operation applied to the snapshot. If a conflict exists, process 1900 moves to operation 1950, which displays a conflict resolution dialog. The conflict resolution dialog is configured to provide a manual selection of one of the two conflicting operations. The selected operation is applied to the distributed data structure, and the second of the two operations is canceled. If decision operation 1940 determines that there is no conflict, process 1900 moves from decision operation 1940 to end operation 1955.
[0188] Figure 20 This is a flowchart illustrating an exemplary process for accessing a distributed data structure. In some aspects, process 2000 may be executed by a device running an instance of a collaborative module. For example, process 2000 may be executed by one or more devices, namely devices 102a and / or 102b. In some aspects, the execution of the following... Figure 20 Instances of one or more of the collaborative modules discussed can run on the server side of the implementation, for example, on a computer that also runs Synchronization Service 106. In some aspects, the following section discusses… Figure 20The process 2000 discussed can be executed by a cooperative module instance, such as any of the cooperative module instances 105a-b. In some aspects, instructions stored in memory (e.g., instruction 2324 hereinafter) can configure hardware processing circuitry (e.g., processor 2302 hereinafter) to perform one or more of the functions discussed below. Figure 20 In the discussion, the device performing process 2000 may be referred to as an "execution device". Note that in various embodiments, the above at least refers to Figure 15-19 One or more of the functions discussed in 21 and / or 22 may be included in process 2000.
[0189] After starting operation 2005, process 2000 moves to operation 2010. Operation 2010 joins a collaboration session. The collaboration session provides access to a distributed data structure. Joining a collaboration session may include opening a network connection to a service such as synchronization service 106. Joining the collaboration session may also include identifying the distributed data structure that the collaboration session will provide access to to the service. In some aspects, the distributed data structure may be in the form of a file. In some aspects, the file may be stored on a stable storage device accessible via a network. In some aspects, the distributed data structure may be identified via a Uniform Resource Locator (URL).
[0190] In operation 2015, a message identifying a sequentially ordered operation on the distributed data structure is received. In some aspects, the message is received from synchronization service 106. In other aspects, the message may be received from one of the other collaborative module instances participating in the collaborative session, utilizing a peer-to-peer protocol for communication. Each identified operation has pending confirmation. Sequentially ordered operations define a collaboration window for the distributed data structure. In other words, the collaboration window represents an operation on the distributed data structure initiated by one of the collaborative module instances participating in the collaboration, but which has not yet received confirmation from all the collaborative module instances participating in the collaboration. Therefore, when the collaboration is considered as a whole (regarding all the collaborative module instances participating in the collaboration), the collaboration window represents an operation that is still "in progress".
[0191] As mentioned above Figure 2AAs discussed in message 200, an operation sequence number (e.g., via field 206) and a maximum sequence number for synchronization or acknowledgment operations (e.g., via field 214) can be indicated. These two values provide an indication of the sequentially ordered operations of the distributed data structure. For example, if the maximum sequence number for an acknowledgment operation is N, and the operation sequence number is N+C (where C is a constant), then there are C sequentially ordered operations identified by messages with pending acknowledgments. This is because when all devices acknowledge the operation, the maximum sequence number for the acknowledgment operation (e.g., 214) advances the operation sequence number (e.g., 206) "forward". New operations will advance the operation sequence number as it is assigned to the cooperating module instance participating in the collaboration.
[0192] In Operation 2020, a first version of the distributed data structure is represented as including each operation in a sequentially ordered sequence. By representing the distributed data structure as including each operation in the sequentially ordered sequence, subsequent operations on the first version are based on the result of each operation in the sequentially ordered sequence. This can be achieved in some ways by storing or queuing information defining the sequentially ordered operations. For example, some embodiments may use a queue or tree data structure to store pending (unconfirmed) operations. In these embodiments, once the sequentially ordered operations are properly stored or queued, the distributed data structure is considered to "represent" those operations.
[0193] In some aspects, once an operation is confirmed by all collaborating module instances participating in the collaboration, the operation can be applied to the distributed data structure. In other words, the data contained in the distributed data structure may be irrevocably modified by the operation. Details regarding how to represent operations on the distributed data structure vary depending on the embodiment. In some embodiments, a garbage collection process or other process asynchronous to the collaboration on the distributed data structure can apply the operation to the data contained in the distributed data structure. For example, as described above, a garbage collection process can operate on operations confirmed by all collaborating module instances participating in the collaboration. In some aspects, the distributed data structure can be modified based on the operation, based on the message received in operation 2015.
[0194] In operation 2025, a second version of the distributed data structure is represented as including a first operation following the last operation in the sequentially ordered operations. The first operation is initiated by the execution device. As discussed above, some implementations may store information defining pending operations. Therefore, in these embodiments, information defining the first operation may be stored, queued, or otherwise recorded, such as the type of operation (e.g., field 208), the location or scope within the distributed data structure to which the operation is applied (e.g., field 210), the version of the distributed data structure to which the operation is applied (e.g., field 204), and the cooperating module instance that initiated the operation (e.g., field 202). By representing a second version of the distributed data structure, subsequent operations performed on the distributed data structure take into account the result of the first operation. By representing a first operation on the distributed data structure, operations following the first operation depend on the result of the first operation (to a certain extent, subsequent operations depend on the result of the first operation).
[0195] In operation 2030, a notification message is transmitted. The notification message indicates that the first operation is applied to a first version of the distributed data structure. In some aspects, the notification message includes the above-mentioned... Figure 2A One or more fields are described. One or more values of the fields are specific to the first operation. In an embodiment utilizing a centralized service for operation serialization, the notification may be transmitted to synchronization service 106. In other embodiments utilizing a peer-to-peer protocol for serialization, the notification may be transmitted to another instance of a collaboration module participating in the collaboration.
[0196] Some aspects of process 2000 include receiving multiple notifications. Each of these notifications is for a corresponding operation. As with the notifications above, each notification indicates a unique sequence number assigned to the corresponding operation. Each notification also identifies the collaborative module instance that initiated the operation. In some aspects, there may be a one-to-one mapping between collaborative module instances and client devices, and therefore, the identifier of a collaborative module instance may be synonymous with the identifier of a specific, physically different computing device running a particular instance of the collaborative module. The unique sequence number assigned to each of the operations indicates the order in which the operation is applied to the distributed data structure. One or more of the multiple notifications received by process 2000 can be used for operations initiated by the execution device. For example, as discussed above, when a collaborative module instance initiates a new operation, it can send values corresponding to one or more fields described above with respect to message 200 to other collaborative module instances participating in the collaboration. Since a sequence number has not yet been assigned to the operation, it can send a predetermined value indicating it. Once a sequence number is assigned, a notification is provided to the collaborative module instance that initially initiated the new operation. The notification will indicate the identifier of the receiving collaboration module instance (for example, in field 202).
[0197] In some embodiments, process 2000 may include initiating a second operation on the distributed data structure and subsequently receiving notification of an additional operation. The notification may also indicate (e.g., via a sequence number assigned to the additional operation) that the additional operation will be applied to the distributed data structure prior to the second operation. Alternatively, if the execution device receives a notification assigning a sequence number to the second operation and subsequently receives a second notification assigning a second sequence number to the additional operation, the second operation is applied to the distributed data structure prior to the additional operation.
[0198] Some aspects of process 2000 include resolving edit conflicts between two operations on the distributed data structure. The edit conflict is resolved based on the relative order of two sequence numbers, each of which is assigned to one of the conflicting operations. For example, an edit conflict between two edit operations might cause a later-ordered insertion to appear before an earlier-ordered insertion in the distributed data structure. Edit conflicts between two delete operations are resolved by performing delete operations in an order consistent with the order of their assigned sequence numbers.
[0199] Some aspects of process 2000 include receiving an instruction for an updated collaboration window that excludes some operations in a series of operations; and performing garbage collection on the excluded operations based on the received instruction.
[0200] After operation 2030, process 2000 moves to end operation 2035.
[0201] Figure 21 This is a flowchart of an exemplary process for accessing a distributed data structure. In some aspects, process 2100 may be executed by a device running a collaboration module instance. For example, process 2100 may be executed by one or more devices, 102a and / or 102b. In some aspects, the collaboration module instance may also run on the server side of the implementation, such as on a computer that also runs synchronization service 106. In some aspects, the following refers to... Figure 21 The process 2100 discussed can be executed by a cooperative module instance, such as any of the cooperative module instances 105a-b. In some aspects, instructions stored in memory (e.g., instruction 2324 hereinafter) can configure hardware processing circuitry (e.g., processor 2302 hereinafter) to perform one or more of the functions discussed below. Figure 21 In the discussion, the device performing process 2100 may be referred to as an "execution device". Note that in various embodiments, the above at least refers to Figure 15-20 and / or Figure 22 One or more of the functions discussed may be included in process 2100.
[0202] After initiating operation 2105, process 2100 moves to operation 2110. Operation 2110 joins a collaboration session. The collaboration session provides access to a distributed data structure. Joining a collaboration session may include opening a network connection to a service such as synchronization service 106. Joining the collaboration session may also include identifying to the service the distributed data structure that the collaboration session will provide access to. In some aspects, the distributed data structure may be in the form of a file. In some aspects, the file may be stored on a stable storage device accessible via a network. In some aspects, the distributed data structure may be identified via a Uniform Resource Locator (URL).
[0203] In operation 2120, a notification is received. In at least some embodiments, the notification may be in the form of a network message. The notification may be received from synchronization service 106. In some other embodiments where communication is conducted between collaboration module instances in the collaboration session using a peer-to-peer protocol, the notification may be received from one of the other collaboration module instances participating in the collaboration session.
[0204] The notification indicates an operation performed by a remote device. In other words, the operation is performed by a device other than the execution device. The operation may be, for example, an insertion, deletion, or commenting operation on a sequence data structure such as a text string, rich test string, or stream. The operation may be indicated via a sequence number assigned to the operation.
[0205] The notification also indicates the position of the operation relative to the distributed data structure. For example, the position may include an offset from the beginning data point or byte of the distributed data structure to the position in the distributed data structure where the insertion, deletion, or comment operation is to be applied, in bytes or words in various respects. In the case of an insertion operation, the position indicates the location where new data will be inserted into the distributed data structure, wherein data from the previous version of the distributed data structure after the position is located after the newly inserted data. In the case of a comment operation, the position indicates which part of the distributed data structure is to be commented. In the case of a removal operation, the position indicates the data in the distributed data structure to be removed. For example, the position may indicate a range of data locations to be removed from the distributed data structure by a removal operation.
[0206] The notification may also indicate the version of the distributed data structure to which the operation is applied. For example, as discussed above, the highest operation sequence number applied to a particular distributed data structure can define the version of that particular distributed data structure. This version or sequence number information can be obtained via methods described above. Figure 2A The message discussed, which includes one or more fields of message 200, is transmitted to a collaboration module instance participating in the collaboration, such as a collaboration module instance running on the execution device.
[0207] In operation 2130, the minimum length of a portion of the distributed data structure represented by the first node of the tree is identified. The tree represents the distributed data structure. The first node referred to in operation 2130 can be any node in the tree, including the root node or nodes below the root node. In some aspects, the first node can be a leaf node of the tree.
[0208] As discussed above, in embodiments where the distributed data structure is represented as a tree, some portions of the distributed data structure can be synchronized across all collaborating module instances participating in the collaboration. Therefore, these portions are constant for all these collaborating module instances, and the length of these portions can be determined. This length is the minimum length mentioned in operation 2130 and is consistent with the discussion of minimum lengths throughout this disclosure.
[0209] In some embodiments, the notification may also indicate the minimum sequence number of the operation to be confirmed or synchronized with all other clients participating in the collaboration. For example, as described above regarding... Figure 2A The notification discussed may include field 214.
[0210] In operation 2140, a partial length specific to the second device and represented by the first node is determined. As discussed above, when participating devices apply operations to a distributed data structure, the distributed data structure evolves by incorporating sequential versions of each of those operations. Therefore, two participating collaborative module instances may have slightly different versions of the distributed data structure until a specific edit has been delivered to both collaborative module instances. The minimum length determined in operation 2140 takes into account these differences in the versions of the distributed data structure across the collaborative module instances. Operation 2140 determines a partial length represented by the first node of the version of the distributed data structure to which the second device operation is applied. As discussed above, this partial length can be specific not only to the collaborative module instance performing the operation but also to the version of the distributed data structure to which a specific operation is applied. In some aspects, Equation 1 may be used to determine the partial length. In some embodiments, the partial length may have been previously determined before operation 2140 is performed, and the first node is updated to reflect the partial length.
[0211] In operation 2145, the size of a portion of the distributed data structure represented by the first node is determined. The size is based on a minimum length and a portion length. In some aspects, the size is the sum of the minimum length and the portion length.
[0212] In operation 2150, it is determined that the size of the portion is greater than or equal to the position of the operation within the distributed data structure. As discussed above, in embodiments where the distributed data structure is represented as a tree, specific nodes of the tree represent specific portions of the distributed data structure. Procedure 2100 describes how embodiments identify which specific nodes represent a portion of the tree to which an operation is being applied, based on the length of the distributed data structure represented by each node, by searching the nodes of the tree.
[0213] Although operation 2150 indicates that the location of the second device operation is at or below the first node, in other examples, the size may be smaller than the location, indicating that other child nodes of the first node need to be searched to identify nodes representing the portion of the distributed data structure including the location. This process can continue across the sibling nodes of the first node until the appropriate branch of the tree is identified. The process can then be performed at lower levels in the tree to gradually narrow down the portion of the distributed data structure until a leaf node representing that portion is identified.
[0214] In operation 2160, the leaf node is determined based on the determination in operation 2150. Since operation 2150 determines that the portion of the distributed data structure including the location is represented by the first node, operation 2160 can proceed deeper in the tree (to the child nodes of the first node) to gradually narrow the search range until the leaf node representing the location is identified.
[0215] In operation 2170, the second device operation is represented in the tree based on the identified leaf nodes. Representing the operation can include various items. If the operation is an insertion or removal operation, the identified leaf node can be split, if necessary, to provide a single leaf node to include a portion of the distributed data structure affected by the operation. For example, if an insertion operation is performed, the leaf node can be split into three nodes at the location: a first node representing a portion of the distributed data structure prior to the inserted data, a second leaf node representing the inserted data, and a third leaf node representing any remaining data from the identified leaf node. For a removal operation, the leaf node can be updated to indicate a sequence number for the operation. The identified leaf node can also be split for a removal operation such that the removed data is represented by a single leaf node, or at least the unremoved data is not represented in the same leaf node as the deleted data. For a comment operation, the leaf node can be updated to include the comment information. The leaf node can also be appropriately split, if necessary, to represent operations in the tree.
[0216] After operation 2170 is completed, process 2100 moves to end operation 2175.
[0217] Figure 22 This is a flowchart of an exemplary process for accessing a distributed data structure. In some aspects, process 2200 may be executed by a device running a collaboration module instance. For example, process 2200 may be executed by one or more devices, 102a and / or 102b. In some aspects, the collaboration module instance may also run on the server side of the implementation, such as on a computer that also runs synchronization service 106. In some aspects, the following refers to... Figure 22 The process 2200 discussed can be executed by a cooperative module instance, such as any of the cooperative module instances I05a-b. In some aspects, instructions stored in memory (e.g., instruction 2324 hereinafter) can configure hardware processing circuitry (e.g., processor 2302 hereinafter) to perform one or more of the functions discussed below. Figure 22In the discussion, the device that performs the process 2200 can be referred to as the "execution device".
[0218] Note that in various embodiments, the above refers at least to 15-19 and / or Figure 20 and / or Figure 21 and / or Figure 24 One or more of the functions discussed may be included in process 2200.
[0219] After initiating operation 2202, process 2200 moves to operation 2205. In operation 2205, the execution device joins a collaboration. The collaboration provides access to a distributed data structure. Joining a collaboration session may include opening a network connection to a service such as synchronization service 106. Joining the collaboration session may also include identifying to the service the distributed data structure that the collaboration session will provide access to. In some aspects, the distributed data structure may be in the form of a file. In some aspects, the file may be stored on a stable storage device accessible via a network. In some aspects, the distributed data structure may be identified via a Uniform Resource Locator (URL).
[0220] In operation 2210, instructions for multiple corresponding operations on the serialization of a distributed data structure are received. In some aspects, each of these instructions may be a message from the serialization service, including those mentioned above regarding message 200 and... Figure 2A One or more of the fields discussed. The indication also indicates the device that initiated each operation in the operation (e.g., field 202 of message 200). Instead of indicating a device, the indication may alternatively indicate a collaborative module instance, as discussed above.
[0221] In operation 2215, the leaf nodes of the tree represent the results of multiple operations on the distributed data structure (e.g., via 320) and the operations on said distributed data structure (e.g., at least via 322). As discussed above, for example, regarding Figure 4-13 Operations can be represented in the leaf nodes of the tree, and operations can be selectively applied to the data of the distributed data structure based on the originating device (e.g., a cooperative module instance), the version of the DDS to which the operation is applied by the originating device, and the sequence number assigned to the operation.
[0222] In operation 2220, a specific length of the originating device (or originating cooperative module instance) is represented for a portion of the distributed data structure represented by the leaf nodes below the non-leaf nodes. In various embodiments, operation 2230 may be performed on each non-leaf node in the tree, or at least on the non-leaf nodes between the leaf nodes and the root node of the tree. As discussed above, partial length information is contained in the non-leaf nodes of the tree. A minimum length value, representing the minimum length of the DOS represented by the leaf nodes below the non-leaf nodes, may also be represented in the non-leaf nodes.
[0223] In operation 2225, additional instructions for additional operations performed by a specific originating device are received. Instructions are also received regarding the location within the distributed data structure where the operation is applied (e.g., the location to insert, the range of locations to remove, or the location to comment). In some aspects, operation 2225 includes receiving a message that includes the above information regarding... Figure 2A One or more of the fields discussed, wherein the message provides the indication.
[0224] In operation 2230, a leaf node of the tree representing a portion of the distributed data structure to which the additional operation is applied is identified. For example, if the operation is an insertion into a string at offset X, operation 2230 identifies a leaf node representing offset X in the string. To identify the leaf node, operation 2230 relies on one or more partial length values from the partial length values discussed above, specific to the device (or cooperative module instance) initiating the operation. The leaf node is further identified based on the version of the distributed data structure to which the additional operation is applied by the specific originating device.
[0225] In operation 2235, the leaf node is modified based on the additional operation. As discussed above, some exemplary modifications include splitting the leaf node into multiple nodes to facilitate data insertion. For a deletion operation, the leaf node can be modified by recording the sequence number assigned to the deletion operation. In some embodiments, the leaf node can also be split to provide leaf nodes dedicated to the deleted data. If the operation is a comment operation, the comment data can be copied to the leaf node or otherwise associated with the leaf node. (Examples of text in bold strings that are comments).
[0226] As discussed above, process 2200 may include receiving a message from a serialization service (or peer-to-peer protocol) indicating one or more fields in message 200, as described above regarding... Figure 2A The above is about... Figure 2BThe messages discussed define a collaboration window for operations that have not yet been acknowledged by all collaborating participants (collaboration module instances or devices). Process 2200 may also include applying maintenance procedures to the leaf nodes of the tree representing the operations below the collaboration window. As discussed above, these maintenance operations can include combining leaf nodes when those leaf nodes represent continuous data of the DDS, or completely removing nodes (e.g., if data is deleted by an operation below the collaboration window, then the leaf node representing that data is removed from the tree). After operation 2235, process 2200 moves to end operation 2238.
[0227] Figure 23 A block diagram of an exemplary machine 2300 is illustrated, on which one or more of the techniques (e.g., methods) discussed herein can be executed. In alternative embodiments, machine 2300 may operate as a standalone device or may be connected (e.g., networked) to other machines. In a networked deployment, machine 2300 may operate as a server machine, a client machine, or both in a server-client network environment. In the example, machine 2300 may act as a peer-to-peer (P2P) (or other distributed) network environment. Machine 2300 may be a personal computer (PC), tablet PC, set-top box (STB), personal digital assistant (PDA), mobile phone, smartphone, network device, network router, switch or bridge, server computer, database, conference room equipment, or any machine capable of executing instructions (sequentially or otherwise) specifying actions to be taken by a designated machine. In various embodiments, machine 2300 may execute the above description with respect to Figures 1-22 or the following description with respect to Figures 1-22. Figure 24 One or more processes described herein. Furthermore, although only a single machine is illustrated, the term "machine" should also be considered as any collection of machines that individually or jointly execute a set (or more sets) of instructions to perform any one or more of the methods discussed herein, such as cloud computing, Software as a Service (SaaS), and other computer cluster configurations.
[0228] As described herein, examples may include logic or multiple components, modules, or mechanisms (hereinafter referred to as "modules") or on which operations may be performed. A module is a tangible entity (e.g., hardware) capable of performing the specified operations and configured or arranged in a particular manner. In the examples, circuitry may be arranged in a specified manner (e.g., internally or relative to an external entity such as other circuitry) as a module. In the examples, all or part of one or more computer systems (e.g., standalone, client, or server computer systems) or one or more hardware processors may be configured as modules by firmware or software (e.g., instructions, application portions, or applications) that operate to perform the specified operations. In the examples, the software may reside on a machine-readable medium. In the examples, when the software is executed by the underlying hardware of the module, it causes the hardware to perform the specified operations.
[0229] Therefore, the term "module" is understood to encompass tangible entities that are physically constructed, specifically configured (e.g., hardwired), or temporarily (e.g., provisionally) configured (e.g., programmed) to operate in a particular manner or perform any of the operations described herein. Considering an example where modules are temporarily configured, it is not necessary to instantiate each module at any given time. For example, in the case where the modules include a general-purpose hardware processor configured using software, the general-purpose hardware processor can be configured as distinct modules at different times. The software can accordingly configure the hardware processor, for example, to constitute a particular module at one time and different modules at different times.
[0230] Machine (e.g., computer system) 2300 may include a hardware processor 2302 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a hardware processor core, or any combination thereof), main memory 2304, and static memory 2306, some or all of which may communicate with each other via interconnect 2308 (e.g., a bus). Machine 2300 may also include a display unit 2310, an alphanumeric input device 2312 (e.g., a keyboard), and a user interface (UI) navigation device 2314 (e.g., a mouse). In the example, the display unit 2310, the input device 2312, and the UI navigation device 2314 may be a touchscreen display. Machine 2300 may additionally include a storage device (e.g., a drive unit) 2316, a signal generation device 2318 (e.g., a speaker), a network interface device 2320, and one or more sensors 2321, such as a global positioning system (GPS) sensor, a compass, an accelerometer, or other sensors. Machine 2300 may include output controller 2328, such as serial (e.g., Universal Serial Bus (USB), parallel, or other wired or wireless (e.g., infrared (IR), near field communication (NFC), etc.) connections, to communicate with or control one or more peripheral devices (e.g., printers, card readers, etc.).
[0231] Storage device 2316 may include machine-readable medium 2322 thereon storing one or more sets of data structures or instructions 2324 (e.g., software) embodying or utilized by any one or more of the technologies or functions described herein. Instructions 2324 may also reside wholly or at least partially in main memory 2304, static memory 2306, or hardware processor 2302 during execution by machine 2300. In the example, one or any combination of hardware processor 2303, main memory 2304, static memory 2306, or storage device 2316 may constitute a machine-readable medium.
[0232] Although machine-readable medium 2322 is illustrated as a single medium, the term "machine-readable medium" can include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) configured to store one or more instructions 2324.
[0233] The term "machine-readable medium" can include any medium capable of storing, encoding, or carrying instructions executable by machine 2300 and causing machine 2300 to perform any one or more of the technologies disclosed herein, or a data structure capable of storing, encoding, or carrying data used by or associated with such instructions. Examples of non-limiting machine-readable media can include solid-state memory as well as optical and magnetic media. Specific examples of machine-readable media can include: non-volatile memory, such as semiconductor memory devices (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM)) and flash memory devices; disks, such as internal hard disks and removable disks; magneto-optical disks; random access memory (RAM); solid-state drives (SSDs); and CD-ROM and DVD-ROM disks. In some examples, machine-readable media can include non-transitory machine-readable media. In some examples, machine-readable media can include machine-readable media that do not transiently propagate signals.
[0234] Instruction 2324 can also be sent or received on communication network 2326 via network interface device 2320 using a transmission medium. Machine 2300 can communicate with one or more other machines using any of a variety of transmission protocols (e.g., Frame Relay, Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), etc.). Exemplary communication networks may include local area networks (LANs), wide area networks (WANs), packet data networks (e.g., the Internet), mobile phone networks (e.g., cellular networks), conventional telephone (POTS) networks, and wireless data networks (e.g., the Institute of Electrical and Electronics Engineers (IEEE) 802.11 standard series, referred to as...). The IEEE 802.16 standard series is known as This includes standards such as the IEEE 802.15.4 series, the Long Term Evolution (LTE) series, the Universal Mobile Telecommunications System (UMTS) series, and peer-to-peer (P2P) networks. In the example, network interface device 2320 may include one or more physical jacks (e.g., Ethernet, coaxial, or telephone jacks) or one or more antennas to connect to communication network 2326. In the example, network interface device 2320 may include multiple antennas to perform wireless communication using at least one of single-input multiple-output (SIMO), multiple-input multiple-output (MIMO), or multiple-input single-output (MISO) technologies. In some examples, network interface device 2320 may use multi-user MIMO technology for wireless communication.
[0235] Figure 24This is a flowchart of an exemplary process that can be implemented by a serialization service. For example, process 2400 can be executed by synchronization service 106. In some aspects, instructions stored in memory (e.g., instruction 2324 above) can configure hardware processing circuitry (e.g., processor 2302 above) to perform one or more of the functions discussed below. Figure 24 In the discussion, the device performing process 2200 may be referred to as an "execution device". Note that in various embodiments, the above at least refers to Figure 14-19 and / or Figure 20 and / or Figure 21 and / or Figure 22 One or more of the functions discussed may be included in process 2400.
[0236] After initiating operation 2402, process 2400 moves to operation 2405, which establishes collaborative sessions with multiple devices. These collaborative sessions provide access to distributed data structures.
[0237] In operation 2410, a message is received. This message requests the allocation of a sequence number for the operation. For example, operation 2410 may receive a message including the information mentioned above regarding... Figure 2A A message containing one or more of the fields discussed in message 200. For example, the message may indicate the cooperating module instance or device initiating the operation (e.g., field 202) and request that a sequence number be assigned to the operation indicated in the message (e.g., an operation that may be indicated by one or more of fields 208, 210, 212). In some aspects as described above, the request may be indicated via a predetermined value (e.g., -1) included in the operation sequence number field 206.
[0238] In some aspects, the received message can be decoded to identify that the sequence number indicated in the message is set to a predetermined value, wherein the predetermined value indicates that the sequence number needs to be assigned to an operation (e.g., the operation sequence number field 206 can be set to a predetermined value, such as -1). This determination can then be made based on the identifier that the message is requesting sequence number assignment.
[0239] In some aspects, the received message indicates the version number of the distributed data structure associated with the operation (e.g., via field 204).
[0240] In operation 2415, a sequence number is assigned to the operation. The operation is executed in response to receiving the message in operation 2410. The sequence number can be, for example, as described above regarding... Figure 2B and / or Figure 2CThe sequence number can be assigned as described in any of the descriptions. For example, the sequence number can be assigned to define the order of operations initiated by various collaborative module instances or devices participating in the collaborative session.
[0241] In operation 2420, a first notification indicative of the operation and an assigned sequence number is transmitted to each of the plurality of devices. For example, the first notification may include the information above regarding... Figure 2A And one or more of the fields discussed in message 200. As discussed above, the execution device may notify each of the collaborative participants (device and / or collaborative module instances) via a single broadcast message, multiple multicast messages, or an individual unicast message transmitted separately to each participant. Since the first notification may include... Figure 2A The notification can also notify each of the plurality of devices of the current maximum sequence number (e.g., reference sequence number) of the operation that has been confirmed by all devices.
[0242] The first notification can be generated to further indicate the version associated with the operation. Additionally, any information related to the operation can also be provided in the first notification. For example, the first notification may include values from any field of message 200 that are consistent with the values included in the message received in operation 2410. (Refer to the above) Figure 2D As described, in some cases, the version number associated with a first operation received from a collaborating participant can be adjusted based on other locally initiated operations at that participant. Specifically, the number of operations without assigned sequence numbers at the collaborating participant when the first operation is executed is related to how the version number is adjusted. The version number is adjusted to account for these operations (with pending sequence number assignments). As discussed above, if the first operation is associated with a first version, but the initiating participant has N other pending operations before initiating the first operation, the first version can be adjusted before initiating the first operation at the collaborating participant to reflect that these N pending operations have been applied to the distributed data structure. In some aspects, the first version can be adjusted to first version + N by using higher sequence numbers to represent higher-order operations and versions. In some aspects, N can be determined by first determining the number of operations initiated by the participant from the version of the distributed data structure specified in the message of operation 2410 (excluding operations in the message of 2410). N can be set to this number of operations. This is discussed above regarding... Figure 2D It has been described.
[0243] Operation 2425 determines that each of the multiple devices (the participants in the collaboration are device and / or collaboration module instances) has acknowledged the operation. Operation 2425 can make this determination by monitoring messages received from each of the multiple devices. Specifically, it can monitor the messages mentioned above regarding... Figure 2A The value of the maximum sequence number field 214 of message 200 under discussion. When the value of this field in a specific message from a specific device (or collaboration module instance) indicates that the value is equal to or higher than the sequence number assigned in operation 2415 in the order of operations, operation 2425 determines that the specific device has acknowledged the operation. Operation 2425 then monitors the values received from each of the multiple devices (or collaboration module instances). When all devices indicate the same value, the determination of operation 2425 is complete.
[0244] In operation 2430, a second notification is transmitted to each of the plurality of devices. The second notification indicates the determination of operation 2425. The second notification may include the information above regarding... Figure 2A The message in message 200 refers to one or more fields among the fields discussed. In some aspects, by... Figure 2A The value of field 214 in the message is set to a value that is at least as high as the assigned sequence number in the operation order to indicate the determination.
[0245] Although process 2400 is described above as assigning a single sequence number to a single operation, those skilled in the art will understand that process 2400 can operate iteratively to assign multiple sequence numbers to multiple operations. Operations initiated by various collaborating participants (devices and / or collaborating module instances) can be received in an interleaved manner, and process 2400 can assign sequence numbers to these interleaved operations in an order consistent with the order in which the operations were received.
[0246] Process 2400 may also include generating a snapshot of the distributed data structure. (See above regarding...) Figure 2E As discussed, snapshots can merge the results of all operations on a distributed data structure up to a specific sequence number. Once those results have been applied to the distributed data structure, the snapshot represents the value of that distributed data structure. When a new participant joins the collaboration, the snapshot can be provided to that new participant. Alternatively, any operation assigned a sequence number but not included in the snapshot can also be provided to the new participant. This allows the new participant to apply those operations to the snapshot and synchronize its local version of the distributed data structure with the local versions of the other participants in the collaboration. After operation 2430, process 2400 moves to end operation 2435.
[0247] As described herein, examples may include logic or multiple components, modules, or mechanisms, or on which operations can be performed. A module is a tangible entity (e.g., hardware) capable of performing the specified operations and configured or arranged in a particular manner. In the examples, circuitry may be arranged as a module in the manner referred to (e.g., internally or relative to an external entity such as other circuitry). In the examples, all or part of one or more computer systems (e.g., standalone, client, or server computer systems) or one or more hardware processors may be configured as modules by firmware or software (e.g., instructions, application portions, or applications) to operate and perform the specified operations. In the examples, the software may reside on a machine-readable medium. In the examples, when the software is executed by the underlying hardware of the module, it causes the hardware to perform the specified operations.
[0248] Example 1 is a method performed by a first device, comprising: receiving a notification of a second device operation on a distributed data structure, the second device operation indicating a location within the distributed data structure associated with the second device operation; identifying a first node of a tree, the first node representing a portion of the distributed data structure; first determining a minimum length of the represented portion based on the first node; identifying a second device-specific portion length of the represented portion based on the first node; second determining a size of the represented portion based on the minimum length and the portion length; third determining, based on the size, that the location within the distributed data structure is contained within the represented portion; identifying a leaf node below the first node in the tree based on the third determination; and indicating a second device operation on the distributed data structure based on the identified leaf node.
[0249] In Example 2, the subject of Example 1 may optionally include: receiving a notification of a third device operation on the distributed data structure, the third device operation indicating a second location within the distributed data structure for the third device operation; identifying a second portion length of the represented portion, specific to the third device; fourth determining a second size of the represented portion, specific to the third device, based on the minimum length and the second portion length; fifth determining, based on the size, that the second location within the distributed data structure is not included in the represented portion; identifying a second leaf node based on the fifth determination, representing the second portion of the distributed data structure that includes the second location; and indicating the third device operation on the distributed data structure based on the identified second leaf node.
[0250] In Example 3, the subject of Example 2 may optionally include: identifying the sibling node of the first node in response to the fifth determination, wherein the second leaf node is a child node of the sibling node.
[0251] In Example 4, the subject matter of any one or more of Examples 1-3 may optionally include: wherein the notification further indicates a version of the distributed data structure associated with the operation of the second device, wherein the determination of the portion length is also based on the version.
[0252] In Example 5, the subject matter of any one or more of Examples 1-4 may optionally include: wherein the second device operation includes splitting the identified leaf node into two or more leaf nodes, each of the two or more leaf nodes representing a unique portion of the distributed data structure.
[0253] In Example 6, the subject of Example 5 may optionally include: assigning the unique portion to each of the two or more leaf nodes based on the location of the position relative to the represented portion.
[0254] In Example 7, the subject matter of any one or more of Examples 4-6 may optionally include: updating the partial length information for the second device at each node in the tree between the identified leaf nodes and the root of the tree, based on the representation.
[0255] In Example 8, the subject of Example 7 may optionally include: wherein the update of the partial length information includes: associating the updated partial length with the version of the distributed data structure.
[0256] In Example 9, the subject of Example 8 may optionally include: wherein updating the partial length information includes: attaching the association between the version and the updated length of the portion represented by the modified leaf node in the distributed data structure to the partial length information at the updated node.
[0257] In Example 10, the subject matter of any one or more of Examples 1-9 may optionally include: identifying a specific device operation; identifying a specific version of the distributed data structure associated with the specific device operation; representing the specific device operation in the tree by modifying the leaf nodes of the tree; and updating partial length information for the specific device at each node in the tree between the leaf nodes and the root node, wherein the update of the partial length information at each node is based on the difference between the minimum length of the portion represented by the node and the length of the portion represented at the specific device.
[0258] In Example 11, the subject of Example 10 may optionally include: determining a first set of leaf nodes below a specific node in the tree, representing an operation with an assigned sequence number greater than that of the specific version; determining a second set of leaf nodes below the specific node, representing an operation initiated by the specific device with an operation sequence number greater than that of the specific version; and updating device-specific partial length information at the specific node based on the counts of the first set of leaf nodes and the second set of leaf nodes.
[0259] Example 12 is a non-transitory computer-readable storage medium including instructions that, when executed, configure hardware processing circuitry to perform operations including: receiving notification of a second device operation on a distributed data structure, the second device operation indicating a location within the distributed data structure associated with the second device operation; identifying a first node of a tree, the first node representing a portion of the distributed data structure; first determining a minimum length of the represented portion based on the first node; identifying a second device-specific portion length of the represented portion based on the first node; second determining a size of the represented portion based on the minimum length and the portion length; third determining, based on the size, that the location within the distributed data structure is contained within the represented portion; identifying a leaf node below the first node in the tree based on the third determination; and indicating a second device operation on the distributed data structure based on the identified leaf node.
[0260] In Example 13, the subject of Example 12 may optionally include: receiving a notification of a third device operation on the distributed data structure, the third device operation indicating a second location within the distributed data structure for the third device operation; identifying a second portion length of the represented portion, specific to the third device; fourth determining a second size of the represented portion, specific to the third device, based on the minimum length and the second portion length; fifth determining, based on the size, that the second location within the distributed data structure is not included in the represented portion; identifying a second leaf node based on the fifth determination, the second leaf node representing a second portion of the distributed data structure that includes the second location; and representing the third device operation on the distributed data structure based on the identified second leaf node.
[0261] In Example 14, the subject of Example 13 may optionally include: identifying the sibling node of the first node in response to the fifth determination, wherein the second leaf node is a child node of the sibling node.
[0262] In Example 15, the subject matter of any one or more of Examples 12-14 may optionally include: wherein the notification further indicates a version of the distributed data structure associated with the operation of the second device, wherein the determination of the portion length is also based on the version.
[0263] In Example 16, the subject matter of any one or more of Examples 12-15 may optionally include: wherein the second device operation includes splitting the identified leaf node into two or more leaf nodes, each of the two or more leaf nodes representing a unique portion of the distributed data structure.
[0264] In Example 17, the subject of Example 16 may optionally include: assigning the unique portion to each of the two or more leaf nodes based on the location of the position relative to the represented portion.
[0265] In Example 18, the subject matter of any one or more of Examples 12-17 may optionally include: updating the partial length information for the second device at each node in the tree between the identified leaf nodes and the root of the tree, based on the representation.
[0266] In Example 19, the subject of Example 18 may optionally include: wherein the update of the partial length information includes: associating the updated partial length with the version of the distributed data structure.
[0267] In Example 20, the subject of Example 19 may optionally include: wherein updating the partial length information includes: attaching the association between the version and the updated length of the portion represented by the modified leaf node in the distributed data structure to the partial length information at the updated node.
[0268] In Example 21, the subject matter of any one or more of Examples 12-20 may optionally include: identifying a specific device operation; identifying a specific version of the distributed data structure associated with the specific device operation; representing the specific device operation in the tree by modifying the leaf nodes of the tree; and updating partial length information for the specific device at each node in the tree between the leaf nodes and the root node, wherein the update of the partial length information at each node is based on the difference between the minimum length of the portion represented by the node and the length of the portion represented at the specific device.
[0269] In Example 22, the subject of Example 21 may optionally include: determining a first set of leaf nodes below a specific node in the tree, representing an operation with an assigned sequence number greater than that of the specific version; determining a second set of leaf nodes below the specific node, representing an operation initiated by the specific device with an operation sequence number greater than that of the specific version; and updating device-specific partial length information at the specific node based on the counts of the first set of leaf nodes and the second set of leaf nodes.
[0270] Example 23 is an apparatus comprising: unit for receiving notification of a second device operation on a distributed data structure, the second device operation indicating a location within the distributed data structure associated with the second device operation; unit for identifying a first node of a tree, the first node representing a portion of the distributed data structure; unit for first determining a minimum length of the represented portion based on the first node; unit for identifying a second device-specific portion length of the represented portion based on the first node; unit for second determining a size of the represented portion based on the minimum length and the portion length; unit for third determining, based on the size, that the location within the distributed data structure is included in the represented portion; unit for identifying a leaf node below the first node in the tree based on the third determination; and unit for indicating a second device operation on the distributed data structure based on the identified leaf node.
[0271] In Example 24, the subject matter of Example 23 may optionally include: a unit for receiving notification of a third device operation on the distributed data structure, the third device operation indicating a second location within the distributed data structure for the third device operation; a unit for identifying a second portion length specific to the third device of the represented portion; a unit for fourth determining a second size of the represented portion specific to the third device based on the minimum length and the second portion length; a unit for fifth determining, based on the size, that the second location within the distributed data structure is not included in the represented portion; a unit for identifying a second leaf node in the distributed data structure that includes the second portion of the second location based on the fifth determination; and a unit for indicating the third device operation on the distributed data structure based on the identified second leaf node.
[0272] In Example 25, the subject of Example 24 may optionally include: a unit for identifying a sibling node of the first node in response to the fifth determination, wherein the second leaf node is a child node of the sibling node.
[0273] In Example 26, the subject matter of any one or more of Examples 23-25 may optionally include: wherein the notification further indicates a version of the distributed data structure associated with the operation of the second device, wherein the unit for determining the portion length is configured to further base the portion length on the version.
[0274] In Example 27, the subject matter of any one or more of Examples 23-26 may optionally include: wherein the unit for representing the operation of the second device is configured to split the identified leaf node into two or more leaf nodes, each of the two or more leaf nodes representing a unique portion of the distributed data structure.
[0275] In Example 28, the subject of Example 27 may optionally include: a unit for assigning the unique portion to each of the two or more leaf nodes based on the positioning of the location relative to the represented portion.
[0276] In Example 29, the subject matter of any one or more of Examples 23-28 may optionally include: a unit for updating, based on the representation, partial length information for the second device at each node in the tree between the identified leaf nodes and the root.
[0277] In Example 30, the subject of Example 29 may optionally include: wherein the unit for updating the partial length information is configured to associate the updated partial length with the version of the distributed data structure.
[0278] In Example 31, the subject of Example 30 may optionally include: wherein the unit for updating the partial length information is configured to: attach the association between the version and the updated length of the portion represented by the modified leaf node in the distributed data structure to the partial length information at the updated node.
[0279] In Example 32, the subject matter of any one or more of Examples 23-31 may optionally include: a unit for identifying a specific device operation; a unit for identifying a specific version of the distributed data structure associated with the specific device operation; a unit for representing the specific device operation in the tree by modifying the leaf nodes of the tree; and a unit for updating partial length information for the specific device at each node in the tree between the leaf nodes and the root node, wherein the unit for updating the partial length information at each node is configured such that the update is based on the difference between the minimum length of the portion represented by the node and the length of the portion represented at the specific device.
[0280] In Example 33, the subject matter of Example 32 may optionally include: a unit for determining a first set of leaf nodes below a specific node in the tree, representing an operation with an assigned sequence number greater than that of the specific version; a unit for determining a second set of leaf nodes below the specific node, representing an operation initiated by the specific device with an operation sequence number greater than that of the specific version; and a unit for updating the device-specific partial length information at the specific node based on the counts of the first set of leaf nodes and the second set of leaf nodes.
[0281] Example 34 is a first device comprising: hardware processing circuitry; an electronic memory storing instructions, which, when executed, configure the hardware processing circuitry to perform operations including: receiving notification of a second device operation on a distributed data structure, the second device operation indicating a location within the distributed data structure associated with the second device operation; identifying a first node of a tree, the first node representing a portion of the distributed data structure; first determining a minimum length of the represented portion based on the first node; identifying a second device-specific portion length of the represented portion based on the first node; second determining a size of the represented portion based on the minimum length and the portion length; third determining, based on the size, that the location within the distributed data structure is contained within the represented portion; identifying a leaf node below the first node in the tree based on the third determination; and indicating a second device operation on the distributed data structure based on the identified leaf node.
[0282] In Example 35, the subject of Example 34 may optionally include the following operations: receiving a notification of a third device operation on the distributed data structure, the third device operation indicating a second location within the distributed data structure for the third device operation; identifying a second portion length of the represented portion, specific to the third device; fourth determining a second size of the represented portion, specific to the third device, based on the minimum length and the second portion length; fifth determining, based on the size, that the second location within the distributed data structure is not included in the represented portion; identifying a second leaf node based on the fifth determination, the second leaf node representing a second portion of the distributed data structure that includes the second location; and indicating the third device operation on the distributed data structure based on the identified second leaf node.
[0283] In Example 36, the subject of Example 35 may optionally include the following operation: in response to the fifth determination, identifying the sibling node of the first node, wherein the second leaf node is a child node of the sibling node.
[0284] In Example 37, the subject matter of any one or more of Examples 34-36 may optionally include: wherein the notification further indicates a version of the distributed data structure associated with the operation of the second device, wherein the determination of the portion length is also based on the version.
[0285] In Example 38, the subject matter of any one or more of Examples 34-37 may optionally include: wherein the second device operation includes splitting the identified leaf node into two or more leaf nodes, each of the two or more leaf nodes representing a unique portion of the distributed data structure.
[0286] In Example 39, the subject of Example 38 may optionally include the operation of assigning the unique portion to each of the two or more leaf nodes based on the location of the position relative to the represented portion.
[0287] In Example 40, the subject matter of any one or more of Examples 34-39 may optionally include the following operation: updating the partial length information for the second device at each node in the tree between the identified leaf nodes and the root in the tree, based on the representation.
[0288] In Example 41, the subject of Example 40 may optionally include: wherein the update of the partial length information includes: associating the updated partial length with the version of the distributed data structure.
[0289] In Example 42, the subject of Example 41 may optionally include: wherein updating the partial length information includes: attaching the association between the version and the updated length of the portion represented by the modified leaf node in the distributed data structure to the partial length information at the updated node.
[0290] In Example 43, the subject matter of any one or more of Examples 34-42 may optionally include the following operations: identifying a specific device operation; identifying a specific version of the distributed data structure associated with the specific device operation; representing the specific device operation in the tree by modifying the leaf nodes of the tree; and updating partial length information for the specific device at each node in the tree between the leaf nodes and the root node, wherein the update of the partial length information at each node is based on the difference between the minimum length of the portion represented by the node and the length of the portion represented at the specific device.
[0291] In Example 44, the subject of Example 43 may optionally include the following operations: determining a first set of leaf nodes below a specific node in the tree, representing an operation with an assigned sequence number greater than that of the specific version; determining a second set of leaf nodes below the specific node, representing an operation initiated by the specific device with an operation sequence number greater than that of the specific version; and updating device-specific partial length information at the specific node based on the counts of the first set of leaf nodes and the second set of leaf nodes.
[0292] Example 45 is a method performed by a first device, comprising: joining a collaboration that provides access to a distributed data structure; receiving multiple indications of a plurality of corresponding operations of a serialized distributed data structure, each indication further indicating a corresponding originating device associated with each of the corresponding operations; representing the serialized plurality of operations and the resulting corresponding portions of the distributed data structure in leaf nodes of a tree data structure; representing, in non-leaf nodes of the tree, and based on a specific operation of the serialized plurality of operations, an originating device-specific length of a portion of the distributed data structure represented by a leaf node below the non-leaf node; receiving additional indications of additional operations of the distributed data structure, and further indications of a specific originating device associated with the additional operations, the additional indications further indicating a position within the distributed data structure associated with the additional operations; determining a leaf node of the tree associated with the position based on the originating device-specific length of the portion and the specific originating device associated with the additional operations; and modifying the determined leaf node based on the additional operations.
[0293] In Example 46, the subject of Example 45 may optionally include: receiving an indication of a specific version of the distributed data structure associated with the specific operation for the specific operation, wherein the origin device-specific length of the portion is based on the specific version.
[0294] In Example 47, the subject matter of any one or more of Examples 45-46 may optionally include: receiving an instruction for the operation of the serialized plurality of operations confirmed by all participants in the collaboration; and performing maintenance on the tree based on the instruction.
[0295] In Example 48, the subject of any one or more of Examples 45-47 may optionally include: initiating a local operation on the distributed data structure, representing the local operation in the tree; and transmitting an indication of the operation to a serialization service.
[0296] In Example 49, the subject of any one or more of Examples 45-48 may optionally include: generating a message indicating that a sequence number associated with the local operation has not been assigned, and a version of the distributed data structure associated with the local operation.
[0297] In Example 50, the subject of Example 49 may optionally include: receiving a second message indicating that a sequence number is assigned to the local operation, and updating the tree based on the assigned sequence number.
[0298] In Example 51, the subject matter of any one or more of Examples 45-50 may optionally include: receiving an indication of a second operation associated with a second originating device and a second position associated with the second operation; searching the tree based on a specific length of the second originating device of the portion; identifying a second portion of the tree representing the second position in the distributed data structure based on the result of the search; and representing the second operation in the tree based on the second portion.
[0299] In Example 52, the subject of Example 51 may optionally include: receiving an indication of a second version of the distributed data structure associated with the second operation, wherein the search is also based on the second version.
[0300] Example 53 is a non-transitory computer-readable medium including instructions that, when executed by the hardware processing circuitry of a device, configure the device to perform operations including: joining a cooperation that provides access to a distributed data structure; receiving multiple indications of a plurality of corresponding operations serialized to the distributed data structure, each indication further indicating a corresponding origin device associated with each of the corresponding operations; representing the serialized plurality of operations and the resulting corresponding portions of the distributed data structure in leaf nodes of a tree data structure; representing, in non-leaf nodes of the tree and based on a specific operation of the serialized plurality of operations, an origin device-specific length of a portion of the distributed data structure represented by a leaf node below the non-leaf node; receiving additional indications of additional operations of the distributed data structure, and further indications of a specific origin device associated with the additional operations, the additional indications further indicating a position within the distributed data structure associated with the additional operations; determining a leaf node of the tree associated with the position based on the origin device-specific length of the portion and the specific origin device associated with the additional operations; and modifying the determined leaf node based on the additional operations.
[0301] In Example 54, the subject of Example 53 may optionally include: for the specific operation, receiving an indication of a specific version of the distributed data structure associated with the specific operation, wherein the origin device-specific length of the portion is based on the specific version.
[0302] In Example 55, the subject matter of any one or more of Examples 53-54 may optionally include: receiving an instruction for an operation among the serialized plurality of operations confirmed by all participants in the collaboration; and performing maintenance on the tree based on the instruction.
[0303] In Example 56, the subject of any one or more of Examples 53-55 may optionally include: initiating a local operation on the distributed data structure, representing the local operation in the tree; and transmitting an indication of the operation to a serialization service.
[0304] In Example 57, the subject of any one or more of Examples 53-56 may optionally include: generating a message indicating that a sequence number associated with the local operation has not been assigned, and a version of the distributed data structure associated with the local operation.
[0305] In Example 58, the subject of Example 57 may optionally include: receiving a second message indicating that a sequence number has been assigned to the local operation, and updating the tree based on the assigned sequence number.
[0306] In Example 59, the subject matter of any one or more of Examples 53-58 may optionally include: receiving an indication of a second operation associated with a second originating device and a second position associated with the second operation; searching the tree based on a specific length of the second originating device of the portion; identifying a second portion of the tree representing the second position in the distributed data structure based on the result of the search; and representing the second operation in the tree based on the second portion.
[0307] In Example 60, the subject of Example 59 may optionally include: receiving an indication of a second version of the distributed data structure associated with the second operation, wherein the search is also based on the second version.
[0308] Example 61 is an apparatus comprising: a unit for joining a collaboration that provides access to a distributed data structure; a unit for receiving a plurality of indications of a plurality of corresponding operations of a serialized distributed data structure, each indication further indicating a corresponding origin device associated with each of the corresponding operations; a unit for representing, in leaf nodes of a tree data structure, the serialized plurality of operations and a resulting corresponding portion of the distributed data structure; a unit for representing, in non-leaf nodes of the tree, and based on a specific operation of the serialized plurality of operations, a specific length of the origin device of a portion of the distributed data structure represented by a leaf node below the non-leaf node; a unit for receiving additional indications of additional operations of the distributed data structure, and further indications of a specific origin device associated with the additional operations, the additional indications further indicating a position within the distributed data structure associated with the additional operations; a unit for determining a leaf node of the tree associated with the position based on the specific length of the origin device of the portion and the specific origin device associated with the additional operations; and a unit for modifying the determined leaf node based on the additional operations.
[0309] In Example 62, the subject matter of Example 61 may optionally include: a unit for receiving an indication of a specific version of the distributed data structure associated with the specific operation for the specific operation, wherein the origin device-specific length of the portion is based on the specific version.
[0310] In Example 63, the subject matter of any one or more of Examples 61-62 may optionally include: a unit for receiving an indication of an operation among the serialized plurality of operations confirmed by all participants in the collaboration; and a unit for performing maintenance on the tree based on the indication.
[0311] In Example 64, the subject matter of any one or more of Examples 61-63 may optionally include: a unit for initiating local operations on the distributed data structure, representing the local operations in the tree; and a unit for transmitting an indication of the operation to a serialization service.
[0312] In Example 65, the subject of any one or more of Examples 61-64 may optionally include a unit for: generating a message indicating that a sequence number associated with the local operation has not been assigned, and a version of the distributed data structure associated with the local operation.
[0313] In Example 66, the subject of Example 65 may optionally include: a unit for receiving a second message indicating that a sequence number has been assigned to the local operation, and a unit for updating the tree based on the assigned sequence number.
[0314] In Example 67, the subject matter of any one or more of Examples 61-66 may optionally include: a unit for receiving an indication of a second operation associated with a second originating device and a second position associated with the second operation; a unit for searching the tree based on a specific length of the second originating device of the portion; a unit for identifying a second portion of the tree representing the second position in the distributed data structure based on the result of the search; and a unit for representing the second operation in the tree based on the second portion.
[0315] In Example 68, the subject of Example 67 may optionally include: a unit for receiving an indication of a second version of the distributed data structure associated with the second operation, wherein the search is also based on the second version.
[0316] Example 69 is a device comprising: hardware processing circuitry; an electronic hardware memory storing instructions that, when executed, configure the hardware processing circuitry to perform operations including: joining a cooperation that provides access to a distributed data structure; receiving a plurality of indications of a plurality of corresponding operations of a serialized distributed data structure, each indication further indicating a corresponding origin device associated with each corresponding operation of the corresponding operations; representing the serialized plurality of operations and corresponding portions of the distributed data structure in leaf nodes of a tree data structure; representing, in non-leaf nodes of the tree, and based on a specific operation of the serialized plurality of operations, a specific length of the origin device of a portion of the distributed data structure represented by a leaf node below the non-leaf node; receiving additional indications of additional operations of the distributed data structure, and further indications of a specific origin device associated with the additional operations, the additional indications further indicating a position within the distributed data structure associated with the additional operations; determining a leaf node of the tree associated with the position based on the specific length of the origin device of the portion and the specific origin device associated with the additional operations; and modifying the determined leaf node based on the additional operations.
[0317] In Example 70, the subject of Example 69 may optionally include: receiving an indication of a specific version of the distributed data structure associated with the specific operation for the specific operation, wherein the origin device-specific length of the portion is based on the specific version.
[0318] In Example 71, the subject matter of any one or more of Examples 69-70 may optionally include: receiving an instruction for an operation among the serialized plurality of operations confirmed by all participants in the collaboration; and performing maintenance on the tree based on the instruction.
[0319] In Example 72, the subject matter of any one or more of Examples 69-71 may optionally include: initiating a local operation on a distributed data structure, representing the local operation in the tree; and transmitting an indication of the operation to a serialization service.
[0320] In Example 73, the subject matter of any one or more of Examples 69-72 may optionally include: generating a message indicating that a sequence number associated with the local operation has not been assigned, and a version of the distributed data structure associated with the local operation.
[0321] In Example 74, the subject of Example 73 may optionally include: receiving a second message indicating that a sequence number has been assigned to the local operation, and updating the tree based on the assigned sequence number.
[0322] In Example 75, the subject matter of any one or more of Examples 69-74 may optionally include: receiving an indication of a second operation associated with a second originating device and a second position associated with the second operation; searching the tree based on a specific length of the second originating device of the portion; identifying a second portion of the tree representing the second position in the distributed data structure based on the result of the search; and representing the second operation in the tree based on the second portion.
[0323] In Example 76, the subject of Example 75 may optionally include: receiving an indication of a second version of the distributed data structure associated with the second operation, wherein the search is also based on the second version.
[0324] Therefore, the term "module" is understood to encompass tangible entities that are physically constructed, specifically configured (e.g., hardwired), or temporarily (e.g., provisionally) configured (e.g., programmed) to operate in a particular manner or perform any of the operations described herein. Consider the example of temporarily configured modules, where it is not necessary to instantiate each module at any given time. For example, in the case where modules include general-purpose hardware processors configured using software, the general-purpose hardware processors can be configured as distinct modules at different times. The software can accordingly configure the hardware processors, for example, to constitute a particular module at one time and different modules at different time instances.
[0325] Various embodiments can be implemented entirely or partially in software and / or firmware. This software and / or firmware may take the form of instructions contained in or on a non-transitory computer-readable storage medium. Those instructions can then be read and executed by one or more processors to perform the operations described herein. The instructions can be in any suitable form, such as, but not limited to: source code, compiled code, interpreted code, executable code, static code, dynamic code, etc. Such computer-readable medium may include any tangible non-transitory medium for storing information in a form readable by one or more computers, such as, but not limited to, read-only memory (ROM); random access memory (RAM); disk storage media; optical storage media; flash memory; etc.
Claims
1. The first device, comprising: Hardware processing circuitry; An electronic memory that stores instructions, which, when executed, configure the hardware processing circuitry to perform operations including the following: Receive a notification of a second device operation on a distributed data structure, the notification of the second device operation indicating the location within the distributed data structure associated with the second device operation, the sequence number assigned by the synchronization service, and the version of the distributed data structure on which the second device operation takes effect; The first node of the identifier tree, where the first node represents a part of the distributed data structure; Based on the first node, the minimum length of the represented portion is determined first, the minimum length representing the length of the synchronization data in the distributed data structure, the synchronization data coming from the operations of multiple participants that have taken effect on the distributed data structure in the order of their assigned sequence numbers; The partial length of the represented portion, which is specific to the version of the distributed data structure at the second device, is identified based on the first node. The size of the represented portion is determined secondly based on the minimum length and the partial length; Based on the size, it is thirdly determined that the position within the distributed data structure is included in the represented portion; The leaf node below the first node in the tree is identified based on the third determination; as well as The second device operation on the distributed data structure is represented by the identified leaf nodes.
2. The first device according to claim 1, further comprising: Receive notification of a third device operation on the distributed data structure, wherein the third device operation indicates a second location within the distributed data structure for the third device operation; The length of the second portion, specific to the third device, that identifies the part being identified; A fourth, specific second dimension of the represented portion, relating to the third device, is determined based on the minimum length and the second portion length. Based on the size, it is determined that the second position within the distributed data structure is not included in the represented portion; The second leaf node is identified based on the fifth determination, and the second leaf node represents the second part of the distributed data structure that includes the second position; as well as The operation of the third device on the distributed data structure is represented by the identified second leaf node.
3. The first device according to claim 2, further comprising: In response to the fifth determination, the sibling node of the first node is identified, wherein the second leaf node is a child node of the sibling node.
4. The first device according to claim 1, wherein, The second device operation includes splitting the identified leaf node into two or more leaf nodes, each of the two or more leaf nodes representing a unique part of the distributed data structure.
5. The first device according to claim 4, further comprising: Based on the location of the position relative to the represented portion, the unique portion is assigned to each of the two or more leaf nodes.
6. The first device according to claim 5, further comprising: The partial length information for the second device at each node in the tree, between the identified leaf node and the root, is updated based on the representation.
7. The first device according to claim 6, wherein, The update of the partial length information includes associating the updated partial length with the version of the distributed data structure.
8. The first device according to claim 7, wherein, Updating the partial length information includes attaching the association between the version and the updated length of the portion represented by the updated node in the distributed data structure to the partial length information at the updated node.
9. The first device according to claim 1, further comprising: Identify specific device operations; Identify a specific version of the distributed data structure associated with the operation of the specific device; The specific device operation in the tree is represented by modifying the leaf nodes of the tree; as well as Update the partial length information for a specific device at each node in the tree between the leaf node and the root node, wherein the update of the partial length information at each node is based on the difference between the minimum length of the portion represented by the node and the length of the portion represented at the specific device.
10. The first device according to claim 9, wherein the operation further comprises: Identify the first group of leaf nodes below a specific node in the tree, where the first group of leaf nodes represents an operation with a sequence number greater than the assigned number of the specific version; Determine the second group of leaf nodes below the specific node, where the second group of leaf nodes represents an operation initiated by the specific device and having an operation sequence number greater than the specific version; as well as The partial length information at the specific node specific to the specific device is updated based on the counts of the first group of leaf nodes and the second group of leaf nodes.
11. A computer-readable storage medium including instructions that, when executed, configure hardware processing circuitry of a device to perform operations including: Receive a notification of a second device operation on a distributed data structure, the notification of the second device operation indicating the location within the distributed data structure associated with the second device operation, the sequence number assigned by the synchronization service, and the version of the distributed data structure on which the second device operation takes effect; The first node of the identifier tree, where the first node represents a part of the distributed data structure; Based on the first node, the minimum length of the represented portion is determined first, the minimum length representing the length of the synchronization data in the distributed data structure, the synchronization data coming from the operations of multiple participants that have taken effect on the distributed data structure in the order of their assigned sequence numbers; The partial length of the represented portion, which is specific to the version of the distributed data structure at the second device, is identified based on the first node. The size of the represented portion is determined secondly based on the minimum length and the partial length; Based on the size, it is thirdly determined that the position within the distributed data structure is included in the represented portion; The leaf node below the first node in the tree is identified based on the third determination; as well as The second device operation on the distributed data structure is represented by the identified leaf nodes.
12. The computer-readable storage medium of claim 11, further comprising: Receive a notification of a third device operation on the distributed data structure, the notification indicating a second location within the distributed data structure for the third device operation; The length of the second portion, specific to the third device, that identifies the part being identified; A fourth, specific second dimension of the represented portion, relating to the third device, is determined based on the minimum length and the second portion length. Based on the size, it is determined that the second position within the distributed data structure is not included in the represented portion; The second leaf node is identified based on the fifth determination, and the second leaf node represents the second part of the distributed data structure that includes the second position; as well as The operation of the third device on the distributed data structure is represented by the identified second leaf node.
13. The computer-readable storage medium of claim 12, further comprising: In response to the fifth determination, the sibling node of the first node is identified, wherein the second leaf node is a child node of the sibling node.
Citation Information
Patent Citations
Web-based collaborative document review system
US20130283147A1