Keyword rights distribution

CN116057529BActive Publication Date: 2026-09-25SALESFORCE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180058834.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-23
Filing Date
2021-09-08
Publication Date
2026-09-25
Estimated Expiration
2041-09-08

AI Technical Summary

Benefits of technology

[0004]现代数据库系统通常实现使用户能够以有组织的方式存储能被有效访问和操纵的信息集合体的管理系统。在一些情况下,这些管理系统对具有多个层级的日志结构合并树(Log-structured Merge Tree,LSM树)进行维护,每个层级将数据库记录中的信息存储为键值对。LSM树通常包括两种高层级组件:内存缓存和持久性存储。在操作期间,数据库系统接收事务请求以处理包括将数据库记录写入持久性存储的事务。数据库系统首先将数据库记录写入到内存缓存,然后将它们刷新到持久性存储。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116057529B_ABST
    Figure CN116057529B_ABST
Patent Text Reader

Abstract

This application discloses techniques related to distribution of database key privileges. A database system can distribute first privilege information to a plurality of database nodes, the first privilege information identifying distribution of key range privileges to some of the plurality of database nodes. A given key range privilege distributed to a given database node can allow the database node to write to records whose keys fall within a key range associated with the given key range privilege. The database system can receive, from a first database node, a request for a first key range privilege provided to a second database node. The database system can modify the first privilege information to derive second privilege information that provides the first key range privilege to the first database node instead of the second database node. The database system can distribute the second privilege information to some of the plurality of database nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology Technical Field

[0002] The publicly available text generally concerns database systems, and more specifically, the distribution of database keyword permissions between database nodes.

[0003] Related technical descriptions

[0004] Modern database systems typically implement management systems that enable users to store and efficiently access and manipulate collections of information in an organized manner. In some cases, these management systems maintain a Log-structured Merge Tree (LSM tree) with multiple levels, each level storing information from database records as key-value pairs. An LSM tree typically includes two high-level components: an in-memory cache and persistent storage. During operation, the database system receives transaction requests to process transactions that include writing database records to persistent storage. The database system first writes the database records to the in-memory cache and then flushes them to persistent storage. Attached Figure Description

[0005] Figure 1 This is a block diagram illustrating exemplary elements of a system including a permission orchestrator node and worker nodes according to some embodiments.

[0006] Figure 2 This is a block diagram illustrating exemplary elements of a worker node that relinquishes keyword permissions according to some embodiments.

[0007] Figure 3 This is a block diagram illustrating exemplary elements of a permissions orchestrator node according to some embodiments.

[0008] Figure 4 This is a block diagram illustrating exemplary elements of a worker node that obtains keyword permissions according to some embodiments.

[0009] Figures 5 to 6 This is a flowchart illustrating an exemplary method related to the distribution of keyword permissions according to some embodiments.

[0010] Figure 7 This is a block diagram illustrating exemplary elements of a multitenant system according to some embodiments.

[0011] Figure 8 This is a block diagram illustrating the elements of a computer system according to some embodiments. Detailed Implementation

[0012] As described above, modern database systems can maintain an LSM tree that includes database records with key-value pairs. In many cases, the database system comprises a single active database node (and several standby database nodes) responsible for writing database records to the LSM tree in a persistent storage component. However, in some cases, the database system comprises multiple active database nodes writing database records to the LSM tree. Multiple active database nodes may share a common persistent storage but have their own separate memory caches. In such an implementation, database records written by an active database node to its separate memory cache are not visible to other active database nodes until those records are flushed to the common persistent storage.

[0013] In implementations with multiple active database nodes, there is a possibility that two or more database nodes may write records for the same database key at approximately the same time if the database nodes are not restricted in some way. (This situation is rarely considered in implementations using a single active database node, as there is a local lock manager to ensure that changes to local transactions are coordinated and do not conflict.) Therefore, in some multi-active-node implementations, one of the active database nodes may be granted database key permissions for a specific database key, which allows the database node to write records for that database key. However, other active database nodes are not granted database key permissions and therefore cannot write records for that specific database key. But in some cases, it may be desirable to redistribute database key permissions from the owner active database node to another active database node. For example, if the owner database node is overwhelmed with other database work and there are pending transactions involving writing records for a specific database key, it may be desirable to offload those pending transactions to another database node. Among other things, this public text addresses the technical problem of ensuring that multiple database nodes do not write records for the same key at relatively equal times, while also allowing the redistribution of database key permissions to meet the needs of the database system.

[0014] More specifically, the disclosed text describes various techniques for orchestrating the distribution of database keyword permissions among multiple database nodes in a database system. In the various embodiments described below, the database system includes a permission orchestrator node that orchestrates the distribution of keyword permissions among worker nodes of a database capable of writing database records to the database system. During database system operation, a first worker node may receive a request to perform a transaction involving writing a record for a specific database keyword (e.g., database keyword "XYZ"). In various embodiments, the first worker node first accesses permission information from the permission orchestrator node (if the first worker node has not previously obtained such permission information) and determines whether it has permission to write a database record for the specific database keyword. In some cases, the permission information may indicate that the first worker node has already been granted the relevant keyword permission, and therefore the first worker node can write the record. The first worker node may also ensure that no other record for the specific database keyword has been committed within a certain period of time before it writes the record for the transaction. In other cases, the permission information may indicate that the relevant keyword permission has not yet been granted or has already been granted to another worker node. Subsequently, the first worker node makes a request to the permission orchestrator node for the relevant keyword permission. In some cases, the first worker node can request keyword-scope permissions that include the relevant database keyword permissions (e.g., the keyword scope "XY" that includes the database keyword "XYZ").

[0015] In various embodiments, upon receiving a request from a given worker node, the permission orchestrator node first determines whether the relevant keyword permission is held by another worker node. If the relevant keyword permission is not held, the permission orchestrator node can generate updated permission information that provides the relevant keyword permission to the first worker node. The permission orchestrator node can then distribute the permission information to the worker nodes of the database system. In some embodiments, the permission orchestrator node can provide a keyword permission range that includes the relevant keyword permission (this may be referred to as a "keyword range permission"). As an example, the permission orchestrator node can provide keyword permissions for the keyword range "XY" (which includes any keyword starting with "XY", including the specific database keyword "XYZ"). However, in some embodiments, if a second worker node holds the relevant keyword permission, the permission orchestrator node sends a request to the second worker node to relinquish the relevant keyword permission.

[0016] In various embodiments, upon receiving an abandon request, the second worker node determines whether any active transaction at the second worker node has locked the relevant keyword permissions for a record commit. If the active transaction intends to commit a record for a database keyword corresponding to the keyword permissions, the transaction can lock the keyword permissions. Locked keyword permissions cannot be used by another active transaction, nor can they be returned to the permissions orchestrator node while the keyword permissions are locked / held. If the relevant keyword permissions are held by an active transaction running on the second worker node, the second worker node can wait until the transaction is complete before returning the keyword permissions. In some embodiments, the second worker node can abandon a set of keyword permissions that includes the relevant keyword permissions. In some embodiments, the second worker node also returns historical information identifying the transaction commit number (“XCN”) associated with the abandoned keyword permissions. The XCN can indicate when the latest database record was committed for the abandoned keyword permissions (or the set of keyword permissions that includes the abandoned keyword permissions). Imagine an instance where the latest committed database record for the keyword range “XY” is committed with XCN “501”. Therefore, the second worker node can return historical information indicating that the maximum XCN for the revoked keyword permission "XYZ" is "501". In various embodiments, after ensuring that the relevant keyword permission has been revoked by the second worker node, the permission orchestrator node provides the relevant keyword permission to the first worker node by updating the permission information and distributing it to the worker nodes. The updated permission information may include historical information returned by the second worker node, allowing other worker nodes to determine whether they can commit records for a specific database keyword.

[0017] Upon receiving relevant keyword permissions, the first worker node can use the permission information to determine whether a record has been committed for the relevant keyword permissions during a certain time period—for example, whether the second worker node committed a record for the relevant keyword permissions when it possessed them. If it appears that a record may have been committed, the first worker node can communicate with the second worker node to determine whether the record has actually been committed. If the record has been committed, the first worker node can abort its transaction to ensure that the two transactions do not conflict; otherwise, the first worker node writes the record for the relevant keyword permissions to its memory cache. In various embodiments, when relinquishing keyword permissions (e.g., in response to receiving a relinquishment request), the first worker node notifies the permission orchestrator node about the record write and commit, making it known to other worker nodes when keyword-related permissions are acquired. In some embodiments, relevant keyword permissions identify record commits. Therefore, keyword permissions can simultaneously be a locking and parallel protection mechanism, as well as an enumeration of the most recently updated position.

[0018] In this way, database keyword permissions can be distributed and redistributed among worker nodes via the permission orchestrator node. This can lead to a "contest" scenario where worker nodes "compete" for database keyword permissions in order to write database records for system transactions. That is, a worker node can acquire specific keyword permissions, write a record with those specific permissions, relinquish those permissions, and then later acquire them again for another record write. Therefore, specific keyword permissions can be "pulled back and forth" between worker nodes attempting to write records using those permissions.

[0019] These techniques can be advantageous because they allow database systems to be implemented with multiple active database nodes while ensuring that multiple database nodes do not write records to the same database key at nearly the same time, thus violating transaction isolation and potentially introducing anomalies visible to the application. That is, by assigning key permissions to at most one database node at a time, other database nodes are prevented from writing database records with the key associated with that key permission. These techniques can be further advantageous because they allow key permissions to be moved between active database nodes, so that one active database node does not have to handle all the work associated with the database key. This can make transaction processing faster compared to existing database system implementations. References will now be made to... Figure 1 Let's begin by discussing exemplary applications of these technologies.

[0020] Now go to Figure 1 A block diagram of system 100 is shown. System 100 includes a collection of components that can be implemented by hardware or a combination of hardware and software routines. In the illustrated embodiment, system 100 includes a database 110 having an LSM file 115, worker nodes 120A and 120B, and a permission orchestrator node 150. As further shown, worker node 120 and permission orchestrator node 150 include permission information 140, which defines keyword range permissions 145A and 145B. Also as shown, worker node 120 includes a corresponding memory cache 130 storing a database record 132 associated with keyword 134. In some embodiments, system 100 is implemented in a different manner than that shown. For example, permission orchestrator node 150 may also implement the functionality of worker node 120 in addition to orchestrating the distribution of keyword range permissions 145. Furthermore, while techniques for public text are discussed in conjunction with LSM trees, these techniques can be applied to other types of database implementations in which multiple nodes are writing and committing records to the database.

[0021] In various embodiments, system 100 implements platform services (e.g., Customer Relationship Management (CRM) platform services) that allow users of the service to develop, run, and manage applications. System 100 may be a multi-tenant system providing various functionalities to multiple users / tenants hosted by a multi-tenant system. Therefore, system 100 can execute software routines from various different users (e.g., providers and tenants of system 100), and provide code, web pages, and other data to users, databases, and other entities associated with system 100. For example, as shown, system 100 includes worker nodes 120 that can store, manipulate, and retrieve data from LSM files 115 of database 110 on behalf of users of system 100.

[0022] In various embodiments, database 110 is a collection of information organized in a manner that allows access, storage, and manipulation of information. Therefore, database 110 may include supporting software that allows worker nodes 120 to manipulate (e.g., access, store, etc.) the information stored in database 110. In some embodiments, database 110 is implemented by a single storage device or multiple storage devices connected together on a network (e.g., a Storage Area Network (SAN)) and configured to redundantly store information to prevent data loss. Since storage devices can persistently store data, database 110 can be used as persistent storage. In various embodiments, database records 132 written by one worker node 120 to LSM file 115 can be accessed by other worker nodes 120. LSM file 115 may be stored as part of a Log Structured Merge Tree (LSM tree) implemented at database 110.

[0023] In various embodiments, the LSM tree is a data structure that stores LSM files 115 in an organized manner using a hierarchical scheme. The LSM tree may include two high-level components: an in-memory component implemented at memory cache 130 and an on-disk component implemented at database 110. In some embodiments, memory cache 130 is considered separate from the LSM tree. In various embodiments, worker nodes 120 first write database records 132 to their memory cache 130. When cache 130 is full and / or at a specific point in time, worker nodes 120 may flush their database records 132 to database 110. As part of flushing database records 132, in various embodiments, worker nodes 120 write database records 132 to a new set of LSM files 115 at database 110.

[0024] In various embodiments, LSM file 115 is a collection of database records 132. Database records 132 may be key-value pairs containing data and corresponding database keys 134 that can be used to look up database records. For example, database records 132 may correspond to data rows in a database table, where database records 132 specify values ​​for one or more attributes associated with the database table. In various embodiments, file 115 is associated with one or more database key ranges defined by the keys 134 of the database records 132 included in file 115. Imagine that file 115 stores instances of three database records 132 associated with the keys 134 “XYA”, “XYW”, and “XYZ”, respectively. The three keys 134 span the database key range XYA to XYZ, therefore LSM file 115 can be associated with this database key range.

[0025] In various embodiments, worker node 120 is hardware, software, or a combination thereof capable of providing database services (e.g., data storage, data retrieval, and / or data manipulation). These database services can be provided to other components within system 100 and / or components outside system 100. For example, worker node 120A can receive a database transaction request to perform one or more database tasks—this request may be received from an application server attempting to access a set of database records 132. The database transaction request may specify an SQL SELECT command to select one or more rows from one or more database tables. The content of a row can be qualified in database record 132, so worker node 120A can return one or more database records 132 corresponding to the selected one or more table rows. In various cases, the database transaction request may instruct database node 120 to write one or more database records 132 to an LSM tree. In various embodiments, worker node 120 first writes those database records to its memory cache 130 before flushing the database records 132 to database 110.

[0026] In various embodiments, memory cache 130 is a buffer in the worker node 120's memory (e.g., random access memory) for storing data. HBase TM Storage (HBase) TMMemstore is an instance of memory cache 130. As described above, worker node 120 can first write database record 132 to its memory cache 130. In some cases, the latest / latest version of a row in a database table can be found in database record 132 stored in memory cache 130. However, in various embodiments, the database record 132 written to worker node 120's memory cache 130 is not visible to other worker nodes 120. That is, other worker nodes 120 do not know what information is stored in worker node 120's memory cache 130 without being asked. Since one worker node 120 may not know the database record 132 written by another worker node 120, in various embodiments, to prevent database record conflicts, worker node 120 is provided with permission information 140 that controls which records 132 can be written by a given worker node 120. Therefore, permission information 140 can prevent two or more worker nodes 120 from writing database record 132 for the same database key 134 within a specific time interval, thereby preventing worker node 120 from flushing conflicting database record 132 to database 110.

[0027] In various embodiments, permission information 140 is information that identifies keyword range permissions 145 and the set of owners corresponding to those keyword range permissions. For example, as shown, worker node 120A is granted keyword range permission 145A (shown in a solid box at worker node 120A), while worker node 120B is granted keyword range permission 145B (shown in a solid box at worker node 120B). In various embodiments, keyword range permission 145 indicates the range of keyword permissions corresponding to the range of database keywords 134 for which the worker node 120 is permitted to write database record 132. As an example, keyword range permission 145A could indicate keyword permission for keyword 134A. Thus, as shown, worker node 120A could write database record 132A to its memory cache 130. In various embodiments, keyword permissions are granted to at most one worker node 120 at any given time. Consider an instance where keyword range permission 145B indicates keyword permission for keyword 134B. Although worker node 120B is granted permission for the keyword, worker node 120A is not allowed to write to record 132 containing keyword 134B. In various embodiments, to allow writing to record 132 for a specific keyword 134, worker node 120 may issue a permission request 112 to permission orchestrator node 150 specifying the specific keyword 134. In various cases, permission request 112 may specify multiple keywords 134 (e.g., a keyword range).

[0028] In various embodiments, the permission orchestrator node 150 facilitates the distribution of keyword permissions among worker nodes 120 and ensures that at most one worker node 120 has ownership of the keyword permissions. In various embodiments, as part of facilitating the distribution of keyword permissions, the permission orchestrator node 150 may update permission information 140 in response to receiving permission request 112. As shown, the permission orchestrator node 150 receives permission request 112 from worker node 120B. In response to receiving permission request 112, the permission orchestrator node 150 may first determine whether keyword permissions for the requested database keyword 134 have already been granted. If not, the permission orchestrator node 150 may update permission information 140 to grant keyword permissions to worker node 120B, and may notify worker node 120B of this update via permission response 114, and notify other worker nodes 120 via permission information indication 156. In various embodiments, all worker nodes 120, including worker node 120B, are notified via permission information indication 156.

[0029] If keyword permissions have already been provided (or keyword range permissions 145 if requested), the permissions orchestrator node 150 can identify its associated worker node 120 and issue a relinquishment request 152 to worker node 120. As an example, worker node 120B can issue a permissions request 112 for keyword permissions associated with keyword 134A. The permissions orchestrator node 150 can determine that the keyword permissions are part of keyword range permissions 145A that have already been provided to worker node 120A. Therefore, the permissions orchestrator node 150 can issue a relinquishment request 152 to worker node 120A. In various embodiments, when relinquishing keyword permissions, worker node 120 ensures that the keyword permissions to be relinquished are not in use. Worker node 120 can then relinquish the keyword permissions and notify permissions orchestrator node 150 via a relinquishment response 154. In some cases, worker node 120 may include historical information in the abandon response 154, which indicates whether record 132 may have been committed using the abandoned keyword permission (or in some cases using the abandoned keyword range permission 145).

[0030] In response to receiving the abandonment response 154, the permission orchestrator node 150 can update the permission information 140 to provide keyword permissions to worker node 120B, and can notify worker node 120B of this update via permission response 114, and notify other worker nodes 120 via permission information indication 156. Therefore, by updating the permission information 140 in response to permission request 112 and propagating the updated permission information 140 to worker nodes 120, the permission orchestrator node 150 can distribute and redistribute ownership of keyword permissions to worker nodes 120.

[0031] Turn now Figure 2 A block diagram of an exemplary layout associated with a worker node 120 that relinquishes keyword permissions is shown. In the illustrated embodiment, worker node 120A includes a memory cache 130, permission information 140, and a database application 200. As further shown, the database application 200 is processing an active transaction 210 holding lock 215 and a committed transaction 220 with an associated transaction commit number (XCN) 225. Also as shown, permission information 140 includes keyword range permissions 145A to 145C with keyword permission 205, and corresponding history information 230 specifying XCN 225. In some embodiments, worker node 120A is implemented in a manner different from that shown. As an example, database application 200 may process multiple active transactions 210 and multiple committed transactions 220.

[0032] In various embodiments, database application 200 is a set of program instructions executable to manage database 110, including managing the LSM tree built around database 110. Therefore, database application 200 can process database transactions to read records 132 from and write records 132 to the LSM tree. In various embodiments, to aid in processing database transactions, database application 200 can maintain metadata describing the structural layout of the LSM tree, including where file 115 is stored in database 110 and which records 132 can be included in which files 115. In various embodiments, the metadata includes a trie corresponding to file 115. Database application 200 can use the metadata to perform faster and more efficient keyword range lookups as part of processing database transactions.

[0033] As discussed, database application 200 can receive requests to perform transactions to read and write database records 132. Upon receiving a transaction request, database application 200 can initiate an active transaction 210 based on the received transaction request. In various embodiments, active transaction 210 refers to an ongoing transaction in which database application 200 is writing database records 132 to memory cache 130. When a transaction is active transaction 210, the database records 132 written for that transaction have not yet been committed and may not be readable / accessible by any entity other than the worker nodes 120 that wrote them. In various cases, database application 200 may decide to roll back active transaction 210, thus flushing the database records 132 written for active transaction 210 instead of committing them. For example, if database application 200 does not allow based on a set of criteria (e.g., combined with...). Figure 4 (For a more detailed discussion of time period analysis) If a specific database record 132 is written, the database application 200 can roll back the active transaction 210 that includes the write to database record 132. In some cases, the database application 200 may decide to roll back only a sub-part (or sub-transaction) of the active transaction 210.

[0034] In various embodiments, in order to write database record 132 when processing a given activity transaction 210, database application 200 determines whether worker node 120A has the appropriate permissions 205 and acquires locks 215 for those permissions 205 before writing database record 132. Imagine that the activity transaction 210 shown involves writing an instance of database record 132 with a key 134 (referred to as "keyword 134(XYT)") having a value of "XYT". Database application 200 may first check permission information 140 to determine whether the permission key 205 for key 134(XYT) has been provided to worker node 120A. In various embodiments, key permission 205 identifies the associated key 134 and the owner of key 134. As shown, key scope permissions 145A and 145C (shown in solid boxes) have been provided to worker node 120A, while key scope permission 145B (shown in dashed boxes) has not yet been provided to worker node 120A. Therefore, in the illustrated embodiment, the database application 200 determines that keyword permission 205A identifies worker node 120A as the owner of keyword 134(XYT); however, in other cases, worker node 120A may not be the owner, and therefore may have to request ownership of keyword 134(XYT) – in conjunction with Figure 4 Examples of claims to ownership are discussed in more detail.

[0035] In various embodiments, after determining that the required key permission 205 has been provided to the worker node 120, the worker node 120 acquires a lock 215 for the key 134 when writing the database record 132 with the associated key 134. Continuing the foregoing example, the worker node 120A may acquire the lock 215(XYT) for the key 134(XYT). In various embodiments, the lock 215 prevents another entity (e.g., another worker node 120) from writing the database record 132 with the associated key 134, and may further prevent the associated key permission 205 from being revoked and re-provided to another entity. Once the lock 215 is acquired, the worker node 120 can write the database record 132 for the corresponding key 134. As described in connection with Figure 4 discussed in more detail, before writing the database record 132, the worker node 120 may further determine based on XCN 225 whether another worker node 120 has submitted the database record 132 for the same database key 134 within a specific time period.

[0036] After processing an active transaction 210 (e.g., after writing all database records 132 requested for the transaction), the worker node 120 may commit the transaction, resulting in a committed transaction 220. In various embodiments, as part of the commit process, the worker node 120 marks each database record 132 of the transaction with a transaction commit number (XCN) 225. As shown, the XCN 225 of the committed transaction 220 is T501 (referred to as "XCN 225(T501)"). Accordingly, each record 132 associated with the committed transaction 220 may include metadata identifying XCN 225(T501). In various embodiments, XCN 225 is a monotonically increasing value, and thus can indicate a time period. That is, during operation, the system 100 may periodically increment the database system XCN 225 assigned to a transaction at commit time. Since the database system XCN 225 is periodically incremented, two committed transactions 220 can be associated with different XCN 225. For example, a first committed transaction 220 may be assigned XCN 225(T501), while a second committed transaction 220 may be assigned XCN 225(T412). Since the database system XCN 225 is incrementing, the worker node 120 (or another entity) can determine that the second committed transaction 220 occurred temporally before the first committed transaction 220 (T412 < T501). In various embodiments, the committed database records 132 are made available to other worker nodes 120 upon request and are eventually written to the file 115 at the database 110.

[0037] In various embodiments, as part of the commit process, worker node 120 also updates history information 230 based on the XCN 225 of the transaction being committed. In various embodiments, for keyword permission 205 and / or keyword range permission 145, history information 230 identifies the XCN 225 associated with the most recent record commit involving keyword permission 205 / keyword range permission 145—the XCN 225 is referred to as the “maximum XCN” or “latest XCN” of keyword permission 205 / keyword range permission 145. Imagine that database record 132 with keyword 134 (XYZ) is an instance committed for the illustrated committed transaction 220. Since worker node 120A holds keyword permission 205B, and no other worker node 120 has write permission for keyword 134 (XYZ), and worker node 120A holds keyword permission 205B, database record 132 is the most recently committed record 132 for keyword 134 (XYZ). Therefore, worker node 120A can update historical information 230 to associate keyword permission 205B (or keyword range permission 145A, which includes keyword permission 205B) with XCN 225 (T501).

[0038] Historical information 230 can be updated in different ways. In various cases, worker node 120 can update historical information 230 to associate each key permission 205 used in a given transaction at worker node 120 with the XCN 225 of that given transaction. In some cases, the set of key permissions 205 can be grouped and provided as key range permissions 145, and a portion of historical information 230 can be linked to key range permissions 145. Thus, worker node 120 can update historical information 230 to associate key range permissions 145 with XCN 225—this can be referred to as a “keyword range XCN.” As shown, key range permission 145C is associated with XCN 225 (T412). Therefore, the latest committed database record 132 for key range permission 145C appears in committed transaction 220 with XCN 225 (T412). However, other records 132 for keyword range permission 145C can be associated with different, smaller XCN 225, and therefore potentially committed by another worker node 120 at an earlier point in time. (As combined...) Figure 4 In more detail, worker node 120 can use historical information 230 to determine whether to abort record writing.

[0039] During operation, worker node 120A may receive a relinquishment request 152 to relinquish one or more keyword permissions 205 (or keyword range permissions 145). The relinquishment request 152 may be received from permission orchestrator node 150 and may identify the one or more keyword permissions 205 to be relinquished. In various embodiments, in response to receiving the relinquishment request 152, worker node 120A first prevents any new transaction from acquiring the lock 215 of the requested keyword permission 205. Worker node 120A may then determine whether any active transaction 210 holds the lock 215 of the requested keyword permission 205. If no lock 215 is associated with those keyword permissions 205, worker node 120A may send a relinquishment response 154 back to permission orchestrator node 150. In some cases, worker node 120A may relinquish the keyword permission 205 but retain the keyword permission 205 for the keyword range that includes the relinquished keyword permission 105. For example, worker node 120A may relinquish keyword permission 205 "XYZ" but retain the remaining keyword permissions 205 for the keyword range "XY". In various embodiments, the relinquishment response 154 includes an indication that the requested keyword permission 205 has been relinquished and therefore will not be used in a transaction unless the keyword permission 205 is re-provided to worker node 120A. In some cases, worker node 120 may relinquish a set of keyword range permissions 145 that includes one or more requested keyword permissions 205, so the relinquishment response 154 may include an indication of the relinquished keyword range permissions 145. For example, the relinquishment request 152 may specify keyword permission 205B, but worker node 120A may decide to relinquish the entire keyword range permissions 145A. The relinquishment response 154 may also include history information 230, which provides an indication of which relinquished keyword permissions 205 were used in commit record 132 when the keyword permissions 205 were provided to worker node 120.

[0040] In some cases, worker node 120A may determine that an active transaction 210 holds a lock 215 for the requested keyword permission 205. For example, a give-up request 152 may identify keyword permission 205A that has been acquired by the active transaction 210 in the illustrated embodiment. In some embodiments, worker node 120A waits for the relevant active transaction 210 to commit or abort. Thereafter, worker node 120A may provide a give-up response 154 to orchestrator node 150 for the requested keyword permission 205. If worker node 120 intends to give up keyword scope permission 145, but a lock 215 exists on non-requested keyword permissions 205 within keyword scope permission 145, worker node 120 may retain the locked keyword permission 205 but give up the rest of keyword scope permission 145. As an example, worker node 120A may receive a give-up request 152 identifying keyword permission 205B. Worker node 120A may decide to return keyword scope permission 145A. However, because keyword permission 205A (which is not the requested keyword permission) is locked by active transaction 210, worker node 120A can return all keyword permissions 205 in the keyword range permission 145A except for keyword permission 205A. After no lock 215 is held on keyword permission 205A, worker node 120A can return keyword permission 205A.

[0041] Now go to Figure 3 A block diagram of an exemplary permission orchestrator node 150 is shown. In the illustrated embodiment, the permission orchestrator node 150 includes a permission engine 300 with permission information 140. In some embodiments, the permission orchestrator node 150 is implemented differently than shown—for example, the permission orchestrator node 150 may serve as a worker node 120, and thus also include a database application 200 and a memory cache 130.

[0042] As shown in the figure, the permission orchestrator node 150 receives a permission request 112 from worker node 120B. Permission request 112 may identify one or more keyword permissions 205 or keyword range permissions 145 that worker node 120B is attempting to acquire. For example, permission request 112 from worker node 120B may identify keyword range permission 145A. In response to receiving permission request 112, permission engine 300 may process permission request 112 and return a permission response 114 indicating whether the requester has received the requested keyword permission 205.

[0043] In various embodiments, the permission engine 300 is a collection of executable software routines for facilitating the granting and relinquishing of keyword permissions 205 among worker nodes 120. In various embodiments, in response to receiving a permission request 112, the permission engine 300 first determines whether the requested keyword permission 205 has already been granted to worker nodes 120. As shown, the permission engine 300 stores permission information 140, so the permission engine 300 can consult the permission information 140 to determine whether the requested keyword permission 205 has been granted. If those keyword permissions 205 have not yet been granted, then the permission engine 300 can generate updated permission information 140, which assigns the requested keyword permission 205 to the worker node 120 making the request. The permission engine 300 can then distribute the updated permission information 140 to worker nodes 120. In some embodiments, the permission engine 300 publishes a permission information indication 156 to worker nodes 120 including the updated permission information 140. In other embodiments, the permission information indication 156 indicates that the permission information 140 has been updated. Therefore, worker node 120 can retrieve updated permission information 140 from permission orchestrator node 150 in response to receiving permission information instruction 156.

[0044] If the requested keyword permission 205 has already been provided to worker node 120, permission engine 300 may issue a relinquishment request 152 to worker node 120. For example, permission engine 300 may determine from a first version of permission information 140 (e.g., permission information 140A) that keyword scope permission 145A has been provided to worker node 120A (as shown). Therefore, as described above, permission engine 300 may send a relinquishment request 152 to worker node 120A to relinquish keyword scope permission 145A. In various embodiments, in response to receiving a relinquishment response 154 indicating that the requested keyword permission 205 has been relinquished, permission engine 300 updates the first version of permission information 140 to a second version (e.g., permission information 140B), wherein the requested keyword permission 205 has been assigned to the worker node 120 that is making the request. As shown in the figure, the permission engine 300 updates the permission information 140A, in which worker node 120A is the owner of keyword range permission 145A, to permission information 140B, in which worker node 120B is the owner of keyword range permission 145A. Then, the permission engine 300 can issue a permission information instruction 156 for the updated permission information 140. In some embodiments, the permission engine 300 issues a permission response 114, including the updated permission information 140, to the requesting worker node 120.

[0045] In various embodiments, when the permission information 140 is updated to modify the ownership of keyword permission 205, the permission engine 300 also updates the historical information 230 included in the permission information 140. As described above, the abandonment response 154 may identify XCN 225 for those keyword permissions 205 or keyword range permissions 145 used in the transaction at the abandoning worker node 120. The XCN 225 provided for keyword permission 205 (or keyword range permission 145) may represent the most recent time period of commit record 132 for the corresponding keyword 134 throughout the system 100. Therefore, the permission engine 300 may update the historical information 230 to include the XCN 225 identified in the received abandonment response 154.

[0046] Turn now Figure 4 A block diagram of an exemplary layout associated with a worker node 120 acquiring keyword permission 205 is shown. In the illustrated embodiment, worker node 120B includes a memory cache 130, permission information 140, and a database application 200. As shown, database application 200 includes active transactions 210 associated with lock 215 (XYZ) and snapshot XCN 410 (T432). Also as shown, permission information 140 includes keyword range permissions 145A to 145C and associated history information 230. In some embodiments, worker node 120B is implemented in a manner different from that shown. For example, database application 200 may handle multiple active transactions 210 and multiple committed transactions 220.

[0047] As discussed, database application 200 can receive requests to perform transactions to read and write database records 132. Upon receiving a transaction request, database application 200 can initiate a new active transaction 210 based on the received transaction request. While processing active transaction 210, database application 200 can write various records 132 to memory cache 130. When writing a specific record 132, database application 200 can first determine whether worker node 120B has been granted appropriate permissions 205 based on permission information 140. As previously mentioned, if appropriate permissions 205 have not been granted to worker node 120B, then worker node 120B can request permissions 205. Once worker node 120B has appropriate permissions 205, database application 200 can acquire the lock 215 for permission 205 before writing to database record 132. However, in various cases, before writing record 132, database application 200 can ensure that no other record 132 is committed after the time corresponding to snapshot XCN 410, the snapshot XCN 410 identifier allowing viewing of the status of system 100 of active transaction 210.

[0048] In various embodiments, the snapshot XCN 410 is a value of the latest XCN 225 indicating that the corresponding database record 132 thereof can be read by the corresponding active transaction 210. As shown in the figure, a snapshot XCN 410 (T432) is allocated to the active transaction 210. Therefore, the active transaction 210 can read committed database records 132 whose XCN 225 is less than or equal to "T432". (In some cases, only database records 132 with XCN less than the snapshot XCN 410 can be read.) For example, the active transaction 210 can read, from the database 110, the database record 132 stamped with XCN 225 (T230) where T230 < T432.

[0049] In various embodiments, in order to ensure the integrity of data stored in the system 100, if another database record 132 for a specific key 134 has been committed with an XCN 225 greater than the snapshot XCN 410 allocated to the corresponding active transaction 210, the database application 200 will not write a database record 132 for the specific key 134. Therefore, in various embodiments, for an active transaction 210, the database application 200 determines, based on historical information 230, whether a database record 132 for a specific key 134 has been committed with an XCN 225 greater than the snapshot XCN 410 (T432). For example, the active transaction 210 may involve writing a database record 132 for the key 134 (XYZ). Before writing the database record 132, the database application 200 may check the historical information 230 to determine the XCN 225 associated with the key right 205A. In some embodiments, the historical information 230 identifies the corresponding XCN 225 for each key right 205. If the XCN 225 associated with the key right 205A is greater than the snapshot XCN 410 (T432), the database application 200 may abort writing the database record 132 for the key 134 (XYZ). However, if the XCN 225 is not greater than the snapshot XCN 410 (T432), the database application 200 may write the database record 132.

[0050] In some embodiments, historical information 230 identifies XCN 225 for the entire keyword range permission 145. As shown, XCN 225 (T501) is associated with keyword range permission 145A, which includes keyword permission 205A. In response to determining that XCN 225 (T501) for keyword range permission 145A is greater than snapshot XCN 410 (T432), database application 200 may determine whether keyword range permission 205A itself is associated with a smaller XCN 225—even though XCN 225 (T501) is associated with the entire keyword range permission 145A, it may have already been added to historical information as a result of submitting database record 132 associated with a keyword 134 (e.g., XYA) different from keyword 134XYZ. In various embodiments, to determine XCN 225 for a specific keyword permission 205, worker node 120B sends XCN request 420 to worker node 120. In some embodiments, historical information 230 identifies the last worker node 120 associated with keyword permission 205, so worker node 120B can send XCN request 420 only to the identified worker node 120. In some cases, database application 200 can receive XCN response 425, which can identify the XCN 225 for keyword permission 205 identified in XCN request 420. If the identified XCN 225 is greater than snapshot XCN 410 (T432), database application 200 can abort writing database record 132 for keyword 134 (XYZ). However, if XCN 225 is not greater than snapshot XCN 410 (T432), database application 200 can write database record 132.

[0051] Now go to Figure 5 A flowchart of method 500 is shown. Method 500 is an embodiment of a method performed by a database system (e.g., system 100) for orchestrating the distribution of key-scope permissions (e.g., key-scope permissions 145) among multiple database nodes (e.g., worker nodes 120) of the database system. Method 500 can be performed by executing program instructions stored on a non-transitory computer-readable medium. In some embodiments, method 500 includes more or fewer steps than shown. For example, method 500 may include the step of the database system generating first permission information.

[0052] Method 500 begins at step 510, where the database system distributes first permission information (e.g., permission information 140A) to multiple database nodes. The first permission information may identify the distribution of key-range permissions to some of the multiple database nodes. A given key-range permission distributed to a given database node may allow the given database node to write records (e.g., database record 132) whose key (e.g., key 134) falls within the key range associated with the given key-range permission. In some cases, the first permission information may provide a second key-range permission to a second database node, which includes the first key-range permission.

[0053] In step 520, the database system receives from a first database node (e.g., worker node 120B) a request (e.g., permission request 112) for a first key-scope permission provided to a second database node (e.g., worker node 120A). In various embodiments, the database system sends a relinquishment request to the second database node to relinquish the first key-scope permission (e.g., relinquishment request 152). The second database node may relinquish the first key-scope permission in response to determining that it is not used in the set of active transactions (e.g., active transaction 210) occurring at the second database node. The database system may receive from the second database node an indication that the first key-scope permission has been relinquished (e.g., relinquishment response 154). In some cases, the second database node may relinquish the first key-scope permission but retain the remainder of it. In some cases, the indication may specify a transaction commit number (e.g., XCN 225), which indicates the time interval at which the latest record was committed for the first key-scope permission.

[0054] In step 530, the database system modifies the first permission information to derive second permission information (e.g., permission information 140B), which provides the first key-scope permission to the first database node instead of the second database node. In various embodiments, the second permission information is stored in a trie data structure comprising multiple branches, wherein a specific branch among the multiple branches corresponds to the first key-scope permission.

[0055] In step 540, the database system distributes the second permission information to some of the multiple database nodes. Distributing the second permission information may include notifying the multiple database nodes (e.g., via permission information indication 156) of the second permission information, and returning the second permission information in response to receiving a request for the second permission information from some of the multiple database nodes.

[0056] In various scenarios, the second permission information can identify a keyword-range transaction commit number, which indicates a first time interval when the latest record was committed for the first keyword-range permission. In various embodiments, the first database node determines whether a second time interval associated with the first database node (e.g., corresponding to snapshot XCN 410) occurred after the first time interval. In response to determining that the second time interval occurred after the first time interval, the first database node can write a database record for a specific keyword associated with the first keyword-range permission. In response to determining that the second time interval did not occur after the first time interval, the first database node can retrieve the record transaction commit number for the specific keyword from the database node that committed the latest record for that specific keyword. The first database node can prevent record writing for the specific keyword in response to determining that the record transaction commit number indicates a time interval preceding the second time interval associated with the first database node. The first database node can write a record for the specific keyword in response to determining that the record transaction commit number indicates a time interval preceding the second time interval associated with the first database node.

[0057] Now go to Figure 6 A flowchart of method 600 is shown. Method 600 is an embodiment of a method performed by a database system (e.g., system 100) for orchestrating the distribution of key-scope permissions (e.g., key-scope permissions 145) among multiple database nodes (e.g., worker nodes 120) of the database system. Method 600 can be performed by executing program instructions stored on a non-transitory computer-readable medium. In some embodiments, method 600 includes more or fewer steps than shown. For example, method 600 may include the step of the database system generating first permission information.

[0058] Method 600 begins at step 610, where a permission orchestrator database node (e.g., permission orchestrator node 150) provides a first key-range permission (e.g., key-range permission 145) to a first worker database node of the database system (e.g., worker node 120A). The first key-range permission allows writing to records (e.g., database record 132) whose keywords (e.g., keyword 134) fall within the first key-range range associated with the first key-range permission. In step 620, the permission orchestrator database node receives a permission request (e.g., permission request 112) from a second worker database node of the database system (e.g., worker node 120B) for a second key-range permission associated with a second key-range encompassed by the first key-range.

[0059] In step 630, in response to receiving a permission request, the permission orchestrator database node causes the first worker database node to relinquish at least a portion of the permissions for the first key range. This relinquishment may include sending a relinquishment request (e.g., relinquishment request 152) to the first worker database node to relinquish permissions associated with the second key range. The first worker database node may, in response to receiving the relinquishment request, prevent transactions from using the keywords associated with the second key range. In various cases, the first worker database node may determine that a record associated with a keyword falling within the second key range has been written for an ongoing transaction (e.g., active transaction 210). The first worker database node may commit the ongoing transaction. After committing the ongoing transaction, the first worker database node may return an indication to the permission orchestrator database node that a portion of the permissions for the first key range associated with the second key range has been relinquished (e.g., relinquishment response 154).

[0060] In step 640, after the first worker database node relinquishes at least a portion of the first key range permissions, the permission orchestrator database node provides the second key range permissions to the second worker database node. In various embodiments, providing the second key range permissions to the second worker database node includes the permission orchestrator database node providing the second worker database node with historical information (e.g., historical information 230) indicating one or more writes performed by the first worker database node. The second worker node can determine, based on the historical information, whether the first worker database node committed a record with a specific keyword falling within the second key range during a specific time interval. In response to determining that the first worker database node committed a record with the specific keyword during the specific time interval, the second worker database node can abort a portion of the transaction involving writing the record with the specific keyword.

[0061] Exemplary multi-tenant database system

[0062] Turn now Figure 7 An exemplary multi-tenant database system (MTS) 700 in which various technologies for implementing public text are shown—for example, system 100 could be an MTS 700. Figure 7In this embodiment, MTS 700 includes a database platform 710, an application platform 720, and a network interface 730 connected to a network 740. Also as shown, the database platform 710 includes a data storage 712 and a collection of database servers 714A to 714N that interact with the data storage 712, while the application platform 720 includes a collection of application servers 722A to 722N with a corresponding environment 724. In the illustrated embodiment, MTS 700 is connected to various user systems 750A to 750N via network 740. The disclosed multitenant system is included for illustrative purposes and is not intended to limit the scope of the disclosed text. In other embodiments, the techniques disclosed herein are implemented in a non-multitenant environment, such as a client / server environment, a cloud computing environment, a cluster of computers, etc.

[0063] In various embodiments, MTS 700 is a collection of computer systems that together provide various services to users (or “tenants”) interacting with MTS 700. In some embodiments, MTS 700 implements a Customer Relationship Management (CRM) system that provides tenants (e.g., companies, government agencies, etc.) with mechanisms to manage their relationships and interactions with customers and prospects. For example, MTS 700 enables tenants to store customer contact information (e.g., customer websites, email addresses, phone numbers, and social media data), identify sales opportunities, record service issues, and manage marketing campaigns. Furthermore, MTS 700 enables tenants to identify how they communicate with customers, what customers have purchased, when customers last purchased items, and what customers paid for. To provide the services of the CRM system and / or other services, as shown in the figure, MTS 700 includes a database platform 710 and an application platform 720.

[0064] In various embodiments, database platform 710 is a combination of hardware components and software routines that implement database services for storing and managing data (including tenant data) of MTS 700. As shown, database platform 710 includes data storage 712. In various embodiments, data storage 712 includes a collection of storage devices (e.g., solid-state drives, hard disk drives, etc.) connected together on a network (e.g., storage area network (SAN)) and configured to redundantly store data to prevent data loss. In various embodiments, data storage 712 is used to implement a database (e.g., database 110) comprising a collection of information organized in a manner that allows access to, storage of, and manipulation of information. Data storage 712 can implement a single database, a distributed database, a collection of distributed databases, a database with redundant online or offline backups or other redundancies, etc. As part of implementing the database, data storage 712 can store a file (e.g., file 115) comprising one or more database records, each record having a corresponding data payload (e.g., values ​​of fields in a database table) and metadata (e.g., key values, timestamps, table identifiers of tables associated with the records, tenant identifiers of tenants associated with the records, etc.).

[0065] In various embodiments, database records may correspond to rows of tables. Tables typically contain one or more data categories logically arranged as columns or fields in a viewable schema. Therefore, each record in a table may contain a data instance for each category, qualified by fields. For example, a database may include a table describing customers, with fields for basic contact information such as name, address, phone number, fax number, etc. Therefore, records for this table may include values ​​for each field in the table (e.g., a name for the name field). Another table may describe purchase orders, including fields for information such as customer, product, sales price, date, etc. In various embodiments, standard entity tables are provided for all tenants, such as tables for account, contact, lead, and opportunity data, each containing pre-qualified fields. The MTS 700 may store database records for one or more tenants in the same table—that is, tenants may share a single table. Therefore, in various embodiments, database records include a tenant identifier indicating the owner of the database record. Thus, a tenant's data remains secure and separate from other tenants' data, so that one tenant cannot access another tenant's data unless that data is explicitly shared.

[0066] In some embodiments, the data stored at data store 712 is organized as part of a Log Structure Merge Tree (LSM tree). An LSM tree typically comprises two high-level components: a memory cache and persistent storage. In operation, database server 714 may first write database records to the local memory cache, and then flush these records to persistent storage (e.g., data store 712). As part of flushing database records, database server 714 may write database records to new files included in the “top” level of the LSM tree. Over time, as database records move down the levels of the LSM tree, database server 714 may rewrite database records to new files included in lower levels. In various implementations, as database records age and move down the LSM tree, they move to increasingly slower storage devices (e.g., from solid-state drives to hard disk drives) of data store 712.

[0067] When database server 714 wants to access a database record for a specific key, it can traverse different levels of the LSM tree for a file that potentially contains database records with that specific key. If database server 714 determines that the file may contain relevant database records, it can retrieve the file from data storage 712 into its own storage. Database server 714 can then examine the retrieved file to find database records with the specific key. In various embodiments, database records are immutable once written to data storage 712. Therefore, if database server 714 wants to modify the value of a row in a table (which can be identified from the accessed database record), it writes the new database record to the top level of the LSM tree. Over time, database records are merged down the LSM tree hierarchy. Thus, the LSM tree can store various database records for a database key, where older records for that key are located at lower levels of the LSM tree compared to newer records.

[0068] In various embodiments, database server 714 is a hardware element, software routine, or a combination thereof capable of providing database services (e.g., data storage, data retrieval, and / or data manipulation). Database server 714 may correspond to worker node 120. Such database services can be provided by database server 714 to components within MTS 700 (e.g., application server 722) and components outside MTS 700. As an example, database server 714 may receive database transaction requests from application server 722 that requests to write data to or read data from data storage 712. Database transaction requests may specify an SQL SELECT command to select one or more rows from one or more database tables. The content of the rows can be qualified in the database records, so database server 714 can locate and return one or more database records corresponding to the selected one or more table rows. In various cases, database transaction requests may instruct database server 714 to write one or more database records against an LSM tree—database server 714 maintains the LSM tree implemented on database platform 710. In some embodiments, database server 714 implements a relational database management system (RDMS) or an object-oriented database management system (OODBMS) that facilitates the storage and retrieval of information for data storage 712. In various cases, database servers 714 can communicate with each other to facilitate transaction processing. For example, database server 714A can communicate with database server 714N to determine whether database server 714N has written database records to its memory cache for a specific key.

[0069] In various embodiments, application platform 720 is a combination of hardware components and software routines that implement and execute CRM software applications, providing and receiving relevant data, code, forms, web pages, and other information from user system 750, and storing relevant data, objects, web page content, and other tenant information via database platform 710. In various embodiments, to facilitate these services, application platform 720 communicates with database platform 710 to store, access, and manipulate data. In some instances, application platform 720 may communicate with database platform 710 via different network connections. For example, one application server 722 may be connected via a local area network, while another application server 722 may be connected via a direct network link. Transmission Control Protocol (TCP / IP) and Internet Protocol (TCP / IP) are exemplary protocols for communication between application platform 770 and database platform 710; however, it will be apparent to those skilled in the art that other transport protocols may be used depending on the network interconnection used.

[0070] In various embodiments, application server 722 is a hardware element, software routine, or combination thereof capable of providing services to application platform 720, including processing requests received from tenants of MTS 700. In various embodiments, application server 722 can create environment 724 that can be used for various purposes, such as providing developers with the functionality to develop, execute, and manage applications (e.g., business logic). Data can be transferred from another environment 724 and / or from database platform 710 to environment 724. In some cases, environment 724 cannot access data from other environments 724 unless such data is explicitly shared. In some embodiments, multiple environments 724 may be associated with a single tenant.

[0071] Application platform 720 can provide user system 750 with access to multiple different hosted (standard and / or custom) applications, including CRM applications and / or applications developed by tenants. In various embodiments, application platform 720 can manage application creation, application testing, application storage in database objects at data store 712, application execution in environment 724 (e.g., a virtual machine in process space), or any combination thereof. In some embodiments, application platform 720 can add and remove application server 722 from the server pool at any time for any reason, even if the user and / or organization may not have server affinity for a particular application server 722. In some embodiments, an interface system (not shown) implementing load balancing functionality (e.g., an F5 Big-IP load balancer) is located between application server 722 and user system 750 and is configured to distribute requests to application server 722. In some embodiments, the load balancer uses a least-connections algorithm to route user requests to application server 722. Other instances of load balancing algorithms, such as round-robin scheduling algorithms and observed response times, may also be used. For example, in some embodiments, three consecutive requests from the same user may hit three different servers 722, while three requests from different users may hit the same server 722.

[0072] In some embodiments, the MTS 700 provides security mechanisms such as encryption to keep each tenant's data separate unless the data is shared. If more than one server 714 or 722 is used, they can be located very close to each other (e.g., in a cluster of servers located in a single building or campus), or they can be distributed in locations far apart from each other (e.g., one or more servers 714 are located in city A, while one or more servers 722 are located in city B). Therefore, the MTS 700 can include one or more logically and / or physically connected servers, either locally or distributed across one or more geographic locations.

[0073] One or more users (e.g., via user system 750) can interact with MTS 700 via network 740. User system 750 may correspond to, for example, a tenant of MTS 700, a provider of MTS 700 (e.g., an administrator), or a third party. Each user system 750 can be a desktop PC, workstation, laptop computer, PDA, cellular phone, or any device that supports Wireless Access Protocol (WAP) or any other computing device capable of direct or indirect connection to the Internet or other network connections. User system 750 may include dedicated hardware configured to connect to MTS 700 via network 740. User system 750 can perform interactions with MTS 700, HTTP clients (e.g., browser programs such as Microsoft's Internet Explorer), and other communication methods. TM Browsers, Netscape's Navigator TM A browser, Opera browser, or a browser that enables WAP in the case of a cellular phone, PDA, or other wireless device, or a graphical user interface (GUI) corresponding to both, allows users of user system 750 (e.g., subscribers of a CRM system) to access, process, and view information and pages available to them by MTS 700 via network 740. Each user system 750 may include one or more user interface devices, such as a keyboard, mouse, touchscreen, pen, etc., for interacting with the graphical user interface (GUI) provided by a browser on a monitor screen, LCD display, etc., in conjunction with pages, forms, and other information provided by MTS 700 or other systems or servers. As described above, the disclosed embodiments are suitable for use with the Internet, which refers to a specific globally interconnected network. However, it should be understood that other networks may be used instead of the Internet, such as intranets, extranets, virtual private networks (VPNs), non-TCP / IP-based networks, any LAN or WAN, etc.

[0074] Because users of User System 750 can have different capabilities, the capabilities of a specific User System 750 can be defined as one or more permission levels associated with the current user. For example, when a salesperson is interacting with MTS 700 using a specific User System 750, User System 750 may have capabilities assigned to that salesperson (e.g., user privileges). However, when an administrator is interacting with MTS 700 using the same User System 750, User System 750 may have capabilities assigned to that administrator (e.g., administrative privileges). In systems with a hierarchical role model, a user with one permission level may have access to applications, data, and database information accessible to users with lower permission levels, but may not have access to certain applications, database information, and data accessible to users with higher permission levels. Therefore, depending on the user's security or permission level, different users may have different capabilities in accessing and modifying application and database information. There may also be some data structures managed by MTS 700 at the tenant level, while other data structures are managed at the user level.

[0075] In some embodiments, the user system 750 and its components can be configured using an application (such as a browser) that includes computer code executable on one or more processing elements. Similarly, in some embodiments, the MTS 700 (and other MTS instances, where more than one exists) and its components can be configured by an operator using an application that includes computer code executable on a processing element. Therefore, the various operations described herein can be performed by executing program instructions stored on a non-transitory computer-readable medium and executed by a processing element. The program instructions can be stored on a non-volatile medium such as a hard disk, or in any other well-known volatile or non-volatile storage medium or device, such as ROM or RAM, or provided on any medium capable of bootstrapping program code, such as optical disc (CD) media, digital versatile disc (DVD) media, floppy disks, etc. Furthermore, the entire program code or portions thereof can be transferred and downloaded from a software source (e.g., via the Internet) or from another server, as is well known, or transmitted using any well-known communication medium and protocol (e.g., TCP / IP, HTTP, HTTPS, Ethernet, etc.) via any other conventional network connection (e.g., extranet, VPN, LAN, etc.). It will also be understood that the computer code used to implement the aspects of the disclosed embodiments can be implemented in any programming language that can be executed on a server or server system, such as C, C++, HTML, Java, JavaScript, or any other scripting language, such as VBScript.

[0076] Network 740 can be a LAN (Local Area Network), WAN (Wide Area Network), wireless network, point-to-point network, star network, token ring network, hub network, or any other suitable configuration. The globally interconnected network, often referred to as the "Internet" with a capital "I," is an example of a TCP / IP (Transmission Control Protocol and Internet Protocol) network. However, it should be understood that the disclosed embodiments can utilize any of a variety of other types of networks.

[0077] User system 750 can communicate with MTS 700 using TCP / IP and other common Internet protocols such as HTTP, FTP, AFS, and WAP at higher network levels. For example, when using HTTP, user system 750 can include an HTTP client, commonly referred to as a "browser," for sending and receiving HTTP messages from the HTTP server at MTS 700. Such a server can be implemented as the sole network interface between MTS 700 and network 740, but other technologies can be used or alternatively. In some implementations, the interface between MTS 700 and network 740 includes load-sharing functionality, such as a round-robin scheduling algorithm HTTP request dispatcher, to balance the load and distribute incoming HTTP requests evenly across multiple servers.

[0078] In various embodiments, user system 750 communicates with application server 722 to request and update system-level and tenant-level data from MTS 700, which may require one or more queries to data store 712. In some embodiments, MTS 700 automatically generates one or more SQL statements (SQL queries) designed to access the required information. In some cases, user system 750 may generate a request with a specific format corresponding to at least a portion of MTS 700. As an example, user system 750 may use object notation describing an object-relational mapping (e.g., a JavaScript object notation mapping) of specified multiple objects to request the movement of data objects into a specific environment 724.

[0079] Exemplary computer system

[0080] Now go to Figure 8 A block diagram of an exemplary computer system 800 is shown, which may implement system 100, database 110 and / or worker node 120, MTS 700 and / or user system 750. Computer system 800 includes a processor subsystem 880 connected to system memory 820 and I / O interface 840 via interconnect 860 (e.g., system bus). I / O interface 840 is connected to one or more I / O devices 850. Although for convenience... Figure 8 The diagram shows a single computer system 800, but system 800 can also be implemented as two or more computer systems operating together.

[0081] Processor subsystem 880 may include one or more processors or processing units. In various embodiments of computer system 800, multiple instances of processor subsystem 880 may be coupled to interconnect 860. In various embodiments, processor subsystem 880 (or each processor unit within 880) may include cache or other forms of onboard memory.

[0082] System memory 820 may be used to store program instructions that can be executed by processor subsystem 880 to cause system 800 to perform the various operations described herein. System memory 820 may be implemented using different physical memory media, such as hard disk storage, floppy disk storage, removable disk storage, flash memory, random access memory (RAM-SRAM, EDO RAM, SDRAM, DDR SDRAM, RAMBUS RAM, etc.), read-only memory (PROM, EEPROM, etc.), etc. The memory in computer system 800 is not limited to main memory, such as memory 820. Instead, computer system 800 may also include other forms of memory, such as buffer memory in processor subsystem 880 and auxiliary memory on I / O device 850 (e.g., hard disk drive, storage array, etc.). In some embodiments, these other forms of memory may also store program instructions that can be executed by processor subsystem 880. In some embodiments, program instructions that implement database application 200 when executed may be included / stored within system memory 820.

[0083] According to various embodiments, I / O interface 840 can be any of various types of interfaces configured to connect to and communicate with other devices. In one embodiment, I / O interface 840 is a bridge chip (e.g., a southbridge) from a front end to one or more back end buses. I / O interface 840 can be connected to one or more I / O devices 850 via one or more corresponding buses or other interfaces. Examples of I / O devices 850 include storage devices (hard disk drives, optical disk drives, removable flash drives, storage arrays, SANs, or their associated controllers), network interface devices (e.g., to a local area network or wide area network), or other devices (e.g., graphics, user interface devices, etc.). In one embodiment, computer system 800 is connected to a network (e.g., configured to communicate via WiFi, Bluetooth, Ethernet, etc.) via network interface device 850.

[0084] Implementations of the subject matter of this application include, but are not limited to, the following examples 1 to 20.

[0085] 1. A method comprising:

[0086] The database system distributes first permission information to multiple database nodes of the database system, wherein the first permission information identifies the distribution of keyword range permissions to some of the multiple database nodes, and wherein the given keyword range permission distributed to a given database node allows the given database node to write records whose keywords fall within the keyword range associated with the given keyword range permission.

[0087] The database system receives a request from the first database node for the first key-scope permissions granted to the second database node;

[0088] The database system modifies the first permission information to derive the second permission information, which grants the first key-scope permissions to the first database node instead of the second database node; and

[0089] The database system distributes secondary permission information to some of the multiple database nodes.

[0090] 2. The method described in Example 1 further includes:

[0091] Before modifying the first permission information, the database system:

[0092] Send a request to the second database node to relinquish the first key scope permission, wherein the second database node is operable to relinquish the first key scope permission in response to determining that it has not been used in the set of active transactions performed at the second database node; and

[0093] Receive an indication from the second database node that the permissions for the first key scope have been revoked.

[0094] 3. According to the method described in Example 2, the first permission information provides the second keyword range permission to the second database node, the second keyword range permission is a superset of the first keyword range permission, and the second database node can operate to abandon the first keyword range permission but retain the remainder of the second keyword range permission.

[0095] 4. The method according to Example 2, wherein the instruction specifies a transaction commit number associated with the most recently committed record for the first key range permissions, and wherein the second permission information identifies the transaction commit number to the first database node, wherein the first database node is operable to determine whether to write a specific record based on the transaction commit number.

[0096] 5. The method according to Example 1, wherein the second permission information limits the keyword range transaction commit number associated with the most recently committed record for the first keyword range permission, and wherein the first database node is operable to:

[0097] Determine whether the transaction commit number associated with the first database node is greater than the key range transaction commit number; and

[0098] In response to determining that the transaction commit number is greater than the key range transaction commit number, a record is written for the specific key associated with the first key range permission.

[0099] 6. The method according to Example 1, wherein the second permission information limits the keyword range transaction commit number associated with the most recently committed record for the first keyword range permission, and wherein the first database node is operable to:

[0100] Determine whether the transaction commit number associated with the first database node is greater than the key range transaction commit number; and

[0101] In response to determining that the transaction commit number is not greater than the key range transaction commit number, retrieve the record transaction commit number for the specific key associated with the first key range permission from the database node that committed the latest commit.

[0102] 7. According to the method described in Example 6, the first database node is capable of operating as follows:

[0103] In response to determining that the transaction commit number associated with the first database node is not greater than the record transaction commit number, writes to records for a specific key are prevented.

[0104] 8. The method according to Example 6, wherein the first database node is operable to:

[0105] In response to determining that the transaction commit number associated with the first database node is greater than the record transaction commit number, a record is written for the specific key.

[0106] 9. According to the method described in Example 1, the second permission information is stored in a trie data structure including multiple branches, and a specific branch in the multiple branches corresponds to the first key range permission.

[0107] 10. The method according to Example 9, wherein distributing the second permission information to some of the multiple database nodes includes:

[0108] Notify multiple database nodes of the second permission information;

[0109] Receive information requests for second-level permission information from some of multiple database nodes; and

[0110] It returns a trie data structure in response to an information request.

[0111] 11. A non-transitory computer-readable medium storing program instructions thereon, the program instructions being executable by a database system to cause the database system to perform operations, including:

[0112] Distribute first permission information to multiple database nodes of a database system, wherein the first permission information identifies the distribution of keyword range permissions to some of the multiple database nodes, and wherein a given keyword range permission distributed to a given database node allows the given database node to write records whose keywords fall within the keyword range associated with the given keyword range permission.

[0113] Receive permission requests from the first database node for the first key-scope permissions granted to the second database node;

[0114] Modify the first permission information to export the second permission information, which grants the first key-scope permissions to the first database node instead of the second database node; and

[0115] Distribute secondary permission information to some of the multiple database nodes.

[0116] 12. The medium according to Example 11, wherein the operation further includes:

[0117] Send a request to the second database node to relinquish at least the first key scope permissions; and

[0118] Before modifying the first permission information, receive an indication from the second database node that the first key scope permission and the second key scope permission have been abandoned.

[0119] 13. The medium according to Example 12, wherein the first permission information is modified such that a superset keyword range permission is provided to the first database node, and wherein the superset keyword range permission includes the first keyword range permission and a second keyword range permission that the first database node did not request in the permission request.

[0120] 14. The medium according to Example 12, wherein a designated second database node submits records for keywords falling within a keyword range associated with a first keyword range permission during a specific time period.

[0121] 15. The medium according to Example 11, wherein the first permission information is modified such that the second permission information includes an indication of whether another database node submitted a record with a keyword falling within the keyword range associated with the first keyword range permission during a specific time period.

[0122] 16. A method comprising:

[0123] The permission orchestrator database node of the database system provides a first key range permission to the first worker database node of the database system, wherein the first key range permission allows writing to records whose keywords fall within the first key range associated with the first key range permission;

[0124] The permission orchestrator database node receives permission requests for the second key scope permissions from the second worker database node of the database system. The second key scope permissions are associated with the second key scope contained in the first key scope.

[0125] In response to receiving a permission request, the permission orchestrator database node causes the first worker database node to relinquish at least a portion of the first key-scope permissions; and

[0126] After the first worker database node relinquishes at least a portion of its first key-scope permissions, the permission orchestrator database node grants the second worker database node second key-scope permissions.

[0127] 17. The method according to Example 16, wherein causing includes sending a relinquishment request to a first worker database node to relinquish permissions associated with a second key scope, and wherein the method further includes:

[0128] The first worker database node determines that records have been written for the ongoing transaction and are associated with keywords falling within the range of the second key;

[0129] The ongoing transaction is submitted by the first worker database node; and

[0130] After committing the ongoing transaction, the first worker database node returns an indication to the permission orchestrator database node that a portion of the permissions associated with the second key scope has been relinquished.

[0131] 18. The method according to Example 17 further includes:

[0132] In response to receiving a drop request, the first worker database node prevents the transaction from using the key associated with the second key range.

[0133] 19. The method according to Example 16, wherein providing the second key scope permission to the second worker database node includes providing the second worker database node with historical information indicating one or more writes performed by the first worker database node, and wherein the method further includes:

[0134] The second worker database node determines, based on historical information, whether the first worker database node submitted a record with a specific keyword falling within the range of the second keyword during a specific time interval.

[0135] 20. The method according to Example 19 further includes:

[0136] In response to determining that the first worker database node committed a record with a specific key during a specific time interval, the second worker database node aborts part of the transaction involving writing the record with the specific key.

[0137] The disclosed text includes references to "embodiments," which are non-limiting implementations of the disclosed concepts. References to "embodiments," "an embodiment," "a particular embodiment," "some embodiments," "various embodiments," etc., do not necessarily refer to the same embodiment. A large number of possible embodiments are contemplated, including specific embodiments described in detail, as well as modifications or substitutions falling within the spirit or scope of the disclosed text. Not all embodiments are required to embody any or all of the potential advantages described herein.

[0138] Unless otherwise stated, the specific embodiments are not intended to limit the scope of the claims drafted based on the disclosed text to the disclosed form, even where only a single instance is described in conjunction with a particular feature. Therefore, the disclosed embodiments are illustrative rather than restrictive unless otherwise stated. This application is intended to cover substitutions, modifications, and equivalents that will be apparent to those skilled in the art who benefit from the disclosed text.

[0139] Specific features, structures, or characteristics may be combined in any suitable manner consistent with the published text. Therefore, the published text is intended (explicitly or implicitly) to include any feature or combination of features disclosed herein, or any generalization thereof. Thus, during the examination of this application (or an application claiming priority thereto), new claims may be made for any such combination of features. In particular, with reference to the appended claims, features of dependent claims may be combined with features of independent claims, and features in the individual independent claims may be combined in any suitable manner, not merely in the specific combinations listed in the appended claims.

[0140] For example, although the appended dependent claims are drafted such that each dependent claim is subordinate to a single other claim, additional dependencies are also contemplated, including: claim 5 may be subordinate to any of the preceding claims; claim 6 may be subordinate to any of the preceding claims; claim 9 may be subordinate to any of the preceding claims; claim 15 may be subordinate to any one of claims 11 to 14; and claim 19 may be subordinate to any one of claims 16 to 18. Where appropriate, it is also contemplated that a claim drafted in one statutory type (e.g., apparatus) may imply a corresponding claim in another statutory type (e.g., method).

[0141] Because the published text is a legal document, various terms and phrases are subject to administrative and judicial interpretation. This announcement is made so that the following paragraphs, along with the definitions provided throughout the publication, will be used to determine how claims drafted based on the published text should be interpreted.

[0142] Unless the context clearly specifies otherwise, references to the singular forms such as "a," "an," and "the" mean "one or more." Therefore, reference to "an item" in a claim does not exclude other instances of that item.

[0143] The word “can” is used in this article in the sense of authority (i.e., having the potential to be able to) rather than in the sense of mandatory (i.e., having to).

[0144] The terms “contains” and “includes” and their forms are open-ended, meaning “including but not limited to”.

[0145] When the term "or" is used in conjunction with a list of options in public text, it is generally understood to be used in an inclusive sense unless the context otherwise specifies. Thus, the statement "x or y" is equivalent to "x or y, or both," which covers x but not y, covers y but not x, and covers both x and y. On the other hand, phrases like "x or y, but not both" clearly indicate that "or" is used in an exclusive sense.

[0146] The statement “w, x, y, or z, or any combination thereof” or “...at least one of w, x, y, and z” is intended to cover all possibilities involving a single element up to the total number of elements in the set. For example, given the set [w, x, y, z], these phrases cover any single element in the set (e.g., w but not x, y, or z), any two elements (e.g., w and x, but not y or z), any three elements (e.g., w, x, and y, but not z), and all four elements. Therefore, the phrase “...at least one of w, x, y, and z” refers to at least one element of the set [w, x, y, z], thus covering all possible combinations in the list of options. This phrase should not be interpreted as requiring at least one instance of w, at least one instance of x, at least one instance of y, and at least one instance of z.

[0147] In public text, various “labels” can be nouns. Unless the context otherwise specifies, different labels used for a feature (e.g., “first circuit,” “second circuit,” “specific circuit,” “given circuit,” etc.) refer to different instances of that feature. Unless otherwise stated, the labels “first,” “second,” and “third” do not imply any kind of ordering (e.g., spatial, temporal, logical, etc.) when applied to a particular feature.

[0148] In public texts, different entities (which may be referred to differently as “units,” “circuits,” other components, etc.) can be described or said to be “configured” to perform one or more tasks or operations. The expression—[entity] configured to [perform one or more tasks]—is used here to refer to a structure (i.e., a physical thing). More specifically, this expression is used to indicate that the structure is arranged to perform one or more tasks during operation. Even if the structure is not currently being operated, it can be said that the structure is “configured” to perform certain tasks. Therefore, an entity described or stated as “configured” to perform a certain task refers to a physical thing, such as a device, a circuit, a memory storing program instructions executable to perform that task, etc.

[0149] The term "configured as" does not mean "configurable as". For example, an unprogrammed FPGA is not considered "configured as" to perform certain specific functions. However, such an unprogrammed FPGA can be "configurable as" to perform that function.

[0150] The phrase "based on" is used to describe one or more factors that influence a determination. This term does not exclude other factors that may influence the determination. That is, a determination can be based solely on the specified factor or on the specified factor along with other unspecified factors. Consider the phrase "A is determined based on B." This phrase specifies that B is a factor used to determine A or influences the determination of A. This phrase does not exclude that the determination of A may also be based on other factors, such as C. This phrase is also intended to cover embodiments where A is determined solely based on B. As used herein, the phrase "based on" is synonymous with the phrase "at least partially based on."

[0151] The phrase "in response to" describes one or more factors that cause an effect. This phrase does not exclude the possibility that other factors may influence or otherwise cause this effect. That is, the effect can be in response to only the specified factor or in response to the specified factor along with other unspecified factors. Consider the phrase "in response to B, A is performed." This phrase specifies that B is the factor that causes A to be performed. This phrase does not exclude the possibility that A can also be in response to some other factor, such as C. This phrase is also intended to cover embodiments where A is performed only in response to B.

[0152] In the published text, various “modules” operable to perform specified functions are illustrated in the accompanying drawings and described in detail above. As used herein, a “module” refers to software and / or hardware operable to perform a specified set of operations. A module can refer to a set of software instructions executable by a computer system to perform the set of operations. A module can also refer to hardware configured to perform the set of operations. Hardware modules can constitute general-purpose hardware and a non-transitory computer-readable medium storing program instructions, or special-purpose hardware, such as a custom ASIC. Thus, a module described as “executable” to perform an operation refers to a software module, while a module described as “configured” to perform an operation refers to a hardware module. A module described as “operable” to perform an operation refers to a software module, a hardware module, or a combination thereof. Furthermore, any discussion of modules referred to herein as “executable” to perform certain operations should be understood to mean that, in other embodiments, these operations may be implemented by a hardware module “configured” to perform an operation, and vice versa.

Claims

1. A method for keyword permission distribution, comprising: The database system distributes first permission information to multiple database nodes of the database system, wherein the first permission information identifies the distribution of keyword range permissions to some of the multiple database nodes, and wherein a given keyword range permission distributed to a given database node allows the given database node to write records whose keywords fall within the keyword range associated with the given keyword range permission; The database system receives a request from the first database node for the first key range permissions provided to the second database node; The database system modifies the first permission information to derive second permission information, which provides the first keyword range permission to the first database node instead of the second database node. and The database system distributes the second permission information to some of the plurality of database nodes, wherein the second permission information defines the keyword-range transaction commit number associated with the most recently committed record for the first keyword-range permission, and wherein the first database node is operable to: Determine whether the transaction commit number associated with the first database node is greater than the key range transaction commit number; and In response to determining that the transaction commit number is greater than the keyword range transaction commit number, a record is written for the specific keyword associated with the first keyword range permission.

2. The method according to claim 1, further comprising: Before modifying the first permission information, the database system: Send a request to the second database node to relinquish the first keyword scope permission, wherein the second database node is operable to relinquish the first keyword scope permission in response to determining that the first keyword scope permission has not been used in the set of active transactions performed at the second database node; and Receive an indication from the second database node that the permission for the first keyword range has been revoked.

3. The method according to claim 2, wherein the first permission information provides a second keyword range permission to the second database node, the second keyword range permission being a superset of the first keyword range permission, and wherein the second database node is operable to abandon the first keyword range permission but retain the remainder of the second keyword range permission.

4. The method of claim 2 or 3, wherein the indication specifies the keyword range transaction commit number associated with the record most recently committed for the first keyword range permission.

5. The method of claim 1, wherein the second permission information defines a keyword range transaction commit number associated with the record most recently committed for the first keyword range permission, and wherein the first database node is operable to: Determine whether the second transaction commit number associated with the first database node is greater than the key range transaction commit number; and In response to determining that the second transaction commit number is not greater than the keyword range transaction commit number, the record transaction commit number for the second specific keyword associated with the first keyword range permission is retrieved from the database node that committed the latest commit.

6. The method of claim 5, wherein the first database node is operable to: In response to determining that the second transaction commit number associated with the first database node is not greater than the record transaction commit number, record writing for the second specific key is blocked.

7. The method of claim 5, wherein the first database node is operable to: In response to determining that the second transaction commit number associated with the first database node is greater than the record transaction commit number, a record is written for the second specific key.

8. The method according to claim 1, wherein the second permission information is stored in a trie data structure including multiple branches, and a specific branch among the multiple branches corresponds to the first keyword range permission.

9. The method of claim 8, wherein distributing the second permission information to some of the plurality of database nodes comprises: The second permission information is notified to the plurality of database nodes; Receive information requests for the second permission information from some of the plurality of database nodes; and The trie data structure is returned in response to the information request.

10. A computer-readable medium storing program instructions thereon, the program instructions being executable by a database system to cause the database system to perform operations including: Distribute first permission information to multiple database nodes of the database system, wherein the first permission information identifies the distribution of keyword range permissions to some of the multiple database nodes, and wherein a given keyword range permission distributed to a given database node allows the given database node to write records whose keywords fall within the keyword range associated with the given keyword range permission; Receive permission requests from the first database node for the first key-scope permissions granted to the second database node; Modify the first permission information to export the second permission information, which grants the first keyword scope permission to the first database node instead of the second database node. and Distribute the second permission information to some of the plurality of database nodes, wherein the second permission information defines a keyword-range transaction commit number associated with the most recently committed record for the first keyword-range permission, and wherein the first database node is operable to: Determine whether the transaction commit number associated with the first database node is greater than the key range transaction commit number; and In response to determining that the transaction commit number is greater than the keyword range transaction commit number, a record is written for the specific keyword associated with the first keyword range permission.

11. The medium of claim 10, wherein the operation further comprises: Send a request to the second database node to relinquish at least the permissions for the first keyword scope; and Before modifying the first permission information, receive an indication from the second database node that the first keyword range permission and the second keyword range permission have been abandoned.

12. The medium of claim 11, wherein the modification of the first permission information is made such that a superset keyword range permission is provided to the first database node, and wherein the superset keyword range permission includes the first keyword range permission and the second keyword range permission that the first database node did not request in the permission request.

13. The medium of claim 10, wherein the modification of the first permission information is made such that the second permission information includes an indication of whether another database node submitted a record having a keyword falling within the keyword range associated with the first keyword range permission during a specific time period.

14. A computer system, comprising: At least one processor; and A memory having program instructions stored thereon, the program instructions being executable by the at least one processor to perform the method of any one of claims 1 to 9.

Citation Information

Patent Citations

  • Database system, node and method

    CN111241590A

  • Data access control method and device

    CN111460506A