Generate a hash tree for the database schema
By creating a hash tree in the database system and comparing the root hash value, the data corruption problem caused by database schema changes is solved, and the correct input and processing of data under different schemas is achieved.
Patent Information
- Application Number
- CN202080071836.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-18
- Filing Date
- 2020-12-14
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2040-12-14
AI Technical Summary
In database systems, changes in database schema can lead to data corruption and system crashes, especially in a multi-tenant model where data cannot be entered correctly when the database schema is different from the one when it was created.
By creating a hash tree, the hash function is used to map information in the database schema to the root hash value, thereby determining the unique identifier of the database schema. Compare the root hash values of the two database schemas to tell if they are different and prevent data from being entered under different schemas.
It effectively prevents data corruption and system crashes caused by database schema differences, ensuring the correct input and processing of data between different database schemas.
Smart Images

Figure CN114600094B_ABST
Abstract
Description
Background Art Technical Field
[0002] The disclosure relates generally to database schemas, and more particularly, to creating a hash tree (or hierarchy of hash values) having a root hash value that identifies a database schema.
[0003] Related technical description
[0004] Modern database systems typically implement management systems that enable users to store a collection of information in an organized manner that can be efficiently accessed and manipulated. In many cases, these management systems maintain relational databases that are structured to identify relationships between pieces of information. Information stored in a relational database is typically organized as a collection of tables, each table consisting of columns and rows, where each column defines a grouping of information. In the context of relational databases, the structure of the tables (e.g., the columns that form the tables) and how they are related are specified in a database schema, which provides a logical grouping of objects in the database, such as tables, views, stored procedures, and the like. Throughout the life of a database, the database schema may undergo various changes as additional tables are added and old tables are structurally altered or deleted. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] Figure 1 is a block diagram illustrating example elements of a system capable of creating a hash tree having a root hash value that can be used to identify a database schema, according to some embodiments.
[0006] Figure 2 is a block diagram illustrating exemplary elements of a snapshot generation engine capable of generating snapshots including hash trees according to some embodiments.
[0007] Figure 3A is a block diagram illustrating an example unit of a hash tree having a root hash value that can be used to identify a database schema according to some embodiments.
[0008] Figure 3B is a block diagram illustrating exemplary units of two different hash trees according to some embodiments.
[0009] Figure 4 is a block diagram illustrating example units including example data that may be hashed to create a hash tree having a root hash value according to some embodiments.
[0010] Figure 5 is a block diagram illustrating example elements of an input engine that may be used to create a state associated with a snapshot according to some embodiments.
[0011] Figure 6 and Figure 7 is a flow chart illustrating an exemplary method involving creating a snapshot including a root hash value according to some embodiments.
[0012] Figure 8 is a block diagram illustrating an example computer system according to some embodiments.
[0013] The disclosure includes references to "one embodiment" or "an embodiment." The appearance of the phrase "in one embodiment" or "in an embodiment" does not necessarily refer to the same embodiment. The particular features, structures, or characteristics may be combined in any suitable manner consistent with the disclosure.
[0014] In the disclosure, different entities (which may be variously referred to as "units," "circuits," other components, etc.) may be described or referred to as being "configured" to perform one or more tasks or operations. The expression "[the entity] is configured to [perform one or more tasks]" is used herein to refer to a structure (i.e., a physical thing, such as an electronic circuit). More specifically, the expression is used to indicate that the structure is arranged to perform one or more tasks during operation. Even if the structure is not currently operating, it can be said that the structure is "configured to" perform certain tasks. "A network interface configured to communicate over a network" is intended to cover, for example, an integrated circuit having circuits that perform that function during operation, even if the integrated circuit in question is not currently in use (e.g., a power source is not connected to the integrated circuit). Thus, an entity described or recited as "configured to" perform a certain task refers to a physical thing, such as a device, a circuit, a memory storing program instructions executable to perform the task, etc. This phrase is not used herein to refer to an intangible thing. Therefore, "configured to" is not used herein to refer to a software entity such as an application programming interface (API).
[0015] The term "configured to" does not mean "configurable to." For example, an unprogrammed FPGA would not be considered "configured to" perform a particular function, although it may be "configurable to" perform that function and may be "configured to" perform that function after being programmed.
[0016] As used herein, the terms "first," "second," and the like are used as labels for the nouns preceding them and do not imply any type of ordering (e.g., spatial, temporal, logical, etc.) unless otherwise specified. For example, in a processor having eight processing cores, the terms "first" and "second" processing cores may be used to refer to any two of the eight processing cores. In other words, for example, the first processing core and the second processing core are not limited to processing cores 0 and 1.
[0017] As used herein, the term "based on" is used to describe one or more factors that influence a decision. This term does not exclude other factors that may influence the decision. That is, a decision may be based only on specified factors or on specified factors and other unspecified factors. Consider the phrase "determining A based on B." This phrase specifies that B is a factor used to determine A or influence the decision of A. This phrase does not exclude that the decision of A may also be based on some other factor, such as C. This phrase is also intended to cover embodiments in which A is determined based only on B. As used herein, the phrase "based on" is therefore synonymous with the phrase "based at least in part on." DETAILED DESCRIPTION
[0018] Many database systems are capable of creating database snapshots of the databases they manage. The term "database snapshot" is used herein in accordance with its common and customary meaning in the art, referring to a database structure that saves the state of a portion or all of a database at the point in time when the database snapshot is created. Therefore, a database snapshot can be used to recreate a portion or all of a database corresponding to the state in which the database snapshot was created. However, in some implementations, a database snapshot does not store certain information about a database, such as its database schema. When recreating a portion or all of a database, a database system can use an already existing database schema. However, if the database schema is different from the database schema at the point in time when the snapshot was created, the database system will likely encounter problems such as data corruption and system crashes.
[0019] For example, in a multi-tenant model, a database system stores data for multiple tenants in a database. At some point, the database system can generate a tenant database snapshot that captures all the data for a particular tenant at that point in time. This database snapshot can be used, for example, to recreate the tenant in a sandbox environment or on a different database as part of a tenant migration service. If the database schema of the sandbox environment or the other database is different (for example, some tables include additional columns), the tenant's data will not be able to be entered under that database schema without data corruption or other undesirable issues. Therefore, to avoid these undesirable issues, it may be necessary to determine whether the two database schemas are different
[0020] The disclosure describes various techniques for determining whether database schemas are different and identifying how different database schemas are different. In various embodiments described below, a database system creates a snapshot for a collection of data stored in a database managed by the database system. In some cases, the database system may create a snapshot in response to receiving a request from another system such as an application server or a user device. In various embodiments, as part of creating a snapshot, the database system applies a hash function to different parts of the information specified in a database schema that defines a database hierarchy, in which a database object (e.g., a table) includes attributes / columns, which themselves include properties (e.g., data types). The term "hash function" is used herein in accordance with its common and customary meaning in the art, and refers to a function that can be used to map data of any size to a value of a fixed size. In various embodiments, the database system hashes one or more characteristics of each attribute together (i.e., by applying a hash function) to derive the first layer of the hash value hierarchy. The term "hash value" is used herein in accordance with its common and customary meaning in the art, and refers to a value derived by applying a hash function to data. The database system can group the first layer of hash values associated with the same database table and apply a hash function to each group to derive a second layer of the hash value hierarchy. The process can then continue (e.g., by hashing the second layer of hash values together) until a root hash value is finally derived. Thus, in some embodiments, the root hash value can represent the entire database architecture. In various embodiments, the database system includes a hierarchy of hash values of the created snapshots.
[0021] In various embodiments, when recreating the state identified by the snapshot (e.g., causing the database to store the set of data captured by the snapshot when the snapshot was created), the database system generates a second hierarchy of hash values based on the current database schema of the database in which the snapshot data is stored. Thereafter, the database system determines whether the database schema associated with the snapshot is different from the current database schema by comparing the root hash values of the two hierarchies. In the event that the two root hash values are different, the database system can prevent the data associated with the snapshot from being stored in the database because the two database schemas are different. In various embodiments, the database system compares corresponding branches (or portions) of the two hierarchies to identify at least one different branch. The database system can return a response identifying the branch to the user so that the user understands specifically how the two database schemas differ.
[0022] These techniques advantageously allow a system to determine if two database schemas are different. Thus, the system can prevent data from being entered under different database schemas, thereby preventing problems such as data corruption. For example, if a new database schema is missing a column for a particular table defined in another database schema, then when the system attempts to enter that data from the other database schema into the new database schema, the data associated with that column will be lost. The techniques described in the disclosure can also be used to identify differences between database schemas. This allows a user to know how to update a database schema so that data can be entered correctly under that schema without corruption. Reference will now be made to Figure 1 We begin by discussing exemplary applications of these techniques.
[0023] Now refer to Figure 1 , a block diagram of system 100 is shown. System 100 includes a collection of components that can be implemented by hardware or a combination of hardware and software programs. In the illustrated embodiment, system 100 includes a database 110 having a database architecture 115 and a database system 120 having a snapshot 130, which is a snapshot of database 110. As further shown, snapshot 130 includes a hash tree 140 having a root hash value 145. In some embodiments, system 100 can be implemented in a manner different from that shown. As an example, system 100 can include multiple database systems 120 that interact with database 110.
[0024] In various embodiments, system 100 implements a platform service that allows users of the service to develop, run, and manage applications. As an example, system 100 can be a multi-tenant system that provides various functions to multiple users / tenants hosted by a multi-tenant system. Therefore, system 100 can execute software programs from various different users (e.g., providers and tenants of system 100), and provide code, web pages, and other data to users, databases (e.g., database 110), and other entities associated with system 100. As shown, system 100 includes database system 120, which interacts with database 110 to store and access data of users associated with system 100.
[0025] In various embodiments, database 110 is an aggregate of information organized in a manner that allows access, storage, and manipulation of information. Therefore, database 110 may include support software that allows database system 120 to perform operations (e.g., access, storage, etc.) on information in database 110. Database 110 may be implemented by a single storage device or multiple storage devices connected together on a network (e.g., a storage attachment network) and configured to store information redundantly to prevent data loss. In some embodiments, database 110 is implemented using a merge tree (LSM tree) with a multi-layer log structure. One or more of the layers may include database records that are cached in a memory buffer before being written to the layer maintained on the disk storage medium. These database records may correspond to rows in a database table defined by database architecture 115.
[0026] In various embodiments, database schema 115 is a collection of information that defines how data is organized in database 110 and how subsets of the data are related. That is, database schema 115 may be a blueprint that provides a logical grouping of objects within database 110, such as database tables, views, stored procedures, etc. Figure 2 As discussed in more detail, in various embodiments, database schema 115 (or a collection of database schemas 115) defines database tables maintained for database 110, including the attributes that make up those database tables. For example, database schema 115 may define a "customer" database table having attributes (or columns) such as first name, last name, email address, etc. Each of these attributes may be associated with a set of properties defined by database schema 115. For example, a "first name" attribute may be associated with a data type of "char".
[0027] In various embodiments, the database schema 115 defines relationships between objects (e.g., database tables) in the database 110. As an example, the database schema 115 may define a foreign key for a database table that links the database table to another database table by serving as a reference to the primary key of the other database table. The database schema 115 may also define a primary key for a database table that may be used to uniquely identify each record of the database table. In some embodiments, the database schema 115 defines additional metadata, such as database triggers that are executed when certain events related to database objects defined by the database schema 115 occur. As discussed in more detail below, the database system 120 uses various information defined in the database schema 115 (e.g., database tables, primary keys, etc.) to generate a hash tree 140.
[0028] In various embodiments, database system 120 performs various operations to manage database 110, including data storage, data retrieval, and data manipulation. In this way, database system 120 can provide database services to other systems that wish to interact with database 110. Such database services can include accessing and storing data maintained for the database table of database 110. As an example, an application server can issue a transaction request to database system 120, which includes specifying that a specific database record is written to database 110 for a transaction of a specific database table. As shown, database system 120 can receive snapshot request 125 to create snapshot 130. Snapshot request 125 can be received from another component of user or system 100, and can specify a specific tenant of system 100 for which snapshot 130 is to be created. In some instances, database system 120 can create snapshot 130 at periodic time intervals or in response to certain events (e.g., after a batch of records are written to database 110) except receiving snapshot request 125.
[0029] In various embodiments, snapshot 130 is a database construct that saves the state of part or all of database 110 at a specific point in time (typically, when database snapshot 130 is created). In some embodiments, snapshot 130 can be a tenant snapshot that saves the state associated with the corresponding tenant about database 110. For example, snapshot 130 can identify part or all of the tenant's data stored in database 110 at a specific point in time, rather than data associated with other tenants. Therefore, in various embodiments, snapshot 130 can be used to rebuild part or all of the data of database 110 that existed when snapshot 130 was created. Continuing with the previous example, database system 120 can restore database 110 to the state when it stored the tenant data that existed at the specific point in time when snapshot 130 was created. Due to this ability, snapshot 130 can be used to repair database 110 if data corruption occurs. Snapshot 130 can also be used for testing purposes-for example, entering the tenant's data into a database 110 associated with a test environment.
[0030] However, in various embodiments, the snapshot 130 does not identify the state of the database schema 115. That is, the snapshot 130 does not identify the metadata defined by the database schema 115. Therefore, when using the snapshot 130 to recreate the saved state, the database system 120 can use the database schema 115 that already existed when the snapshot 130 was created (that is, the current version). In some instances, the snapshot 130 can be stored for a long time. During this time, the database schema 115 can be changed-for example, new attributes can be added to an existing database table. In other instances, the snapshot 130 can be used to recreate the saved state for a database other than the database for which the snapshot 130 was created. For example, as part of a tenant migration service, the tenant's data can be moved from a first database to a second database. Therefore, the snapshot 130 can be created based on the first database. However, the second database 110 can have a database schema 115 different from the first database. Therefore, in any instance, if the database schema 115 being used is different from the database schema associated with the snapshot 130, the database system 120 may not be able to recreate the state saved by the snapshot 130.
[0031] In view of these shortcomings, the inventors of the present invention have realized that it is necessary to detect not only whether two database schemas are different, but also the specific parts of the schemas that are different. To solve this problem, the inventors have proposed to associate a schema hash tree with the database schema 115. The schema hash tree includes a root hash value that uniquely identifies a specific database schema 115, and can be compared with other root hash values to determine whether those other root hash values identify the same specific database schema 115 or other database schemas 115. Figure 1 As shown, the snapshot 130 of the database 110 includes a hash tree 140, so that when the snapshot 130 is used (for example, the tenant's data is input into the database 110), the database system 120 can determine whether the currently existing database schema 115 is different from the database schema for which the snapshot 130 was created.
[0032] In various embodiments, the hash tree 140 is a hierarchy of hash values that includes a root hash value 145 corresponding to the database schema 115. By comparing the hash value of the hash tree 140 with the hash value of another hash tree 140 corresponding to another database schema 115, the hash tree 140 can be used to determine whether and how its corresponding database schema 115 is different from the other database schema 115. As an example, if the root hash values 145 of the two hash trees 140 are different values, then their corresponding database schemas 115 are different. In various embodiments, if the two hash trees 140 are different, the database system 120 traverses the branches of each hash tree 140 and compares the corresponding hash values between each hash tree 140. (For examples of differences between branches, see Figure 3B Based on the difference in the hash values, the database system 120 can identify the specific parts that are different between the two database schemas 115. For example, the database system 120 can identify that the second database schema 115 defines additional attributes for the database tables defined by the two database schemas 115.
[0033] In various embodiments, the database system 120 creates a hash tree 140 by applying a hash function to various parts of the metadata defined in the database schema 115. In some embodiments, the database system 120 hashes the characteristics of each attribute together to create the initial layer of the hash tree 140. If the snapshot 130 is a tenant snapshot of a specific tenant of the database 110, the database system 120 can only hash the attributes of those tables associated with the specific tenant. Then, the database system 120 can apply the hash function to a set of hash values of the initial layer, where these hash values correspond to the same database table / object. As a result, the second layer of the hash tree 140 is created. The database system 120 can apply a hash function that receives all the hash values of the second layer as input to generate a root hash value 145. In various embodiments, after creating the hash tree 140, the database system 120 includes the created hash tree 140 with the snapshot 130. Therefore, when using snapshot 130, database system 120 can use hash tree 140 to determine whether current database schema 115 is different from the database schema 115 associated with snapshot 130 described above. In this way, database system 120 can prevent data corruption and other problems when using snapshot 130.
[0034] Now refer to Figure 2, a block diagram of an exemplary database architecture 115 and an exemplary database system 120 is shown. In the illustrated embodiment, the database architecture 115 includes a database table 210 having attributes 220, and the attributes 220 include characteristics 230. Also as shown, the database system 120 includes a snapshot generation engine 240 having a hash function 245. In some embodiments, the database architecture 115 and / or the database system 120 can be implemented in a manner different from that shown. As an example, the snapshot generation engine 240 can include multiple hash functions 245.
[0035] As previously mentioned, in various embodiments, the database schema 115 is a collection of information that provides structure to the data stored in the database 110. As shown, the database schema 115 defines database tables 210, attributes 220, and properties 230. In various embodiments, this information can be used to generate a hash tree 140 that uniquely identifies the database schema 115.
[0036] The database schema 115 defines a set of database tables 210 for organizing the data of the database 110. In various embodiments, the database tables 210 are composed of columns and rows, wherein each column defines a field of information and each row provides a record that may include a value for the field. In various embodiments, the columns of the database tables 210 correspond to the attributes 220 defined by the database schema 115. For example, a "product" database table 210 may include a "price" attribute 220, which is a column of the product database table 210. The product database table 210 may include other attributes 220, such as a "product ID" attribute 220, a "product name" attribute 220, a "product quantity" attribute 220, and the like. Thus, the records of the product database table 210 may define a product ID value, a product name value, a price value, and a product quantity value.
[0037] As further shown, the attribute 220 is associated with a set of characteristics 230. In various embodiments, the characteristics 230 identify characteristics of its corresponding attribute 220. For example, the characteristics 230 may include the name of the attribute (e.g., "product name"), the data type of the attribute (e.g., char), the maximum length of the data type (e.g., 20 characters), whether null values are accepted, whether the attribute is used as a primary key, etc. The characteristics 230 may be expressed in a text-based format that allows a hash function to be applied to some or all of the characteristics of a particular attribute 220.
[0038] In various embodiments, the snapshot generation engine 240 is a collection of software programs executable to create a snapshot 130 including a hash tree 140. When creating the hash tree 140, the snapshot generation engine 240 may initially retrieve the database schema 115 from the database 110 or another storage area. In various embodiments, to create the hash tree 140, the snapshot generation engine 240 provides one or more characteristics 230 of the attribute 220 to a hash function 245 for each attribute 220 to derive a hash value. (In various embodiments, the hash function 245 is a non-cryptographic hash function suitable for generating hash values. The hash function 245 may be, for example, a "MurmurHash" function.) The bottom layer of the hash tree 140 may include a collection of hash values, each hash value corresponding to a corresponding attribute 220 defined in the database schema 115.
[0039] After generating the bottom layer of hash tree 140, snapshot generation engine 240 can group each hash value of the bottom layer so that a given group includes a hash value associated with a corresponding database table 210. In some embodiments, snapshot generation engine 240 provides the hash value of the group to hash function 245 as input for each group to derive the hash value. The next layer of hash tree 140 can include a set of hash values, each hash value corresponding to a corresponding database table 210. In various embodiments, snapshot generation engine 240 provides the hash value of the next layer to hash function 245 as input to derive a root hash value 145. After generating snapshot 130, snapshot generation engine 240 can merge hash tree 140 into snapshot 130. As an example, hash tree 140 can be stored in the "fingerprint" field of snapshot 130. In some embodiments, only root hash value 145 can be included in snapshot 130.
[0040] Reference now Figure 3A , a block diagram of an exemplary hash tree 140 is shown. In the illustrated embodiment, the hash tree 140 includes a first layer having attribute hash values 310A-E, a second layer having database table hash values 320A and 320B, and a third layer having a root hash value 145. As further shown, the attribute hash values 310A-E are based on the characteristics 230A-O. In some embodiments, the hash tree 140 may be implemented in a manner different from that shown. As an example, the hash tree 140 may not include a second layer of hash values.
[0041] In various embodiments, attribute hash value 310 is a hash value corresponding to attribute 220. Attribute hash value 310 can be generated by applying hash function 245 (or a set of hash functions 245) to one or more characteristics 230 associated with the corresponding attribute 220 of the attribute hash value. As shown, attribute hash value 310A is generated by hashing characteristics 230A-C together. In various embodiments, specific characteristics 230 are not provided as inputs to hash function 245. For example, human-readable features such as attribute names or the order of object creation may not be inputs to hash function 245. Therefore, in this way, hash tree 140 can be insensitive to human-readable features.
[0042] In various embodiments, table hash value 320 is a hash value corresponding to database table 210. Table hash value 320 can be generated by applying hash function 245 (or a set of hash functions 245) to attribute hash value 310A of associated database table 210 corresponding to table hash value. As shown, table hash value 320A is generated by hashing attribute hash values 310A-C together. In various embodiments, root hash value 145 is a hash value that uniquely corresponds to database architecture 115. Root hash value 145 can be generated by applying hash function 245 (or a set of hash functions 245) to all table hash values 320 of appropriate hash tree 140. As shown, root hash value 145 is generated by hashing table hash values 320A and 320B together.
[0043] Because the top and middle levels of the hash tree 140 incorporate the hash values of their children as parameters, the database system 120 can identify not only whether a mismatch has occurred, but also where the mismatch has occurred. For example, if the characteristic 230A is modified, this will cause the attribute hash value 310A to be different, which in turn will cause the table hash value 320A to be different, and therefore cause the root hash value 145 to be different. Therefore, the database system 120 can identify that a particular branch includes a hash value that is different from the corresponding branch of another hash tree 140. In this way, the database system 120 can traverse to the end of the particular branch to identify which attribute 220 is different. In some cases, the database system 120 can identify branches that do not exist in other database schemas 115.
[0044] Now refer to Figure 3B, a block diagram of two exemplary hash trees 140 is shown. In the illustrated embodiment, hash tree 140A is different from hash tree 140B. Although hash trees 140A and 140B both include the same attribute hash value 310B and table hash value 320B as shown, hash tree 140A includes attribute hash value 310A, table hash value 320A, and root hash value 145A that are different from hash tree 140B, and hash tree 140B includes attribute hash value 310C, table hash value 320C, and root hash value 145B.
[0045] Database system 120 can first compare root hash value 145A of hash tree 140A with root hash value 145B of hash tree 140B. Since value 2843 of root hash value 145A is different from value 3492 of root hash value 145B, database architecture 115 of hash tree 140A is different from database architecture 115 of hash tree 140B. Database system 120 can determine how two database architectures 115 are different by analyzing corresponding branches (e.g., branch 330A and branch 330B) of two hash trees 140. As shown in the figure, attribute hash value 310A is generated by hashing characteristic 230A, and corresponding attribute hash value 310C is generated by hashing characteristic 230D. This causes table hash value 320A to be different from table hash value 320C, and root hash value 145A is different from root hash value 145B. As a result, as shown in the figure, branch 330A is different from branch 330B. After database system 120 determines that branch 330A is different from branch 330B, database system 120 can further determine what caused the difference. In this way, database system 120 can identify that feature 230A is different from feature 230C. Database system 120 can provide this information to users of system 100 for evaluation.
[0046] Now refer to Figure 4 , a block diagram of an example of a portion of a hash tree 140 stored in JavaScript Object Notation (JSON) format is shown. The JSON format is used as an example; the hash tree 140 can be stored in other hierarchical formats. The portion of the hash tree 140 shown corresponds to a database table 210. As shown, the database table 210 includes three attributes 220A-C. Therefore, three attribute hash values 310A-C are generated based on the characteristics 230 of these three attributes 220. The database table 210 can be associated with an index. In various embodiments, the characteristics of the index are hashed to form an index hash value 410. The index hash value 410 can be provided as an input to a hash function 245 together with an appropriate attribute hash value 310 to generate a table hash value 320. As an example, the index hash value 410 and the attribute hash values 310A-C can be hashed together to generate a table hash value 320.
[0047] Now refer to Figure 5 , a block diagram of an exemplary database system 120 capable of using snapshots 130 to store data in database 110 is shown. In the illustrated embodiment, database system 120 includes input engine 510. In some embodiments, database system 120 can be implemented in a manner different from that shown. For example, input engine 510 can generate hash tree 140B instead of receiving hash tree 140B.
[0048] In various embodiments, the input engine 510 is a collection of software programs that can be executed to store data in the database 110 using the received snapshot 130. For example, the input engine 510 can deserialize the snapshot 130 and use it to form a tenant - for example, by recreating the tenant's data within the database 110. As shown, the input engine 510 can receive a snapshot 130 having a hash tree 140A created based on the database schema 115A. The snapshot 130 can be received as part of an input request to input data into the database 110. In response to the request, the input engine 510 can cause the creation of a hash tree 140B based on the database schema 115B - for example, by communicating with the snapshot generation engine 240.
[0049] After hash tree 140B is created, input engine 510 can compare the root hash value 145 of hash tree 140A with the root hash value 145 of hash tree 140B. If the two root hash values 145 match, input engine 510 can continue to input the data identified by snapshot 130 into database 110. If the two root hash values 145 do not match, input engine 510 may not input the identified data into database 110. Input engine 510 can return a reply indicating the success or failure of the data identified by input snapshot 130. In some embodiments, input engine 510 can traverse both hash trees 140A and 140B to identify what hash values between the two hash trees are different. As part of the returned reply, input engine 510 can indicate how hash trees 140A and 140B are different--for example, by identifying which hash values are different.
[0050] Now refer to Figure 6, a flow chart of method 600 is shown. Method 600 is an embodiment of a method performed by a computer system (e.g., database system 120) for creating a snapshot (e.g., snapshot 130), the snapshot (e.g., snapshot 130) including a root hash value (e.g., root hash value 145), the root hash value identifying a first database schema (e.g., database schema 115) of a database (e.g., database 110). In various cases, method 600 can be performed by executing a set of program instructions stored on a non-transitory computer-readable medium. In some embodiments, method 600 includes more or fewer steps - for example, method 600 may include a step of the computer system storing the data set (captured in the snapshot) under the second database schema in response to determining that the root hash value of the first database schema matches the root hash value of the second database schema.
[0051] Method 600 begins at step 610, where a computer system receives a request (e.g., snapshot request 125) to create a snapshot for a data set stored in a database having a first database schema. The first database schema may define a plurality of database tables (e.g., database table 210). A given database table may be associated with a set of attributes (e.g., attributes 220). A given attribute is associated with a set of properties (e.g., properties 230). At step 620, in response to receiving the request, the computer system creates a snapshot for the data set.
[0052] At step 622, as part of creating the snapshot, the computer system generates a first hierarchy of hash values (e.g., hash tree 140) based on the first database schema, the first hierarchy of hash values including a first root hash value for the first database schema. As part of generating the first hierarchy of hash values, the computer system may apply a hash function (e.g., hash function 245) to one or more of a set of characteristics of a given attribute for a given attribute to derive hash values (e.g., attribute hash values 310) that form a portion of the first hierarchy of hash values. The particular set of hash values included in the first hierarchy of hash values may correspond to a particular database table in the plurality of database tables. As part of generating the first hierarchy of hash values, the computer system may apply the hash function to the particular set of hash values to derive hash values (e.g., database table hash values 320) that form a portion of a second hierarchy of hash values. The computer system may apply the hash function to the hash values included in the second hierarchy of hash values to derive the first root hash value.
[0053] At step 624, as part of creating the snapshot, the computer system includes a first hierarchy of hash values with the snapshot. The first hierarchy of hash values can be used to determine whether the first database schema is different from the second database schema. The snapshot can be used to enter the data set into a database associated with the sandbox environment.
[0054] In some embodiments, the computer system receives a second request to input the data set into a database having a second database schema based on the snapshot. In response to the second request, the computer system may generate a second hierarchy of hash values, the second hierarchy of hash values including a second root hash value for the second database schema. The computer system may compare the first root hash value to the second root hash value to determine whether the first database schema is different from the second database schema. In response to determining that the first database schema is different from the second database schema, the computer system may prevent the data set from being input into the database having the second database schema. The computer system may return a response to the second request that identifies at least one difference between the first database schema and the second database schema.
[0055] In some cases, the computer system may identify at least one different hash value between the first layer of the first series of hash values and the first layer of the second series of hash values. The at least one different hash value may indicate at least one difference between the first database schema and the second database schema. The second database schema may define attributes for a particular database table, while the first database schema does not define attributes for the particular database table.
[0056] Now refer to Figure 7 , a flowchart of method 700 is shown. Method 700 is an embodiment of a method for creating a snapshot (e.g., snapshot 130) performed by a computer system (e.g., database system 120), the snapshot having a root hash value (e.g., root hash value 145) of a database schema (e.g., database schema 115) identifying a database (e.g., database 110). In various cases, method 700 can be performed by executing a set of program instructions stored on a non-transitory computer-readable medium. In some embodiments, method 700 includes more or fewer steps-for example, method 700 may include a step in which a computer system sends a specific tenant snapshot to another computer system, the other computer system managing a second database having a second database schema. The other computer system is capable of generating a second hash tree based on the second database schema, and comparing the second root hash value of the second hash tree with the first root hash value included in the specific tenant snapshot to determine whether the first database schema is different from the second database schema.
[0057] Method 700 begins at step 710, where a computer system maintains a first database schema for a database storing data for a plurality of tenants of the computer system. The first database schema may include metadata defining a plurality of database tables (e.g., database table 210). As part of defining the plurality of database tables, the metadata may define a set of attributes (e.g., attribute 220) and a set of properties (e.g., property 230) for each of the attributes.
[0058] At step 720, the computer system receives a request (eg, snapshot request 125) to create a tenant snapshot (eg, snapshot 130) of data stored in a database. The tenant snapshot may be associated with a particular one of a plurality of tenants of the computer system.
[0059] At step 730, the computer system creates a tenant-specific snapshot that identifies the data of the specific tenant but does not identify metadata of the first database schema. In some cases, a specific database table among the plurality of database tables may store a portion of the data of the specific tenant and data of a second specific tenant among the plurality of tenants. Thus, the tenant-specific snapshot may identify a portion of the data of the specific tenant but does not identify the data of the second specific tenant stored in the specific database table.
[0060] At step 732 , as part of creating the specific tenant snapshot, the computer system generates a first hash tree (eg, hash tree 140 ) based on the first database schema, the first hash tree including a first root hash value corresponding to the first database schema.
[0061] At step 734, as part of creating the specific tenant snapshot, the computer system includes including a first hash tree with the specific tenant snapshot. The first root hash value can be used to determine whether the first database schema is different from the second database schema.
[0062] In various embodiments, the computer system may receive a request for causing a second database (which may be the database mentioned in step 710) to store data of a specific tenant identified by a specific tenant snapshot. The second database may store data under a third database architecture. Therefore, the computer system may determine whether the first database architecture is different from the third database architecture. As part of the determination process, the computer system may generate a second hash tree based on the third database architecture, the second hash tree including a second root hash value corresponding to the third database architecture. The computer system may compare the second root hash value with the first root hash value. In response to determining that the first database architecture is different from the third database architecture, the computer system may prevent the data of a specific tenant from being stored in the second database. In some embodiments, the computer system may traverse one or more branches of the first hash tree and one or more corresponding branches of the second hash tree to identify a difference set of hash values between the first hash tree and the second hash tree.
[0063] Exemplary Computer System
[0064] Now refer to Figure 8 , a block diagram of an exemplary computer system 800 in which system 100, database 110, and / or database system 120 may be implemented is shown. Computer system 800 includes a processor subsystem 880 coupled to system memory 820 and I / O interface 840 via an interconnect 860 (e.g., a system bus). I / O interface 840 is coupled to one or more I / O devices 850. Computer system 800 may be any of a variety of types of devices, including, but not limited to: a server system, a personal computer system, a desktop computer, a laptop or notebook computer, a mainframe computer system, a tablet computer, a handheld computer, a workstation, a network computer, a consumer device such as a mobile phone, a music player, or a personal digital assistant (PDA). Although for convenience, the computer system 800 is not limited to a server system, a personal computer system, a desktop computer, a laptop or notebook computer, a mainframe computer system, a tablet computer, a handheld computer, a workstation, a network computer, a consumer device such as a mobile phone, a music player, or a personal digital assistant (PDA). Figure 8 A single computer system 800 is shown in FIG. 8 , but system 800 may also be implemented as two or more computer systems operating together.
[0065] Processor subsystem 880 may include one or more processors or processing units. In various embodiments of computer system 800, multiple instances of processor subsystem 880 may be coupled to interconnect 860. In various embodiments, processor subsystem 880 (or each processor unit within 880) may include cache or other forms of onboard memory.
[0066] The system memory 820 can be used to store operational instructions that can be executed by the processor subsystem 880 to enable the system 800 to perform various operations described herein. The system memory 820 can be implemented using different physical memory media, such as hard disk memory, floppy disk memory, removable disk memory, flash memory, random access memory (RAM-SRAM, Edo RAM, SDRAM, DDR SDRAM, RAMBUS RAM, etc.), read-only memory (PROM, EEPROM, etc.), etc. The memory in the computer system 800 is not limited to main memory such as memory 820. On the contrary, the computer system 800 can also include other forms of memory, such as cache memory in the processor subsystem 880 and auxiliary memory (e.g., hard disk drive, storage array, etc.) on the I / O device 850. In some embodiments, these other forms of memory can also store program instructions that can be executed by the processor subsystem 880. In some embodiments, program instructions that implement the snapshot generation engine 240, hash function 245, and input engine 510 when executed can be included / stored in the system memory 820.
[0067] According to various embodiments, the EO interface 840 can be any of various types of interfaces configured to be connected to other devices and communicate with other devices. In one embodiment, the I / O interface 840 is a bridge chip (e.g., a south bridge) from the front end to one or more back-end buses. The EO interface 840 can be connected to one or more EO devices 850 via one or more corresponding buses or other interfaces. Examples of I / O devices 850 include storage devices (hard drives, optical drives, removable flash drives, storage arrays, SANs or their associated controllers), network interface devices (e.g., to local area networks or wide area networks) or other devices (e.g., graphics, user interface devices, etc.). In one embodiment, the computer system 800 is connected to a network via the network interface device 850 (e.g., configured to communicate via WiFi, Bluetooth, Ethernet, etc.).
[0068] Implementation of the subject matter of the present application includes but is not limited to the following embodiments 1 to 20.
[0069] 1. A method comprising:
[0070] receiving, by a computer system, a request to create a snapshot for a set of data stored in a database having a first database schema;
[0071] In response to receiving the request, the computer system creates a snapshot for the data set, wherein the creating includes:
[0072] generating a first hierarchy of hash values based on the first database schema, the first hierarchy of hash values comprising a first root hash value for the first database schema; and
[0073] A first hierarchy of hash values is included with the snapshot, wherein the first hierarchy of hash values can be used to determine whether the first database schema differs from the second database schema.
[0074] 2. The method according to embodiment 1, wherein the first database schema defines a plurality of database tables, a given database table is associated with a set of attributes, and a given attribute is associated with a set of characteristics.
[0075] 3. According to the method of embodiment 2, the first layer of generating the hash value includes:
[0076] For a given attribute, a hash function is applied to one or more of the set of characteristics of the given attribute to derive a hash value that forms part of a first layer of hash values.
[0077] 4. The method according to embodiment 3, wherein the specific set of hash values included in the first layer of hash values corresponds to a specific database table among the plurality of database tables, and wherein the first layer of generating hash values comprises:
[0078] A hash function is applied to a particular set of hash values to derive a hash value that forms part of a second layer of hash values.
[0079] 5. The method according to embodiment 4, wherein the first layer of generating the hash value comprises:
[0080] A hash function is applied to the hash values included in the second layer of hash values to derive a first root hash value.
[0081] 6. The method according to embodiment 1 further includes:
[0082] receiving, by the computer system, a second request to input the data set into a database having a second database schema based on the snapshot;
[0083] In response to the second request, the computer system generates a second hierarchy of hash values, the second hierarchy of hash values including a second root hash value for the second database schema; and the computer system compares the first root hash value and the second root hash value to determine whether the first database schema is different from the second database schema.
[0084] 7. The method according to embodiment 6, further comprising:
[0085] In response to determining that the first database schema is different from the second database schema, the computer system prevents the data set from being imported into a database having the second database schema; and
[0086] A response to the second request is returned by the computer system, the response identifying at least one difference between the first database schema and the second database schema.
[0087] 8. The method according to embodiment 7, wherein returning a response to the second request comprises:
[0088] At least one different hash value is identified, by a computer system, between a first layer of a first series of hash values and a first layer of a second series of hash values, wherein the first layer of the first series of hash values includes hash values of characteristics associated with attributes of a plurality of database tables defined in a first database schema, and wherein the at least one different hash value indicates at least one difference between the first database schema and the second database schema.
[0089] 9. The method of embodiment 7, wherein the second database schema defines attributes for a specific database table, and the first database schema does not define attributes for the specific database table.
[0090] 10. The method of embodiment 1, wherein the snapshot can be used to input the data set into a database associated with the sandbox environment.
[0091] 11. A non-transitory computer readable medium having program instructions stored thereon, the program instructions enabling a computer system to perform operations, the operations comprising:
[0092] receiving a request to create a snapshot for a collection of data stored in a database having a first database schema;
[0093] In response to receiving the request, creating a snapshot for the data set, wherein the creating includes:
[0094] generating a hierarchy of hash values based on the first database schema, the hierarchy of hash values including a root hash value for the first database schema; and
[0095] A hierarchy of hash values is included with the snapshot, wherein the hierarchy of hash values can be used to determine whether the first database schema is different from the second database schema.
[0096] 12. The non-transitory computer-readable medium of embodiment 11, wherein the first database schema defines a plurality of database tables, a given database table includes a set of attributes, a given attribute is associated with a set of characteristics, and wherein the hierarchy for generating the hash value includes:
[0097] For each attribute in the set of attributes of a database table included in the plurality of database tables, a hash function is applied to the set of characteristics associated with the attribute to produce a hash value that forms a portion of a first layer of hash values of the hierarchy.
[0098] 13. The non-transitory computer-readable medium of embodiment 12, wherein the hash values of the first layer of hash values correspond to two or more sets of hash values, and wherein generating the layer system of hash values comprises:
[0099] For each of the two or more sets of hash values, a hash function is applied to the set of hash values to produce a hash value that forms a portion of a second layer of the hierarchy of hash values.
[0100] 14. The non-transitory computer-readable medium according to embodiment 13, wherein the hierarchy of generating the hash value comprises:
[0101] The hash function is applied to the second layer of the series of hash values to produce a root hash value.
[0102] 15. The non-transitory computer readable medium of embodiment 11, wherein the operations further comprise:
[0103] receiving a request to store a data set under a second database schema; determining whether the first database schema and the second database schema are different by comparing a root hash value of the first database schema to a root hash value of the second database schema; and
[0104] In response to determining that the root hash value of the first database schema matches the root hash value of the second database schema, the data set is stored under the second database schema.
[0105] 16. A method comprising:
[0106] maintaining, by a computer system, a first database schema for a database for storing data for a plurality of tenants of the computer system, wherein the first database schema includes metadata defining a plurality of database tables;
[0107] A request is received by a computer system for creating a tenant snapshot of a database, wherein the tenant snapshot is associated with a specific tenant among a plurality of tenants; and the computer system is created, by the computer system, the specific tenant snapshot that identifies data of the specific tenant but does not identify metadata of a first database schema, wherein the creating includes:
[0108] generating a first hash tree based on the first database schema, the first hash tree comprising a first root hash value corresponding to the first database schema; and
[0109] A first hash tree is included with the particular tenant snapshot, wherein a first root hash value can be used to determine whether the first database schema is different from the second database schema.
[0110] 17. The method according to embodiment 16, further comprising:
[0111] receiving, by a computer system, a request for a second database to store data for a specific tenant identified by the specific tenant snapshot, wherein the second database stores the data under a third database schema; and
[0112] Determining, by a computer system, whether the first database schema is different from the third database schema includes:
[0113] generating a second hash tree based on the third database schema, the second hash tree including a second root hash value corresponding to the third database schema; and
[0114] comparing the second root hash value to the first root hash value; and in response to determining that the first database schema is different from the third database schema, the computer system blocks the data of the particular tenant from being stored in the second database.
[0115] 18. The method according to embodiment 17, further comprising:
[0116] A computer system traverses one or more branches of the first hash tree and one or more corresponding branches of the second hash tree to identify a set of difference hash values between the first hash tree and the second hash tree, wherein the set of difference hash values can be used to identify a difference between the first database schema and the third database schema.
[0117] 19. The method according to embodiment 16, further comprising:
[0118] The computer system sends the specific tenant snapshot to a second computer system that manages a second database including a second database schema, wherein the second computer system is capable of generating a second hash tree using the second database schema and comparing a second root hash value of the second hash tree with the first root hash value to determine whether the first database schema is different from the second database schema.
[0119] 20. A method according to embodiment 16, wherein a specific database table among multiple database tables stores a portion of the data of a specific tenant and data of a second specific tenant among multiple tenants, and wherein the specific tenant snapshot identifies a portion of the data of the specific tenant but does not identify the data of the second specific tenant stored in the specific database table.
[0120] Although specific embodiments have been described above, even in the case where only a single embodiment is described with respect to a particular feature, these embodiments are not intended to limit the scope of the disclosure. Unless otherwise stated, the examples of features provided in the disclosure are illustrative rather than restrictive. The above description is intended to encompass substitutions, modifications, and equivalents that are obvious to those skilled in the art having the benefit of the disclosure.
[0121] The scope of the disclosure includes any feature or combination of features disclosed herein (explicitly or implicitly), or any generalization thereof, whether or not it solves any or all of the problems raised herein. Accordingly, new claims may be formulated during the prosecution of the present application (or an application claiming priority thereto) to any such combination of features. In particular, features of dependent claims may be combined with features of the independent claims with reference to the appended claims, and features in individual independent claims may be combined in any appropriate manner and not merely in the specific combinations listed in the appended claims.
Claims
1. A method comprising: Receiving, by a computer system, a request to create a snapshot for a set of data stored in a database having a first database schema, the first database schema including metadata defining a set of database tables; In response to receiving the request, the computer system creates the snapshot for the data set, wherein the creating comprises: generating the snapshot of the data set, wherein the snapshot does not identify the metadata included in the first database schema; generating, based on the metadata of the first database schema, a first hierarchy of hash values, the first hierarchy of hash values comprising a first root hash value for the first database schema, wherein the first hierarchy of hash values comprises a first layer of hash values derived from characteristics of attributes of the set of database tables; and A first hierarchy of hash values is included with the snapshot, wherein the first hierarchy of hash values enables a system to determine whether the set of data of the snapshot can be entered into another set of database tables defined by a second database schema.
2. The method of claim 1, wherein a given database table is associated with a set of attributes, and wherein a given attribute is associated with a set of properties.
3. The method according to claim 2, wherein generating the first hierarchy of the hash value comprises: For the given attribute, a hash function is applied to one or more of the set of characteristics of the given attribute to derive an attribute hash value, the attribute hash value forming part of a first layer of hash values.
4. The method according to claim 3, wherein the specific attribute hash value set included in the first layer of the hash value corresponds to a specific database table in the database table set, wherein generating the first layer of the hash value comprises: A hash function is applied to the particular set of attribute hash values to derive a table hash value that forms part of a second layer of hash values.
5. The method of claim 4, wherein the particular database table is associated with an index, wherein characteristics of the index are hashed to form an index hash value, and wherein the index hash value is hashed together with the particular set of attribute hash values to form the table hash value.
6. The method according to claim 4, wherein generating the first hierarchy of hash values comprises: A hash function is applied to the table hash values included in the second layer of hash values to derive the first root hash value.
7. The method according to any one of claims 1 to 5, further comprising: receiving, by the computer system, a second request to input the data set into a database having the second database schema based on the snapshot; In response to the second request, the computer system generates a second hierarchy of hash values, the second hierarchy of hash values including a second root hash value for the second database schema; and The first root hash value is compared to the second root hash value by the computer system to determine whether the first database schema is different from the second database schema.
8. The method according to claim 7, further comprising: In response to determining that the first database schema is different from the second database schema, the computer system prevents the data set from being entered into the database having the second database schema; and A response to the second request is returned by the computer system, the response identifying at least one difference between the first database schema and the second database schema.
9. The method of claim 8, wherein returning the response to the second request comprises: At least one different hash value between a first layer of the first series of hash values and a first layer of the second series of hash values is identified by the computer system, and wherein the at least one different hash value indicates at least one difference between the first database schema and the second database schema.
10. A computer system comprising: at least one processor; and A memory having program instructions executable by the at least one processor to perform the method according to any one of claims 1 to 9 stored therein.
11. A non-transitory computer readable medium having program instructions stored thereon, the program instructions being capable of causing a computer system to perform operations including: receiving a request to create a snapshot for a set of data stored in a database having a first database schema, the first database schema including metadata defining a set of database tables; In response to receiving the request, creating the snapshot for the data set, wherein the creating comprises: generating the snapshot of the data set, wherein the snapshot does not identify the metadata included in the first database schema; generating, based on the metadata of the first database schema, a hierarchy of hash values including a root hash value for the first database schema, wherein the hierarchy of hash values includes a first layer of hash values of characteristics derived from attributes of the set of database tables; and The hierarchy of hash values is included with the snapshot, wherein the hierarchy of hash values enables a system to determine whether the set of data of the snapshot can be entered into another set of database tables defined by a second database schema.
12. The non-transitory computer-readable medium of claim 11, wherein a given database table includes a set of attributes, wherein a given attribute is associated with a set of characteristics, wherein generating the hierarchy of hash values includes: For each attribute in the attribute set of database tables included in the set of database tables, a hash function is applied to the set of characteristics associated with the attribute to produce a hash value that forms part of a first layer of hash values of the hierarchy.
13. The non-transitory computer-readable medium of claim 12, wherein the hash values in the first layer of hash values correspond to two or more sets of hash values, and wherein generating the hierarchy of hash values comprises: For each of the two or more sets of hash values, a hash function is applied to the set of hash values to produce a hash value that forms a portion of a second layer of the hierarchy of hash values.
14. The non-transitory computer-readable medium of claim 13, wherein generating the hierarchy of hash values comprises: A hash function is applied to a second level of the hash values of the hierarchy to produce the root hash value.
15. The non-transitory computer readable medium of any one of claims 11 to 14, wherein the operations further comprise: receiving a request to store the data set in the another set of database tables defined by the second database schema; determining whether the first database schema and the second database schema are different by comparing a root hash value of the first database schema with a root hash value of the second database schema; and In response to determining that the root hash value of the first database schema matches the root hash value of the second database schema, storing the set of data in the other set of database tables.
Citation Information
Patent Citations
System and method for managing deduplication using checkpoints in a file storage system
CN104641365A
Creating match cohorts and exchanging protected data using blockchain
CN110582987A