Data storage method, data query method and related device
Patent Information
- Application Number
- CN202210969865.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-12
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-08-12
AI Technical Summary
如此,数据查询的复杂性较大,数据查询效率较低
[0017]第六方面,本申请提供一种计算机可读存储介质,当存储介质中的指令由电子设备的处理器执行时,使得电子设备能够执行如第一方面或第二方面所提到的方法。
Smart Images

Figure CN116126850B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of database technology, and in particular to a data storage method, a data query method, and related apparatus. Background Technology
[0002] As businesses grow, the volume of user data and business data becomes enormous. When faced with large amounts of data storage, businesses often use database sharding and table partitioning based on time, region, and identity ID.
[0003] In some scenarios, when large amounts of data are stored using database sharding and table partitioning, data queries require cross-database and cross-table searches, followed by aggregation of the query results from each database and table. This results in significant complexity and low efficiency in data querying. Summary of the Invention
[0004] This application provides a data storage method, a data query method, and related apparatus, which reduces the complexity of data query and improves the efficiency of data query.
[0005] In a first aspect, embodiments of this application provide a data storage method, comprising: determining at least one sub-database in a source database that stores first application data generated by a target application, and at least one sub-table in the at least one sub-database that stores the first application data, wherein the source database includes multiple sub-databases, each sub-database includes multiple sub-tables, and the source database is used to store application data generated by different applications in multiple sub-tables of the multiple sub-databases; obtaining first target data of the target application in at least one sub-table in at least one sub-database of the source database, wherein the first target data is at least a portion of the data in the first application data; determining, according to a preset mapping relationship between database tables in the source database and database tables in the target database, a target single database corresponding to at least one sub-database in the target database, and a target single table corresponding to at least one sub-table in the target single database, wherein the target database includes at least one single database and at least one single table; and synchronizing the first target data to the target single table in the target single database for storage.
[0006] In a second aspect, embodiments of this application provide a data query method, comprising: receiving a data query request for a target application, wherein the data requested in the data query request is data stored according to the data storage method mentioned in the first aspect; in response to the data query request, querying first target data of the target application from a target single table of a target single database in a target database, wherein the first target data is data in at least one table of at least one sub-database in a source database, the source database stores first application data generated by the target application, the source database includes multiple sub-databases, each sub-database includes multiple tables, the source database is used to store application data generated by different applications in multiple tables of multiple sub-databases, the first target data is synchronized from the source database to the target database, the target single database has a mapping relationship with at least one sub-database, the target single table has a mapping relationship with at least one table, and the target database includes at least one single database and at least one single table.
[0007] Thirdly, embodiments of this application provide a data storage device, including:
[0008] The determination module is used to determine at least one sub-database in the source database that stores the first application data generated by the target application, and at least one sub-table in the at least one sub-database that stores the first application data. The source database includes multiple sub-databases, and each sub-database includes multiple sub-tables. The source database is used to store application data generated by different applications in multiple sub-tables of multiple sub-databases.
[0009] The acquisition module is used to acquire the first target data of the target application in at least one table of at least one sub-database in the source database, wherein the first target data is at least a portion of the data in the first application data;
[0010] The determination module is also used to determine, based on the preset mapping relationship between the database tables in the source database and the database tables in the target database, at least one sub-database of the source database corresponds to a target single database in the target database, and at least one sub-table corresponds to a target single table in the target single database. The target database includes at least one single database and at least one single table.
[0011] The storage module is used to synchronize the first target data to the target single table in the target single database for storage.
[0012] Fourthly, embodiments of this application provide a data query device, including:
[0013] The receiving module is used to receive data query requests for the target application. The data query request requests that the data be stored in accordance with the data storage method mentioned in the first aspect above.
[0014] The query module is used to respond to the data query request and query the first target data of the target application from the target single table of the target single database of the target database. The first target data is synchronized from the source database to the target database. The source database includes multiple sub-databases, each sub-database includes multiple sub-tables, and the first target data is stored in at least one sub-table of at least one sub-database in the source database. The target single database has a mapping relationship with the at least one sub-database, and the target single table has a mapping relationship with the at least one sub-table. The target database includes at least one single database and at least one single table.
[0015] Fifthly, this application provides an electronic device, comprising:
[0016] A processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute instructions to implement the methods mentioned in the first or second aspect.
[0017] Sixthly, this application provides a computer-readable storage medium that, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform the methods mentioned in the first or second aspect.
[0018] As can be seen, synchronizing the data of the target application stored in the sharded and partitioned tables of the source database to the target single database and single table in the target database means that the data of the target application in the sharded and partitioned tables of the source database is synchronized to the same database and table in the target database for storage. By storing the target application's data in the target database in a single-database, single-table manner, subsequent queries for the target application's data can be performed directly from that target database. Because the target application's data is stored in the same database and table, it can be retrieved directly from that database and table, eliminating the need for cross-database and cross-table queries and aggregation in the source database. This reduces the complexity of data queries and improves data query efficiency. Attached Figure Description
[0019] The accompanying drawings, which are included to provide a further understanding of this specification and form part of this specification, illustrate exemplary embodiments and are used to explain this specification, but do not constitute an undue limitation thereof. In the drawings:
[0020] Figure 1 This is a schematic diagram of the data storage structure in the related technologies provided in the embodiments of this application;
[0021] Figure 2 A flowchart illustrating a data storage method provided in an embodiment of this application;
[0022] Figure 3This application provides a schematic diagram of a data synchronization structure.
[0023] Figure 4 A flowchart illustrating a data query method provided in an embodiment of this application;
[0024] Figure 5 This is a schematic diagram of a data query structure provided in an embodiment of this application;
[0025] Figure 6 This is a schematic diagram of the structure of a data storage device provided in an embodiment of this application;
[0026] Figure 7 This is a schematic diagram of the structure of a data query device provided in an embodiment of this application;
[0027] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.
[0029] The terms "first," "second," etc., used in this specification and claims are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, in this specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0030] As mentioned earlier, large datasets are typically stored using database sharding and table partitioning. Common sharding rules include partitioning by time, ID, and region. For example, in an application, after generating a large amount of data, the data is stored in different databases and tables. During queries, the sharding rules are applied to different databases and tables, and the results are then aggregated. For instance, ... Figure 1As shown, the target application generates a large amount of data, which is stored in three database shards: Database 1, Database 2, and Database 3. Each shard contains multiple tables. Querying data from the target application requires cross-database, cross-table queries from the shards in Database 1, Database 2, and Database 3, followed by aggregation to obtain the final query result. This cross-database, cross-table query and aggregation method is not only complex but also inefficient, and improper operation can even result in no query results.
[0031] Therefore, this application aims to provide a data storage scheme and a subsequent query scheme for data stored based on the data storage scheme. For the data storage method, the following steps are taken: First, a source database is determined to contain at least one sub-database storing first application data generated by a target application, and at least one sub-table storing the first application data within that sub-database. The source database includes multiple sub-databases, each containing multiple sub-tables. The source database stores application data generated by different applications in multiple sub-tables within these sub-databases. First target data is obtained from at least one sub-table in at least one sub-database of the source database. The first target data is at least a portion of the first application data. Based on a preset mapping relationship between tables in the source database and tables in the target database, a target single database corresponding to at least one sub-database in the source database and a target single table corresponding to at least one sub-table in the target single database are determined. The target database includes at least one single database and at least one single table. The first target data is then synchronized to the target single table in the target single database for storage. In this way, the data stored in the target application's sharded and partitioned tables in the source database is synchronized to the target single database and single table in the target database. That is, the data of the target application in the source database is synchronized to the same database and table in the target database for storage. By storing the target application's data in the target database in a single-database, single-table manner, subsequent queries for the target application's data can be performed directly from that target database. Because the target application's data is stored in the same database and table, it can be retrieved directly from that database and table, eliminating the need for cross-database and cross-table queries and aggregation in the source database. This reduces the complexity of data queries and improves data query efficiency.
[0032] For the data query method, the data to be queried is the data stored according to the data storage method described above. First, a data query request for the target application is received. In response to the data query request, the first target data of the target application is retrieved from the target table of the target single database in the target database. The first target data is data from at least one table in at least one sub-database of the source database. The source database stores the first application data generated by the target application. The source database includes multiple sub-databases, each containing multiple tables. The source database is used to store application data generated by different applications in multiple tables across multiple sub-databases. The first target data is synchronized from the source database to the target database. The target single database has a mapping relationship with at least one sub-database, and the target single table has a mapping relationship with at least one table. The target database includes at least one single database and at least one single table. In this way, the data of the target application can be directly queried from the same database and table in the target database, eliminating the need for cross-database and cross-table queries and aggregation from the source database, reducing the complexity of data queries and improving data query efficiency.
[0033] It should be understood that the data storage methods provided in the embodiments of this application can all be executed by electronic devices or by software installed in electronic devices, specifically by terminal devices or server devices.
[0034] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0035] Please refer to Figure 2 This is a flowchart illustrating a data storage method according to an embodiment of this specification, applied to an electronic device. The method may include:
[0036] Step S201: Determine at least one sub-database in the source database that stores the first application data generated by the target application, and at least one sub-table in the at least one sub-database that stores the first application data.
[0037] The source database consists of multiple shards, and each shard contains multiple tables. The source database is used to store application data generated by different applications in multiple tables of multiple shards.
[0038] Specifically, the source database can be a relational database capable of database sharding and table partitioning. It uses a relational model to organize data, storing it in the database in the form of rows and columns. When the data volume becomes too large, it is necessary to expand into multiple databases, and at least one database needs to be expanded into multiple tables to store the data. When expanding databases and tables, they can be expanded according to time, region, and ID. A table refers to any one of the multiple tables in a database shard, which consists of rows and columns. A database shard refers to any one of the multiple databases in the source database. A group of tables forms a database shard. For example, the source database type can be an Oracle database, a MySQL database, a DB2 database, a PostgreSQL database, a Microsoft SQL Server database, etc.
[0039] The target application is one that generates a large amount of application data and requires database sharding and table partitioning to store this application data.
[0040] Step S203: Obtain the first target data of the target application in at least one table of at least one sub-database in the source database.
[0041] The first target data is at least a portion of the data in the first application data. For example, it may be a part of the data in the first application data or all of the data in the first application data, i.e., part of the application data or all of the application data, without any limitation here.
[0042] Step S205: Based on the preset mapping relationship between the database tables in the source database and the database tables in the target database, determine the target single database corresponding to at least one sub-database of the source database, and the target single table corresponding to at least one sub-table in the target single database.
[0043] The target database includes at least one single database and at least one single table. The target single database can be any one of the at least one single databases in the target database, and the target single table can be any one of the at least one single table in the target single database.
[0044] Specifically, as mentioned above Figure 1As described, for the target application, the data it generates will be stored in at least one shard and at least one table in the source database. For the shards and tables storing the target application's data, a mapping relationship can be established between at least one shard and at least one table corresponding to the target application in the source database and the databases and tables in the target database. This mapping relationship is then saved through a Distributed World Wide Web (Web) Operational Data Store (ODS). Specifically, if the source database stores the target application's data in one shard, it maps to one database in the target database; if the source database stores the target application's data in one table, it maps to one table in the target database; if the source database stores the target application's data in multiple shards, it maps to one database in the target database; and if the source database stores the target application's data in multiple tables, it maps to one table in the target database. It is worth noting that a single table in the target database mapped to by the same target application has a corresponding relationship with a single database in the target database; that is, a single table in the target database mapped to by at least one shard of the same target application belongs to a single database in the target database mapped to by at least one shard of the same target application.
[0045] More specifically, when establishing mapping relationships, mapping relationships can be established according to the names of databases and tables, or according to the types of databases and tables. In the case of establishing mapping relationships according to the names of database tables in the source database and the names of database tables in the target database, in one possible implementation, determining the target single database in the target database corresponding to at least one sub-database and the target single table in the target database corresponding to at least one sub-table, based on the preset mapping relationship between database tables in the source database and database tables in the target database, includes: obtaining the database name of at least one sub-database and the table name of at least one sub-table; determining the target database name of the single database in the target database corresponding to the database name of at least one sub-database and the target table name of the single table in the target single database corresponding to the table name of at least one sub-table, based on the preset mapping relationship between database tables in the source database and database tables in the target database; determining the target single database based on the target database name; and determining the target single table based on the target table name. Specifically, each shard and table in the source database has a unique database name and table name, and the databases and tables in the target database also have corresponding database names and table names. The names of the shards and tables in the source database are mapped to the names of the databases and tables in the target database. It is worth noting that the names of multiple shards in the source database for the same target application are mapped to the name of a single database in the target database, and the names of multiple shards in the source database for the same target application are mapped to the name of a single table in the target database. Furthermore, the table mapped to by the target application in the target database belongs to that single database. Thus, because the names of the databases and tables are unique, establishing the mapping relationship between the source and target databases based on database and table names is more reliable, resulting in higher accuracy and reliability of data synchronization.
[0046] After establishing the mapping relationship between the source database tables and the target database tables, the target database is determined according to this mapping relationship, corresponding to at least one shard of the target application's data and at least one target table of the target application's data. The target database can be a distributed database, such as TiDB. That is, at least one shard of the target application in the source database corresponds to a database in TiDB, and at least one table of the target application in the source database corresponds to a table in TiDB.
[0047] Step S207: Synchronize the first target data to the target single table in the target single database for storage.
[0048] Specifically, synchronizing the first target data to the target table within the target database refers to transferring the first target data from the source database to the target database's target table and database. The method for synchronizing the first target data to the target table within the target database varies depending on the type of the first target data. Specifically: If the first target data is full data, the DM tool can be used to synchronize it all at once to the target table within the target database's target table. DM is a data migration tool that supports full migration and incremental synchronization, allows filtering of tables and operations, and supports merging and migrating sharded databases and tables. If the first target data is incremental data, the Otter tool can be used to synchronize it to the target table within the target database's target table. Otter is an open-source database incremental data synchronization tool that can be used for incremental data synchronization scenarios. It supports sharding and table merging, mapping table fields between the source and target databases, and ignoring deletion events in the source database, preventing deletion events from being synchronized to the target database. Furthermore, during the process of synchronizing the first target data to the target table in the target database, the structure of the sharded tables in the source database may change. Therefore, it is necessary to synchronize the modification operations of the tables in the source database to the tables in the target database, so as to modify the structure of the tables in the target database to correspond with the results of the tables in the source database, thereby ensuring the consistency of the table structure between the source database and the target database, and further ensuring the stability of data synchronization from the source database to the target database.
[0049] For example, taking TiDB, a distributed database as the target database, as an example... Figure 3 As shown, the data generated by the target application is stored in three databases: Database 1, Database 2, and Database 3. There are three ways to synchronize the data from these three databases to the TiDB distributed database. The first method is to use the DM tool to synchronize all data from these three databases to the TiDB distributed database. The second method is to use the Otter tool to synchronize incremental data from these three databases to the TiDB distributed database. The third method is to synchronize the DDL data from these three databases to the TiDB distributed database. DDL data refers to modification statements that modify the table structure of the tables in the three databases of the source database, including adding table columns, deleting table columns, deleting tables, and adding tables.
[0050] In this embodiment, the target application's data is stored in the target database in a single-database, single-table manner. When querying the target application's data later, it can be queried directly from the target database. Since the target application's data is stored in the same database and table, the target application's data can be queried directly from that database and table without having to query and aggregate data across databases and tables in the source database, thus reducing the complexity of data query and improving the efficiency of data query.
[0051] In one possible implementation, the first target data includes incremental data. Synchronizing the first target data to a target table in the target single database for storage includes: determining a first table field of at least one sub-table storing the incremental data, and a second table field of the target single table, wherein the first table field and the second table field are table fields that indicate the same target table header in different forms; and synchronizing at least a portion of the incremental data corresponding to the first table field to the table column corresponding to the second table field in the target single table according to a preset mapping relationship between the first table field and the second table field.
[0052] Specifically, for tables in the source and target databases, each table consists of at least one table field. A table field refers to the name of the data in that table column, and this name is called the target table header. For example, for a table that stores product information, the table fields include, but are not limited to, price, color, model, etc. The first table field refers to the target table header of the sub-table in the source database that stores incremental data, and the second table field refers to the target table header of the target single table in the target database that stores data. The data stored in the table column of the target table header corresponding to the first table field can be at least a portion of the data in the incremental data.
[0053] For incremental data, the meaning of the first table field in the column containing data stored in at least one sub-table in the source database and the second table field in a target single table that has a mapping relationship with the at least one sub-table in the target database may be the same, but the first table field and the second table field may be presented in different forms. For example, if price is used as the target table header, the first table field may be the Chinese price, and the second table field may be the English price (price). To avoid the inconsistency between the first table field in the source database and the second table field in the target database, which represent the same meaning, this embodiment establishes a mapping relationship between the first table field and the second table field in the source database that represent the same meaning but have different forms. Specifically, the Otter tool can be used to establish the mapping relationship between the first table field and the second table field. During subsequent data synchronization, the forms of table fields representing the same meaning can be unified. That is, during the data synchronization process, the data corresponding to the first table field in the source database is synchronized to the table column corresponding to the second table field in the target database, thereby ensuring the consistency and stability of the synchronized data between the source database and the target database. Furthermore, during the process of synchronizing the first target data to the target table in the target database, if the first table field in the incremental data of the source database and the second table field in the target table, representing the same meaning, are different, the first table field can be modified to the second table field or vice versa to unify the first and second table fields in the two databases. Then, at least a portion of the data corresponding to the first table field in the source database is synchronized to the column of the target table in the target database, thereby ensuring the consistency and stability of the synchronized data between the source and target databases. It is worth noting that during the process of synchronizing at least a portion of the incremental data from the source database to the target database, if at least a portion of the incremental data already exists in the target database, this portion of data in the target database is directly overwritten to avoid duplicate storage and unnecessary resource consumption in the target database.
[0054] In one possible implementation, synchronizing the first target data to a target single table in the target single database for storage includes: receiving modification statements that modify the table structure of at least one sub-table in the source database and the table structure of the target single table, wherein the table structure includes the number of rows and columns of the table; modifying the table structure of at least one sub-table in the source database and the table structure of the target single table in the target database according to the modification statements; determining the first target data from the at least one sub-table after the table structure modification; and synchronizing the first target data to the target single table after the table structure modification.
[0055] Specifically, modification statements refer to SQL statements that modify the table structure. These include statements that delete tables, create tables, add columns, and delete columns. Based on these modification statements, columns are added, columns are deleted, or tables are removed from at least one sub-table in the source database and the target single table in the target database. The first target data is then determined from the modified sub-table and synchronized to the modified target single table. Therefore, if the table structure of at least one sub-table in the source database is inconsistent with the target single table in the target database, data synchronization will be interrupted. Thus, it is necessary to uniformly modify the results of the tables in both the source and target databases to ensure consistency between the table structures of at least one sub-table in the source database and the target single table in the target database, achieving continuous data synchronization and guaranteeing its continuity and stability.
[0056] During data synchronization, if the number of columns in the target database is greater than the total number of columns in at least one sub-table in the source database, data synchronization can continue. If the number of columns in the target database is less than the total number of columns in at least one sub-table in the source database, data synchronization may be interrupted. Therefore, when a modification statement indicates adding a column, at least one column is first added to the target table in the target database, and then at least one column is added to at least one sub-table in the source database. Conversely, when a modification statement indicates deleting a column, at least one column is first deleted from at least one table in the source database, and then at least one column is deleted from the target table in the target database. This ensures that the number of columns in the target table in the target database is greater than the total number of columns in at least one sub-table in the source database, guaranteeing the continuity of data synchronization. The total number of columns in at least one sub-table in the source database refers to the total number of columns in all sub-tables in the source database corresponding to the same target application.
[0057] In one possible implementation, after synchronizing the first target data to the target single table in the target single database for storage, the method further includes: receiving a first input to at least a portion of the data in the first target data in the source database; in response to the first input, deleting at least a portion of the data in the first target data in the source database; and retaining at least a portion of the data in the first target data in the target database according to a pre-configured synchronization rule, wherein the synchronization rule indicates that the first input in the source database is not synchronized to the target database.
[0058] Specifically, the first input refers to a deletion request that removes at least a portion of the target data from the source database. A pre-configured synchronization rule allows the deletion operation of at least a portion of the target data in the source database to not be synchronized to the target database. That is, deleting data in the source database does not correspondingly delete data in the target database. This synchronization rule is configured and stored by Web ODS. In this way, when data is deleted from the source database, the data can be continuously stored in the target database, ensuring the stability of data storage in the target database.
[0059] Following the above scheme, the target application's data is stored in the target database. Data queries for the target application can then be performed directly from the target database. Since the target application's data is stored in the same database and table, the complexity of data queries is reduced, and the efficiency of data queries is improved. The following section will combine... Figure 4 This application provides a data query method based on an embodiment.
[0060] Please refer to Figure 4 This is a flowchart illustrating a data query method provided in one embodiment of this specification, applied to an electronic device. The method may include:
[0061] Step S401: Receive a data query request for the target application.
[0062] Specifically, a data query request refers to a request to query data from the target database for the target application. In other words, the data requested in the data query request is data stored according to the data storage method mentioned in the above embodiments. For example, such as... Figure 5 As shown in the above embodiment, the full data, incremental data, and table structure modification data of the target application app1 are synchronized from database shard 1, database shard 2, and database shard 3 to the target database TiDB. When querying data of the target application app1, the target application app1 can directly initiate a data query request to the target database TiDB.
[0063] Step S403: In response to the data query request, query the first target data of the target application from the target single table of the target single database of the target database.
[0064] Specifically, the first target data of the target application is queried from the target table of the target single database in the target database. The first target data is synchronized from the source database to the target database. The source database includes multiple sub-databases, and each sub-database includes multiple sub-tables. The first target data is data stored in at least one sub-table in at least one sub-database of the source database. The target single database has a mapping relationship with the at least one sub-database, and the target single table has a mapping relationship with the at least one sub-table. The target database includes at least one single database and at least one single table.
[0065] Specifically, the data query method provided in this application embodiment queries the data of the target application from the target database. The data in the target database is the data stored in accordance with the data storage method mentioned in the above embodiment. Therefore, the data query method provided in this application embodiment and the data storage method mentioned in the above embodiment can be referred to each other where they are the same or similar. This application embodiment will not be described again here.
[0066] Specifically, when the data query request type is an OLAP query request, the Tiflash component is used to query the target data from the target database.
[0067] Specifically, OLAP queries refer to queries that require computation. Adding a Tiflash component can help with the computation. The Tiflash component will perform columnar storage on the data to be queried again, and increase the number of storage nodes to share the computation task, thereby improving the response speed to OLAP query requests.
[0068] The technical solution disclosed in this application first receives a data query request for a target application. In response to the data query request, it queries the first target data of the target application from a target table in a target single database within the target database. The first target data is at least a portion of the data from at least one table in at least one sub-database in the source database. This at least portion of the data is synchronized from the source database to the target database. The target single database has a mapping relationship with at least one sub-database, and the target single table has a mapping relationship with at least one sub-table. The target single database includes the target single table. In this way, the data of the target application can be directly queried from the same database and table in the target database, eliminating the need for cross-database and cross-table data queries and aggregation from the source database, reducing the complexity of data queries and improving data query efficiency.
[0069] In addition, with the above Figure 2 Corresponding to the data storage method shown, this application also provides a data storage device. Figure 6 This is a schematic diagram of the structure of a data storage device 600 provided in an embodiment of this application, including:
[0070] The determining module 601 is used to determine at least one sub-database in the source database that stores the first application data generated by the target application, and at least one sub-table in the at least one sub-database that stores the first application data. The source database includes multiple sub-databases, and each sub-database includes multiple sub-tables. The source database is used to store application data generated by different applications in multiple sub-tables of multiple sub-databases.
[0071] The acquisition module 602 is used to acquire the first target data of the target application in at least one table of at least one sub-database in the source database, wherein the first target data is at least a portion of the data in the first application data;
[0072] The determining module 601 is further configured to determine, based on the preset mapping relationship between the database tables in the source database and the database tables in the target database, at least one sub-database of the source database corresponds to a target single database in the target database, and at least one sub-table corresponds to a target single table in the target single database. The target database includes at least one single database and at least one single table.
[0073] Storage module 603 is used to synchronize the first target data to the target single table in the target single database for storage.
[0074] In one possible implementation, the determining module 601 is further configured to obtain the database name of at least one sub-database and the table name of at least one sub-table; based on a preset mapping relationship between databases and tables in the source database and databases and tables in the target database, determine the target database name of the single database corresponding to the database name of at least one sub-database, and the target table name of the single table in the target single database corresponding to the table name of at least one sub-table; determine the target single database based on the target database name, and determine the target single table based on the target table name.
[0075] In one possible implementation, the first target data includes incremental data. The storage module 603 is further configured to determine a first table field of at least one sub-table storing the incremental data, and a second table field of the target single table, wherein the first table field and the second table field are table fields that indicate the same target table header in different forms. According to a preset mapping relationship between the first table field and the second table field, at least a portion of the incremental data corresponding to the first table field is synchronized to the table column corresponding to the second table field in the target single table for storage.
[0076] In one possible implementation, the storage module 603 is further configured to receive modification statements that modify the table structure of at least one sub-table in the source database and the table structure of the target single table, wherein the table structure includes the number of rows and columns of the table; modify the table structure of at least one sub-table in the source database and the table structure of the target single table in the target database according to the modification statements; determine the first target data from the at least one sub-table whose table structure has been modified, and synchronize the first target data to the target single table whose table structure has been modified.
[0077] In one possible implementation, storage module 603 is further configured to, when a modification statement indicates the addition of a table column, first add at least one table column to a target single table in the target database, and then add at least one table column to at least one table in the source database.
[0078] In one possible implementation, storage module 603 is further configured to, when a modification statement indicates that a table column should be deleted, first delete at least one table column of at least one table in the source database, and then delete at least one table column of the target single table in the target database.
[0079] In one possible implementation, the system further includes: a receiving module for receiving a first input to at least a portion of the data in the first target data in the source database; a deletion module for deleting at least a portion of the data in the first target data in the source database in response to the first input; and a retention module for retaining at least a portion of the data in the first target data in the target database according to a pre-configured synchronization rule, wherein the synchronization rule indicates that the first input in the source database is not synchronized to the target database.
[0080] Obviously, the data storage device in the embodiments of this application can be used as described above. Figure 2 The entity executing the data storage method shown can therefore realize the data storage method in Figure 2 The functions achieved are the same and have the same beneficial effects. Since the principles are the same, they will not be described in detail here.
[0081] In addition, with the above Figure 4 Corresponding to the data query method shown, this application also provides a data query device. Figure 7 This is a schematic diagram of the structure of a data query device 700 provided in an embodiment of this application, including:
[0082] The receiving module 701 is used to receive a data query request for the target application. The data query request is used to request data stored in accordance with the data storage method mentioned in the above embodiments.
[0083] The query module 702 is used to respond to a data query request and query the first target data of the target application from the target single table of the target single database of the target database. The first target data is synchronized from the source database to the target database. The source database includes multiple sub-databases, and each sub-database includes multiple sub-tables. The first target data is data stored in at least one sub-table of at least one sub-database in the source database. The target single database has a mapping relationship with the at least one sub-database, and the target single table has a mapping relationship with the at least one sub-table. The target database includes at least one single database and at least one single table.
[0084] In one possible implementation, the query module 702 is also used to query target data from the target database using the Tiflash component when the type of the data query request is an OLAP query request.
[0085] The data query device can be used as the above Figure 4 The execution entity of the data query method shown is therefore able to realize the data query method in Figure 4 The functions achieved are the same and have the same beneficial effects. Since the principles are the same, they will not be described in detail here.
[0086] Figure 8 This is a schematic diagram of the structure of an electronic device according to one embodiment of this specification. Please refer to it. Figure 8 At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and memory. The memory may include main memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for other business operations.
[0087] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 8 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0088] Memory is used to store programs. Specifically, programs may include program code, which includes computer operation instructions. Memory may include main memory and non-volatile memory, and provides instructions and data to the processor.
[0089] The processor reads the corresponding computer program from non-volatile memory into main memory and then executes it, forming a data storage device or data retrieval device at the logical level. The processor executes the program stored in memory and specifically performs the following operations:
[0090] The process involves identifying at least one shard in the source database that stores first application data generated by the target application, and at least one shard table in the at least one shard that stores the first application data. The source database comprises multiple shards, each containing multiple shard tables. The source database stores application data generated by different applications in multiple shard tables across these shards. The process also involves obtaining first target data from at least one shard table in the at least one shard of the source database, where the first target data is at least a portion of the first application data. Based on a pre-defined mapping relationship between tables in the source database and tables in the target database, the process identifies the target single database corresponding to the at least one shard in the source database, and the target single table corresponding to the at least one shard table in the target single database. The target database includes at least one single database and at least one single table. Finally, the process synchronizes the first target data to the target single table in the target single database for storage.
[0091] Alternatively, a data query request for the target application is received, wherein the data to be queried is data stored according to the data storage method mentioned in the above embodiments; in response to the data query request, the first target data of the target application is queried from the target single table of the target single database of the target database. The first target data is synchronized from the source database to the target database. The source database includes multiple sub-databases, each sub-database includes multiple sub-tables, and the first target data is data stored in at least one sub-table of at least one sub-database in the source database. The target single database has a mapping relationship with the at least one sub-database, and the target single table has a mapping relationship with the at least one sub-table. The target database includes at least one single database and at least one single table.
[0092] The above is as described in this instruction manual. Figure 1 The method performed by the data storage device disclosed in the illustrated embodiments or as described in this specification Figure 6The data storage device disclosed in the illustrated embodiments can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0093] It should be understood that the electronic device in the embodiments of this application can implement the data storage method in Figure 1 The functions or data storage devices of the illustrated embodiments are in Figure 6 The embodiments shown have the same function. Since the principles are the same and they have the same beneficial effects, the embodiments of this application will not be described again here.
[0094] It should be understood that the electronic device in the embodiments of this application can implement the data query method in Figure 4 The functional or data query device of the illustrated embodiment is in Figure 7 The embodiments shown have the same function. Since the principles are the same and they have the same beneficial effects, the embodiments of this application will not be described again here.
[0095] Of course, in addition to software implementation, the electronic device described in this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software. In other words, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0096] This application also proposes a computer-readable storage medium that stores one or more programs, the programs including instructions that, when executed by a portable electronic device including multiple applications, enable the portable electronic device to perform... Figure 1 or Figure 4 The method of the illustrated embodiment is specifically used to perform the following operations:
[0097] The process involves identifying at least one shard in the source database that stores first application data generated by the target application, and at least one shard table in the at least one shard that stores the first application data. The source database comprises multiple shards, each containing multiple shard tables. The source database stores application data generated by different applications in multiple shard tables across these shards. The process also involves obtaining first target data from at least one shard table in the at least one shard of the source database, where the first target data is at least a portion of the first application data. Based on a pre-defined mapping relationship between tables in the source database and tables in the target database, the process identifies the target single database corresponding to the at least one shard in the source database, and the target single table corresponding to the at least one shard table in the target single database. The target database includes at least one single database and at least one single table. Finally, the process synchronizes the first target data to the target single table in the target single database for storage.
[0098] Alternatively, a data query request for the target application is received, wherein the data to be queried is data stored according to the data storage method mentioned in the above embodiments; in response to the data query request, the first target data of the target application is queried from the target single table of the target single database of the target database. The first target data is synchronized from the source database to the target database. The source database includes multiple sub-databases, each sub-database includes multiple sub-tables, and the first target data is data stored in at least one sub-table of at least one sub-database in the source database. The target single database has a mapping relationship with the at least one sub-database, and the target single table has a mapping relationship with the at least one sub-table. The target database includes at least one single database and at least one single table.
[0099] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0100] In summary, the above are merely preferred embodiments of this specification and are not intended to limit the scope of protection of this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of protection of this specification.
[0101] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0102] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0103] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0104] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
Claims
1. A data storage method, characterized in that, include: The source database is identified as having at least one sub-database storing first application data generated by the target application, and at least one sub-table storing the first application data in the at least one sub-database. The source database includes multiple sub-databases, each sub-database including multiple sub-tables. The source database is used to store application data generated by different applications in the multiple sub-databases and multiple sub-tables. Obtain first target data from at least one table in at least one sub-database of the source database for the target application, wherein the first target data is at least a portion of the data in the first application data; Based on the preset mapping relationship between the database tables in the source database and the database tables in the target database, determine the target single database corresponding to at least one sub-database of the source database in the target database, and the target single table corresponding to the at least one sub-table in the target single database. The target database includes at least one single database and at least one single table. The first target data is synchronized to the target single table in the target single database for storage. During the synchronization of the first target data to the target single table, if a modification statement indicates that a column should be added, at least one column should be added to the target single table in the target database first, and then at least one column should be added to the at least one sub-table in the source database. If a modification statement indicates that a column should be deleted, at least one column of the at least one sub-table in the source database should be deleted first, and then at least one column of the target single table in the target database should be deleted. The modification statement refers to a statement that modifies the table structure of at least one sub-table in the source database.
2. The data storage method according to claim 1, characterized in that, The step of determining the target single database corresponding to at least one sub-database of the source database in the target database, and the target single table corresponding to the at least one sub-table in the target single database, based on the preset mapping relationship between the database tables in the source database and the database tables in the target database, includes: Obtain the database name of the at least one sub-database and the table name of the at least one sub-table; Based on the preset mapping relationship between the database tables in the source database and the database tables in the target database, determine the target database name of the single database corresponding to the database name of the at least one sub-database in the target database, and the target table name of the single table in the target single database corresponding to the table name of the at least one sub-table. The target single database is determined based on the target database name, and the target single table is determined based on the target table name.
3. The data storage method according to claim 1, characterized in that, The first target data includes incremental data, and the step of synchronizing the first target data to the target single table in the target single database for storage includes: Determine the first table field of the sub-table storing the incremental data in the at least one sub-table, and the second table field of the target single table, wherein the first table field and the second table field are table fields that indicate the same target table header in different forms; Based on the preset mapping relationship between the first table field and the second table field, at least a portion of the incremental data corresponding to the first table field is synchronized to the table column corresponding to the second table field in the target single table for storage.
4. The data storage method according to claim 1, characterized in that, The step of synchronizing the first target data to the target single table in the target single database for storage includes: Receive modification statements that modify the table structure of at least one sub-table in the source database and the table structure of the target single table, wherein the table structure includes the number of rows and columns of the table; Modify the table structure of at least one sub-table in the source database and the table structure of the target single table in the target database according to the modification statement; The first target data is determined from at least one sub-table after the table structure is modified, and the first target data is synchronized to the target single table after the table structure is modified.
5. The data storage method according to claim 1, characterized in that, After synchronizing the first target data to the target single table in the target single database for storage, the method further includes: Receive a first input of at least a portion of the data in the first target data in the source database; In response to the first input, at least a portion of the data in the first target data in the source database is deleted; According to a pre-configured synchronization rule, at least a portion of the data in the first target data in the target database is retained, and the synchronization rule indicates that the first input in the source database is not synchronized to the target database.
6. A data query method, characterized in that, include: Receive a data query request for a target application, wherein the data query request is for data stored in accordance with the data storage method according to any one of claims 1-5; In response to the data query request, the first target data of the target application is queried from the target single table of the target single database in the target database. The first target data is synchronized from the source database to the target database. The source database includes multiple sub-databases, and each sub-database includes multiple sub-tables. The first target data is data stored in at least one sub-table in at least one sub-database of the source database. The target single database has a mapping relationship with the at least one sub-database, and the target single table has a mapping relationship with the at least one sub-table. The target database includes at least one single database and at least one single table.
7. The data query method according to claim 6, characterized in that, The step of querying the first target data of the target application from the target single table of the target single database in response to the data query request includes: When the data query request is of type OLAP query request, the target data is queried from the target database using the Tiflash component.
8. A data storage device, characterized in that, include: The determination module is used to determine at least one sub-database in the source database that stores first application data generated by the target application, and at least one sub-table in the at least one sub-database that stores the first application data. The source database includes multiple sub-databases, and each sub-database includes multiple sub-tables. The source database is used to store application data generated by different applications in the multiple sub-databases and multiple sub-tables. The acquisition module is used to acquire first target data of the target application in at least one sub-table in at least one sub-database of the source database, wherein the first target data is at least a portion of the data in the first application data; The determining module is further configured to determine, based on a preset mapping relationship between the database tables in the source database and the database tables in the target database, a target single database corresponding to at least one sub-database of the source database in the target database, and a target single table corresponding to at least one sub-table in the target single database, wherein the target database includes at least one single database and at least one single table. A storage module is used to synchronize the first target data to the target single table in the target single database for storage; during the synchronization of the first target data to the target single table, if a modification statement is received indicating the addition of a table column, at least one table column is first added to the target single table in the target database, and then at least one table column is added to the at least one sub-table in the source database; if a modification statement is received indicating the deletion of a table column, at least one table column in the at least one sub-table in the source database is first deleted, and then at least one table column in the target single table in the target database is deleted; the modification statement refers to a statement that modifies the table structure of at least one sub-table in the source database.
9. A data query device, characterized in that, include: A receiving module is configured to receive a data query request for a target application, wherein the data requested in the data query request is data stored according to the data storage method described in any one of claims 1-5; The query module is used to respond to the data query request and query the first target data of the target application from the target single table of the target single database of the target database. The first target data is synchronized from the source database to the target database. The source database includes multiple sub-databases, and each sub-database includes multiple sub-tables. The first target data is data stored in at least one sub-table of at least one sub-database in the source database. The target single database has a mapping relationship with the at least one sub-database, and the target single table has a mapping relationship with the at least one sub-table. The target database includes at least one single database and at least one single table.
10. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the data storage method as described in any one of claims 1 to 5 or the data query method as described in claim 6 or 7.
11. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the data storage method as described in any one of claims 1 to 5 or the data query method as described in claim 6 or 7.
Citation Information
Patent Citations
Data query method and device and server
CN110309174A
Device, system and method for aggregating MySQL into PostgreSQL database and storage medium
CN110597891A
Data query method and device, server and readable storage medium
CN111767303A