Data Processing Method, Apparatus, Database System, Electronic Device and Storage Medium

By generating distributed transactions in the distributed MySQL sharding middleware, ensuring data consistency between data tables and external index tables, the problem of data inconsistency when the distributed MySQL sharding middleware is solved, and the atomicity and strong consistency of data processing are achieved.

CN112395284BActive Publication Date: 2025-07-01ALIBABA GROUP HOLDING LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201910754776.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-08-15
Publication Date
2025-07-01
Estimated Expiration
2039-08-15

AI Technical Summary

Technical Problem

When the distributed MySQL sharding middleware updates the global secondary index table, the data of the storage engine is not exposed to the public, resulting in the middleware being unable to obtain the data update of the main table, and can only use the same storage format as the main table, resulting in inconsistent data.

Method used

By generating distributed transactions, including pre-commit tasks to data tables and external index tables, ensuring transactions are submitted only when the task is executed successfully, thus ensuring data consistency between data tables and external index tables.

Benefits of technology

The data consistency in the data table and the external index table at any time is achieved, and the problem of inconsistency between the external index table and the data table data in the prior art is solved, and the error search results are obtained by users when searching data through the external index table.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112395284B_ABST
    Figure CN112395284B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a data processing method, apparatus, database system, electronic device, and storage medium. The data processing method includes: receiving a data processing request for a database, and generating a corresponding distributed transaction, where the distributed transaction includes a first pre-commit task indicating a data table and a second pre-commit task indicating an external index table corresponding to the data table; receiving first execution status information of the first pre-commit task and second execution status information of the second pre-commit task returned by the database; if both the first execution status information and the second execution status information indicate that the task execution is successful, submitting the distributed transaction to complete the data processing of the data table and the external index table. This data processing method can ensure data consistency between the data table and the external index table.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of computer technologies, and in particular, to a data processing method, apparatus, database system, electronic device, and storage medium. Background Art

[0002] With the development of technologies, databases are widely used for data record storage, query, and analysis. During the process of using a database for data storage, as the amount of data increases, the data processing capacity of a single device running the database becomes a bottleneck. To solve this problem, a database sharding and table partitioning strategy (such as the sharding algorithm) is adopted, that is, the database and / or data tables are split into multiple shards, and different shards can be a separate instance and configured on different devices.

[0003] In the existing databases adopting the above database sharding and table partitioning strategy, a global secondary index table needs to be used during operation. The global secondary index table is a kind of auxiliary index, which can be a non-clustered index and is an index table that users can create according to specific requirements.

[0004] For an existing database without database sharding and table partitioning (for example, a MySQL database), after data processing (such as inserting data, deleting data, or updating data) such as updating data, the method for updating the global secondary index table is that while the storage engine updates the records in the main table for storing data in the database, the same data is also updated to the global secondary index table. However, in the scenario of updating the global secondary index table by a distributed MySQL sharding middleware, this method for updating the global secondary index table has the following problems:

[0005] First, since the data of the storage engine is not exposed externally, the distributed MySQL sharding middleware cannot obtain the data written by the storage engine to the main table, and thus cannot update this data to the global secondary index table.

[0006] Second, since only one data source can be written, the global secondary index table can only use the same storage format as the main table. Summary of the Invention

[0007] In view of this, embodiments of the present invention provide a data processing solution to solve some or all of the above problems.

[0008] According to a first aspect of an embodiment of the present invention, there is provided a data processing method, which includes: receiving a data processing request for a database, and generating a corresponding distributed transaction, where the distributed transaction includes a first pre-submission task indicating a data table and a second pre-submission task indicating an external index table corresponding to the data table; receiving first execution status information of the first pre-submission task and second execution status information of the second pre-submission task returned by the database; if both the first execution status information and the second execution status information indicate successful task execution, then submitting the distributed transaction to complete data processing of the data table and the external index table.

[0009] According to a second aspect of an embodiment of the present invention, there is provided a data processing device, which includes: a transaction generation module, configured to receive a data processing request for a database and generate a corresponding distributed transaction, where the distributed transaction includes a first pre-submission task indicating a data table and a second pre-submission task indicating an external index table corresponding to the data table; an information receiving module, configured to receive first execution status information of the first pre-submission task and second execution status information of the second pre-submission task returned by the database; a submission module, configured to, if both the first execution status information and the second execution status information indicate successful task execution, then submit the distributed transaction to complete data processing of the data table and the external index table.

[0010] According to a third aspect of an embodiment of the present invention, there is provided an electronic device, including: a processor, a memory, a communication interface, and a communication bus, where the processor, the memory, and the communication interface complete communication with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the data processing method described in the first aspect.

[0011] According to a fourth aspect of an embodiment of the present invention, there is provided a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the data processing method described in the first aspect.

[0012] According to the data processing solution provided by the embodiments of the present invention, a corresponding distributed transaction is generated according to the received data processing request to instruct the database to process the data table and the external index table, and it is determined whether the data table and the external status table can be executed successfully according to the first execution status information and the second execution status information. Only when both the first execution status information and the second execution status information indicate successful execution, the distributed transaction is submitted to perform data processing on the data table and the external index table, ensuring that the data in the data table and the external index table is consistent at any time, and solving the problem in the prior art that the data in the external index table is inconsistent with the data in the data table, resulting in incorrect retrieval results when users perform data retrieval through the external index table. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For those of ordinary skill in the art, other drawings can also be obtained according to these drawings.

[0014] Figure 1 It is a flowchart of the steps of a data processing method according to Embodiment 1 of the present invention;

[0015] Figure 2 It is a flowchart of the steps of a data processing method according to Embodiment 2 of the present invention;

[0016] Figure 3 It is a flowchart of the steps of a data processing method according to Embodiment 3 of the present invention;

[0017] Figure 4 It is a flowchart of the steps of a data processing method according to Embodiment 4 of the present invention;

[0018] Figure 5 It is a flowchart of the steps of a data processing method according to Embodiment 5 of the present invention;

[0019] Figure 6 It is a structural block diagram of a data processing device according to Embodiment 6 of the present invention;

[0020] Figure 7 It is a structural block diagram of a data processing device according to Embodiment 7 of the present invention;

[0021] Figure 8a It is a structural schematic diagram of a database system according to Embodiment 8 of the present invention;

[0022] Figure 8b It is a schematic diagram of a database system performing a data processing method according to Embodiment 8 of the present invention;

[0023] Figure 8c It is a schematic flow chart of a data processing method executed by a database system according to Embodiment 8 of the present invention;

[0024] Figure 9 It is a schematic structural diagram of an electronic device according to Embodiment 9 of the present invention. Detailed implementation manners

[0025] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art shall fall within the protection scope of the embodiments of the present invention.

[0026] The following further illustrates the specific implementation of the embodiments of the present invention with reference to the accompanying drawings of the embodiments of the present invention.

[0027] Embodiment 1

[0028] Refer to Figure 1 , which shows a step flow chart of a data processing method according to Embodiment 1 of the present invention.

[0029] The data processing method in this embodiment includes the following steps:

[0030] Step S102: Receive a data processing request for indicating a database, and generate a corresponding distributed transaction.

[0031] In this embodiment, the proxy layer (such as DRDS proxy, Distributed Relational Database Services proxy) in the distributed database system is used as the execution subject to illustrate the data processing method provided in the embodiments of the present invention. Among them, the distributed database system can be a scenario based on MySQL sharding, or other types of scenarios.

[0032] In the distributed database of this embodiment, an independent proxy layer (such as DRDS proxy) is set. The proxy layer is a service process added between the client and the database, mainly providing the routing ability of the distributed database for users. An SQL statement of the client will be routed to one or more sub-databases (such as MySQL instances) according to the sharding algorithm of DRDS proxy. Through DRDS proxy, users can easily manage and operate multiple MySQL instances.

[0033] The data processing requests received by the proxy layer may be requests indicating inserting records into a data table, or requests indicating updating records in a data table, or requests indicating deleting records in a data table, etc.

[0034] When the proxy layer receives a data processing request indicating updating records in a data table, for a distributed database system with different sub - table methods for the data table and the external index table, the data processing request will involve data processing on records in multiple sub - databases. To avoid the situation where records in some sub - databases can be successfully updated while records in other sub - databases cannot be successfully updated, resulting in data inconsistency between the external index table and the data table, the proxy layer generates a corresponding distributed transaction according to the data processing request.

[0035] The distributed transaction includes a first pre - commit task and a second pre - commit task. The first pre - commit task is used to indicate processing the data table, and the second pre - commit task is used to indicate processing the external index table corresponding to the data table. This distributed transaction can only be successfully executed when all the pre - commit tasks it includes are successfully executed, thereby ensuring data consistency between the external index table and the data table. The processing corresponds to the processing requested by the aforementioned data processing request, such as a request for inserting records, a request for updating records, a request for deleting records, etc.

[0036] It should be noted that according to the different content in the data processing request, both the first pre - commit task and the second pre - commit task may include one or more branch tasks, and each branch task may correspond to a MySQL instance. Different branch tasks may correspond to the same or different MySQL instances.

[0037] Step S104: Receive the first execution status information of the first pre - commit task and the second execution status information of the second pre - commit task returned by the database.

[0038] The first execution status information is used to indicate whether each MySQL instance involved in the first pre - commit task has successfully executed the first pre - commit task. For example, if the first pre - commit task indicates inserting record A into the data table of a certain MySQL instance, the first execution status information is used to indicate whether record A has been successfully inserted into the data table.

[0039] Similarly, the second execution status information is used to indicate whether each MySQL instance involved in the second pre - commit task has successfully executed the second pre - commit task. For example, if the second pre - commit task indicates inserting record A into the external index table of a certain MySQL instance, the second execution status information is used to indicate whether record A has been successfully inserted into the external index table.

[0040] In practical applications, the first execution status information and the second execution status information can be represented in any appropriate form. For example, 1 indicates successful execution, and 0 indicates failed execution; or, True indicates successful execution, and False indicates failed execution, and so on. The embodiments of the present invention do not limit the specific forms of the first execution status information and the second execution status information.

[0041] After obtaining the first execution status information and the second execution status information, it is possible to determine whether to submit a distributed transaction based on the first execution status information and the second execution status information. For example, if both the first execution status information and the second execution status information indicate that the task execution is successful, it means that the data processing of both the data table and the external index table is successful, and the data consistency of the two tables can be ensured, then step S106 can be executed. On the contrary, if at least one of the first execution status information and the second execution status information indicates that the task execution fails, it means that the data processing of at least one of the data table and the external index table fails, and the data consistency between the two cannot be guaranteed, then it can be indicated not to submit the distributed transaction. Of course, when at least one task execution fails, those skilled in the art can configure and execute any appropriate operations as needed, and this embodiment does not limit this.

[0042] Step S106: If both the first execution status information and the second execution status information indicate that the task execution is successful, then submit the distributed transaction to complete the data processing of the data table and the external index table.

[0043] If both the first execution status information and the second execution status information indicate that the task execution is successful, it means that the distributed transaction can be successfully executed completely. The proxy layer can indicate to submit the distributed transaction so that each MySQL instance involved in the distributed transaction officially executes the task submission to complete the data processing of the data table and the external index table.

[0044] Since the distributed transaction is only submitted when both the first execution status information and the second execution status information indicate that the task execution is successful, that is, the distributed transaction is only submitted when both the data table and the external index table can successfully perform data processing, which ensures the data consistency between the data table and the external index table. And because the data table and the external index table perform data processing simultaneously, while ensuring the atomicity of the data processing of the data table and the external index table, it can also ensure the strong consistency between the data table and the external index table.

[0045] Among them, strong consistency is opposite to eventual consistency in the prior art. Strong consistency means that at any moment, the data table and the external index table are in a completely consistent state. Compared with eventual consistency, which has an inconsistent intermediate state before reaching the final consistency, strong consistency can ensure that users can obtain accurate data when using the database at any time.

[0046] Through this embodiment, a corresponding distributed transaction is generated according to the received data processing request to instruct the database to process the data table and the external index table, and it is determined whether the data table and the external status table can be successfully executed according to the first execution status information and the second execution status information. Only when both the first execution status information and the second execution status information indicate successful execution, the distributed transaction is committed to perform data processing on the data table and the external index table, ensuring that the data in the data table and the external index table are consistent at any time, and solving the problem in the prior art that the data in the external index table is inconsistent with the data in the data table, resulting in incorrect retrieval results when the user retrieves data through the external index table.

[0047] The data processing method of this embodiment can be executed by any suitable electronic device with data processing capabilities, including but not limited to: servers, mobile terminals (such as tablet computers, mobile phones, etc.), and PC machines, etc.

[0048] Embodiment Two

[0049] Refer to Figure 2 , which shows a step flowchart of a data processing method according to Embodiment Two of the present invention.

[0050] The data processing method of this embodiment includes the aforementioned steps S102 to step S106.

[0051] Among them, when the method determines whether to commit a distributed transaction according to the first execution status information and the second execution status information, if both the first execution status information and the second execution status information indicate successful task execution, step S106 is executed; otherwise, if at least one of the first execution status information and the second execution status information indicates failed task execution, the following step S108 is executed.

[0052] Step S108: If at least one of the first execution status information and the second execution status information indicates failed task execution, a rollback message is generated to indicate a rollback operation on the first pre-commit task and / or the second pre-commit task through the rollback message.

[0053] When one of the first execution status information and the second execution status information indicates that a task execution fails. For example, when the first execution status information indicates that the task execution is successful and the second execution status information indicates that the task execution fails, it means that the data processing of the data table is successful and the data processing of the external index table fails. At this time, the data in the data table is inconsistent with the data in the corresponding external index table, and the distributed transaction cannot be committed. A rollback message can be generated and sent to each involved MySQL instance to perform a rollback operation on the first pre-commit task and / or the second pre-commit task, so that the data in the data table and the external index table is consistent after the rollback operation.

[0054] It should be noted that in this embodiment, the external index table can be an external global secondary index table. In other embodiments, the external index table can also be any other appropriate external index table. The external global secondary index table in the distributed database can be split onto multiple MySQL instances, and the sharding method it adopts can be different from the sharding method of the data table.

[0055] Similarly, the data table can be the main table in the database or any other appropriate data table.

[0056] In some cases, the storage structures used by the data table (such as row-store data source or column-store data source) and the external index table may be different. In such a case, the proxy layer can generate pre-commit tasks adapted to the storage structure.

[0057] For example, if the data table is a row-store MySQL data source and the external index table is a column-store PostgreSQL data source, when the proxy layer executes a certain data processing request, it generates a first pre-commit task recognizable by the MySQL data source and a second pre-commit task recognizable by the PostgreSQL data source according to the data processing request, so as to adapt to different storage structures. Thus, the processing of the data processing request becomes more flexible and compatible.

[0058] Through this embodiment, on the basis of ensuring that the data in the data table and the external index table is consistent at any time and solving the problem in the prior art that the data in the external index table is inconsistent with the data in the data table, resulting in incorrect retrieval results when users retrieve data through the external index table, since the proxy layer can generate a first pre-commit task adapted to the storage structure used by the data table and a second pre-commit task adapted to the storage structure used by the external index table, it can adapt to the situation where the external index table and the data table use different storage structures, and improves the flexibility and compatibility of the processing of the data processing request.

[0059] Moreover, when any one of the first pre-submission task and the second pre-submission task fails, a rollback message can be generated to instruct the involved MySQL instance to perform a rollback operation, ensuring strong data consistency between the external index table and the data table, and there is no intermediate state of data inconsistency.

[0060] The data processing method of this embodiment can be executed by any suitable electronic device with data processing capabilities, including but not limited to: servers, mobile terminals (such as tablets, mobile phones, etc.), and PC machines, etc.

[0061] Embodiment III

[0062] Refer to Figure 3 , which shows a step flowchart of a data processing method according to Embodiment III of the present invention.

[0063] The data processing method of this embodiment includes the aforementioned steps S102 to S106. Optionally, it may or may not include step S108 according to needs.

[0064] In this embodiment, the step S102 includes the following sub-steps:

[0065] Sub-step S1021: Obtain the SQL statement in the data processing request.

[0066] In the first case, when the data processing request does not contain a dynamic function, the step S1021 includes: directly using the original SQL statement in the data processing request as the obtained SQL statement.

[0067] In the second case, when the data processing request contains a dynamic function, the sub-step S1021 includes: obtaining the original SQL statement from the data processing request, and replacing the dynamic function in the original SQL statement with a constant; generating a replaced SQL statement according to the replacement result.

[0068] For example, the original SQL statement includes the NOW() function, which is a dynamic function indicating to obtain the current system time. If the value of the NOW() function is not calculated in advance, but independently calculated on the data table and the external index table respectively, it may cause the corresponding record contents in the data table and the external index table to be inconsistent due to different system times during calculation, resulting in data inconsistency between the two. To prevent this situation, when a dynamic function is included, first determine the value of the dynamic function (this value is a constant that will not change), replace the dynamic function in the original SQL statement with this value, and generate a replaced SQL statement. This replaced SQL statement is the SQL statement obtained from the data processing request.

[0069] Optionally, to further improve reliability and ensure data consistency between the data table and the external index table, it is possible to determine whether the primary key of the data table involved in the data processing request is an auto-increment primary key, and perform appropriate operations based on the judgment result.

[0070] For example, in the second case, determine the data table to be processed according to the original SQL statement, determine whether the primary key of the data table is an auto-increment primary key, and whether there is a statement containing an insert operation in the original SQL statement; if it is an auto-increment primary key and there is a statement containing an insert operation, generate a globally unique auto-increment value, and generate a replaced SQL statement based on the auto-increment value and the replacement result.

[0071] An auto-increment primary key is a primary key whose primary key value automatically increases. For the case where the primary key of the data table is an auto-increment primary key and there is a statement containing an insert operation, generate the value of the auto-increment primary key in the proxy layer (i.e., DRDS proxy), ensuring that the value of the auto-increment primary key is a globally unique auto-increment value, thereby preventing the situation where the values of the auto-increment primary keys independently generated by the data table and the external index table may be inconsistent in subsequent steps, and ensuring data consistency.

[0072] When generating the replaced SQL statement based on the replacement result and the auto-increment value, the generated auto-increment value can be added to the SQL statement that has replaced the constant to form the replaced SQL statement.

[0073] For example, the SQL statement that has replaced the constant is: insert into t_primary(name,date)values('a','20190803').

[0074] Among them, t_primary represents the table name of the data table, name and date represent the field names in the data table, a represents the value of the name field, 20190803 represents the value of the date field, and the field name of the primary key of the data table t_primary is "id", and the auto-increment value generated by the proxy layer is "1", then the replaced SQL statement is: insert into t_primary(id,name,date)values(1,'a','20190803').

[0075] Sub-step S1022: Determine whether the SQL statement is an SQL statement containing an insert operation.

[0076] SQL statements include, but are not limited to, insert statements (insert), update statements (update), replace statements (replace), and delete statements (delete), etc. Among them, insert statements and replace statements are both SQL statements containing insert operations.

[0077] Those skilled in the art can determine whether an SQL statement is an SQL statement containing an insert operation in any appropriate manner as needed, and this embodiment does not limit this. For example, a table can be created based on all SQL statements containing insert operations, and it can be determined whether a certain SQL statement is an SQL statement containing an insert operation by determining whether the SQL statement is in the table.

[0078] Sub-step S1023: Determine a query statement according to the judgment result, and generate the corresponding distributed transaction according to the query statement and the SQL statement.

[0079] In the first case, if the judgment result indicates that the SQL statement is an SQL statement containing an insert operation, it means that the SQL statement may be an insert statement, a replace statement, etc. For these statements, new records need to be inserted into the data table and the external index table during execution. Therefore, it is necessary to first determine whether there are conflicting records in the data table and / or the external index table to prevent partial or complete insertion failure due to the existence of conflicting records when inserting new records, thus affecting the data consistency between the data table and the external index table.

[0080] Based on this, in the first case, according to the judgment result, determine a first query statement according to the SQL statement, and the first query statement is used to indicate querying records conflicting with the to-be-processed records indicated by the SQL statement. After determining the query statement, those skilled in the art can generate the corresponding distributed transaction in any appropriate manner as needed.

[0081] In the second case, if the judgment result indicates that the SQL statement is not an SQL statement containing an insert operation, it means that the SQL statement may be an update statement, a delete statement, etc. The objects of these statements are records that already exist in the data table and / or the external index table, and these statements usually contain a conditional clause (where clause).

[0082] Based on this, in the second case, according to the judgment result, determine a second query statement according to the SQL statement, and the second query statement is used to indicate obtaining records in the data table that meet the conditional clause in the SQL statement. After determining the query statement, those skilled in the art can generate the corresponding distributed transaction in any appropriate manner as needed.

[0083] Through this embodiment, it is ensured that the data in the data table and the external index table are consistent at any time, and the problem in the prior art that the external index table has a state inconsistent with the data in the data table, resulting in incorrect retrieval results when the user retrieves data through the external index table, is solved.

[0084] In addition, the original SQL statement obtained from the data processing request is first processed such as dynamic function replacement and generation of the auto-increment value of the auto-increment primary key, which prevents data inconsistency between the data table and the external index table in subsequent processing due to issues such as inconsistent auto-increment values or inconsistent dynamic function calculation results, ensuring reliability. Before generating a distributed transaction, it is determined whether the SQL statement contains an insert operation, and a corresponding query statement is generated according to the determination result. Then, a distributed transaction is generated based on the generated query statement and the SQL statement, improving the adaptability of the solution.

[0085] The data processing method of this embodiment can be executed by any suitable electronic device with data processing capabilities, including but not limited to: servers, mobile terminals (such as tablet computers, mobile phones, etc.), and PC machines, etc.

[0086] Embodiment 4

[0087] Referring to Figure 4 , a step flowchart of a data processing method according to Embodiment 4 of the present invention is shown.

[0088] The data processing method of this embodiment includes the foregoing steps S102 to S106. Optionally, it may include or not include step S108. Among them, step S102 may be implemented in the implementation manner of Embodiment 1 or 3, or implemented in any other suitable implementation manner.

[0089] In this embodiment, when step S102 is implemented in the implementation manner of Embodiment 3, the first implementation manner of sub-step S1023, that is, determining a query statement according to the determination result and generating the corresponding distributed transaction according to the query statement and the SQL statement, is as follows.

[0090] Sub-step S1023 includes: if the determination result indicates an SQL statement containing an insert operation, determining to generate a first query statement according to the SQL statement, where the first query statement is used to indicate querying records conflicting with the to-be-processed records indicated by the SQL statement; obtaining a first query result according to the first query statement; and generating the distributed transaction according to the first query result and the SQL statement.

[0091] As described in Embodiment 3, when the determination result indicates that the SQL statement includes an insert operation, in order to prevent existing records in the data table and / or the external index table from conflicting with new records and affecting the execution of the insert operation, resulting in data inconsistency between the data table and the external index table, it is determined to generate a first query statement for indicating querying conflicting records according to the SQL statement.

[0092] When generating the first query statement, if there is a global unique index in the tables involved in the SQL statement, the generated first query statement is used to query the data table and the external index table. At this time, the first query statement includes branch query statement A and branch query statement B. Branch query statement A can be generated based on the table name of the data table involved in the SQL statement, the primary key of the data table, and other unique keys, and is used to query the conflicting records in the data table. Branch query statement B can be generated based on the table name of the external index table involved in the SQL statement and the global unique index key, and is used to query the conflicting records in the external index table.

[0093] If there is no global unique index, the generated first query statement is only used to query the data table. It can be generated based on the table name of the data table involved in the SQL statement, the primary key of the data table, and other unique keys, and is used to query the conflicting records in the data table. Conflicting records are, for example, records in the data table where the primary key value of a certain or certain existing records is the same as the primary key value of the record to be processed indicated by the SQL statement. Since the primary key is unique, these records are conflicting records.

[0094] By sending the first query statement to each involved MySQL instance for execution, the query results returned by each MySQL instance can be obtained, and their sum is the first query result.

[0095] Based on the first query result and the SQL statement, the distributed transaction can be generated. For those skilled in the art, they can use any appropriate method to generate the distributed transaction, and this embodiment does not limit it.

[0096] The following are several examples of generating the distributed transaction based on the first query result and the SQL statement.

[0097] Among them, the first query result may only include table conflicting records (denoted as query result A), which can be conflicting records obtained by querying based on the primary key of the data table or the primary key and other unique keys. The first query result may also include table conflicting records (denoted as query result A) and index conflicting records (denoted as query result B). Index conflicting records can be conflicting records obtained by querying the external index table based on the global unique index key.

[0098] In Case 1, the SQL statement is a basic insert statement. For example, the SQL statement is "insert into t_primary (id, name, date) values (1, 'a', '20190803')". If the first query result indicates that there are no conflicting records in both the data table and the external index table, a distributed transaction is directly generated according to this SQL statement. The first pre-commit task in this distributed transaction instructs to insert the to-be-processed records indicated by this SQL statement into the data table; the second pre-commit task of this distributed transaction instructs to insert the to-be-processed records indicated by this SQL statement into the external index table.

[0099] If the first query result indicates that there are conflicting records in the data table and / or the external index table, those skilled in the art can configure appropriate operations as needed, such as still generating a distributed transaction according to the SQL statement or performing other operations, which are not limited in this embodiment.

[0100] In Case 2: The SQL statement is an insert statement and includes a first clause (such as the ignore clause) indicating to discard conflicting records. At this time, generating the distributed transaction according to the first query result and the SQL statement includes: If it is determined that the SQL statement is an insert statement including a first clause indicating to discard conflicting records, then according to the to-be-processed records indicated by the SQL statement and the first query result, determine the first records in the to-be-processed records that do not match the first query result; generate a first pre-commit task indicating to insert the first records into the data table, and a second pre-commit task indicating to insert the first records into the external index table; generate the distributed transaction according to the first pre-commit task and the second pre-commit task.

[0101] Among them, the first records are the difference set between the to-be-processed records indicated by the SQL statement and the first query result. In other words, they are records that conflict neither with the records in the data table nor with the records in the external index table.

[0102] Case 3: The SQL statement is an insert statement and includes a second clause (such as the on duplicate key update clause) indicating record update. At this time, generating the distributed transaction according to the first query result and the SQL statement includes: when it is determined that the SQL statement is an insert statement including a second clause indicating record update, obtaining the first records in the to-be-processed records indicated by the SQL statement that do not match the first query result and the second records that match the first query result; generating a first pre-commit task indicating inserting the first records into the data table and using the second records to update the first query result in the data table; and generating a second pre-commit task indicating inserting the first records into the external index table and using the second records to update the first query result in the external index table; generating the distributed transaction according to the generated first pre-commit task and the second pre-commit task.

[0103] Since the second clause indicates inserting the non-conflicting records in the to-be-processed records and using the to-be-processed records to update the conflicting records, when generating the distributed transaction, it is necessary to first determine the first records in the to-be-processed records that do not match the first query result, that is, the non-conflicting records; and determine the records that match the first query result, that is, the conflicting records.

[0104] Among them, if the first query result includes both table conflict records (i.e., query result A) and index table conflict records (i.e., query result B), the records in the first records can be divided into two sets, set X and set Y. The records in set X are the records that do not match query result A, and set Y is the records that do not match query result B; similarly, the second records can be divided into set M and set N. The records in set M are the records that match query result A, and the records in set N are the records that match query result B.

[0105] After determining the first records and the second records, when generating the first pre-commit task according to the first records, the second records and the SQL statement, generate a first pre-commit task indicating inserting the records in set X into the data table and using the records in set Y to update query result A in the data table. When generating the second pre-commit task according to the first records, the second records and the SQL statement, generate a second pre-commit task indicating inserting the records in set M into the external index table and using the records in set N to update query result B in the external index table.

[0106] Case 4: The SQL statement is a replacement statement (such as a replace statement). At this time, according to the first query result and the SQL statement, generating the distributed transaction includes: if it is determined that the SQL statement is a replacement statement for indicating record replacement, then according to the first query result, the SQL statement, and the record to be processed indicated by the SQL statement, generating the first pre-commit task and the second pre-commit task, and generating the distributed transaction according to the first pre-commit task and the second pre-commit task.

[0107] Among them, since the execution of the replacement statement requires first performing a delete operation and then an insert operation, the first pre-commit task instructs to delete the first query result in the data table and insert the record to be processed into the data table; the second pre-commit task instructs to delete the first query result in the external index table and insert the record to be processed into the external index table.

[0108] Specifically, for example, if the first query result includes table conflict records (i.e., query result A) and index conflict records (i.e., query result B), the generated first pre-commit task is used to instruct deleting query result A in the data table and inserting the record to be processed into the data table. The generated second pre-commit task is used to instruct deleting query result B in the external index table and inserting the record to be processed into the external data table.

[0109] Through this embodiment, it is ensured that the data in the data table and the external index table are consistent at any time, solving the problem in the prior art that the external index table has a state inconsistent with the data in the data table, resulting in incorrect retrieval results when users perform data retrieval through the external index table.

[0110] In addition, when generating the distributed transaction, the first pre-commit task and the second pre-commit task that indicate different actions are generated for different SQL statements, further improving the applicability of the solution.

[0111] The data processing method of this embodiment can be executed by any suitable electronic device with data processing capabilities, including but not limited to: servers, mobile terminals (such as tablet computers, mobile phones, etc.), and PC machines, etc.

[0112] Embodiment Five

[0113] Refer to Figure 5 , which shows a step flowchart of a data processing method according to Embodiment Five of the present invention.

[0114] The data processing method of this embodiment includes the foregoing steps S102 to S106. Optionally, it may or may not include step S108. Among them, step S102 may be implemented in the implementation manner of Embodiment 1 or 3, or in any other appropriate implementation manner.

[0115] In this embodiment, when step S102 is implemented in the implementation manner of Embodiment 3, the sub-step S1023, that is, determining a query statement according to the judgment result and generating a corresponding second implementation manner of the distributed transaction according to the query statement and the SQL statement, is as follows:

[0116] In this embodiment, sub-step S1023 includes: if the judgment result indicates that the SQL statement does not include an insert operation, generating a second query statement according to the SQL statement, where the second query statement is used to indicate obtaining records in the data table that meet the conditional clause in the SQL statement; obtaining a second query result according to the second query statement; and generating the distributed transaction according to the second query result and the SQL statement.

[0117] As described in Embodiment 3, when the judgment result indicates that the SQL statement includes an insert operation, it is usually an update statement or a delete statement, etc. Since the operation objects of these statements are existing records in the data table and the external index table, the second query statement can be generated according to the conditional clause (such as the where clause) in the SQL statement to obtain records that meet the conditional clause. For example, the second query statement is a select statement including the conditional clause, and the conditional clause of the select statement is the same as the conditional clause in the SQL statement to find records that meet the conditional clause from the data table. For example, the SQL statement is: UPDATE tb SET name = 'a' WHERE name = 'b'; where tb is the data table name and name is the field name. The second query statement generated according to the conditional clause of this SQL statement is: SELECT id, name FROM tb WHERE name = 'b', where id and name are the field names of the data table and tb is the data table name.

[0118] The following examples several cases of generating the distributed transaction according to the second query result and the SQL statement.

[0119] Case 5: The SQL statement is an update statement. At this time, generating the distributed transaction according to the second query result and the SQL statement includes: when the SQL statement is an update statement, generating a first pre-commit task and a second pre-commit task according to the SQL statement; and generating the distributed transaction according to the first pre-commit task and the second pre-commit task.

[0120] Among them, the first pre-submission task instruction updates the second query result in the data table according to the SQL statement; the second pre-submission task instruction updates the record corresponding to the second query result in the external index table according to the SQL statement. This can ensure that the records in the data table and the external index table are updated synchronously to ensure the data consistency between the two.

[0121] For example, the SQL statement is: UPDATE tb SET name='a' WHERE name='b'. Among them, tb is the data table name and name is the field name.

[0122] Its corresponding second query statement is: SELECT id,name FROM tb WHERE name='b'. Among them, tb is the data table name, and id and name are the field names. Suppose the second query result of this second query statement is: the id values are 1 and 2.

[0123] At this time, the syntax recognizable by the MySQL data source used in the first pre-submission task generated according to the second query result and the SQL statement can be expressed as: UPDATE tb SET tb.name='a' WHERE id IN(1,2), which is used to instruct to update the second query result in the data table.

[0124] If the data source used by the external index table is the same as the data source used by the data table, then in the second pre-submission task, only the data table name tb needs to be modified to the table name of the external index table. If the data source used by the external index table is different from the data source used by the data table, then a representation recognizable by the syntax of the data source of the external index table can be generated to adapt to the situation where the data tables and the external index tables use different data sources.

[0125] Situation six, the SQL statement is a delete statement. At this time, generating the distributed transaction according to the second query result and the SQL statement includes: when the SQL statement is a delete statement, generating the first pre-submission task and the second pre-submission task according to the SQL statement; generating the distributed transaction according to the first pre-submission task and the second pre-submission task.

[0126] Among them, the first pre-submission task instruction deletes the second query result in the data table according to the SQL statement; the second pre-submission task instruction deletes the record corresponding to the second query result in the external index table according to the SQL statement to ensure the data consistency between the data table and the external index table.

[0127] Through this embodiment, the data processing request ensures that the data in the data table and the external index table is consistent at any time, solving the problem in the prior art that the data in the external index table is inconsistent with the data in the data table, resulting in incorrect retrieval results when the user retrieves data through the external index table.

[0128] In addition, by generating a query statement to query the records of the data table, determining the records that need to be processed according to the query results, and generating a first pre-commit task for the data table and a second pre-commit task for the external index table for synchronization according to the query results, the problem of not supporting the external global secondary index of the distributed database system is solved.

[0129] Since the proxy layer generates the first pre-commit task and the second pre-commit task in the distributed transaction according to the data processing request, and converts the SQL statement into a statement recognizable by the target data source (i.e., the external index table) when generating the second pre-commit task, the problem in the prior art that only one data source can be written can be solved, and the problem of only supporting the row-store data source and not supporting the column-store data source can be avoided. The data processing method of this embodiment can be executed by any suitable electronic device with data processing capabilities, including but not limited to: servers, mobile terminals (such as tablet computers, mobile phones, etc.) and PC machines, etc.

[0130] Embodiment Six

[0131] Refer to Figure 6 , which shows a structural block diagram of a data processing device according to Embodiment Six of the present invention.

[0132] The data processing device of this embodiment includes: a transaction generation module 602, configured to receive a data processing request for indicating a database, and generate a corresponding distributed transaction, where the distributed transaction includes a first pre-commit task for indicating a data table and a second pre-commit task for indicating an external index table corresponding to the data table; an information receiving module 604, configured to receive first execution status information of the first pre-commit task and second execution status information of the second pre-commit task returned by the database; and a submission module 606, configured to submit the distributed transaction to complete the data processing of the data table and the external index table if both the first execution status information and the second execution status information indicate that the task execution is successful.

[0133] Through this embodiment, a corresponding distributed transaction is generated according to the received data processing request to instruct the database to process the data table and the external index table, and it is determined whether the data table and the external status table can be successfully executed according to the first execution status information and the second execution status information. Only when both the first execution status information and the second execution status information indicate successful execution, the distributed transaction is committed to perform data processing on the data table and the external index table, ensuring that the data in the data table and the external index table is consistent at any time, and solving the problem in the prior art that the data in the external index table is inconsistent with the data in the data table, resulting in incorrect retrieval results when the user retrieves data through the external index table.

[0134] Embodiment Seven

[0135] Refer to Figure 7 , which shows a structural block diagram of a data processing device according to Embodiment Seven of the present invention.

[0136] The data processing device in this embodiment includes: a transaction generation module 702, configured to receive a data processing request for instructing a database and generate a corresponding distributed transaction, where the distributed transaction includes a first pre-commit task for instructing a data table and a second pre-commit task for instructing an external index table corresponding to the data table; an information receiving module 704, configured to receive first execution status information of the first pre-commit task and second execution status information of the second pre-commit task returned by the database; and a submission module 706, configured to, if both the first execution status information and the second execution status information indicate successful task execution, submit the distributed transaction to complete data processing of the data table and the external index table.

[0137] Optionally, the device further includes: a rollback module 708, configured to, if at least one of the first execution status information and the second execution status information indicates failed task execution, generate a rollback message to instruct a rollback operation on the first pre-commit task and / or the second pre-commit task through the rollback message.

[0138] Optionally, the transaction generation module 702 includes: a first acquisition module 7021, configured to acquire an SQL statement in the data processing request; a first determination module 7022, configured to determine whether the SQL statement is an SQL statement including an insert operation; and a first generation module 7023, configured to determine a query statement according to the determination result and generate a corresponding distributed transaction according to the query statement and the SQL statement.

[0139] Optionally, the first generation module 7023 includes: a second determination module 7023a, configured to determine, if the determination result indicates an SQL statement including an insert operation, to generate a first query statement according to the SQL statement, where the first query statement is used to indicate querying records conflicting with the records to be processed indicated by the SQL statement; a second acquisition module 7023b, configured to obtain a first query result according to the first query statement; and a second generation module 7023c, configured to generate the distributed transaction according to the first query result and the SQL statement.

[0140] Optionally, the second generation module 7023c includes: a third determination module, configured to determine, if it is determined that the SQL statement is an insert statement including a first clause indicating abandonment of conflicting records, the first record in the records to be processed that does not match the first query result according to the records to be processed indicated by the SQL statement and the first query result; a third generation module, configured to generate a first pre-commit task indicating inserting the first record into the data table, and a second pre-commit task indicating inserting the first record into the external index table; and a fourth generation module, configured to generate the distributed transaction according to the first pre-commit task and the second pre-commit task.

[0141] Optionally, the second generation module 7023c includes: a fourth determination module, configured to obtain, when it is determined that the SQL statement is an insert statement including a second clause indicating record update, the first record in the records to be processed indicated by the SQL statement that does not match the first query result and the second record that matches the first query result; a fifth generation module, configured to generate a first pre-commit task indicating inserting the first record into the data table and using the second record to update the first query result in the data table; and a sixth generation module, configured to generate a second pre-commit task indicating inserting the first record into the external index table and using the second record to update the first query result in the external index table; and a seventh generation module, configured to generate the distributed transaction according to the generated first pre-commit task and the second pre-commit task.

[0142] Optionally, the second generation module 7023c includes: an eighth generation module, configured to generate the first pre-commit task and the second pre-commit task according to the first query result, the SQL statement, and the record to be processed indicated by the SQL statement if it is determined that the SQL statement is a replacement statement for indicating record replacement, and generate the distributed transaction according to the first pre-commit task and the second pre-commit task; wherein, the first pre-commit task indicates deleting the first query result in the data table and inserting the record to be processed into the data table; the second pre-commit task indicates deleting the first query result in the external index table and inserting the record to be processed into the external index table.

[0143] Optionally, the first query result includes table conflict records, or the first query result includes table conflict records and index conflict records.

[0144] Optionally, the first generation module 7023 includes: a ninth generation module 7023d, configured to generate a second query statement according to the SQL statement if the judgment result indicates that the SQL statement does not include an insert operation, where the second query statement is used to indicate obtaining records in the data table that meet the conditional clause in the SQL statement; a third obtaining module 7023e, configured to obtain a second query result according to the second query statement; a tenth generation module 7023f, configured to generate the distributed transaction according to the second query result and the SQL statement.

[0145] Optionally, the tenth generation module 7023f includes: an eleventh generation module, configured to generate the first pre-commit task and the second pre-commit task according to the SQL statement when the SQL statement is an update statement; wherein, the first pre-commit task indicates updating the second query result in the data table according to the SQL statement; the second pre-commit task indicates updating the record corresponding to the second query result in the external index table according to the SQL statement; a twelfth generation module, configured to generate the distributed transaction according to the first pre-commit task and the second pre-commit task.

[0146] Optionally, the tenth generation module 7023f includes: a thirteenth generation module, configured to generate the first pre-commit task and the second pre-commit task according to the SQL statement when the SQL statement is a delete statement; wherein, the first pre-commit task indicates deleting the second query result in the data table according to the SQL statement; the second pre-commit task indicates deleting the record corresponding to the second query result in the external index table according to the SQL statement; a fourteenth generation module, configured to generate the distributed transaction according to the first pre-commit task and the second pre-commit task.

[0147] Optionally, the first acquisition module 7021 includes: a replacement module 7021a, configured to acquire an original SQL statement from the data processing request, and replace dynamic functions in the original SQL statement with constants; a fifteenth generation module 7021b, configured to generate a replaced SQL statement according to the replacement result.

[0148] Optionally, the fifteenth generation module 7021b includes: a fifth determination module, configured to determine the data table to be processed according to the original SQL statement, and determine whether the primary key of the data table is an auto-increment primary key, and whether there is a statement including an insert operation in the original SQL statement; a sixteenth generation module, configured to, if it is an auto-increment primary key and there is a statement including an insert operation, generate a globally unique auto-increment value, and generate a replaced SQL statement according to the auto-increment value and the replacement result.

[0149] It should be noted that the foregoing fourth generation module, seventh generation module, eighth generation module, twelfth generation module, and fourteenth generation module may be the same module or multiple different modules.

[0150] The data processing device in this embodiment is used to implement the corresponding data processing methods in the foregoing multiple method embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated herein.

[0151] Embodiment Eight

[0152] Referring to Figure 8a , a schematic structural diagram of a database system according to Embodiment Eight of the present invention is shown.

[0153] As Figure 8a shown, the database system includes an agent layer, a storage layer, and at least one database layer. A plurality of database instances are configured on the at least one database layer, and each of the database instances is connected to at least one of the storage layers. The storage layer is used to store data in the database layer; the agent layer is used to execute operations indicated by the foregoing data processing method to send the first pre-commit task and the second pre-commit task in the distributed transaction to at least one database instance in at least one of the database layers; the database layer receiving the distributed transaction instructs the corresponding database instance to interact with the storage layer according to the first pre-commit task and the second pre-commit task to execute the first pre-commit task and the second pre-commit task; and returns the corresponding first execution status information and second execution status information to the agent layer.

[0154] The proxy layer can be deployed in an independent server or in one or more of the database layers. In this embodiment, taking the proxy layer deployed in an independent server as an example, the proxy layer is also used to obtain data processing requests for the database instances and route the data processing requests to at least one of the database instances according to the database sharding and table partitioning algorithm to manage and operate the multiple database instances.

[0155] There can be one or more database layers, and each database layer can include one or more database instances. It is only necessary to ensure that multiple database instances are configured so as to disperse the data processing pressure. Each database instance is correspondingly connected to a storage layer, and the storage layer is used to store data tables and external index tables.

[0156] It should be noted that the storage layers connected to different database instances can be on the same storage disk or distributed on different storage disks.

[0157] Through this embodiment, it can be ensured that the data in the data table and the external index table are consistent at any time, solving the problem in the prior art that the data in the external index table is inconsistent with the data in the data table in the eventual consistency solution, resulting in incorrect retrieval results when users retrieve data through the external index table.

[0158] The following combines Figure 8b and Figure 8c , and the data processing process is described as follows:

[0159] The proxy layer obtains a data processing request from the client, such as an original SQL statement, and starts a distributed transaction (such as an XA transaction) according to the data processing request.

[0160] According to the original SQL statement, determine the data table to be processed. If necessary, fill in the auto-increment field of the data table to be processed. For example, when the primary key of the data table to be processed is an auto-increment primary key and the original SQL statement contains an insert operation (such as containing an insert statement, etc.), fill in the auto-increment primary key of the data table to be processed. If not, this action can be omitted.

[0161] If there is a dynamic function in the original SQL statement, replace the dynamic function in the original SQL statement with a constant and generate a replaced SQL statement. For example, if the original SQL statement contains the now() function, replace now() with the corresponding current time. If not, this action can be omitted.

[0162] According to the replaced SQL statement, determine whether it contains an insert operation. Among them, SQL statements containing insert operations are, for example, insert statements, replace statements, etc.

[0163] In the first case, if an insert operation is included, conflict records (also known as matching records) are queried in the data table according to the primary key and unique key. And it is confirmed whether there is a global unique index key. If there is, conflict records (also known as matching records) are queried in the external index table according to the global unique index key. If there is no global unique index key, this action can be omitted.

[0164] In the first case, sub-case A, after querying conflict records (there may or may not be conflict records), if the replaced SQL statement contains a replace statement, it is indicated to delete all conflict records in the data table and the external index table, and it is indicated to insert all records indicated by the replaced SQL statement into the data table and the external index table. In the case where the statuses returned by the data table and the external index table both indicate successful processing, the distributed transaction is committed.

[0165] In the first case, sub-case B, after querying conflict records, if the replaced SQL statement contains on duplicate key update, it is indicated to insert the records that do not match the conflict records among all records indicated by the replaced SQL statement into the data table and the external index table, and update all conflict records in the data table and the external index table. In the case where the statuses returned by the data table and the external index table both indicate successful processing, the distributed transaction is committed.

[0166] In the first case, sub-case C, after querying conflict records, if the replaced SQL statement contains ignore, it is indicated to insert the records that do not match the conflict records among all records indicated by the replaced SQL statement into the data table and the external index table; if the replaced SQL statement does not contain ignore, it is indicated to insert all records indicated by the replaced SQL statement into the data table and the external index table. In the case where the statuses returned by the data table and the external index table both indicate successful processing, the distributed transaction is committed.

[0167] In the second case, if an insert operation is not included, matching records are queried in the data table according to the conditional clause in the replaced SQL statement. After querying the matching records, if the replaced SQL statement is an update statement, it is indicated to update the matching records in the data table and the external index table. In the case where the statuses returned by the data table and the external index table both indicate successful processing, the distributed transaction is committed; if the replaced SQL statement is not an update statement, it is indicated to delete the matching records in the data table and the external index table. In the case where the statuses returned by the data table and the external index table both indicate successful processing, the distributed transaction is committed.

[0168] Through the above process, it can be ensured that the data processing of the data table and the external index table is consistent, thus ensuring the consistency of the data between the two.

[0169] Embodiment Nine

[0170] Reference Figure 9 , which shows a schematic structural diagram of an electronic device according to Embodiment 9 of the present invention. The specific implementation of the electronic device in the specific embodiments of the present invention is not limited.

[0171] As Figure 9 shown, the electronic device may include: a processor 802, a communications interface 804, a memory 806, and a communication bus 808.

[0172] Among them:

[0173] The processor 802, the communications interface 804, and the memory 806 communicate with each other through the communication bus 808.

[0174] The communications interface 804 is used to communicate with other electronic devices such as terminal devices or servers.

[0175] The processor 802 is used to execute the program 810, and specifically may execute the relevant steps in the above-mentioned data processing method embodiments.

[0176] Specifically, the program 810 may include program code, and the program code includes computer operation instructions.

[0177] The processor 802 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the electronic device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0178] The memory 806 is used to store the program 810. The memory 806 may include a high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.

[0179] The program 810 can be specifically used to cause the processor 802 to perform the following operations: receive a data processing request for indicating data processing for a database, and generate a corresponding distributed transaction, where the distributed transaction includes a first pre-commit task for indicating data processing for a data table, and a second pre-commit task for indicating data processing for an external index table corresponding to the data table; receive first execution status information of the first pre-commit task and second execution status information of the second pre-commit task returned by the database; if both the first execution status information and the second execution status information indicate that the task execution is successful, then commit the distributed transaction to complete the data processing of the data table and the external index table.

[0180] In a feasible manner, the program 810 is further used to cause the processor 802 to generate a rollback message if at least one of the first execution status information and the second execution status information indicates that the task execution fails, so as to indicate a rollback operation for the first pre-commit task and / or the second pre-commit task through the rollback message.

[0181] In a feasible manner, the program 810 is further used to cause the processor 802 to obtain an SQL statement in the data processing request when receiving a data processing request for indicating data processing for a data table in the database and generating a corresponding distributed transaction; determine whether the SQL statement is an SQL statement including an insert operation; determine a query statement according to the determination result, and generate the corresponding distributed transaction according to the query statement and the SQL statement.

[0182] In a feasible manner, the program 810 is further used to cause the processor 802 to determine a first query statement according to the SQL statement when determining a query statement according to the determination result and generating a corresponding distributed transaction according to the query statement and the SQL statement, where the first query statement is used to indicate querying records conflicting with the to-be-processed records indicated by the SQL statement; obtain a first query result according to the first query statement; and generate the distributed transaction according to the first query result and the SQL statement.

[0183] In a feasible manner, the program 810 is further configured to cause the processor 802, when generating the distributed transaction according to the first query result and the SQL statement, if it is determined that the SQL statement is an insert statement including a first clause indicating to abandon conflicting records, to determine, according to the to-be-processed records indicated by the SQL statement and the first query result, a first record in the to-be-processed records that does not match the first query result; generate a first pre-commit task indicating to insert the first record into the data table, and a second pre-commit task indicating to insert the first record into the external index table; and generate the distributed transaction according to the first pre-commit task and the second pre-commit task.

[0184] In a feasible manner, the program 810 is further configured to cause the processor 802, when generating the distributed transaction according to the first query result and the SQL statement, if it is determined that the SQL statement is an insert statement including a second clause indicating to update records, to obtain a first record in the to-be-processed records indicated by the SQL statement that does not match the first query result and a second record that matches the first query result; generate a first pre-commit task indicating to insert the first record into the data table and use the second record to update the first query result in the data table; and generate a second pre-commit task indicating to insert the first record into the external index table and use the second record to update the first query result in the external index table; and generate the distributed transaction according to the generated first pre-commit task and second pre-commit task.

[0185] In a feasible manner, the program 810 is further configured to cause the processor 802, when generating the distributed transaction according to the first query result and the SQL statement, if it is determined that the SQL statement is a replacement statement used to indicate record replacement, to generate the first pre-commit task and the second pre-commit task according to the first query result, the SQL statement, and the to-be-processed records indicated by the SQL statement, and generate the distributed transaction according to the first pre-commit task and the second pre-commit task; wherein the first pre-commit task indicates to delete the first query result in the data table and insert the to-be-processed records into the data table; and the second pre-commit task indicates to delete the first query result in the external index table and insert the to-be-processed records into the external index table.

[0186] In a feasible manner, the first query result includes table conflict records, or the first query result includes table conflict records and index conflict records.

[0187] In a feasible manner, the program 810 is further configured to cause the processor 802 to determine a query statement according to the judgment result, and generate the corresponding distributed transaction according to the query statement and the SQL statement. If the judgment result indicates that the SQL statement does not contain an insert operation, a second query statement is generated according to the SQL statement, where the second query statement is used to indicate obtaining records in the data table that meet the condition clause in the SQL statement; obtaining a second query result according to the second query statement; and generating the distributed transaction according to the second query result and the SQL statement.

[0188] In a feasible manner, the program 810 is further configured to cause the processor 802 to generate the distributed transaction according to the second query result and the SQL statement. When the SQL statement is an update statement, a first pre-commit task and the second pre-commit task are generated according to the SQL statement; where the first pre-commit task indicates updating the data table according to the SQL statement with the second query result; the second pre-commit task indicates updating the records corresponding to the second query result in the external index table according to the SQL statement; and generating the distributed transaction according to the first pre-commit task and the second pre-commit task.

[0189] In a feasible manner, the program 810 is further configured to cause the processor 802 to generate the distributed transaction according to the second query result and the SQL statement. When the SQL statement is a delete statement, a first pre-commit task and the second pre-commit task are generated according to the SQL statement; where the first pre-commit task indicates deleting the data table according to the SQL statement with the second query result; the second pre-commit task indicates deleting the records corresponding to the second query result in the external index table according to the SQL statement; and generating the distributed transaction according to the first pre-commit task and the second pre-commit task.

[0190] In a feasible manner, the program 810 is further configured to cause the processor 802 to obtain the SQL statement in the data processing request, obtain the original SQL statement from the data processing request, and replace the dynamic function in the original SQL statement with a constant; and generate a replaced SQL statement according to the replacement result.

[0191] In a feasible manner, the program 810 is further configured to cause the processor 802 to determine the data table to be processed according to the original SQL statement when generating the replaced SQL statement according to the replacement result, determine whether the primary key of the data table is an auto-increment primary key, and determine whether there is a statement containing an insert operation in the original SQL statement; if it is an auto-increment primary key and there is a statement containing an insert operation, generate a globally unique auto-increment value, and generate the replaced SQL statement according to the auto-increment value and the replacement result.

[0192] For the specific implementation of each step in the program 810, reference may be made to the corresponding steps and units in the above-described data processing method embodiments, which will not be elaborated herein. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding process descriptions in the foregoing method embodiments, which will not be elaborated herein.

[0193] Through the electronic device of this embodiment, a corresponding distributed transaction is generated according to the received data processing request to instruct the database to process the data table and the external index table, and it is determined whether the data table and the external status table can be successfully executed according to the first execution status information and the second execution status information, and the distributed transaction is submitted to perform data processing on the data table and the external index table only when both the first execution status information and the second execution status information indicate successful execution, ensuring that the data in the data table and the external index table are consistent at any time, and solving the problem in the prior art that the external index table has a state inconsistent with the data in the data table, resulting in incorrect retrieval results when the user retrieves data through the external index table.

[0194] It should be noted that according to the needs of implementation, the various components / steps described in the embodiments of the present invention can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present invention.

[0195] The method according to an embodiment of the present invention can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and to be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the data processing method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the data processing method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the data processing method shown herein.

[0196] Those of ordinary skill in the art can realize that the units and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present invention.

[0197] The above embodiments are only used to illustrate the embodiments of the present invention, rather than to limit the embodiments of the present invention. Those of ordinary skill in the relevant technical field can also make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present invention, and the patent protection scope of the embodiments of the present invention should be defined by the claims.

Claims

1. A data processing method, characterized in that, Including: Receiving a data processing request for a database and generating a corresponding distributed transaction, where the distributed transaction includes a first pre-commit task indicating a data table and a second pre-commit task indicating an external index table corresponding to the data table; Receiving first execution status information of the first pre-commit task and second execution status information of the second pre-commit task returned by the database, where the first execution status information is used to indicate whether each instance involved in the first pre-commit task is successfully executed, and the second execution status information is used to indicate whether each instance involved in the second pre-commit task is successfully executed; If both the first execution status information and the second execution status information indicate successful task execution, then submitting the distributed transaction to complete data processing of the data table and the external index table.

2. The method according to claim 1, characterized in that, The method further includes: If at least one of the first execution status information and the second execution status information indicates failed task execution, then generating a rollback message to indicate a rollback operation on the first pre-commit task and / or the second pre-commit task through the rollback message.

3. The method according to claim 1, characterized in that, The receiving a data processing request for a database and generating a corresponding distributed transaction includes: Obtaining an SQL statement in the data processing request; Determining whether the SQL statement is an SQL statement including an insert operation; Determining a query statement according to the judgment result and generating a corresponding distributed transaction according to the query statement and the SQL statement.

4. The method according to claim 3, wherein The determining a query statement according to the judgment result and generating a corresponding distributed transaction according to the query statement and the SQL statement includes: If the judgment result indicates an SQL statement including an insert operation, then determining a first query statement generated according to the SQL statement, where the first query statement is used to indicate querying records conflicting with the to-be-processed records indicated by the SQL statement; Obtaining a first query result according to the first query statement; Generating the distributed transaction according to the first query result and the SQL statement.

5. The method according to claim 4, wherein The generating the distributed transaction according to the first query result and the SQL statement includes: If it is determined that the SQL statement is an insert statement including a first clause indicating discarding conflicting records, then determining a first record in the to-be-processed records that does not match the first query result according to the to-be-processed records indicated by the SQL statement and the first query result; Generating the first pre-commit task indicating inserting the first record into the data table and the second pre-commit task indicating inserting the first record into the external index table; Generating the distributed transaction according to the first pre-commit task and the second pre-commit task.

6. The method according to claim 4, characterized in that The generating the distributed transaction according to the first query result and the SQL statement includes: If it is determined that the SQL statement is an insert statement including a second clause indicating updating records, then obtaining a first record in the to-be-processed records indicated by the SQL statement that does not match the first query result and a second record that matches the first query result; Generate a first pre-commit task that instructs to insert the first record into the data table and update the first query result in the data table with the second record; and, Generate a second pre-commit task that instructs to insert the first record into the external index table and update the first query result in the external index table with the second record; Generate the distributed transaction according to the generated first pre-commit task and second pre-commit task.

7. The method according to claim 4, characterized in that, The generating the distributed transaction according to the first query result and the SQL statement includes: If it is determined that the SQL statement is a replacement statement for instructing record replacement, generate the first pre-commit task and the second pre-commit task according to the first query result, the SQL statement, and the record to be processed indicated by the SQL statement, and generate the distributed transaction according to the first pre-commit task and the second pre-commit task; Wherein, the first pre-commit task instructs to delete the first query result in the data table and insert the record to be processed into the data table; the second pre-commit task instructs to delete the first query result in the external index table and insert the record to be processed into the external index table.

8. The method according to any one of claims 4 to 7, characterized in that, The first query result includes table conflict records, or the first query result includes table conflict records and index conflict records.

9. The method according to claim 3, wherein The determining the query statement according to the judgment result and generating the corresponding distributed transaction according to the query statement and the SQL statement includes: If the judgment result indicates that the SQL statement does not contain an insert operation, generate a second query statement according to the SQL statement, where the second query statement is used to instruct to obtain the records in the data table that meet the conditional clause in the SQL statement; Obtain a second query result according to the second query statement; Generate the distributed transaction according to the second query result and the SQL statement.

10. The method according to claim 9, characterized in that, The generating the distributed transaction according to the second query result and the SQL statement includes: When the SQL statement is an update statement, generate a first pre-commit task and a second pre-commit task according to the SQL statement; wherein, the first pre-commit task instructs to update the second query result in the data table according to the SQL statement; the second pre-commit task instructs to update the record corresponding to the second query result in the external index table according to the SQL statement; Generate the distributed transaction according to the first pre-commit task and the second pre-commit task.

11. The method according to claim 9, characterized in that, The generating the distributed transaction according to the second query result and the SQL statement includes: When the SQL statement is a delete statement, generate a first pre-commit task and a second pre-commit task according to the SQL statement; wherein, the first pre-commit task instructs to delete the second query result in the data table according to the SQL statement; the second pre-commit task instructs to delete the record corresponding to the second query result in the external index table according to the SQL statement; Generate the distributed transaction according to the first pre-submission task and the second pre-submission task.

12. The method according to claim 3, wherein The obtaining of the SQL statement in the data processing request includes: Obtain the original SQL statement from the data processing request, and replace the dynamic functions in the original SQL statement with constants; Generate the replaced SQL statement according to the replacement result.

13. The method according to claim 12, wherein The generating of the replaced SQL statement according to the replacement result includes: Determine the data table to be processed according to the original SQL statement, and determine whether the primary key of the data table is an auto-increment primary key, and whether there is a statement including an insert operation in the original SQL statement; If it is an auto-increment primary key and there is a statement including an insert operation, generate a globally unique auto-increment value, and generate the replaced SQL statement according to the auto-increment value and the replacement result.

14. A data processing device, characterized in that, Includes: A transaction generation module, configured to receive a data processing request for a database, and generate a corresponding distributed transaction, wherein the distributed transaction includes a first pre-submission task indicating a data table and a second pre-submission task indicating an external index table corresponding to the data table; An information receiving module, configured to receive first execution status information of the first pre-submission task and second execution status information of the second pre-submission task returned by the database, where the first execution status information is used to indicate whether each instance involved in the first pre-submission task is successfully executed, and the second execution status information is used to indicate whether each instance involved in the second pre-submission task is successfully executed; A submission module, configured to, if both the first execution status information and the second execution status information indicate that the task execution is successful, submit the distributed transaction to complete the data processing of the data table and the external index table.

15. A database system, characterized in that, Includes a proxy layer, a storage layer, and at least one database layer, where a plurality of database instances are configured on the at least one database layer, and each of the database instances is connected to at least one of the storage layers, and the storage layer is used to store the data in the database layer; The proxy layer is configured to perform the operations indicated by the method according to any one of claims 1-13, so as to send the first pre-submission task and the second pre-submission task in the distributed transaction to at least one database instance in at least one of the database layers; The database layer receiving the distributed transaction, according to the first pre-submission task and the second pre-submission task, instructs the corresponding database instance to interact with the storage layer to execute the first pre-submission task and the second pre-submission task; And return the corresponding first execution status information and second execution status information to the proxy layer.

16. The database system according to claim 15, characterized in that, The proxy layer is further configured to obtain a data processing request for a database instance, and route the data processing request to at least one of the database instances according to a database sharding and table partitioning algorithm, so as to manage and operate the multiple database instances.

17. An electronic device, comprising: A processor, a memory, a communication interface, and a communication bus, where the processor, the memory, and the communication interface complete communication with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the data processing method described in any one of claims 1-13.

18. A computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the data processing method described in any one of claims 1-13.

Citation Information

Patent Citations

  • System and method for generating extensible file system metadata and file system content processing

    CN1906613A