Cross-engine data processing method and device and computer equipment

By synchronizing the target data by using the data index relationship between the row storage engine and the column storage engine in a distributed database system, the problem of difficult data consistency in distributed databases when processing OLTP and OLAP tasks is solved, and the consistency and efficient synchronization of cross-engine data is achieved.

CN120067223APending Publication Date: 2025-05-30CHINA SOUTHERN POWER GRID COMPANY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510190152.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When existing distributed databases handle OLTP and OLAP tasks, it is difficult to ensure cross-engine consistency of data, especially when there are significant differences in the data model, storage structure, indexing mechanism and transaction processing strategies of different engines.

Method used

By determining the target data to be written into the distributed database system, and updating the data in the first database using the row storage engine, configuring a cross-engine synchronization trigger condition, and synchronizing the target data into the second database according to the data index relationship between the row storage engine and the column storage engine.

Benefits of technology

It ensures data consistency between the row storage engine and the column storage engine. By configuring cross-engine synchronization trigger conditions, the cross-engine synchronization timing of data is accurately controlled, and during the synchronization process, it ensures that the structure and index of the target data in the two storage engines are consistent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067223A_ABST
    Figure CN120067223A_ABST
Patent Text Reader

Abstract

The invention relates to a cross-engine data processing method and device, computer equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: determining target data to be written into a distributed database system; the distributed database system comprises a first database corresponding to a row storage engine and a second database corresponding to a column storage engine; updating the target data to a first database through a row storage engine; determining a cross-engine synchronization triggering condition correspondingly configured for the attribute information of the target data; under the condition that the cross-engine synchronization triggering condition is met, the target data is synchronized into the second database according to the data index relation between the row storage engine and the column storage engine, it can be ensured that the structures and indexes of the target data in the two storage engines are kept consistent, and the data synchronization efficiency is improved. And the consistency of the data between the row storage engine and the column storage engine can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, and particularly to a method, device, computer device, computer-readable storage medium, and computer program product for processing cross-engine data. Background Art

[0002] With the rapid development of information technology, the amount of data has grown explosively. Traditional databases face many challenges in processing large-scale data, such as problems of compatibility, reliability, and insufficient performance. Distributed databases have emerged, but there are still many technical problems in multi-modal data processing, heterogeneous data migration, and operation and maintenance management.

[0003] Existing distributed databases usually need to support multiple data processing engines to meet the needs of different business scenarios. In some applications, it is necessary to process on-line transaction processing (OLTP) and on-line analytical processing (OLAP) tasks simultaneously. OLTP can quickly respond to transaction operations, while OLAP focuses on complex data analysis, which poses higher requirements for the parallel processing ability of the database across engines.

[0004] Due to significant differences in data models, storage structures, index mechanisms, and transaction processing strategies between data processing engines such as OLTP and OLAP, it is difficult to ensure data consistency during the process of cross-engine access and processing of data. Summary of the Invention

[0005] Based on this, in view of the above technical problems, it is necessary to provide a method, device, computer device, computer-readable storage medium, and computer program product for processing cross-engine data that can improve data consistency between different engines.

[0006] In a first aspect, the present application provides a method for processing cross-engine data, the method including:

[0007] Determine target data to be written into a distributed database system; the distributed database system includes a first database corresponding to a row storage engine and a second database corresponding to a column storage engine;

[0008] Update the target data into the first database through the row storage engine;

[0009] Determine a cross-engine synchronization trigger condition configured corresponding to the attribute information of the target data;

[0010] When the cross - engine synchronization trigger condition is met, synchronize the target data to the second database according to the data index relationship between the row - based storage engine and the column - based storage engine.

[0011] In one embodiment, the step of synchronizing the target data to the second database according to the data index relationship between the row - based storage engine and the column - based storage engine includes:

[0012] Determine the first index information of the target data in the first database, and construct second index information corresponding to the first index information in the second database according to the first index information;

[0013] Determine the data index relationship between the first database and the second database based on the first index information and the second index information;

[0014] Synchronize the target data to the second database based on the data index relationship.

[0015] In one embodiment, the step of determining the target data to be written into the distributed database system includes:

[0016] Obtain a data processing request for the distributed database system, and determine the data identifier carried in the data processing request;

[0017] Determine the data to be processed pointed to by the data identifier and the target storage engine for processing the data to be processed;

[0018] Process the data to be processed through the target storage engine to obtain the target data.

[0019] In one embodiment, the step of determining the data to be processed pointed to by the data identifier and the target storage engine for processing the data to be processed includes:

[0020] Determine the intermediate storage engine associated with the data to be processed according to the mapping rule between the data to be processed and each storage engine;

[0021] Determine the execution cost of each intermediate storage engine for executing the data processing request according to the data processing request, the real - time load information of each intermediate storage engine, the distribution information of the data to be processed on each intermediate storage engine, and the object statistics information of the data table to which the data to be processed belongs;

[0022] Determine the target storage engine for executing the data processing request from each intermediate storage engine according to the execution cost.

[0023] In one embodiment, the mapping rule is constructed by the following method:

[0024] Obtain the distributed data in the distributed database system and the data processing rules of each storage engine associated with the distributed data;

[0025] Construct a mapping rule between the distributed data and each of the storage engines according to the distributed data and the data processing rules;

[0026] In the case where at least one of the storage engine and the distributed data changes, update the mapping rule based on at least one of the changed storage engine and the changed distributed data.

[0027] In one embodiment, the method further includes:

[0028] Construct a distributed database system, and deploy a first database and a second database in the distributed database system;

[0029] Determine a source database, obtain source data in the source database, and determine a data conversion rule between the source data and the distributed data according to the data compatibility information between the source database and the distributed database system; the data compatibility information includes at least one of data structure information, data type information, and database object information;

[0030] Convert the source data according to the data conversion rule to obtain the distributed data, and migrate the distributed data to at least one of the first database and the second database of the distributed database system.

[0031] In one embodiment, the method further includes:

[0032] Obtain the operation status data of the distributed database system;

[0033] Based on the operation status data, determine the operation status of the distributed database system; the operation status is used to characterize at least one of the performance metrics, status information, and fault information of the distributed database system.

[0034] In a second aspect, the present application further provides a processing device for cross-engine data, and the device includes:

[0035] A data determination module, configured to determine target data to be written into a distributed database system; the distributed database system includes a first database corresponding to a row storage engine and a second database corresponding to a column storage engine;

[0036] A data update module, configured to update the target data to the first database through the row storage engine;

[0037] A condition determination module, configured to determine a cross-engine synchronization trigger condition configured corresponding to the attribute information of the target data;

[0038] A data synchronization module, configured to synchronize the target data to the second database according to the data index relationship between the row storage engine and the column storage engine when the cross-engine synchronization trigger condition is satisfied.

[0039] In a third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0040] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0041] In a fifth aspect, the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0042] The above method, device, computer device, computer-readable storage medium, and computer program product for processing cross-engine data determine the target data to be written into a distributed database system; the distributed database system includes a first database corresponding to a row storage engine and a second database corresponding to a column storage engine; update the target data to the first database through the row storage engine; determine a cross-engine synchronization trigger condition configured corresponding to the attribute information of the target data; when the cross-engine synchronization trigger condition is satisfied, synchronize the target data to the second database according to the data index relationship between the row storage engine and the column storage engine; by configuring the cross-engine synchronization trigger condition, the timing of cross-engine data synchronization can be precisely controlled, and during the synchronization process, based on the data index relationship between the storage engines, it can ensure that the structure and index of the target data are consistent in the two storage engines, which is beneficial to ensuring the consistency between the row storage engine and the column storage engine. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for describing the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained without creative efforts.

[0044] Figure 1 It is an application environment diagram of the method for processing cross-engine data in an embodiment;

[0045] Figure 2 It is a schematic flowchart of the method for processing cross-engine data in an embodiment;

[0046] Figure 3 It is a schematic flowchart of the steps for synchronizing target data in an embodiment;

[0047] Figure 4 It is a structural block diagram of the device for processing cross-engine data in an embodiment;

[0048] Figure 5 It is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0049] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0050] The method for processing cross-engine data provided by the embodiments of the present application can be applied to an application environment as shown in Figure 1 . Among them, the terminal 102 communicates with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed in the cloud or other network servers. The server 104 determines the target data to be written into the distributed database system according to the data processing request initiated by the user on the terminal 102; the distributed database system includes a first database corresponding to a row storage engine and a second database corresponding to a column storage engine; subsequently, the server 104 controls the row storage engine to update the target data into the first database, and determines the cross-engine synchronization trigger condition configured for the attribute information of the target data; when the server 104 monitors that the cross-engine synchronization trigger condition is met, according to the data index relationship between the row storage engine and the column storage engine, the target data is synchronized into the second database to complete the cross-engine processing of the target data and ensure the consistency of the target data between different storage engines.

[0051] Among them, the terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server 104 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0052] In an exemplary embodiment, as Figure 2 shown, a method for processing cross-engine data is provided. Taking the method applied to the Figure 1 server in it as an example for illustration, it can be understood that this method can also be applied to the Figure 1 terminal 102 in it, and can also be applied to a system including the terminal 102 and the server 104, which is realized through the interaction between the terminal 102 and the server 104. The method of this embodiment includes the following steps 201 to step 204. Among them:

[0053] Step 201, determine the target data to be written into the distributed database system; the distributed database system includes a first database corresponding to a row storage engine and a second database corresponding to a column storage engine.

[0054] Among them, the distributed database system refers to a database system that disperses data storage on multiple independent physical nodes. Each physical node is connected through a network and works together to provide data storage, management, and access functions. In this embodiment, the distributed database system includes a first database corresponding to a row storage engine and a second database corresponding to a column storage engine. Multiple physical nodes are deployed in both the first database and the second database. For example, computing nodes, storage nodes, etc. The computing nodes are used to execute data processing requests, transaction management, data calculation, and other operations, and the storage nodes are used to execute data storage, backup, replication, and other operations.

[0055] Among them, the row storage engine refers to a database engine that stores and processes data by row. Each row of data contains a complete record, that is, all the details of the corresponding columns of that row. The row storage engine is usually suitable for scenarios with frequent insert, update, and delete operations such as online transaction processing (OLTP); for example, order systems, user management systems, etc.

[0056] Among them, the first database refers to the database corresponding to the row storage engine, which is used to store transactional data, such as order information, user operation records, etc. In specific implementation, one or more first databases can be deployed to store transactional data distributively. The transactional data can be distributively stored in each node of the first database in the form of data slices or in different first databases. Moreover, for each data slice of the transactional data, multiple data replicas can be created as needed, and each data replica can also be distributively stored in each node of the first database or in different first databases to cope with high-concurrency application scenarios and improve the data processing efficiency.

[0057] Among them, the column storage engine refers to a database engine that stores and processes data by columns. All values of the same field are stored in each column. The column storage engine is suitable for scenarios such as complex data analysis and aggregation queries in online analytical processing (OLAP); for example, data warehouses, reporting systems, etc.

[0058] Among them, the second database refers to the database corresponding to the column storage engine, which is used to store analytical data, such as sales statistics, user behavior analysis, etc. In specific implementation, similar to the first database, one or more second databases can be deployed to store analytical data distributively. Similarly, the analytical data can be distributively stored in each node of the second database in the form of data slices or in different second databases. Moreover, for each data slice of the analytical data, multiple data replicas can be created as needed, and each data replica can also be distributively stored in each node of the second database or in different second databases to cope with high-concurrency application scenarios and improve the data processing efficiency.

[0059] Among them, the target data refers to the data to be written into the distributed database system. In this embodiment, the target data mainly refers to the data that is written into the first database through the row storage engine and can be synchronized to the second database when the synchronization conditions are met, so as to increase the consistency of data across engines. In specific implementation, the target data can be the data carried in the data processing request initiated by the user through the terminal, the intermediate data generated during the processing, and the final data obtained after the processing is completed, etc.

[0060] Exemplarily, the server determines the target data to be written into the distributed database system according to the data processing request initiated by the terminal.

[0061] Step 202, update the target data to the first database through the row storage engine.

[0062] Among them, updating refers to writing target data into the database. The updating method can be new addition from scratch, new addition based on the original data, or overwriting and modifying the original data according to storage requirements or the storage performance of the first database.

[0063] Exemplarily, the server controls the row storage engine to write the determined target data into the first database of the distributed database.

[0064] Step 203: Determine the cross-engine synchronization trigger condition configured corresponding to the attribute information of the target data.

[0065] Among them, the attribute information refers to the metadata or characteristic information associated with the target data, which is used to describe the generation, update, or access mode of the data. The attribute information may include, but is not limited to, the source, application scenario, security level, integrity, data structure, data timeliness, and the relevance and dependence between data of the target data.

[0066] Among them, the cross-engine synchronization trigger condition refers to the condition pre-configured in the distributed database system and used to determine when to synchronize the target data from the row storage engine to the column storage engine. Specifically, the cross-engine synchronization trigger condition can be one or more of the trigger conditions based on time interval, data volume, data change frequency, etc. When setting multiple cross-engine synchronization trigger conditions, the data synchronization can be performed according to the first trigger mechanism, that is, as long as one of the multiple cross-engine synchronization trigger conditions is satisfied first, the data synchronization can be triggered; or all of the multiple cross-engine synchronization trigger conditions need to be satisfied to trigger the data synchronization; or when a certain cross-engine synchronization trigger condition is satisfied, the monitoring of another cross-engine synchronization trigger condition is started, and the data synchronization is triggered when the other cross-engine synchronization trigger condition is satisfied, etc.

[0067] Among them, the trigger condition based on time interval refers to using the time interval between two data synchronization operations as the condition for triggering synchronization. For example, synchronize once every ten minutes (that is, trigger a data synchronization operation every ten minutes), or synchronize once every hour, etc. The trigger condition based on data volume refers to using the accumulated data volume reaching a certain threshold as the condition for triggering synchronization. For example, synchronize once every 1000 records are accumulated. The trigger condition based on data change frequency refers to using the frequency of the target data being updated or inserted as the condition for triggering synchronization. For example, update 10 times per second, or insert 100 records per minute.

[0068] Exemplarily, the server configures the corresponding cross-engine synchronization trigger condition for the target data according to the attribute information such as the source, application scenario, security level, integrity, data structure, data timeliness, and the relevance and dependence between data of the target data.

[0069] Step 204, when the cross-engine synchronization trigger condition is satisfied, synchronize the target data to the second database according to the data index relationship between the row storage engine and the column storage engine.

[0070] The data index relationship refers to the mapping relationship between the data of the row storage engine and the column storage engine, so as to ensure that the structure and index of the target data in the two storage engines are consistent.

[0071] Among them, data synchronization refers to the process of copying the target data from the first database to the second database based on the data index relationship between the row storage engine and the column storage engine. In an exemplary embodiment, the method of synchronizing the target data can be implemented by using asynchronous replication technology. Specifically, after the data is written into the row storage engine, the message queue mechanism is used to encapsulate the change event of the target data into a message, and when the cross-engine synchronization trigger condition is satisfied, that is, at a preset time interval, a preset update frequency, or when the data volume of the target data reaches a preset number, the message encapsulating the change event of the target data is pushed to the column storage engine.

[0072] In an exemplary embodiment, to accelerate the synchronization process of the target data and ensure that the target data can be copied from the row storage engine to the column storage engine in a timely and accurate manner, a multi-threaded concurrent processing method can be used to synchronize the target data; specifically, when performing data synchronization, the data synchronization task can be divided into multiple synchronization subtasks, and each synchronization subtask is processed by an independent thread, and multiple threads run concurrently to achieve the accelerated synchronization of the target data.

[0073] Exemplarily, after the server determines the target data, it continuously monitors whether the cross-engine synchronization trigger condition for the target data is satisfied. If the cross-engine synchronization trigger condition is satisfied, the target data is synchronized to the second database according to the data index relationship between the row storage engine and the column storage engine. If the cross-engine synchronization trigger condition is not satisfied, the server continues to monitor the satisfaction of the cross-engine synchronization trigger condition until the cross-engine synchronization trigger condition is satisfied, and then synchronizes the target data to the second database according to the data index relationship between the row storage engine and the column storage engine.

[0074] In the above method for processing cross-engine data, the target data to be written into the distributed database system is determined; the distributed database system includes a first database corresponding to the row storage engine and a second database corresponding to the column storage engine; the target data is updated to the first database through the row storage engine; the cross-engine synchronization trigger condition configured corresponding to the attribute information of the target data is determined; when the cross-engine synchronization trigger condition is satisfied, the target data is synchronized to the second database according to the data index relationship between the row storage engine and the column storage engine; by configuring the cross-engine synchronization trigger condition, the timing of cross-engine data synchronization can be precisely controlled, and during the synchronization process, based on the data index relationship between the storage engines, it can be ensured that the structure and index of the target data in the two storage engines are consistent, which is beneficial to ensuring the consistency of data between the row storage engine and the column storage engine.

[0075] In one embodiment, as Figure 3 shown, step 204 includes:

[0076] Step 301, determine the first index information of the target data in the first database, and construct the second index information corresponding to the first index information in the second database according to the first index information.

[0077] Among them, the first index information refers to the index information established for the target data in the first database corresponding to the row storage engine, which is used to quickly locate the target data in the first database; the first index information can be the field information, identification information, etc. of the target data.

[0078] Among them, the second index information refers to the index information constructed or mapped in the second database corresponding to the column storage engine in order to correspond to the first index information in the row storage engine, so as to quickly locate the target data in the second database after data synchronization. And the second index information is constructed based on the first index information. When constructing the second index information, it can be obtained according to the characteristics of the column storage engine and its corresponding second database, so as to ensure that the data can be correctly and efficiently corresponded and accessed between the first database and the second database. Furthermore, when reading data in different databases, it can be ensured that consistent results can be obtained according to the same or related index rules (information), that is, no matter whether transaction queries or analytical queries are performed through the row storage engine or the column storage engine, the consistency of data reading can be guaranteed. Similarly, the second index information can be the field information, identification information, etc. of the target data. When the first index information is field information, the second index information can be the same field information or identification information. In implementation, it only needs to satisfy the mapping correspondence between the first index information and the second index information.

[0079] Exemplarily, when synchronizing target data, the server first determines the first index information of the target data in the first database corresponding to the row storage engine, and constructs the second index information corresponding to the first index information in the second database corresponding to the column storage engine according to the characteristics of the first index information, the column storage engine, and the second database corresponding to the row storage engine.

[0080] Step 302, determine the data index relationship between the first database and the second database based on the first index information and the second index information.

[0081] Among them, the data index relationship refers to the mapping or corresponding relationship between the first index information in the first database corresponding to the row storage engine and the second index information in the second database corresponding to the column storage engine, so as to ensure the consistency and accessibility of data between the two databases.

[0082] Exemplarily, the server determines the data index relationship between the first database corresponding to the row storage engine and the second database corresponding to the column storage engine according to the mapping and corresponding relationship between the first index information and the second index information.

[0083] Step 303, synchronize the target data to the second database based on the data index relationship.

[0084] Among them, data synchronization refers to the process of copying the target data from the first database corresponding to the row storage engine to the second database corresponding to the column storage engine according to the data index relationship.

[0085] Exemplarily, the server synchronizes the target data from the first database to the second database according to the data index relationship between the first database and the second database.

[0086] In this embodiment, by determining the first index information of the target data in the first database and constructing the corresponding second index information in the second database, the data index relationship between the two storage engines can be effectively established, so as to efficiently synchronize the target data to the second database, ensure the consistency and integrity of data between different databases, and thus improve the efficiency and flexibility of data processing.

[0087] In one embodiment, step 201 includes:

[0088] Obtain a data processing request for a distributed database system, and determine the data identifier carried in the data processing request; determine the data to be processed pointed to by the data identifier and the target storage engine for processing the data to be processed; process the data to be processed through the target storage engine to obtain the target data.

[0089] Among them, a data processing request refers to a request for performing specific operations or queries on a distributed database system. The data processing request can be triggered by a user through a terminal, an application deployed or installed on the terminal, or other systems, aiming to obtain, modify, delete, or analyze the data stored in the distributed database system.

[0090] Among them, a data identifier is an identifier used to identify data in a distributed database. The data identifier of each piece of data is unique. The data identifier can be a primary key, a unique index, or other fields that can ensure data uniqueness, so as to specify the specific data to be processed, that is, the data to be processed.

[0091] Among them, the data to be processed refers to the data that needs to be further operated or analyzed according to the data identifier in the data processing request. The data to be processed can be located in any database or node of the distributed database system and is retrieved and extracted through a data access layer or middleware, such as a storage engine (node), a cache node, a computing node, etc.

[0092] Among them, a target storage engine refers to a database engine responsible for processing the data to be processed based on a data processing request. In a distributed database system, it can include multiple types of storage engines, such as a row storage engine, a column storage engine, a graph storage engine, etc. Each storage engine has a specific data organization method, an index mechanism, and a query optimization strategy. By determining the target storage engine, the data processing request can be sent to the appropriate storage engine to process the involved data to be processed, ensuring that the data processing request can be completed efficiently and quickly, thereby improving the overall throughput of the distributed database system.

[0093] Among them, the process of processing the data to be processed by the target storage engine can include but is not limited to multiple steps such as data retrieval, data conversion, data analysis, data writing, etc. During the processing process, the target storage engine uses its own data organization method and index mechanism to accelerate the data access and processing speed, thereby improving the performance of the entire system.

[0094] In an exemplary embodiment, when processing a data processing request through a target storage engine, first, the data processing request is parsed by the target storage engine to obtain a data execution statement and the data identifier of the data to be processed targeted by the data execution statement. The data to be processed exists in the form of data shards, and each data shard has a corresponding data identifier; then, the data execution statement is routed by the target storage engine to the first database, the second database, or a database node corresponding to the data shard according to the data identifier; subsequently, the target storage engine corresponding to the database or database node processes the data shard according to the execution of the data execution statement to obtain target data.

[0095] Understandably, in some other embodiments, such as the operation of writing data, the target data can also be directly obtained from the data to be processed carried in the data processing request (i.e., the write data request). At this time, the data to be processed that needs to be written carried in the write data request can be directly determined as the target data.

[0096] Exemplarily, when a user initiates a data processing request for a distributed database system at a terminal, the server obtains the data processing request and determines the data identifier carried in the data processing request to facilitate further determining the data to be processed pointed to by the data identifier. Then, the server determines the data to be processed according to the data identifier and determines the target storage engine for processing the data to be processed. Finally, the server controls the target storage engine to process the data to be processed to obtain the target data.

[0097] In this embodiment, by identifying and analyzing the data processing request to obtain the data identifier, it is possible to accurately identify and locate the data to be processed and the target storage engine matching the data to be processed according to the data identifier, and then effectively process the data through the target storage engine to quickly obtain the target data, which is beneficial to improving the accuracy and efficiency of data processing in the distributed database system.

[0098] In an alternative embodiment, when there are multiple data processing requests, the currently being processed data processing request is defined as the first data processing request, and the other parallel data processing requests are defined as the second data processing requests. There is at least one second data processing request. The above method further includes the following steps:

[0099] Determine the request types of the first data processing request and the second data processing requests; determine the execution priority of the first data processing request relative to the second data processing requests according to the request types of the first data processing request and the second data processing requests; and based on the execution priority of the first data processing request, process the data to be processed by the target storage engine in the execution round of the first data processing request.

[0100] Among them, the execution priority refers to which data processing request can process the data to be processed first or synchronously process the data to be processed when there are multiple data processing requests that need to process a certain data to be processed. In an exemplary embodiment, during the process of processing the data to be processed based on the execution priority, a metadata lock can be introduced; specifically, when the first data processing request involves changing the metadata during execution, the server will automatically acquire the metadata lock of the metadata. If at this time another second data processing request operates on the data related to the metadata (including insertion, update, etc.), the metadata lock will coordinate the execution priorities between the first data processing request and the second data processing request, making the first data processing request being executed wait until the second data processing request holding the old version of the metadata submits and then execute the first data processing request. However, if both the first data processing request and the second data processing request are operations such as insertion and update, they can be executed synchronously to ensure the coherence and consistency of the data operation process, avoid data operation errors caused by metadata changes, and thus effectively maintain the stability and data integrity of the database.

[0101] In one embodiment, determining the data to be processed pointed to by the data identifier and the target storage engine for processing the data to be processed includes:

[0102] According to the mapping rules between the data to be processed and each storage engine, determine the intermediate storage engine associated with the data to be processed; according to the data processing request, the real-time load information of each intermediate storage engine, the distribution information of the data to be processed on each intermediate storage engine, and the object statistics information of the data table to which the data to be processed belongs, determine the execution cost of each intermediate storage engine for executing the data processing request; determine the target storage engine for executing the data processing request among each intermediate storage engine according to the execution cost.

[0103] Among them, the mapping rule refers to the association relationship between the data to be processed and the storage engine. When establishing the mapping rule from the data to be processed to the storage engine, the internal logical structure and actual access pattern of the data to be processed can be considered. For example, factors such as data type, data structure, data size, and data access mode can be used to determine which storage engine the data to be processed should be processed through, so as to optimize data distribution, improve data access efficiency, and ensure data consistency and availability.

[0104] Taking the Database-level (i.e., database level) mapping as an example, for data tables with tight business logic and often accessed as a whole, they can be completely mapped to a specific storage engine to ensure data integrity and operation consistency. In terms of Column-level (i.e., column level) mapping, for some columns containing a large amount of text data or frequently used for full-text search, they are separately mapped to a storage engine suitable for text processing to optimize query performance. For Data-partitioning (i.e., data partitioning) mapping, according to factors such as the time series of data (e.g., sales data divided by month or quarter) or business regions (e.g., user data in different regions), different partitions are flexibly allocated to different storage engines to achieve refined data management and improve the pertinence and efficiency of data processing.

[0105] Among them, the intermediate storage engine refers to the storage engine associated with the data to be processed. There are usually multiple intermediate storage engines, and the intermediate storage engine is determined based on mapping rules. When determining the intermediate storage engine, for example, for OLTP scenarios, such as real-time scheduling systems in the power industry and monitoring and data acquisition systems in smart grids, since they operate on the data to be processed frequently and focus on atomicity and immediacy, a row storage engine can be selected, making the data extremely efficient when performing row-by-row read and write operations and being able to quickly respond to operation requests such as account balance updates and order status modifications to ensure the efficient operation of transactions. For OLAP scenarios, such as enterprise sales data analysis and market trend prediction, where the data volume is huge and mainly used for complex data analysis calculations, a column storage engine can be selected to effectively improve the calculation speed when performing multi-dimensional data aggregation and grouping calculations by quickly locating and reading relevant column data and provide timely and accurate data support for decision-making.

[0106] Among them, the real-time load information refers to the information used to reflect the current running state and resource usage of each intermediate storage engine, including but not limited to CPU (i.e., Central Processing Unit) usage rate, memory occupancy, I / O (i.e., Input / Output) throughput, current task queue length, etc.

[0107] Among them, the distribution information refers to the distribution of the data to be processed on each database or database node, including but not limited to the quantity, type, distribution uniformity, distribution hotspots, distribution patterns, etc. of the data to be processed.

[0108] Among them, the object statistical information of the data table refers to the detailed information about the data table itself and its internal data structure, including but not limited to the number of rows, number of columns, data type distribution of columns, index usage, data distribution characteristics, etc. of the data table.

[0109] Among them, the execution cost refers to the measurement of resource consumption and time required to execute a data processing request. The execution cost of this embodiment comprehensively considers factors such as the real-time load of the storage engine, the distribution of data, and the statistical information of data tables, and can be used to evaluate the relative efficiency of different storage engines in executing data processing requests.

[0110] Exemplarily, the server determines intermediate storage engines associated with the data to be processed according to the mapping rules between the data to be processed and each storage engine; subsequently, the server determines the execution cost of each intermediate storage engine for executing the data processing request according to the data processing request, the real-time load information of each intermediate storage engine, the distribution information of the data to be processed on each intermediate storage engine, and the object statistical information of the data table to which the data to be processed belongs; finally, the server determines the target storage engine for executing the data processing request among each intermediate storage engine according to the measurement of resource consumption and time required when each intermediate storage engine executes the data processing request.

[0111] In this embodiment, by comprehensively considering the data processing request, the real-time load of the storage engine, the data distribution, and the statistical information of the data table, the execution cost of each intermediate storage engine can be more accurately quantified, and the intermediate storage engine with the lowest execution cost can be objectively and accurately selected to execute the data processing request, so as to optimize resource allocation, reduce processing time, improve the overall data processing efficiency, and can reasonably allocate the workload of each storage engine, avoid resource idleness or overload, and improve the overall utilization rate of storage resources.

[0112] In one embodiment, the mapping rule in the above steps is constructed by the following method:

[0113] Obtain the distributed data in the distributed database system and the data processing rules of each storage engine associated with the distributed data; construct the mapping rules between the distributed data and each storage engine according to the distributed data and the data processing rules; when at least one of the storage engine and the distributed data changes, update the mapping rules based on at least one of the changed storage engine and the changed distributed data.

[0114] Among them, the distributed data refers to the data stored in the distributed database system, and the distributed data is scattered and stored in multiple databases or multiple database nodes of the distributed database. The distributed data can include structured data (such as tables in a relational database), semi-structured data (such as JSON documents, that is, JavaScript Object Notation, or XML documents, that is, Extensible Markup Language), and unstructured data (such as text, images, or video files).

[0115] Among them, the data processing rules refer to the rules for how each storage engine processes distributed data, which may include but are not limited to the storage format of distributed data, index strategies, data compression methods, data partitioning strategies, etc.

[0116] Among them, the changed storage engine refers to the storage engine after changes in configuration, performance, or function. For changes in the storage engine, it may include but is not limited to adding a new storage engine, removing an old storage engine, upgrading the hardware or software of the storage engine, adjusting the parameter settings of the storage engine, etc.

[0117] Among them, the changed distributed data refers to the distributed data after changes in structure, content, or access mode. For changes in distributed data, it may include but is not limited to data addition, deletion, modification, reorganization of data tables, adjustment of data partitions, etc.

[0118] Among them, the update mapping rule refers to the process of adjusting or re - determining the mapping rule according to the new situation when the storage engine or distributed data changes, which involves re - evaluating the storage requirements of data, selecting a new storage engine, adjusting the data distribution strategy, etc., to ensure that the distributed data can continue to be processed efficiently by the appropriate storage engine and meet the performance and reliability requirements of the system.

[0119] Exemplarily, the server obtains the distributed data stored in each database or database node from the distributed database system and determines the data processing rules of each storage engine associated with each distributed data; subsequently, the server constructs a mapping rule between the distributed data and each storage engine according to the obtained distributed data and the determined corresponding data processing rules, so as to allocate a suitable storage engine for the data to be processed (i.e., distributed data) during the data processing process. In addition, the server also monitors the changes of the storage engine and distributed data in real - time, so that in the case of changes in at least one of the storage engine and distributed data, the server updates the mapping rule based on at least one of the changed storage engine and the changed distributed data.

[0120] In this embodiment, by obtaining and constructing the mapping rule between the distributed data and each storage engine, the storage and access paths of the distributed data can be optimized, the data retrieval and processing speed can be accelerated, and the overall system operation efficiency can be improved. At the same time, when the storage engine or distributed data changes, the mapping rule is dynamically updated, with strong adaptability, facilitating system expansion and upgrade. Without large - scale reconstruction, it can cope with data growth or storage strategy adjustment, ensuring that the distributed data can always be processed by the appropriate storage engine to meet the continuously changing business requirements.

[0121] In one embodiment, the above - mentioned method further includes:

[0122] Build a distributed database system, and deploy a first database and a second database in the distributed database system; determine a source database, obtain source data in the source database, and determine a data conversion rule from the source data to the distributed data according to the data compatibility information between the source database and the distributed database; convert the source data according to the data conversion rule to obtain distributed data, and migrate the distributed data to at least one of the first database and the second database in the distributed database system.

[0123] Among them, the source database refers to a database that contains the original data that needs to be migrated or synchronized to the distributed database system. The source database can be any type of database system that has data compatibility with the distributed database, such as relational database systems like MySQL (My Structured Query Language, that is, Structured Query Language), Oracle, or MongoDB (a database based on distributed file storage).

[0124] Among them, the source data refers to the original data stored in the source database. The source data can be structured data (such as table data in a relational database), semi-structured data (such as JSON or XML documents), or unstructured data (such as text, image, or video files).

[0125] Among them, the data compatibility information refers to the information about the differences in data structures and data types between the source database and the distributed database. The data compatibility information includes at least one of data structure information (such as table structure, field definition, field name, data type, primary key and foreign key constraints, etc.), data type information (such as detailed definitions and compatibilities of types such as integers, floating-point numbers, strings, dates, etc.), and database object information (such as tables, views, indexes, etc.). By carefully comparing the characteristics of the source database and the distributed database, the compatibility between the two can be accurately evaluated, the possible differences and potential migration risk points can be determined, so as to generate a detailed and comprehensive data conversion rule according to the compatibility evaluation results, and then obtain a data migration plan. For the situation where the data types of the source database and the distributed database system do not match, the data type conversion strategy can be accurately determined (such as converting certain specific data types in the source database to equivalent or suitable data types in the distributed database system). During the data migration process, an appropriate synchronization method can also be selected according to business requirements and data update frequency, such as a one-time full-volume migration or a regular incremental migration.

[0126] Among them, the data conversion rule refers to the rule for converting source data into the format and type acceptable to the distributed database system, including but not limited to data type conversion (such as converting a string to an integer), data structure adjustment (such as expanding a nested document into a flat table structure), or data value mapping (such as mapping a status code in an old system to a status description in a new system).

[0127] For example, for the conversion rule of SQL statements, convert the SQL syntax in the source database into a syntax form that the distributed database system can recognize and execute efficiently, determine the sequence and time nodes of migration, generate a data migration plan, and then, according to the data migration plan, write corresponding migration scripts and tools for performing operations such as data conversion, synchronization, and SQL statement conversion. Finally, before migration, test and verify the migration scripts and tools to ensure their correctness and reliability.

[0128] Among them, data migration refers to the process of transferring source data from the source database to the distributed database system, so as to accurately convert the source data into distributed data and store the distributed data in the appropriate location of the distributed database system.

[0129] Exemplarily, the server constructs a distributed database system and deploys a first database and a second database in the distributed database system; subsequently, the server determines the source database according to the data to be migrated and obtains the source data to be migrated from the source database. Then, the server determines the data conversion rule between the source data and the distributed data according to at least one of the data structure information, data type information, and database object information between the source database and the distributed database; and finally, converts the source data according to the data conversion rule to obtain distributed data, and migrates the distributed data to at least one of the first database and the second database in the distributed database system.

[0130] In this embodiment, determining the data conversion rule based on the data compatibility between the source database and the distributed database can ensure the integrity and accuracy of the source data during the migration process, which helps to achieve smooth data migration and integration between different database systems.

[0131] In one embodiment, the above method further includes:

[0132] Obtain the operation status data of the distributed database system; based on the operation status data, determine the operation status of the distributed database system; the operation status is used to characterize at least one of the performance metrics, status information, and fault information of the distributed database system.

[0133] Among them, the running state data refers to various parameters reflecting the current running status of the distributed database system, and the running state data may include at least one of the system's performance metrics, status information, and fault information.

[0134] Among them, the performance metrics refer to various quantitative data for measuring the performance of the distributed database system. The performance metrics include but are not limited to at least one of throughput, response time, CPU usage, memory occupancy, disk I / O rate, etc., which are used to evaluate the system's processing capacity, resource utilization, and overall performance. By monitoring and analyzing the performance metrics, performance bottlenecks can be identified, resource allocation can be optimized, and system efficiency can be improved.

[0135] Among them, the status information refers to various descriptive information about the current state of the distributed database system. The status information includes but is not limited to the running state of nodes (such as the online / offline state of nodes), the start and stop states of services, and the number of database connections, so as to determine the overall architecture of the distributed data system, the relationships between components, and the current working state.

[0136] Among them, the fault information refers to information about faults or abnormal events that occur in the distributed database system. The fault information includes but is not limited to the time of fault occurrence, location (such as which node, which component), fault type (such as hardware faults, software errors, node downtime, insufficient disk space, network connection interruption, and other abnormal situations), and the scope of influence of the fault, etc. It is crucial for quickly locating problems, taking emergency measures, and restoring the system to normal operation.

[0137] Among them, the running state refers to the current situation presented by the distributed database system based on its running state data. The running state can be normal, abnormal, or in a certain specific mode. By analyzing and interpreting the running state data, the running state of the system can be determined, and corresponding measures can be taken to optimize performance, handle faults, or adjust system configuration.

[0138] Exemplarily, the server obtains the running state data of the distributed database system; based on the running state data, determines the running state of the distributed database system; the running state is used to characterize at least one of the performance metrics, status information, and fault information of the distributed database system.

[0139] In this embodiment, by obtaining the running state data of the distributed database system and determining the running state of the system based on the running state data, it can reflect the key contents such as the performance metrics, status information, and fault information of the distributed database system in real time and accurately, so that system administrators or automated management systems can comprehensively understand the current situation of the database system, and timely discover potential performance bottlenecks, abnormal states, or fault problems.

[0140] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0141] Based on the same inventive concept, an embodiment of the present application also provides a cross-engine data processing device for implementing the above-mentioned cross-engine data processing method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the cross-engine data processing device provided below can refer to the limitations on the cross-engine data processing method in the above text, and will not be repeated here.

[0142] In an exemplary embodiment, as Figure 4 shown, a cross-engine data processing device is provided, including: a data determination module 401, a data update module 402, a condition determination module 403, and a data synchronization module 404, where:

[0143] The data determination module 401 is configured to determine target data to be written into the distributed database system; the distributed database system includes a first database corresponding to the row storage engine and a second database corresponding to the column storage engine.

[0144] The data update module 402 is configured to update the target data into the first database through the row storage engine.

[0145] The condition determination module 403 is configured to determine a cross-engine synchronization trigger condition configured for the attribute information of the target data.

[0146] The data synchronization module 404 is configured to synchronize the target data into the second database according to the data index relationship between the row storage engine and the column storage engine when the cross-engine synchronization trigger condition is satisfied.

[0147] In an optional embodiment, the data determination module 401 is further configured to obtain a data processing request for the distributed database system, and determine the data identifier carried in the data processing request; determine the to-be-processed data pointed to by the data identifier and the target storage engine for processing the to-be-processed data; and process the to-be-processed data through the target storage engine to obtain the target data.

[0148] In an optional embodiment, the data determination module 401 is further configured to determine an intermediate storage engine associated with the to-be-processed data according to the mapping rule between the to-be-processed data and each storage engine; determine the execution cost of each intermediate storage engine for executing the data processing request according to the data processing request, the real-time load information of each intermediate storage engine, the distribution information of the to-be-processed data on each intermediate storage engine, and the object statistical information of the data table to which the to-be-processed data belongs; and determine the target storage engine for executing the data processing request from each intermediate storage engine according to the execution cost.

[0149] In an optional embodiment, the data determination module 401 is further configured to obtain the distributed data in the distributed database system and the data processing rules of each storage engine associated with the distributed data; construct a mapping rule between the distributed data and each storage engine according to the distributed data and the data processing rules; and update the mapping rule based on at least one of the changed storage engine and the changed distributed data when at least one of the storage engine and the distributed data changes.

[0150] In an optional embodiment, the data synchronization module 404 is further configured to determine the first index information of the target data in the first database, and construct the second index information corresponding to the first index information in the second database according to the first index information; determine the data index relationship between the first database and the second database based on the first index information and the second index information; and synchronize the target data to the second database based on the data index relationship.

[0151] In an optional embodiment, the above device further includes a data migration module, configured to construct a distributed database system, and deploy the first database and the second database in the distributed database system; determine the source database, obtain the source data in the source database, and determine the data conversion rule from the source data to the distributed data according to the data compatibility information between the source database and the distributed database system; the data compatibility information includes at least one of data structure information, data type information, and database object information; convert the source data into distributed data according to the data conversion rule, and migrate the distributed data into at least one of the first database and the second database in the distributed database system.

[0152] In an alternative embodiment, the above device further includes a status monitoring module, configured to obtain the operation status data of the distributed database system; determine the operation status of the distributed database system based on the operation status data; and the operation status is used to characterize at least one of the performance metrics, status information, and fault information of the distributed database system.

[0153] Each module in the above cross-engine data processing device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.

[0154] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as data to be processed, target data, distributed database systems, each storage engine, each database, index information, and data identifiers. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a cross-engine data processing method.

[0155] Those skilled in the art can understand that Figure 5 the structure shown in

[0156] is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component layout.

[0157] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the cross-engine data processing method in the above embodiment are implemented.

[0158] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps of the cross-engine data processing method in the above embodiment are implemented.

[0159] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0160] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0161] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded in the present application.

[0162] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for processing cross-engine data, characterized in that: The method comprises: Determine target data to be written into a distributed database system; the distributed database system includes a first database corresponding to a row storage engine and a second database corresponding to a column storage engine; Updating the target data into the first database through the row storage engine; Determine a cross-engine synchronization trigger condition configured corresponding to the attribute information of the target data; When the cross-engine synchronization trigger condition is met, the target data is synchronized to the second database according to the data index relationship between the row storage engine and the column storage engine.

2. The method according to claim 1, characterized in that The step of synchronizing the target data to the second database according to the data index relationship between the row storage engine and the column storage engine includes: Determine first index information of the target data in the first database, and construct second index information corresponding to the first index information in the second database according to the first index information; Determine a data index relationship between the first database and the second database based on the first index information and the second index information; Based on the data index relationship, the target data is synchronized to the second database.

3. The method according to claim 1, characterized in that The step of determining target data to be written into the distributed database system includes: Obtaining a data processing request for the distributed database system, and determining a data identifier carried in the data processing request; Determine the data to be processed pointed to by the data identifier and a target storage engine for processing the data to be processed; The target storage engine processes the data to be processed to obtain target data.

4. The method according to claim 3, characterized in that: The step of determining the data to be processed pointed to by the data identifier and a target storage engine for processing the data to be processed includes: Determine the intermediate storage engine associated with the data to be processed according to the mapping rules between the data to be processed and each storage engine; Determine the execution cost of each intermediate storage engine executing the data processing request according to the data processing request, the real-time load information of each intermediate storage engine, the distribution information of the data to be processed on each intermediate storage engine, and the object statistics information of the data table to which the data to be processed belongs; A target storage engine for executing the data processing request is determined in each intermediate storage engine according to the execution cost.

5. The method according to claim 4, characterized in that The mapping rule is constructed by the following method: Obtaining distributed data in a distributed database system and data processing rules of each storage engine associated with the distributed data; According to the distributed data and the data processing rules, construct a mapping rule between the distributed data and each of the storage engines; In the case where at least one of the storage engine and the distributed data changes, the mapping rule is updated based on at least one of the changed storage engine and the changed distributed data.

6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Building a distributed database system, and deploying a first database and a second database in the distributed database system; Determine a source database, obtain source data in the source database, and determine a data conversion rule between the source data and the distributed data according to data compatibility information between the source database and the distributed database system; the data compatibility information includes at least one of data structure information, data type information, and database object information; The source data is converted according to the data conversion rule to obtain the distributed data, and the distributed data is migrated to at least one of the first database and the second database of the distributed database system.

7. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Acquiring operating status data of the distributed database system; Based on the operation status data, the operation status of the distributed database system is determined; the operation status is used to characterize at least one of a performance indicator, status information and fault information of the distributed database system.

8. A cross-engine data processing device, characterized in that: The device comprises: A data determination module, used to determine target data to be written into a distributed database system; the distributed database system includes a first database corresponding to a row storage engine and a second database corresponding to a column storage engine; A data update module, used to update the target data into the first database through the row storage engine; A condition determination module, used to determine a cross-engine synchronization trigger condition configured corresponding to the attribute information of the target data; A data synchronization module is used to synchronize the target data to the second database according to the data index relationship between the row storage engine and the column storage engine when the cross-engine synchronization trigger condition is met.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Multi-modal data storage method and device, equipment, medium and product

    CN120950005A

  • Block chain decentralized data security storage method and system

    CN120974544A