A method for implementing a distributed data warehouse for real-time data processing

By employing a distributed data warehouse approach and leveraging the collaborative work of a transaction manager, data shard receiver, and writer, the challenge of real-time data processing in the era of big data has been solved. This approach enables data writing and querying within seconds, improving the real-time performance and parallelism of data processing.

CN120578719BActive Publication Date: 2025-12-02GUANGZHOU YAXIN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511072939.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-12-02
Estimated Expiration
2045-08-01

AI Technical Summary

Technical Problem

How to efficiently and accurately process rapidly growing real-time data in the era of big data, and achieve data writing and instant querying functions in a very short time.

Method used

A real-time data processing distributed data warehouse method is adopted, which realizes data writing and instant querying in a very short time through the collaborative work of transaction manager, data shard receiver and data shard writer. The process includes the transaction manager distributing tasks to data shard receiver, data shard receiver sending the key of the consumed data and the consumed data to data shard writer, data shard writer determining the writing information and sending it to transaction manager, data version manager publishing data version, and transaction manager updating task status in real time.

Benefits of technology

It enables data writing and querying within seconds, ensuring that data is instantly visible in the data warehouse, improving the real-time performance and parallelism of data processing, optimizing the processing flow of the data shard writer, and reducing the interruption of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120578719B_ABST
    Figure CN120578719B_ABST
Patent Text Reader

Abstract

This application discloses a method for implementing a distributed data warehouse for real-time data processing, relating to the field of data processing technology. The method comprises: a transaction manager distributing tasks to a data shard receiver; the data shard receiver sending consumption information to the transaction manager; the data shard receiver first sending the key of the consumed data to a data shard writer, and then sending the consumed data to the data shard writer; the data shard writer determining the writing information based on the consumed data and the key; the data shard writer sending the writing information to the transaction manager; the transaction manager sending a notification message to a data version manager; the data version manager publishing a data version and obtaining the publication information; the data version manager sending the publication information to the transaction manager; and the transaction manager updating the transaction status of the tasks based on the consumption information, the writing information, and the publication information. This method enables second-level data visibility in a distributed data warehouse.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method for implementing a distributed data warehouse for real-time data processing. Background Technology

[0002] Real-time data processing technology plays a crucial role in today's information society. With the advent of the big data era, data is being generated at an increasingly faster speed and on a larger scale, making efficient and accurate processing of this data a major challenge for enterprises and organizations. Summary of the Invention

[0003] This application provides a method for implementing a distributed data warehouse for real-time data processing, aiming to enable data to be written and queried instantly in a very short time, ensuring that data can be queried immediately after being written within seconds.

[0004] To achieve the above objectives, this application provides the following technical solution:

[0005] A method for implementing a real-time data processing distributed data warehouse, applied in a data warehouse management system including a transaction manager, a data shard receiver, a data shard writer, and a data version manager, the method comprising:

[0006] The transaction manager distributes tasks to the data shard receiver so that while the data shard receiver is processing any one task, another task is queued for processing; the task is used to consume one second of data from a designated message caching system.

[0007] The data sharding receiver sends the consumption information of the task to the transaction manager;

[0008] The data sharding receiver first sends the key of the consumed data obtained from processing the task to the data sharding writer, and then sends the consumed data to the data sharding writer.

[0009] The data shard writer determines the writing information of the task based on the consumed data and the corresponding key;

[0010] The data shard writer sends the write information to the transaction manager;

[0011] Based on the write information, the transaction manager sends a corresponding notification message to the data version manager;

[0012] The data version manager responds to the notification message and publishes a data version to obtain the release information of the task;

[0013] The data version manager sends the release information to the transaction manager;

[0014] The transaction manager updates the transaction status of the task in real time based on the task's consumption, write, and publish information.

[0015] Optionally, the transaction manager distributes tasks to the data shard receiver such that while the data shard receiver is processing any one task, another task is queued for processing, including:

[0016] The transaction manager responds to the data write request sent by the client and generates a corresponding job; the job is used to obtain the data to be written from the specified message caching system in order to complete the consumption of the data to be written.

[0017] The transaction manager splits the job into multiple tasks and initializes the transaction status of the multiple tasks.

[0018] The transaction manager distributes multiple tasks to multiple data shard receivers such that while each data shard receiver is processing any one task, another task is queued for processing.

[0019] Optionally, the method further includes:

[0020] After determining that the task has been distributed to the data shard receiver, the transaction manager updates the transaction status of the task to "sent and ready to be consumed".

[0021] Optionally, the data sharding receiver sends the consumption information of the task to the transaction manager, including:

[0022] The data fragment receiver will sequentially save at least one task to the task queue.

[0023] The data fragment receiver processes the tasks in the task queue sequentially to obtain the processing result of each task;

[0024] If the processing result of the task contains consumption data, the data fragment receiver determines that the task has been processed successfully and sends the corresponding consumption information to the transaction manager.

[0025] Optionally, the method further includes:

[0026] If the processing result of the task is empty, the data shard receiver determines that the task processing has failed and sends a corresponding consumption failure report to the transaction manager;

[0027] Based on the consumption failure report of the task, the transaction manager reinitializes the task and resends the reinitialized task to the data shard receiver.

[0028] Optionally, the data shard writer determines the write information of the task based on the consumed data and the corresponding key, including:

[0029] After obtaining the key, the data shard writer loads the disk data corresponding to the key into memory;

[0030] After obtaining the consumed data, the data shard writer merges and sorts the consumed data with the disk data to obtain the merge sort result.

[0031] The data shard writer performs data write-to-disk operation on the consumed data according to the merge sort result, so as to obtain the corresponding write-to-disk result;

[0032] If the disk write result is successful, the data shard writer determines that the consumed data has been successfully written to the data warehouse and sends the corresponding write information to the transaction manager.

[0033] Optionally, the method further includes:

[0034] If the disk write result is a failure, the data shard writer determines that the consumption data write has failed and sends a corresponding write failure report to the transaction manager;

[0035] Based on the write failure report, the transaction manager sends the corresponding write task to the data shard receiver;

[0036] Based on the write task, the data sharding receiver resends the consumed data and the corresponding key to the data sharding writer, so that the data sharding writer rewrites the consumed data to disk.

[0037] Optionally, the transaction manager updates the transaction status of the task in real time based on the task's consumption information, write information, and publish information, including:

[0038] After obtaining the consumption information of the task, the transaction manager updates the transaction status of the task to "consumed".

[0039] After obtaining the write information of the task, the transaction manager updates the transaction status of the task to committed;

[0040] After obtaining the task's publication information, the transaction manager updates the task's transaction status to complete.

[0041] Optionally, the method further includes:

[0042] The transaction manager records the execution time of each stage of the task in real time; each stage includes at least a consumption stage and a writing stage; the consumption stage includes the process from the distribution of the task until the consumption information of the task is obtained; the writing stage includes the process from the acquisition of the consumption information of the task until the writing information of the task is obtained.

[0043] If the execution time of the consumption phase exceeds the specified time, the transaction manager reinitializes the task, sends the reinitialized task to the data shard receiver, and resets the execution time of the consumption phase.

[0044] If the execution time of the write phase exceeds the specified time, the transaction manager reinitializes the write task, sends the write task to the data shard receiver, and resets the execution time of the write phase; the write task is used to trigger the data shard receiver to send the consumed data and the corresponding key to the data shard writer.

[0045] A data warehouse management system, comprising:

[0046] Transaction manager, data shard receiver, data shard writer, and data version manager;

[0047] The transaction manager is used to: distribute tasks to the data shard receiver so that while the data shard receiver is processing any task, another task is waiting to be processed; the task is used to consume one second of data from a specified message caching system; send corresponding notification messages to the data version manager based on write information; and update the transaction status of the task in real time based on the consumption information, write information, and publication information of the task.

[0048] The data sharding receiver is used to: send the consumption information of the task to the transaction manager; first send the key of the consumption data obtained from processing the task to the data sharding writer, and then send the consumption data to the data sharding writer;

[0049] The data shard writer is used to: determine the write information of the task based on the consumed data and the corresponding key; and send the write information to the transaction manager.

[0050] The data version manager is used to: respond to the notification message, publish a data version to obtain the release information of the task; and send the release information to the transaction manager.

[0051] The technical solution provided in this application involves a transaction manager distributing tasks to a data shard receiver. The data shard receiver sends consumption information to the transaction manager. The data shard receiver first sends the key of the consumed data to the data shard writer, and then sends the consumed data to the data shard writer. The data shard writer determines the write information based on the consumed data and the corresponding key. The data shard writer sends the write information to the transaction manager. The transaction manager sends a notification message to the data version manager. The data version manager publishes the data version to obtain the publication information. The data version manager sends the publication information to the transaction manager. The transaction manager updates the transaction status of the task in real time based on the consumption information, write information, and publication information. This application enables data writing and instant querying in a very short time, ensuring that data can be queried in the data warehouse immediately after being written within seconds. Attached Figure Description

[0052] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0053] Figure 1 A schematic diagram of the architecture of a data warehouse management system provided in an embodiment of this application;

[0054] Figure 2 A flowchart illustrating a method for implementing a distributed data warehouse for real-time data processing, provided in an embodiment of this application;

[0055] Figure 3 A flowchart illustrating a task transaction status update process provided in an embodiment of this application;

[0056] Figure 4 A flowchart illustrating a task processing procedure provided in an embodiment of this application;

[0057] Figure 5 A flowchart illustrating the process of writing consumer data to disk, provided as an embodiment of this application;

[0058] Figure 6 A flowchart illustrating another implementation method for a real-time data processing distributed data warehouse provided in this application embodiment;

[0059] Figure 7 This is a flowchart illustrating the operational logic of a data warehouse management system provided in an embodiment of this application. Detailed Implementation

[0060] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0061] In this application, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0062] like Figure 1 The diagram shown is an architectural schematic of a data warehouse management system provided in an embodiment of this application, including the following modules: transaction manager 100, data shard receiver 200, data shard writer 300, and data version manager 400.

[0063] In some examples, the data shard receiver, data shard writer, and data version manager reside in the same instance, meaning they all run within the same instance.

[0064] It should be noted that the data warehouse management system shown in the embodiments of this application can be a distributed system (specifically, a distributed data warehouse). An instance can be understood as a computer node in a distributed system. Since a distributed system includes multiple computer nodes, the number of instances can be multiple. Correspondingly, the number of data shard receivers, data shard writers, and data version managers can each be multiple.

[0065] In some examples, the transaction manager can run within any instance of the data warehouse management system, or as a separate instance.

[0066] The functions of the transaction manager can be summarized as follows: receiving data write requests from clients, generating corresponding jobs based on these requests; splitting jobs into multiple tasks and distributing these tasks to multiple data shard receivers; maintaining the transaction status of each task, including initialization, sent and awaiting consumption, consumed, committed, and completed; metadata management, i.e., driving the data version manager to update data versions and recording the data version merging path; maintaining the task queue for each data shard receiver, ensuring that when each data shard receiver executes any task, there is another task waiting in the queue for processing, and distributing an appropriate number of tasks to each data shard receiver based on its load capacity.

[0067] The functions of a data shard receiver can be summarized as follows: receiving tasks sent by the transaction manager; maintaining its own task queue; reporting task consumption information to the transaction manager, which includes the task's consumption data offset attribute (used to represent the data position of the consumption data in the specified message caching system, the offset attribute includes start offset and end offset, the start offset represents the starting position of the consumption data, and the end offset represents the ending position of the consumption data), and the task's corresponding write number (representing the number of the data shard writer responsible for writing the consumption data to disk); sending the consumption data key to all data shard writers; and sending the consumption data to the corresponding data shard writer (the corresponding data shard writer can be determined based on the consumption data key).

[0068] The functions of a data shard writer can be summarized as follows: receiving keys sent by multiple data shard receivers and loading the disk data corresponding to the keys into memory; receiving consumer data sent by multiple data shard receivers and using a preset data storage model (such as the LSM-Tree storage model) to write the consumer data to disk (i.e., writing the consumer data into the system); and sending task write information to the transaction manager, which includes the task number, the instance number to which it belongs, the write result (i.e., the specific write location), the write shard number (i.e., the shard where the write location is located), and the write batch number (i.e., the number of times the data shard writer executes the write task).

[0069] The functions of the data version manager can be summarized as follows: receiving notification messages from the transaction manager corresponding to the task; publishing the data version based on the notification message to obtain the task's publication information; sending the publication information to the transaction manager and determining the data version merge path.

[0070] A job can be understood as a real-time data processing task corresponding to a data write request submitted by a client. This real-time data processing task is persistent; as long as the user does not cancel the job, it will continue to generate tasks to trigger data consumption. Data consumption refers to reading data from a specified message caching system (such as Kafka, which can be understood as an external message caching component where real-time data is first sent), processing the read data, and saving it to the local system (which can be understood as a data warehouse).

[0071] In some examples, if the number of tasks is greater than the number of data fragment receivers, each data fragment receiver can process multiple tasks in parallel. If the number of tasks is less than the number of data fragment receivers, tasks are randomly distributed to data fragment receivers with higher load capacity, and tasks may be temporarily withheld from data fragment receivers with lower load capacity.

[0072] The so-called data version can be understood as the data version of this system (i.e., distributed data warehouse) (e.g., 1.2.3881). Since the data shard writer uses the LSM-Tree model to store data, a data version will be generated each time the task is successfully processed. The LSM-Tree model can be used to merge the various data versions into the latest data version to achieve path merging of data versions.

[0073] In some examples, the transaction manager generates a notification message for the task based on the write information obtained and sends the notification message to all data version managers in the system so that all data version managers can synchronously publish data versions. During the data version publication process, the LSM-Tree model is used to merge the various data versions into the latest data version to achieve data version path merging and obtain the corresponding publication information.

[0074] Optionally, a method for implementing real-time data processing based on the interaction between the transaction manager, data shard receiver, data shard writer, and data version manager can be found in [reference needed]. Figure 2 The steps shown are accompanied by corresponding explanations.

[0075] S201: The transaction manager distributes tasks to the data shard receiver so that while the data shard receiver is processing any one task, another task is queued for processing.

[0076] The task is used to consume one second of data from the specified message caching system.

[0077] It should be noted that the transaction manager distributes tasks to the data shard receiver so that while the data shard receiver is processing any task, another task is waiting in the queue for processing. This avoids the data shard receiver having to wait for a new task to be sent after processing a task, allowing the data shard receiver to process the next task immediately after completing one task. This makes the data consumption process of the data shard receiver uninterrupted (current open source data consumption is intermittent), effectively enhancing the real-time performance of the system's data processing.

[0078] Optionally, the transaction manager distributes tasks to the data shard receiver so that while the data shard receiver is processing one task, there is another task waiting in the queue for processing. See also: Figure 3 The steps described above, along with their corresponding explanations.

[0079] S202: The data fragment receiver sends the task consumption information to the transaction manager.

[0080] The consumption information of a task is obtained by the data sharding receiver processing the task. Specifically, the data sharding receiver consumes one second of data from the specified message cache system according to the task to obtain the consumed data. Based on the offset attribute of the consumed data and the write number corresponding to the task, the consumption information of the task is generated.

[0081] It is important to note that if the data shard receiver cannot consume data from the specified message cache system based on the task, that is, if the data shard receiver does not obtain the data to be consumed, the data shard receiver will not generate consumption information for the task, nor will it be able to send consumption information to the transaction manager.

[0082] Optionally, the implementation process of the data shard receiver sending task consumption information to the transaction manager can be found in [reference needed]. Figure 4 The steps shown are accompanied by corresponding explanations.

[0083] S203: The data fragmentation receiver first sends the key of the consumed data obtained from the processing task to the data fragmentation writer, and then sends the consumed data to the data fragmentation writer.

[0084] The key of the consumption data is essentially the primary key of the consumption data. This primary key uniquely identifies the consumption data, and when the specified message caching system (and the system shown in the embodiments of this application) stores multiple consumption data, they are all sorted according to the primary key.

[0085] It is important to note that if the system contains multiple data fragment writers, the data fragment receiver will send the key of the consumed data to each of the multiple data fragment writers respectively.

[0086] S204: The data shard writer determines the writing information of the task based on the consumed data and the corresponding key.

[0087] Since the data sharding receiver sends the key to the consumer data before sending the consumer data, the data sharding writer has enough time to load the key's disk data into memory during this process, thus saving the writing time of the consumer data and improving the real-time performance of the system's data processing.

[0088] Optionally, the data shard writer determines the task's write information based on the consumed data and the corresponding key. See [link to relevant documentation]. Figure 5 The steps shown are accompanied by corresponding explanations.

[0089] S205: The data shard writer sends the write information to the transaction manager.

[0090] S206: The transaction manager sends a corresponding notification message to the data version manager based on the write information.

[0091] S207: The data version manager responds to the notification message, publishes the data version, and obtains the task's release information.

[0092] The implementation process of data version release is common knowledge familiar to those skilled in the art, and will not be elaborated here.

[0093] It should be noted that the data version manager responds to notification messages and publishes data versions to obtain task release information. The consumption data corresponding to the task can then be queried, meaning the consumption data is visible.

[0094] S208: The data version manager sends the release information to the transaction manager.

[0095] S209: The transaction manager updates the transaction status of a task in real time based on the task's consumption, write, and publish information.

[0096] The transaction manager updates the transaction status of tasks in real time based on their consumption, writing, and publishing information. It can monitor the status of tasks at each stage and take corresponding measures to improve data processing efficiency.

[0097] Optionally, after obtaining the consumption information of a task, the transaction manager updates the transaction status of the task to "consumed"; after obtaining the write information of a task, the transaction manager updates the transaction status of the task to "committed"; and after obtaining the publish information of a task, the transaction manager updates the transaction status of the task to "completed".

[0098] Optionally, the transaction manager records the execution time of each stage of a task in real time. Each stage includes at least a consumption stage and a write stage. The consumption stage includes the process from task distribution until the task's consumption information is obtained. The write stage includes the process from obtaining the task's consumption information until the task's write information is obtained.

[0099] If the execution time of the consumption phase exceeds the specified time, the transaction manager reinitializes the task, sends the reinitialized task to the data shard receiver, and resets the execution time of the consumption phase.

[0100] If the execution time of the write phase exceeds the specified time, the transaction manager reinitializes the write task, sends the write task to the data shard sink, and resets the execution time of the write phase. This write task is used to trigger the data shard sink to send the consumed data and the corresponding key to the data shard writer.

[0101] In a possible implementation, the transaction manager can maintain a task list that records the execution time of each task at each stage, and persist the transaction status of each task and the task list to the system log. The task can be recovered through the log in the event of an abnormal system restart.

[0102] In some examples, the system shown in this application embodiment needs to preset corresponding system interfaces in order to implement the processes shown in S201-S209 above. The system interfaces include external interfaces and internal interfaces. The external interfaces include: interface 1 (start / stop real-time write job) and interface 2 (query the status of real-time write job and the transaction status of task). The internal interfaces include interface 3 (transaction manager sends task to data shard receiver), interface 4 (data shard receiver sends key to data shard writer), interface 5 (data shard receiver sends consumed data to data shard writer), interface 6 (data shard receiver reports consumption information to transaction manager), interface 7 (data shard writer reports write information to transaction manager), interface 8 (transaction manager sends notification message to data version manager), and interface 9 (data version manager reports release information to transaction manager).

[0103] In some examples, the processes shown in S201-S209 can be summarized as follows in specific use cases: Figure 6The process is as follows: Transaction manager issues a task to data shard receiver; data shard receiver consumes one second of data based on the task and obtains the consumption result; data shard receiver reports the consumption result to transaction manager; transaction manager transitions the task's transaction status to "consumed"; data shard receiver sends the data key to data shard writer; data shard writer loads the disk data into memory based on the key; data shard writer receives data sent by data shard receiver; data shard writer performs a merge operation based on the data and disk data; data shard writer writes the data to disk; data shard writer reports the write result to transaction manager; transaction manager transitions the task's transaction status to "committed"; transaction manager notifies data version manager to publish a data version; data version manager publishes the data version; data version manager sends the publication result to transaction manager; transaction manager transitions the task's transaction status to "completed," confirming the data is queryable.

[0104] In some examples, for application scenarios where the system contains multiple instances, the processes shown in S201-S209 can be summarized as follows in specific use cases: Figure 7 The process is as follows: 1. The user initiates a real-time data write task; 2. The transaction manager distributes the task to the data shard receivers of each instance; 3. The data shard receivers receive one second of data from Kafka; 4. The data shard receivers report the consumption information to the transaction manager; 5. The transaction manager issues new consumption tasks to the data shard receivers; 6. The data shard receivers package the key and send it to the relevant instances; 7. The data shard writer preloads the disk data into memory based on the received key; 8. The data shard receivers distribute the data to different instances based on the key; 9. The data shard writer performs a merge operation based on the key and writes the data to disk; 10. The data shard writer reports the write information to the transaction manager; 11. The transaction manager notifies the version manager to publish the data; 12. The data version manager merges the data versions and makes this write visible; 13. The data version manager notifies the transaction manager that the data version has been successfully published.

[0105] Combination Figures 3-5The method shown in this application embodiment has the following advantages: 1. It can enhance the parallelism of data processing and improve the real-time performance of data processing, so that the data shard receiver no longer waits for a single task scheduling, but immediately consumes the data of the next task after consuming a segment of data; 2. It refines the status and management granularity of real-time processing tasks, allowing tasks to be retried at a lower cost when errors or anomalies occur; 3. It optimizes the processing and reporting relationship of the data shard writer, enabling the data shard writer to directly report writing information to the transaction manager, and the transaction manager clearly understands the transaction status of each task in the consumption and writing chain, which can better achieve real-time data processing; 4. After data consumption is completed, the key is sent to the data shard writer in advance, thereby saving the waiting time for consumed data to be read from disk to memory; 5. It can truly achieve second-level data visibility, and data can be queried in the data warehouse of this system within 1-2 seconds after being written to the Kafka system.

[0106] The processes shown in S201-S209 above enable data to be written and queried in a very short time, ensuring that data can be queried in the data warehouse after being written within seconds.

[0107] like Figure 3 The diagram shown is a flowchart illustrating a task transaction status update process provided in an embodiment of this application, including the following steps.

[0108] S301: The transaction manager responds to the data write request sent by the client and generates the corresponding job.

[0109] The job is used to obtain the data to be written from the specified message caching system in order to consume the data to be written.

[0110] In some examples, a data write request includes specifying the brokerserver of a message caching system (such as Kafka), the topic, the partition, authentication ticket information (used to access the Kafka system), and the associated database table for data writing. The brokerserver can be understood as a Kafka server; the Kafka system consists of multiple servers. A topic is a thread for messages (specifically, data within the Kafka system), essentially a name for a type of message. For example, if a set of location signaling messages is sent to Kafka, the corresponding topic must be specified. When the transaction manager reads the message from the Kafka system, it needs to use the corresponding topic to retrieve the signaling messages for that location. A partition refers to a partition of a topic. Because the data in a topic can be quite large, it is split into multiple partitions, distributing the data across multiple servers. Each server corresponds to one partition, allowing data to be retrieved from each partition simultaneously, achieving parallel processing of multiple data streams.

[0111] S302: The transaction manager splits the job into multiple tasks and initializes the transaction state of the multiple tasks.

[0112] The transaction manager can split a job into multiple tasks based on the number of partitions in the Kafka system (each partition can be set to have one second of data, meaning data can be read within one second). Each task is used to consume data from a corresponding partition.

[0113] It should be noted that if the transaction manager initializes the transaction state of multiple tasks, then the transaction state of each task can be recorded as initialization.

[0114] S303: The transaction manager distributes multiple tasks to multiple data shard receivers so that while each data shard receiver is processing one task, another task is queued for processing.

[0115] When the transaction manager distributes a task for the first time, each data shard receiver will distribute two tasks. After the transaction manager receives the consumption information sent by the data shard receiver, it will send a new task to the data shard receiver. This ensures that while each data shard receiver is processing any task, there is another task waiting in the queue for processing, thus achieving uninterrupted data consumption and parallel data processing.

[0116] S304: After determining that the task has been distributed to the data fragment receiver, the transaction manager updates the transaction status of the task to "sent and ready to be consumed".

[0117] The transaction manager updates the transaction status of a task to "sent and ready for consumption" and can monitor the consumption status (i.e., processing status) of a task after it has been split in real time.

[0118] The processes shown in S301-S304 above allow the transaction manager to refine the task granularity. The transaction manager can see the consumption status after the specific task is split, which increases the concurrency and refines the management granularity. When a certain stage of a task (such as distribution failure) fails, it can be retried at a single point without having to start over on a large scale.

[0119] like Figure 4 The diagram shown is a flowchart of a task processing procedure provided in an embodiment of this application, including the following steps.

[0120] S401: The data fragment receiver will save at least one acquired task to the task queue in sequence.

[0121] The data shard receiver can sequentially save at least one obtained task to the task queue according to the task timestamp (i.e., generation time) from earliest to latest.

[0122] In some examples, the data fragment receiver can also save at least one acquired task to the task queue in ascending order of task number.

[0123] S402: The data fragment receiver processes the tasks in the task queue sequentially to obtain the processing result of each task.

[0124] The data fragment receiver processes the current task first, and only after obtaining the processing result of the current task will it retrieve the next task from the task queue for processing.

[0125] S403: If the processing result of the task contains consumption data, the data fragment receiver determines that the task processing was successful and sends the corresponding consumption information to the transaction manager.

[0126] The data shard receiver generates task consumption information based on the consumption data shown in the task processing results and sends the consumption information to the transaction manager.

[0127] S404: If the task processing result is empty, the data fragment receiver determines that the task processing has failed and sends the corresponding consumption failure report to the transaction manager.

[0128] If the processing result of a task is empty, the data shard receiver generates a task consumption failure report, which may include the task number and the instance number to which the data shard receiver belongs.

[0129] S405: Based on the task consumption failure report, the transaction manager reinitializes the task and resends the reinitialized task to the data shard receiver.

[0130] After receiving a task consumption failure report, the transaction manager resends the reinitialized task to the data shard receiver so that the data shard receiver can reprocess the task.

[0131] The processes described in S401-S405 above can reinitialize the task in the event of task processing failure, and resend the reinitialized task to the data fragment receiver to avoid missing data consumption. Only the failed task is reprocessed, without the need for large-scale data consumption, which effectively improves the data processing efficiency of the system.

[0132] like Figure 5 The diagram shown is a flowchart illustrating a process for writing consumer data to disk according to an embodiment of this application, including the following steps.

[0133] S501: After obtaining the key, the data shard writer loads the disk data corresponding to the key into memory.

[0134] The data shard writer retrieves the corresponding disk data from the hardware storage resource area of ​​the system based on the key, and loads the disk data into memory.

[0135] S502: After obtaining the consumed data, the data shard writer merges and sorts the consumed data with the disk data to obtain the merge sort result.

[0136] The so-called merge sort is the merge operation, which merges the consumer data and disk data based on their respective timestamps, so that data with earlier timestamps are ranked before data with later timestamps.

[0137] S503: The data shard writer writes the consumed data to disk based on the merge sort result to obtain the corresponding disk write result.

[0138] The process of writing consumer data to disk essentially involves writing the consumer data into this system (i.e., the distributed data warehouse).

[0139] S504: If the disk write result is successful, the data shard writer confirms that the consumed data has been successfully written to the data warehouse and sends the corresponding write information to the transaction manager.

[0140] If the disk write result is successful, the data shard writer generates write information based on the task number, the instance number to which the data shard writer belongs, the write location of the consumed data, the shard number to which the consumed data is written, and the write batch number (i.e., the number of times the data shard writer executes the write task).

[0141] S505: If the disk write result fails, the data shard writer determines that the data consumption write has failed and sends the corresponding write failure report to the transaction manager.

[0142] The write failure report includes the task number, the instance number to which the data shard writer belongs, and the data shard writer number.

[0143] S506: The transaction manager sends the corresponding write task to the data shard receiver based on the write failure report.

[0144] Upon receiving a write failure report, the transaction manager generates a corresponding write task based on the report and sends the task to the corresponding data shard receiver according to the instance number shown in the report.

[0145] S507: The data fragment receiver resends the consumed data and the corresponding key to the data fragment writer based on the write task, so that the data fragment writer can rewrite the consumed data to disk.

[0146] After receiving a write task, the data fragment receiver sends the consumption data and the corresponding key to the corresponding data fragment writer according to the number of the data fragment writer shown in the write task.

[0147] The processes described in S501-S507 above can reinitialize the write task and resend the write task to the data shard receiver in the event of a data write failure, so as to avoid missing any write data. Only the consumer data that failed to write is rewritten, without the need for large-scale data writing, which effectively improves the data processing efficiency of the system.

[0148] While several specific implementation details are included in the foregoing discussion, these should not be construed as limiting the scope of this application. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0149] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A method for implementing a distributed data warehouse for real-time data processing, characterized in that, The method, applied in a data warehouse management system comprising a transaction manager, a data shard receiver, a data shard writer, and a data version manager, includes: The transaction manager distributes tasks to the data shard receivers, so that while the data shard receivers are processing any one task, another task is queued and waiting to be processed. The task is used to consume one second of data from a designated message caching system. When the transaction manager distributes the task for the first time, each data shard receiver will be distributed two tasks. After the transaction manager receives the consumption information sent by the data shard receiver, it will send a new task to the data shard receiver. The data sharding receiver sends the consumption information of the task to the transaction manager; The data sharding receiver first sends the key of the consumed data obtained from processing the task to the data sharding writer, and then sends the consumed data to the data sharding writer. The data shard writer determines the writing information of the task based on the consumed data and the corresponding key; the writing information is determined based on the disk write result of the consumed data, which is obtained by writing the consumed data to disk based on the merge sort result between the disk data corresponding to the key pre-loaded in memory and the consumed data; The data shard writer sends the write information to the transaction manager; Based on the write information, the transaction manager sends a corresponding notification message to the data version manager; The data version manager responds to the notification message and publishes a data version to obtain the release information of the task; The data version manager sends the release information to the transaction manager; The transaction manager updates the transaction status of the task in real time based on the task's consumption, write, and publish information.

2. The method according to claim 1, characterized in that, The transaction manager distributes tasks to the data shard receiver so that while the data shard receiver is processing any task, another task is queued for processing, including: The transaction manager responds to the data write request sent by the client and generates a corresponding job; the job is used to obtain the data to be written from the specified message caching system in order to complete the consumption of the data to be written. The transaction manager splits the job into multiple tasks and initializes the transaction status of the multiple tasks. The transaction manager distributes multiple tasks to multiple data shard receivers such that while each data shard receiver is processing any one task, another task is queued for processing.

3. The method according to claim 2, characterized in that, The method further includes: After determining that the task has been distributed to the data shard receiver, the transaction manager updates the transaction status of the task to "sent and ready to be consumed".

4. The method according to claim 1, characterized in that, The data sharding receiver sends the consumption information of the task to the transaction manager, including: The data fragment receiver will sequentially save at least one task to the task queue. The data fragment receiver processes the tasks in the task queue sequentially to obtain the processing result of each task; If the processing result of the task contains consumption data, the data fragment receiver determines that the task has been processed successfully and sends the corresponding consumption information to the transaction manager.

5. The method according to claim 4, characterized in that, The method further includes: If the processing result of the task is empty, the data shard receiver determines that the task processing has failed and sends a corresponding consumption failure report to the transaction manager; Based on the consumption failure report of the task, the transaction manager reinitializes the task and resends the reinitialized task to the data shard receiver.

6. The method according to claim 1, characterized in that, The data shard writer determines the write information of the task based on the consumed data and the corresponding key, including: After obtaining the key, the data shard writer loads the disk data corresponding to the key into memory; After obtaining the consumed data, the data shard writer merges and sorts the consumed data with the disk data to obtain the merge sort result. The data shard writer performs data write-to-disk operation on the consumed data according to the merge sort result, so as to obtain the corresponding write-to-disk result; If the disk write result is successful, the data shard writer determines that the consumed data has been successfully written to the data warehouse and sends the corresponding write information to the transaction manager.

7. The method according to claim 6, characterized in that, The method further includes: If the disk write result is a failure, the data shard writer determines that the consumption data write has failed and sends a corresponding write failure report to the transaction manager; Based on the write failure report, the transaction manager sends the corresponding write task to the data shard receiver; Based on the write task, the data sharding receiver resends the consumed data and the corresponding key to the data sharding writer, so that the data sharding writer rewrites the consumed data to disk.

8. The method according to claim 1, characterized in that, The transaction manager updates the transaction status of the task in real time based on the task's consumption, write, and publish information, including: After obtaining the consumption information of the task, the transaction manager updates the transaction status of the task to "consumed". After obtaining the write information of the task, the transaction manager updates the transaction status of the task to committed; After obtaining the task's publication information, the transaction manager updates the task's transaction status to complete.

9. The method according to claim 1, characterized in that, The method further includes: The transaction manager records the execution time of each stage of the task in real time; each stage includes at least a consumption stage and a writing stage; the consumption stage includes the process from the distribution of the task until the consumption information of the task is obtained; the writing stage includes the process from the acquisition of the consumption information of the task until the writing information of the task is obtained. If the execution time of the consumption phase exceeds the specified time, the transaction manager reinitializes the task, sends the reinitialized task to the data shard receiver, and resets the execution time of the consumption phase. If the execution time of the write phase exceeds the specified time, the transaction manager reinitializes the write task, sends the write task to the data shard receiver, and resets the execution time of the write phase; the write task is used to trigger the data shard receiver to send the consumed data and the corresponding key to the data shard writer.

10. A data warehouse management system, characterized in that, include: Transaction manager, data shard receiver, data shard writer, and data version manager; The transaction manager is used to: distribute tasks to the data shard receivers, so that while the data shard receivers are processing any one task, another task is queued and waiting to be processed; the task is used to consume one second of data from a designated message caching system; send a corresponding notification message to the data version manager based on the write information; and update the transaction status of the task in real time based on the consumption information, write information, and publication information of the task; wherein, when the transaction manager distributes the task for the first time, each data shard receiver will distribute two tasks accordingly, and after the transaction manager receives the consumption information sent by the data shard receiver, it will send a new task to the data shard receiver; The data sharding receiver is used to: send the consumption information of the task to the transaction manager; first send the key of the consumption data obtained from processing the task to the data sharding writer, and then send the consumption data to the data sharding writer; The data shard writer is used to: determine the writing information of the task based on the consumed data and the corresponding key; The write information is sent to the transaction manager; the write information is determined based on the disk write result of the consumed data, which is obtained by writing the consumed data to disk based on the merge sort result between the disk data corresponding to the key pre-loaded in memory and the consumed data; The data version manager is used to: respond to the notification message, publish a data version to obtain the release information of the task; and send the release information to the transaction manager.

Citation Information

Patent Citations

  • Transaction processing method, system and device, electronic equipment and storage medium

    CN116578395A