Data concurrent processing method and system and electronic equipment
By independently processing the processing of data operation requests in the storage system for concurrent processing, the delay problem caused by serial processing is solved, and more efficient data processing is achieved.
Patent Information
- Application Number
- CN202311640962.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-01
- Publication Date
- 2025-06-03
AI Technical Summary
In the storage system, when multiple clients perform concurrent data operations on multiple data files, serial processing operation requests lead to large delays, affecting data processing efficiency.
By independent of the processing of data operation requests from the submission process, the data operation requests can be processed concurrently, reducing data processing delays and improving efficiency. The specific method includes obtaining batch operation requests, performing concurrent processing, obtaining data processing results, adding the results to the queue to be submitted, and storing the results in the queue.
Concurrent processing of data operation requests is realized, reducing data processing delays, improving data processing efficiency, and improving the overall performance of the system.
Smart Images

Figure CN120086223A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and more particularly, to a data concurrent processing method, system, and electronic device. Background Art
[0002] In a storage system, if multiple clients perform data operations on multiple data file requests, multiple operation requests will be concurrently issued. However, if each operation request is processed serially, there will be a large delay. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a data concurrent processing method, system, and electronic device, which can separate the processing process of data operation requests from the submission process, that is, enable the submission process to be executed for the data processing results corresponding to the data operation requests, so that the data operation requests can be concurrently processed, reducing data processing latency and improving data processing efficiency.
[0004] In a first aspect, an embodiment of the present invention provides a data concurrent processing method, the method comprising:
[0005] Obtaining a batch operation request, the batch operation request including at least one data operation request;
[0006] Concurrently processing the batch operation request to obtain at least one data processing result;
[0007] Adding each of the data processing results to a to-be-submitted queue;
[0008] Storing each of the data processing results in the to-be-submitted queue.
[0009] Optionally, adding each of the data processing results to the to-be-submitted queue includes:
[0010] Storing each of the data processing results in a corresponding memory;
[0011] Generating a to-be-submitted request corresponding to each of the data processing results;
[0012] Adding each of the to-be-submitted requests to the to-be-submitted queue.
[0013] Optionally, concurrently processing the batch operation request includes:
[0014] Concurrently processing the batch operation request through a predetermined concurrent model.
[0015] Optionally, the concurrent model is a Pipeline model or an RTC model.
[0016] Optionally, obtaining the batch operation request includes:
[0017] Obtain the batch operation request through a predetermined communication network.
[0018] Optionally, the predetermined communication network is a TCP network or a GRPC network.
[0019] Optionally, a data verification field is included in the data packet corresponding to the data operation request.
[0020] Optionally, the data operation request has a corresponding request identifier;
[0021] The batch operation request has a corresponding communication connection identifier, and the communication connection identifier is used to characterize the communication connection for receiving the batch operation request.
[0022] Optionally, the method further includes:
[0023] Determine the feedback messages corresponding to each of the data operation requests;
[0024] Add each of the feedback messages to a reply queue;
[0025] Retrieve at least one feedback message from the reply queue through a corresponding first thread to obtain a message combination packet;
[0026] Feed back the message combination packet to the corresponding client through the corresponding communication connection.
[0027] In a second aspect, an embodiment of the present invention provides a data concurrent processing method, and the method includes:
[0028] Add the received data operation request to a sending queue;
[0029] Retrieve at least one data operation request from the sending queue through a corresponding second thread to obtain a batch operation request;
[0030] Send the batch operation request through the corresponding communication connection.
[0031] Optionally, the retrieving at least one data operation request from the sending queue through the corresponding second thread to obtain a batch operation request includes:
[0032] Obtain the batch operation request in response to the data volume corresponding to the retrieved data operation request reaching a threshold, or the sending queue being empty.
[0033] Optionally, the mapping relationship between the data operation request and the corresponding feedback message is determined through the corresponding request identifier.
[0034] In a third aspect, an embodiment of the present invention provides a data concurrent processing system, and the system includes:
[0035] A server, configured to execute the method described in the first aspect of the embodiments of the present invention; and
[0036] At least one client, configured to execute the method described in the second aspect of the embodiments of the present invention.
[0037] In a fourth aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor, where the memory is used to store one or more computer program instructions, and wherein the one or more computer program instructions are executed by the processor to implement the method described in the first aspect and / or the second aspect of the embodiments of the present invention.
[0038] In a fifth aspect, an embodiment of the present invention provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it implements the method described in the first aspect and / or the second aspect of the embodiments of the present invention.
[0039] In the embodiments of the present invention, by obtaining a batch operation request, performing concurrent processing on the batch operation request, obtaining at least one data processing result, adding each data processing result to a to-be-submitted queue, and storing each data processing result in the to-be-submitted queue. Thus, in this embodiment, the processing process of the data operation request is separated from the submission process, that is, the submission process is performed on the data processing result corresponding to the data operation request, so that the data operation request can be processed concurrently, reducing the data processing delay and improving the data processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Through the following description of the embodiments of the present invention with reference to the drawings, the above and other objects, features, and advantages of the present invention will become clearer. In the drawings:
[0041] Figure 1 is a schematic diagram of a data processing framework of a comparative example;
[0042] Figure 2 is a schematic diagram of a data processing framework of an embodiment of the present invention;
[0043] Figure 3 is a flowchart of a data concurrent processing method of an embodiment of the present invention;
[0044] Figure 4 is a flowchart of a processing method corresponding to a data processing result of an embodiment of the present invention;
[0045] Figure 5 is a flowchart of a message feedback method of an embodiment of the present invention;
[0046] Figure 6It is a flowchart of another data concurrency processing method according to an embodiment of the present invention;
[0047] Figure 7 It is a schematic diagram of a data concurrency processing system according to an embodiment of the present invention;
[0048] Figure 8 It is a schematic diagram of the architecture of data concurrency consistency according to an embodiment of the present invention;
[0049] Figure 9 It is a schematic diagram of a data concurrency processing device according to an embodiment of the present invention;
[0050] Figure 10 It is a schematic diagram of another data concurrency processing device according to an embodiment of the present invention;
[0051] Figure 11 It is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed implementation manners
[0052] The following is a description of the present application based on embodiments, but the present application is not limited to these embodiments. In the following detailed description of the present application, some specific details are described in detail. Those skilled in the art can fully understand the present application without the description of these details. In order to avoid obscuring the essence of the present application, well-known methods, processes, procedures, components, and circuits are not described in detail.
[0053] In addition, those of ordinary skill in the art should understand that the drawings provided herein are all for illustrative purposes and the drawings are not necessarily drawn to scale.
[0054] Unless the context clearly requires otherwise, the words such as "including", "comprising" and the like in the entire application document should be interpreted as the meaning of including rather than exclusive or exhaustive meaning; that is, the meaning of "including but not limited to".
[0055] In the description of the present application, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present application, unless otherwise specified, the meaning of "plurality" is two or more.
[0056] For the solutions described in this specification and embodiments, if they involve personal information processing, they will be processed on the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for performing a contract, etc.), and will only be processed within the specified or agreed scope. If the user refuses to process personal information other than the necessary information required for the basic functions, it will not affect the user's use of the basic functions.
[0057] Figure 1It is a schematic diagram of a data processing framework for a comparative example. As Figure 1 shown, the client 11 sends a data operation request to the server 12. Among them, the server 12 includes a Leader node 121, follower nodes 122 and 123. Among them, the Leader node 121 is determined through election. The Leader node 121 is responsible for managing the entire cluster, synchronizing data to the follower nodes 122 and 123, and ensuring the global consistency of the data. The follower nodes can provide a data query function, and they can send the received data operation requests to the Leader node 121. Further, the server 12 can achieve linearizable reads through the Leader node 121 and the follower nodes 122 and 123, thereby ensuring data consistency in the distributed system. For example, when the follower node 122 needs to read data, it will request the latest ReadIndex from the Leader node 121 to ensure the linearizability of the data read operation.
[0058] The transaction processing process corresponding to the data operation request generally includes a propose operation, an apply operation, and a commit operation. Among them, the propose operation is the proposal stage of the transaction, mainly responsible for determining the operation scope and target of the transaction. If the proposal is successful, the transaction will enter the apply stage operation. The apply operation is the process of applying the transaction, responsible for applying all modifications of the transaction to the database. These modifications can include operations such as data reading, updating, and deletion. The commit operation is the process of committing the transaction, responsible for permanently recording all modifications of the transaction in the database. In the commit operation stage, the database will confirm all the modifications applied in the previous apply operation stage and persistently save them to the disk.
[0059] In this comparative example, in the apply operation stage, the modification operations of the transaction will be gradually applied as the transaction is executed. That is to say, the modification operations of the transaction are executed and applied in the apply operation stage. Specifically, as Figure 1 shown, after the server 12 receives the data operation request, it adds the data operation request to the proposal queue 13 through the propose operation, and then performs the apply operation on the data operation requests in the proposal queue 13. When applying, the OP (operation) operations (that is, the data operations corresponding to the data operation requests x1, x2, etc., such as write operations, etc.) are executed serially, and the obtained KV (Key-Value, Pair) results of the operations are applied to the KV database. Among them, the KV database is a database that stores key-value pair data, where each key uniquely corresponds to a value.
[0060] It can be seen that since apply is serial, the execution of OP operations is also serial. Moreover, the sizes of each OP operation are different, including reads and writes. If one OP operation is slow, the overall apply will be slow, thereby increasing the overall data processing latency. Therefore, the embodiments of the present invention provide a data concurrent processing method, system, and electronic device. By separating the processing process of data operation requests from the submission process, that is, enabling the submission process to execute for the data processing results corresponding to the data operation requests, data operation requests can be processed concurrently, reducing data processing latency and improving data processing efficiency.
[0061] Figure 2 is a schematic diagram of the data processing framework according to the embodiments of the present invention. As Figure 2 shown, the client 21 sends a data operation request to the server 22. Among them, the server 22 includes a Leader node 221, follower nodes 222, and follower node 223. Among them, the Leader node 221 is determined through election. The Leader node 221 is responsible for managing the entire cluster, synchronizing data to the follower nodes 222 and 223, and ensuring the global consistency of the data. The follower nodes can provide a data query function and can send the received data operation requests to the Leader node 221.
[0062] Among them, after the server 22 receives a data operation request, it concurrently executes OP operations (that is, concurrently executes data operations corresponding to data operation requests y1, y2, etc., such as write operations), and stores the KV results obtained from the OP operations in the memory 23. Then, for the KV results in the memory 23, the KV results are added to the proposal queue 24 through a propose operation. For example, the results KV1, KV2, etc. in the memory are added to the proposal queue 24 through a propose operation. Further, an apply operation is performed on the KV results in the proposal queue 24 to apply the KV results to the corresponding KV database 25. It can be seen that, compared with Figure 1 the comparative example in, the embodiments of the present invention separate the processing process of data operation requests from the submission process, that is, enable the submission process to execute for the data processing results corresponding to the data operation requests, enabling data operation requests to be processed concurrently, thereby reducing data processing latency and improving data processing efficiency.
[0063] Figure 3 is a flowchart of a data concurrent processing method according to the embodiments of the present invention. Optionally, the data concurrent processing method of this embodiment can be applied to a server such as a multi-protocol storage system or a meta-service storage system. Further, as Figure 3As shown, the data concurrency processing method corresponding to the data processing framework of the embodiment of the present invention includes the following steps:
[0064] Step S110, obtain a batch operation request. Among them, the batch operation request includes at least one data operation request. Optionally, the data operation request may include file renaming, file creation, file reading and writing, etc. This embodiment does not limit the specific operation type, and it may be any operation in the storage system.
[0065] In an optional implementation manner, a batch operation request combination package in this embodiment may include data operation requests of the same client (that is, different batch operation request combination packages corresponding to different clients), or may include data operation requests of multiple clients (that is, different clients may correspond to the same batch operation request combination package or different batch operation request combination packages). This embodiment does not limit this.
[0066] In an optional implementation manner, this embodiment may obtain a batch operation request through a predetermined communication network. Among them, the service end corresponding to this embodiment is communicatively connected to the client through a predetermined communication network to receive the batch operation request.
[0067] In an optional implementation manner, the predetermined communication network may be a TCP (Transmission Control Protocol) network or a GRPC network. It should be understood that this embodiment is not limited thereto, and other communication networks may be selected based on actual application scenarios to establish a communication connection between the service end and the client.
[0068] GRPC is a high-performance, open-source and general RPC (Remote Procedure Call Protocol) framework, mainly designed for mobile applications and HTTP / 2. It is used for calls between services, and the streaming RPC can easily build a distributed application system and supports multiple languages. Moreover, GRPC uses the Protobuf serialization protocol, which can effectively reduce the amount of transmitted data, reduce the transmission delay, and increase the information density. The Protobuf serialization protocol is a process of converting a data structure or object into a binary string (or vice versa), and is used for functions such as data storage and RPC data exchange.
[0069] Furthermore, if the concurrency intensity of data operations is relatively high, the GRPC network may also have a certain latency. According to the concurrency intensity of the server, embodiments of the present invention can also use a TCP network library to implement network communication to improve system performance. TCP network is a reliable, connection-oriented network protocol that can ensure the order and integrity of data. The TCP network library can provide functions such as creating network connections, sending and receiving data, handling connection closures, flow control, congestion control, error detection, and recovery.
[0070] In an alternative implementation, if embodiments of the present invention use the Golang programming language environment to implement concurrent data processing, the TCP network library can be implemented through the native netpoll package in Golang that encapsulates the epoll package, or the native epoll package. Among them, Golang is a statically typed, compiled, concurrent programming language with garbage collection capabilities.
[0071] Epoll is an IO multiplexing technology in the Linux kernel for processing a large number of file descriptors. It passes results by multiplexing the file descriptor set, rather than having to re-prepare the corresponding file descriptor set before waiting for events each time, which can significantly improve the system CPU utilization rate when there are only a small number of active connections in a large number of concurrent connections.
[0072] Furthermore, in the process of implementing high-performance network communication based on epoll and TCP, when using epoll, the management of connections and the transmission of data can be achieved by detecting events on the TCP connection. At the same time, the TCP connection can also improve efficiency and performance through epoll.
[0073] If TCP network communication is implemented through the netpoll package, two threads (Goroutines) need to be started for each TCP connection to achieve reading and writing. If there are too many TCP connections, more threads will be started, which may cause a certain scheduling pressure. However, the meta-service storage system generally uses long connections, which reduces a certain scheduling pressure.
[0074] If native epoll is used to implement TCP network communication, two threads do not need to be started for each TCP connection. Instead, the thread corresponding to the multiplexed epoll_wait (a function used to wait for file descriptors to be ready) can be used, which reduces the number of started threads. However, native epoll requires a separate thread to perform epoll_wait. Since thread scheduling has no priority, there may be a problem that the thread running epoll_wait cannot be scheduled.
[0075] The embodiments described below mainly take the netpoll package as an example. However, it should be understood that the embodiments of the present invention can select the netpoll package or the native epoll to implement TCP network communication according to the actual application scenario, the advantages and disadvantages of the netpoll package and the native epoll. This embodiment does not limit this.
[0076] Step S120: Perform concurrent processing on the batch operation requests to obtain at least one data processing result. In this embodiment, the data processing result represents the modification result of the operation data corresponding to the data operation request.
[0077] In the embodiments of the present invention, after receiving a combined packet of one or more batch operation requests, concurrent processing is performed on each data operation request to obtain the corresponding data processing result.
[0078] In an optional implementation manner, the batch operation requests are concurrently processed through a predetermined concurrency model. Optionally, the concurrency model can be a Pipeline model or an RTC (run_to_completion) model.
[0079] The concurrent model Pipeline is a model for efficiently processing concurrent tasks in a pipeline manner. Among them, the concurrent model Pipeline adopts a multi-thread + queue mode. When the server receives a data operation request, it sends it to the corresponding request queue. Then, multiple threads fetch tasks from the request queue for processing, and the server returns the feedback result corresponding to the operation request. Among them, receiving the operation request - processing the request - returning the feedback result are all asynchronous. The Pipeline model can control the number of threads through the multi-thread + queue method, thereby ensuring the overall throughput of the system and reducing latency.
[0080] The RTC model is a mode of program or task execution. In the RTC model, each task is regarded as an independent execution unit and is executed sequentially in a certain order, ensuring the execution order and integrity of the program or task. That is to say, in the RTC model, for a certain data operation request, receiving the operation request - processing the request - returning the feedback result are all executed sequentially in the same execution unit, which depends on the asynchronous input and output of the disk, without context switching, Cache Miss (that is, the required data cannot be obtained from the cache and needs to be obtained from the next-level cache or disk), and lock overhead.
[0081] In this embodiment, the embodiments described below mainly take the Pipeline model as an example. However, it should be understood that the embodiments of the present invention can select the Pipeline model or the RTC model to implement the concurrent processing of requests according to the actual application scenario, the advantages and disadvantages of the Pipeline model and the RTC model. This embodiment does not limit this.
[0082] Step S130: Add each data processing result to the queue to be submitted.
[0083] In an alternative implementation, in this embodiment, the data processing results can be directly added to the queue to be submitted (i.e., the proposal queue), that is, the data processing results are added to the proposal queue through a propose operation.
[0084] Figure 4 It is a flowchart of the processing method corresponding to the data processing results of the embodiments of the present invention. In another alternative implementation, as Figure 4 shown, the specific process of adding the data processing results to the queue to be submitted in this embodiment includes the following steps:
[0085] Step S131: Store each data processing result in the corresponding memory. That is to say, after the concurrent processing of each data operation request, the obtained corresponding data processing results (such as KV results) are first stored in the corresponding memory to cache each data processing result, so as to improve the data query efficiency of subsequent transaction processing.
[0086] Step S132: Generate a request to be submitted corresponding to each data processing result. Among them, the request to be submitted can be the data processing result itself, or can include the corresponding data processing result, and this embodiment does not limit this.
[0087] Step S133: Add each request to be submitted to the queue to be submitted.
[0088] In this embodiment, each request to be submitted is added to the queue to be submitted through a propose operation. That is to say, in this embodiment, through a propose operation, after the proposal corresponding to the request to be submitted is passed, the request to be submitted is added to the queue to be submitted.
[0089] Step S140: Store each data processing result in the queue to be submitted.
[0090] Further, in this embodiment, an apply operation is performed on the data processing results (such as KV results) in the queue to be submitted to apply each data processing result, that is, an apply disk write operation is performed. For example, the KV results are applied to the corresponding KV database. Further, after the data processing results are applied to the KV database through the apply operation, a commit operation is performed to persist the data processing results (i.e., the modifications corresponding to each OP operation).
[0091] In an alternative implementation, after performing an apply operation or a commit operation on the data processing results, the embodiments of the present invention return corresponding feedback messages to the corresponding client.
[0092] Figure 5 is a flowchart of the message feedback method according to an embodiment of the present invention. As Figure 5 shown, the message feedback method according to an embodiment of the present invention includes the following steps:
[0093] Step S210, determine the feedback message corresponding to each data operation request. For example, after the rename operation request is processed and written to disk, a feedback message indicating successful rename is generated.
[0094] Step S220, add each feedback message to the reply queue. In this embodiment, a reply queue is maintained to achieve batch reply of feedback messages.
[0095] In an alternative implementation, the server in this embodiment may maintain one reply queue or multiple reply queues. Further, the server in this embodiment may maintain a reply queue for each communication connection.
[0096] In an alternative implementation, the data operation request has a corresponding request identifier, and the batch operation request has a corresponding communication connection identifier. Among them, the communication connection identifier is used to represent the communication connection that receives the batch operation request. Thus, the embodiment of the present invention can implement the mapping between the data operation request and the feedback message through the request identifier and the communication connection identifier.
[0097] Further optionally, the feedback message has the same or corresponding request identifier and communication connection identifier as the data operation request.
[0098] Further, if this embodiment maintains a reply queue for each communication connection, the corresponding communication connection identifier of the feedback message can be judged and the feedback message can be added to the corresponding reply queue. If the reply queues maintained in this embodiment do not correspond one-to-one to the communication connections, the feedback message can be added to the working reply queue or randomly added to a reply queue, and this embodiment does not limit this.
[0099] Step S230, take out at least one feedback message from the reply queue through the corresponding first thread to obtain a message combination packet.
[0100] In this embodiment, the first thread is started to retrieve feedback messages from the reply queue to obtain a message combination packet. In an alternative implementation, if there is a one-to-one correspondence between the reply queue and the communication connection, multiple feedback messages are retrieved in batches from the reply queue each time until the maximum data volume (or the maximum number of messages) of the message combination packet is reached or the corresponding reply queue is empty, and the corresponding message combination packet is obtained. In another alternative implementation, if there is no one-to-one correspondence between the reply queue and the communication connection, the feedback messages are retrieved from the reply queue, and the feedback messages with the same communication connection identifier are added to the same message combination packet until the maximum data volume of the message combination packet is reached or there is no feedback message corresponding to the communication connection identifier in the reply queue, and the final message combination packet is obtained.
[0101] Step S240: Feed back the message combination packet to the corresponding client through the corresponding communication connection. Since the server can be connected to multiple clients through different communication connections, it is necessary to feed back the message combination packet to the correct client through the corresponding communication connection (i.e., the client that sends the corresponding data operation request). Thus, in this embodiment, a one-to-one mapping between the message combination packet and the client is achieved through the communication connection identifier.
[0102] Furthermore, since the feedback message carries a request identifier that is the same as or corresponding to the data operation request, the client can wake up the corresponding request through the request identifier in the feedback message after receiving the message combination packet, so as to achieve a one-to-one mapping between the data operation request and the feedback message.
[0103] In an alternative implementation, in order to avoid problems such as packet disorder (e.g., TCP packet disorder) during the transmission of the data packet over the communication connection, this embodiment also adds a data checksum field to the header of the data packet to overall verify the correctness of the data at the application layer.
[0104] In the embodiment of the present invention, by obtaining batch operation requests, performing concurrent processing on the batch operation requests, obtaining at least one data processing result, adding each data processing result to the queue to be submitted, and storing each data processing result in the queue to be submitted. Thus, in this embodiment, the processing process of the data operation request is separated from the submission process, that is, the submission process is performed on the data processing result corresponding to the data operation request, so that the data operation request can be processed concurrently, reducing the data processing delay and improving the data processing efficiency.
[0105] Figure 6 It is a flowchart of another data concurrent processing method according to an embodiment of the present invention. The data concurrent processing method of this embodiment is applied to the client, such as Figure 6As shown in the figure, the data concurrent processing method of this embodiment includes the following steps:
[0106] Step S310, add the received data operation request to the sending queue. In an alternative implementation, take out the data operation request from the data operation request pool and add it to the sending queue.
[0107] Step S320, obtain a batch operation request by taking out at least one data operation request from the sending queue through the corresponding second thread.
[0108] In an alternative implementation, this embodiment obtains a batch operation request combination package in response to the data volume corresponding to the taken-out data operation request reaching the threshold or the sending queue being empty. Further, this embodiment starts the second thread to take out one or a batch of requests from the corresponding sending queue, and obtains a batch operation request combination package when the data volume or the number of requests corresponding to the taken-out data operation request reaches the threshold or the sending queue is empty.
[0109] Step S330, send the batch operation request through the corresponding communication connection. In this embodiment, as described above, if the client and the server communicate through TCP, then this communication connection is the TCP connection between this client and the server. Optionally, the client and the server can also communicate through a GRPC connection, which will not be described in detail here.
[0110] In an alternative implementation, the client can also receive a message combination package feedback by the server through this communication connection. Among them, the message combination package includes feedback messages corresponding to each data operation request. Further, the mapping relationship between the data operation request and the feedback message in this embodiment is determined by the corresponding request identifier. Specifically, the client can wake up the corresponding data operation request through the request identifier in the header of the feedback message package in the combined message package to determine the processing status of the corresponding data operation request through the feedback message. Among them, the processing status can indicate successful processing (such as successful renaming, etc.) or failed processing (such as file creation failure, etc.).
[0111] The embodiment of the present invention adds the received data operation request to the sending queue, takes out at least one data operation request from the sending queue through the corresponding second thread to obtain a batch operation request, and sends the batch operation request through the corresponding communication connection. Thus, this embodiment can batch and concurrently send data operation requests, so that the server can batch receive data operation requests to concurrently process corresponding operations, improve data processing messages, and reduce latency.
[0112] Figure 7 is a schematic diagram of the data concurrent processing system of the embodiment of the present invention. As Figure 7As shown in the figure, the data concurrency processing system of this embodiment includes at least one client 7 and a server 7'. Among them, the server 7' can be a data storage system or a meta-service system, and this embodiment does not limit this. Further, each client 7 of this embodiment is communicatively connected to the server 7' through TCP. Among them, each client 7 interacts with the server 7' through different TCP connections. In this embodiment, one of the clients 7 is taken as an example to describe the data concurrency processing process.
[0113] In the client 7, first, data operation requests z1, etc. generated are put into the corresponding request pool 71, and one or more data operation requests z1 are taken out from the request pool 71 and put into the sending queue 72.
[0114] In this embodiment, when creating the TCP connection 70 between the client 7 and the server 7', a write thread is started. The write thread of the TCP connection of the client 7 continuously takes batches of data operation request packets from the sending queue 72. When the data volume of the packets taken out from the sending queue 72 reaches the threshold or the sending queue 72 is empty, all the data operation requests taken out in this batch are combined into a combined packet 73 of batch operation requests, and are sent to the server 7' through the corresponding TCP connection 70.
[0115] In this embodiment, the server 7' receives the combined packet 73 of batch operation requests, and starts multiple operation threads to asynchronously and concurrently execute the OP operations corresponding to the respective data operation requests in the batch operation requests, generate corresponding data processing results, and corresponding feedback messages.
[0116] In this embodiment, an apply operation is performed on the data processing results corresponding to each data operation request, that is, each data processing result is added to the proposal queue (i.e., the queue to be submitted) through the propose operation, and an apply operation is performed on each data processing result in the queue to be submitted to apply the corresponding data processing result. For example, the KV result of data processing is applied to the corresponding KV database. The specific processing process is similar to that of Figures 2 - 4 the embodiment described above, and will not be described in detail here.
[0117] Further, the server 7' adds the feedback messages corresponding to each data operation request to the reply queue 74. The write thread of the TCP connection of the server 7' continuously takes out at least one feedback message from the reply queue 74. When the data volume of the packets taken out from the reply queue 74 reaches the threshold or the reply queue 74 is empty, all the feedback messages taken out in this batch are combined into a message combined packet 75, and are sent to the client 7 through the corresponding TCP connection 70. The client 7 receives the message combined packet 75 and asynchronously notifies the request pool 71.
[0118] Since the client 7 and the server 7' send packets in batches, it is necessary to record the mapping relationship between req-resp (i.e., data operation requests and feedback messages) so that the feedback message corresponding to the data operation request can be found according to this mapping relationship. In this embodiment, the mapping between the data operation request and the feedback message is implemented through the message identifier and the communication connection identifier.
[0119] In an alternative implementation, in the client 7, the mapping relationship can be recorded either by using a pointer or by using an incrementing uint64 id. It should be understood that other methods can also be used in this embodiment, such as an incrementing uint32 id, etc. This embodiment does not limit this.
[0120] Further, if the pointer method is used, the value of the pointer can be placed in the header of the data operation request packet, and when generating the feedback message for the data processing request, the value of the pointer is also placed in the header of the feedback message packet. When the client receives the feedback message, the value of the pointer is taken out from the header of the feedback message packet to wake up the corresponding data operation request. Similarly, an incrementing uint64 id can be placed in the header of the data operation request packet so that the data operation request carries the corresponding request identifier to implement the mapping of req-resp.
[0121] Since the server 7' has multiple TCP connections, it is necessary to determine from which TCP connection each data operation request is received to determine that the feedback message corresponding to the data operation request is sent to the correct client 7 through the same TCP connection. In this embodiment, the communication connection identifier is represented by a pointer or an incrementing uint64 id to record from which TCP connection the request is sent, and the reply is made through the same TCP connection. Thus, in this embodiment, the mapping between the data operation request and the feedback message can be implemented through the communication connection identifier and the request identifier, ensuring the accuracy of the message feedback.
[0122] In the embodiment of the present invention, by obtaining batch operation requests, performing concurrent processing on the batch operation requests, obtaining at least one data processing result, adding each data processing result to the queue to be submitted, and storing each data processing result in the queue to be submitted. Thus, in this embodiment, the processing process of the data operation request is separated from the submission process, that is, the data processing result corresponding to the data operation request performs the submission process, so that the data operation request can be processed concurrently, reducing the data processing delay and improving the data processing efficiency.
[0123] In this embodiment, since the OP operation is brought forward before the Raft proposal, this embodiment needs to ensure the correctness of concurrent OP operations in memory; otherwise, problems such as damaged metadata may occur.
[0124] In an alternative implementation, the data concurrent processing method of this embodiment further includes: performing concurrent consistency verification on each data operation request to avoid data operation conflicts. That is to say, this embodiment detects conflicts for each data operation request and applies and stores the data processing results that meet concurrent consistency in the to-be-submitted queue to disk.
[0125] Further optionally, this embodiment performs data overlap detection on each data processing result, and in response to the data overlap detection result indicating overlapping data, determines that there is a transaction conflict.
[0126] Further optionally, this embodiment retries the operation for the data operation request corresponding to the transaction conflict. That is to say, if there is a transaction conflict in the currently to-be-submitted transaction, the currently to-be-submitted transaction is retried. If there is still a conflict after the preset time or the preset number of retries, a message indicating operation failure is fed back to the client. Further, if there is no transaction conflict in the currently to-be-submitted transaction, the currently to-be-submitted transaction (data processing result) is put into the to-be-submitted queue to execute the apply operation, and the data processing result is applied to the corresponding database.
[0127] In an alternative implementation, this embodiment can adopt optimistic transaction conflict detection, such as using the TXN model, or can adopt the method of adding read-write locks to keys to prevent transaction conflicts, such as the method of adding read-write locks to keys similar to Ceph MDS. This embodiment does not limit the way to achieve concurrent consistency in memory.
[0128] Further optionally, this embodiment mainly describes the example of using the in-memory TXN model to ensure the correctness of concurrent OP operations. The TXN model is a logical unit of a database transaction, consisting of a finite sequence of database operations. It represents a transaction or a set of operations and can ensure that these operations are executed atomically in the database. The TXN model can help solve the concurrent control problem and ensure that operations between multiple transactions do not interfere with each other, thus maintaining data consistency and integrity.
[0129] Figure 8 It is a schematic diagram of the architecture of data concurrent consistency in an embodiment of the present invention. As Figure 8 shown, this embodiment takes the process of concurrent consistency processing of transaction TXN as an example for illustration. Among them, transaction TXN represents the commit transaction of the data processing result corresponding to a data operation request.
[0130] In this embodiment, it is necessary to perform conflict detection on the transaction TXN and serialize the propose operation to ensure data concurrency consistency.
[0131] Further optionally, as Figure 8 shown, this embodiment sends one or more data processing results obtained by completing the OP operation to the in-memory KV Buffer 81, and performs conflict detection based on the target operation corresponding to the transaction. The target operation may include Set operation, Delete operation, and / or Add operation, etc. Among them, if the time difference between multiple data processing results is within a predetermined time, conflict detection needs to be performed. Optionally, the time difference may be the time difference of the confirmed commit time, the time difference of being sent to the corresponding memory, or may represent the generation time difference of the data processing results. This embodiment does not limit this. It should be understood that this embodiment does not limit the length of the predetermined time, which can be set according to the data concurrency in the actual application scenario. For example, the more data concurrency, the shorter the predetermined time is set, etc. This embodiment does not limit this. Among them, the Set operation is used to store a value at a specific location in the memory. For example, if the OP operation corresponding to a certain transaction is to create a new data item, the Set operation can store the data item (i.e., the data processing result) created by the OP operation at a specific location in the memory. The Delete operation is used to delete a certain node or data item, and the Add operation is used to add a new node or data item.
[0132] In an optional implementation manner, conflict detection can be performed through the read_TS and commit_TS of the transaction. Among them, read_TS is the timestamp when the transaction starts to execute, indicating the time point when the transaction starts to read data. If the transaction modifies the data after read_TS, this modification is regarded as invalid, that is, there is a conflict. commit_TS is the timestamp when the transaction commits, indicating the time point when the transaction has completed all operations and is ready to commit. If the transaction reads data after commit_TS, this read operation will not be able to obtain the committed modification, that is, there is a conflict. This embodiment can implement conflict detection of each transaction in the memory through the read_TS and commit_TS of the transaction. For example, when a certain transaction performs a data operation on a certain data item, if another transaction also wants to perform a data operation on this data item, it is determined that there is a transaction conflict until a previous transaction commits and updates the commit_TS.
[0133] In this embodiment, in the in-memory KV Buffer 81, conflict detection is performed according to the target operations corresponding to each transaction (such as Set operation, Delete operation, and / or Add operation, etc.), that is, overlapping detection is performed on the data to be changed by each transaction. After the conflict detection of each transaction is successful (that is, there is no transaction conflict situation), the propose operation is executed.
[0134] Furthermore, if there is a conflict between two or more transactions for which conflict detection is performed, the data operation requests corresponding to each conflicting transaction can be retried, or the target request can be determined according to the operation type and / or timestamp of the data operation request, and the target request is retried, and the propose operation is executed for other requests except the target request. For example, the transaction with the earliest start time is determined, the propose operation is executed for the transaction with the earliest start time, and the data operation requests corresponding to other transactions conflicting with it are retried. Furthermore, the specific method of determining the target request that needs to be rolled back according to the operation type can be set based on the characteristics and actual application scenarios of different data operation requests, and this embodiment does not limit this. For example, if the operation types corresponding to the data operation requests with transaction conflicts are read requests and write requests, in order to ensure that the data read is the latest data, the transaction corresponding to the write request can be committed, and the read request is used as the target request to perform a rollback operation.
[0135] For example, for transactions TXN1 and TXN2 that both complete the OP operation and perform corresponding target operations on a certain memory location (that is, there is an overlapping Key), it is determined that there is a transaction conflict between transactions TXN1 and TXN2.
[0136] Further optionally, after the conflict detection of the TXN model is successful, the commit operation and the propose operation are performed asynchronously for each transaction. Since the commit operation and the propose operation are not atomic operations, the order of the commit operation and the propose operation in the same transaction TXN may be reversed, resulting in some unpredictable risks. Therefore, in this embodiment, a locking operation is performed on the conflict detection, Raft propose, and TXN commit in the TXN model to ensure serialization, that is, the serial execution of TXN conflict detection --> Raft propose --> TXN commit is achieved through the locking operation.
[0137] Furthermore, as Figure 8As shown, after successful conflict detection, the data processing results that meet concurrent consistency are added to the to-be-committed queue 82 through the propose operation, and the data processing results in the commit queue 82 are subjected to the apply operation to apply the data processing results to the KV database 83. Specifically, in this embodiment, when executing transaction TXN1, after conflict detection and successful proposal of transaction TXN1, the data processing results of transaction TXN1 in the KV Buffer 81 are added to the to-be-committed queue 82.
[0138] In an alternative implementation, this embodiment can achieve TXN visibility, that is, when multiple transactions operate on the database simultaneously, it ensures that each transaction can correctly see the modification results of other transactions on the database to ensure data consistency.
[0139] Optionally, since the data for the apply operation is the data processing result of the OP operation, in this embodiment, the apply operation and the commit operation can be executed asynchronously. However, since the data modified by a transaction can only be read from the corresponding KV database 83 after the apply is written to disk, when another transaction needs to read the committed data, the previous transaction needs to perform the apply operation first and then the commit operation. That is, in this case, the transaction still needs to perform the apply operation first and then the commit operation. That is, the processing logic of the transaction at this time is: TXN conflict detection --> TXN Raft propose --> TXN Raft apply --> TXN commit. That is, for a transaction, conflict detection is performed first, and then the propose operation, the apply operation, and the commit operation are performed in sequence.
[0140] Furthermore, since the apply operation takes a relatively long time, in order to further improve the throughput, in this embodiment, the data of the apply operation of each transaction can be cached through the in-memory KV Buffer 81. Thus, when executing subsequent transactions, if a subsequent transaction needs to read the data modified by a certain transaction, the Get operation (used to obtain the value at a certain position from the memory) can be preferentially executed to query from the in-memory KV Buffer 81, improving the data query efficiency of subsequent transactions. At this time, the processing logic of the transaction is: TXN conflict detection --> TXN Raft propose --> save to Mem --> TXN commit. That is, for a transaction, first perform conflict detection, and then sequentially perform the propose operation, storage operation (i.e., save to Mem, caching the data of the apply operation to the in-memory KV Buffer 81), and commit operation. Thus, when subsequent transactions read the data of previous transactions, they can preferentially query the relevant data cached in the in-memory KV Buffer 81, improving the data reading efficiency and overall throughput.
[0141] Furthermore, during the execution of a transaction, it may obtain the committed result. At this time, the data of the committed transaction may be in the corresponding in-memory KV Buffer 81 or in the underlying KV database 83. To further improve the reading efficiency, in this embodiment, a corresponding cache Cache 84 can also be set in the memory to store the data corresponding to the transaction in the cache Cache 84 to accelerate the reading efficiency of subsequent transactions. At this time, the transaction processing logic is: TXN conflict detection --> TXN Raft propose --> save to Mem --> save to Cache --> TXN commit. That is, for a transaction, first perform conflict detection, and then sequentially perform the propose operation, storage operation (i.e., save to Mem, caching the data of the apply operation to the in-memory KV Buffer 81), cache operation (i.e., save to Cache, caching the transaction-related data to the cache Cache 84), and commit operation. Thus, when subsequent transactions read the data of previous transactions, they can preferentially query the relevant data cached in the memory or cache, improving the data reading efficiency and overall throughput.
[0142] In this embodiment, the processing process of the data operation request is separated from the submission process, that is, the submission process is executed for the data processing result corresponding to the data operation request, so that the data operation requests can be processed concurrently, reducing the data processing latency and improving the data processing efficiency. And when submitting the data processing result, conflict detection is performed by determining whether there is overlapping data with the data changed by the transaction, ensuring data concurrent consistency.
[0143] Figure 9 It is a schematic diagram of a data concurrent processing device according to an embodiment of the present invention. As Figure 9 shown, the data concurrent processing device 9 according to the embodiment of the present invention includes a request acquisition unit 91, a concurrent processing unit 92, a proposal unit 93, and a submission unit 94.
[0144] The request acquisition unit 91 is configured to acquire a batch operation request, and the batch operation request includes at least one data operation request. The concurrent processing unit 92 is configured to perform concurrent processing on the batch operation request to obtain at least one data processing result. The proposal unit 93 is configured to add each of the data processing results to a to-be-submitted queue. The submission unit 94 is configured to store each of the data processing results in the to-be-submitted queue.
[0145] In an alternative implementation, the proposal unit 93 is further configured to execute:
[0146] Store each of the data processing results in the corresponding memory;
[0147] Generate a to-be-submitted request corresponding to each of the data processing results;
[0148] Add each of the to-be-submitted requests to the to-be-submitted queue.
[0149] In an alternative implementation, the concurrent processing unit 92 is further configured to execute: perform concurrent processing on the batch operation request through a predetermined concurrent model.
[0150] In an alternative implementation, the concurrent model is a Pipeline model or an RTC model.
[0151] In an alternative implementation, the request acquisition unit 91 is further configured to execute: acquire the batch operation request through a predetermined communication network.
[0152] In an alternative implementation, the predetermined communication network is a TCP network or a GRPC network.
[0153] In an alternative implementation, the data packet corresponding to the data operation request includes a data verification field.
[0154] In an alternative implementation, the data operation request has a corresponding request identifier;
[0155] The batch operation request has a corresponding communication connection identifier, and the communication connection identifier is used to characterize the communication connection that receives the batch operation request.
[0156] In an alternative implementation, the data concurrency processing device 9 further includes a message determination unit, a first queue management unit, a combined packet acquisition unit, and a feedback unit.
[0157] The message determination unit is configured to determine the feedback messages corresponding to the respective data operation requests. The first queue management unit is configured to add the respective feedback messages to a reply queue. The combined packet acquisition unit is configured to take out at least one feedback message from the reply queue through a corresponding first thread to obtain a message combined packet. The feedback unit is configured to feedback the message combined packet to the corresponding client through the corresponding communication connection.
[0158] In an embodiment of the present invention, by obtaining a batch operation request, performing concurrent processing on the batch operation request to obtain at least one data processing result, adding each data processing result to a to-be-submitted queue, and storing each data processing result in the to-be-submitted queue. Thus, in this embodiment, the processing process of the data operation request is separated from the submission process, that is, the data processing result corresponding to the data operation request performs the submission process, so that the data operation request can be concurrently processed, reducing the data processing delay and improving the data processing efficiency.
[0159] Figure 10 is a schematic diagram of another data concurrency processing device according to an embodiment of the present invention. As Figure 10 shown, the data concurrency processing device 10 according to an embodiment of the present invention includes a second queue management unit 101, a request determination unit 102, and a request sending unit 103.
[0160] The second queue management unit 101 is configured to add the received data operation requests to a sending queue. The request determination unit 102 is configured to take out at least one data operation request from the sending queue through a corresponding second thread to obtain a batch operation request. The request sending unit 103 is configured to send the batch operation request through the corresponding communication connection.
[0161] In an alternative implementation, the request determination unit 102 is further configured to execute:
[0162] In response to the data volume corresponding to the taken-out data operation request reaching a threshold, or the sending queue being empty, obtaining the batch operation request.
[0163] In an alternative implementation, the mapping relationship between the data operation request and the corresponding feedback message is determined by the corresponding request identifier.
[0164] In the embodiment of the present invention, the received data operation requests are added to the sending queue, at least one data operation request is taken out from the sending queue by a corresponding second thread to obtain a batch operation request, and the batch operation request is sent through a corresponding communication connection. Thus, this embodiment can send data operation requests in batches and concurrently, so that the server can receive data operation requests in batches to process corresponding operations concurrently, improve data processing messages, and reduce latency.
[0165] Figure 11 is a schematic diagram of the electronic device according to the embodiment of the present invention. As Figure 11 shown, the electronic device 110 is a general data processing device, which includes a general computer hardware structure, and at least includes a processor 111 and a memory 112. The processor 111 and the memory 112 are connected through a bus 113. The memory 112 is adapted to store instructions or programs executable by the processor 111. The processor 111 may be an independent microprocessor or a set of one or more microprocessors. Thus, the processor 111 executes the instructions stored in the memory 112 to implement the method flow of the embodiment of the present invention as described above to process data and control other devices. The bus 113 connects the above-mentioned multiple components together and also connects the above-mentioned components to a display controller 114, a display device, and an input / output (I / O) device 115. The input / output (I / O) device 115 may be a mouse, a keyboard, a modem, a network interface, a touch input device, a body sensing input device, a printer, and other devices well known in the art. Typically, the input / output (I / O) device 115 is connected to the system through an input / output (I / O) controller 116.
[0166] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a device (equipment), or a computer program product. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may be implemented as a computer program product on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0167] The present application is described with reference to the flowcharts of the methods, devices (equipment), and computer program products according to the embodiments of the present application. It should be understood that each process in the flowchart can be implemented by computer program instructions.
[0168] These computer program instructions can be stored in a computer-readable memory that directs a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the function specified in one process Figure 1 or functions in a process or processes.
[0169] These computer program instructions may also be provided to a processor of a general purpose computer, special purpose computer, embedded processor or other programmable data processing apparatus to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing apparatus produce means for implementing the function specified in one process Figure 1 or functions in a process or processes.
[0170] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program for a computer to execute some or all of the above method embodiments.
[0171] That is, those skilled in the art can understand that all or part of the steps in implementing the above method embodiments can be completed by specifying relevant hardware through a program, which is stored in a storage medium and includes several instructions to enable a device (which may be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0172] The foregoing are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A data concurrent processing method, characterized in that, the method includes: Obtain a batch operation request, where the batch operation request includes at least one data operation request; Perform concurrent processing on the batch operation request to obtain at least one data processing result; Add each of the data processing results to a pending submission queue; Store each of the data processing results in the pending submission queue.
2. The method according to claim 1, characterized in that, adding each of the data processing results to the pending submission queue includes: Store each of the data processing results in the corresponding memory; Generate a pending submission request corresponding to each of the data processing results; Add each of the pending submission requests to the pending submission queue.
3. The method according to claim 1, characterized in that, performing concurrent processing on the batch operation request includes: Perform concurrent processing on the batch operation request through a predetermined concurrent model.
4. The method according to claim 3, characterized in that, the concurrent model is a Pipeline model or an RTC model.
5. The method according to claim 1, characterized in that, obtaining the batch operation request includes: Obtain the batch operation request through a predetermined communication network.
6. The method according to claim 5, characterized in that, the predetermined communication network is a TCP network or a GRPC network.
7. The method according to claim 1, characterized in that, the data packet corresponding to the data operation request includes a data verification field.
8. The method according to claim 1, characterized in that, the data operation request has a corresponding request identifier; The batch operation request has a corresponding communication connection identifier, and the communication connection identifier is used to characterize the communication connection that receives the batch operation request.
9. The method according to claim 8, characterized in that, the method further includes: Determine the feedback message corresponding to each of the data operation requests; Add each of the feedback messages to a reply queue; Retrieve at least one feedback message from the reply queue through a corresponding first thread to obtain a message combination packet; Feed back the message combination packet to the corresponding client through the corresponding communication connection.
10. A data concurrent processing method, characterized in that, the method includes: Add the received data operation request to a sending queue; Retrieve at least one data operation request from the sending queue through a corresponding second thread to obtain a batch operation request; Send the batch operation request through the corresponding communication connection.
11. The method according to claim 10, characterized in that, retrieving at least one data operation request from the sending queue through the corresponding second thread to obtain a batch operation request includes: Obtain the batch operation request in response to the data volume corresponding to the retrieved data operation request reaching a threshold, or the sending queue being empty.
12. The method according to claim 10, characterized in that, the mapping relationship between the data operation request and the corresponding feedback message is determined through the corresponding request identifier.
13. A data concurrent processing system, characterized in that, the system includes: A server, configured to execute the method according to any one of claims 1-9; and At least one client, configured to execute the method according to any one of claims 10-12.
14. An electronic device, comprising a memory and a processor, wherein, The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1-12.
15. A computer-readable storage medium, wherein, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1-12 is implemented.