Data Processing Method and Device Applied to Distributed System

By using slave node status identifiers in distributed systems to determine the processing path of data read and write requests, the problem that existing databases are difficult to ensure high availability, high performance and strong consistency at the same time, and more efficient data processing and system performance are achieved.

CN112148798BActive Publication Date: 2025-06-10TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011079604.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-10
Publication Date
2025-06-10
Estimated Expiration
2040-10-10

AI Technical Summary

Technical Problem

The read and write strategies of existing databases are difficult to ensure high availability, high performance and strong consistency at the same time, resulting in the inability of the needs of some businesses in actual scenarios.

Method used

By introducing slave node status identifiers in the distributed system, the terminal device can determine whether there is a target slave node in an available state based on these status identifiers, and if there is, a data read request is sent to the target slave node, otherwise it is sent to the master node.

Benefits of technology

It realizes that the requirements of high availability, strong consistency and high performance are better met in distributed systems, ensuring that data writing is not blocked due to unavailable nodes, and the read operation can obtain the written data in a timely manner and make full use of the performance of available nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112148798B_ABST
    Figure CN112148798B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a data processing method and apparatus applied to a distributed system. The distributed system includes a master node and multiple slave nodes. The data processing method includes: obtaining the status identifiers of each slave node from the master node, where the status identifier is used to indicate whether the slave node is in an available state, and the slave nodes in the available state store copies of the data in the master node; determining whether there is a target slave node in the available state in the distributed system according to the status identifiers of each slave node; if there is a target slave node, sending a data read request to the target slave node and receiving the response data for the data read request returned by the target slave node; if there is no target slave node, sending the data read request to the master node and receiving the response data for the data read request returned by the master node. The technical solution of the embodiments of the present application realizes the reading and processing of data, and can meet the requirements of high availability, strong consistency, and high performance for data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer and communication technologies. Specifically, it relates to a data processing method and apparatus applied to a distributed system. Background Art

[0002] A non-relational database (Not Only Structured Query Language, NoSQL) provides users with a very rich set of parameters and schemas to control the read and write strategies of the database. Users can configure the desired database read and write model through combinations of these parameters.

[0003] However, the existing read and write strategies of databases force users to make trade-offs and compromises among high availability, high performance, and strong consistency. For example, to ensure strong consistency in data reading and writing, the master node is selected for reading, but in this case, the performance advantages of multiple nodes are sacrificed; to take advantage of multiple nodes, the slave nodes are selected for reading, but then either consistency or availability has to be sacrificed. Generally speaking, it is very difficult for the existing read and write strategies of databases to simultaneously ensure all three aspects of high availability, high performance, and strong consistency, which also results in the unmet requirements of some services in actual scenarios. Summary of the Invention

[0004] Embodiments of this application provide a data processing method and apparatus applied to a distributed system, which can at least better meet the requirements of high availability, strong consistency, and high performance encountered by some services in the actual environment to a certain extent.

[0005] Other features and advantages of this application will become apparent through the following detailed description, or will be partially learned through the practice of this application.

[0006] According to one aspect of the embodiments of this application, a data processing method applied to a distributed system is provided. The distributed system includes a master node and multiple slave nodes, and the method includes: obtaining the status identifiers of each slave node from the master node, where the status identifier is used to indicate whether the slave node is in an available state, and among them, the slave nodes in the available state store copies of the data in the master node; determining whether there is a target slave node in the available state in the distributed system according to the status identifiers of each slave node; if there is the target slave node, sending a data read request to the target slave node and receiving the response data for the data read request returned by the target slave node; if there is no target slave node, sending the data read request to the master node and receiving the response data for the data read request returned by the master node.

[0007] According to one aspect of the embodiments of the present application, a data processing method applied to a distributed system is provided. The distributed system includes a master node and multiple slave nodes, and the method includes: sending the status identifiers of the respective slave nodes to a terminal device, where the status identifiers are used to indicate whether the slave nodes are in an available state. Among them, the slave nodes in the available state store copies of the data in the master node; receiving a data read request sent by the terminal device, where the data read request is sent by the terminal device when it determines that there is no target slave node in the available state in the distributed system according to the status identifiers of the respective slave nodes; responding to the data read request and returning the response data for the data read request to the terminal device.

[0008] According to one aspect of the embodiments of the present application, a data processing device applied to a distributed system is provided. The distributed system includes a master node and multiple slave nodes, and the device includes: an acquisition unit configured to acquire the status identifiers of the respective slave nodes from the master node, where the status identifiers are used to indicate whether the slave nodes are in an available state. Among them, the slave nodes in the available state store copies of the data in the master node; a determination unit configured to determine whether there is a target slave node in the available state in the distributed system according to the status identifiers of the respective slave nodes; a first sending unit configured to, if the target slave node exists, send a data read request to the target slave node and receive the response data for the data read request returned by the target slave node; a second sending unit configured to, if the target slave node does not exist, send the data read request to the master node and receive the response data for the data read request returned by the master node.

[0009] In some embodiments of the present application, based on the foregoing solution, the device further includes: a write request sending unit configured to send a data write request to the master node so that the master node writes the data specified to be written in the data write request; a notification receiving unit configured to receive a notification message sent by the master node, where the notification message is generated by the master node after receiving a write copy completion message sent by a slave node in the available state.

[0010] In some embodiments of the present application, based on the foregoing solution, the acquisition unit is configured to: acquire a node status view from the master node, where the node status view includes the status identifiers of the respective slave nodes; determine the status identifiers of the respective slave nodes according to the node status view.

[0011] According to one aspect of the embodiments of the present application, a data processing device applied to a distributed system is provided. The distributed system includes a master node and multiple slave nodes. The device includes: a third sending unit configured to send the status identifiers of the respective slave nodes to a terminal device, where the status identifiers are used to indicate whether the slave nodes are in an available state, and among them, the slave nodes in the available state store copies of the data in the master node; a receiving unit configured to receive a data read request sent by the terminal device, where the data read request is sent by the terminal device when it determines that there is no target slave node in the available state in the distributed system according to the status identifiers of the respective slave nodes; and a response unit configured to respond to the data read request and return response data for the data read request to the terminal device.

[0012] In some embodiments of the present application, based on the foregoing solution, the device further includes: a write request receiving unit configured to receive a data write request sent by the terminal device; and a writing unit configured to respond to the data write request, write the data specified to be written in the data write request, and generate an operation log, where the operation log is used to enable the respective slave nodes to determine the data to be written.

[0013] In some embodiments of the present application, based on the foregoing solution, the device further includes: an updating unit configured to update the status identifiers of the respective slave nodes according to the latency times returned after the respective slave nodes pull the operation log, to obtain new status identifiers of the respective slave nodes.

[0014] In some embodiments of the present application, based on the foregoing solution, the device further includes: a notification generating unit configured to generate notification information if a write copy completion message sent by a slave node in the available state is received, where the slave node in the available state is determined according to the new status identifiers of the respective slave nodes; and a notification sending unit configured to send the notification information to the terminal device.

[0015] In some embodiments of the present application, based on the foregoing solution, the updating unit is configured to: if the latency time is not received, determine that the slave node that has not sent the latency time is in an unavailable state; if the latency time is received and the latency time is greater than a preset latency time, determine that the slave node that has sent the latency time is in an unavailable state.

[0016] According to one aspect of the embodiments of the present application, a computer-readable medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, it implements the data processing method applied to a distributed system as described in the foregoing embodiments.

[0017] According to one aspect of the embodiments of the present application, an electronic device is provided, including: one or more processors; a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the data processing method applied to a distributed system as described in the above embodiments.

[0018] According to one aspect of the embodiments of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the data processing method applied to a distributed system provided in the above various alternative embodiments.

[0019] In the technical solutions provided in some embodiments of the present application, target slave nodes in an available state in a distributed system are determined through the status identifiers of each slave node. Then, a data read request is sent to the target slave nodes. After submitting the data read request to the target slave nodes, response data for the data read request returned from the target slave nodes is received. If there are no target slave nodes in an available state in the distributed system, the data read request is sent to the target slave nodes, and then response data for the data read request returned from the master node is received. The technical solutions of the embodiments of the present application have at least the following beneficial effects:

[0020] (1) In the technical solution of the embodiments of the present application, the status identifier is used to indicate whether a slave node is in an available state. The slave nodes in an available state store copies of the data in the master node, indicating that the data write request only needs to be confirmed in the slave nodes in an available state. In other words, the data write request will not be blocked because there are slave nodes in an unavailable state in the distributed system, thus ensuring high availability.

[0021] (2) In addition, when there are target slave nodes in an available state, the data read request is sent to the target slave nodes. Since the slave nodes in an available state store copies of the data in the master node, sending the data read request to the target slave nodes can ensure that the data written previously can be read, guaranteeing strong consistency of read-after-write.

[0022] (3) Furthermore, even if there are nodes in an unavailable state in the distributed system, data can still be read from the slave nodes or the master node in an available state, thus leveraging the performance advantages of the remaining nodes and ensuring high performance.

[0023] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Description of the Drawings

[0024] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application. Obviously, the accompanying drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:

[0025] Figure 1 A schematic diagram showing an exemplary system architecture to which the technical solution of the embodiment of this application can be applied;

[0026] Figure 2 A flowchart showing a data processing method applied to a distributed system according to an embodiment of this application;

[0027] Figure 3 A flowchart showing a data processing method applied to a distributed system according to an embodiment of this application;

[0028] Figure 4 A flowchart showing a data processing method applied to a distributed system according to an embodiment of this application;

[0029] Figure 5 A flowchart showing a data processing method applied to a distributed system according to an embodiment of this application;

[0030] Figure 6 A block diagram showing a data processing device applied to a distributed system according to an embodiment of this application;

[0031] Figure 7 A block diagram showing a data processing device applied to a distributed system according to an embodiment of this application;

[0032] Figure 8 A schematic diagram showing the structure of a computer system of an electronic device suitable for implementing the embodiment of this application. Detailed implementation manners

[0033] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and the concept of the example embodiments will be fully conveyed to those skilled in the art.

[0034] In addition, the described features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.

[0035] It should be noted that the terms used in the specification and claims of the present application and the above-mentioned drawings are only for describing embodiments and are not intended to limit the scope of the present application. It should be understood that when the terms "comprise", "include", "have", etc. are used herein, they specify the presence of the stated features, wholes, steps, operations, elements, components, and / or their groups, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their groups.

[0036] It will be further understood that although the terms "first", "second", "third", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the present invention, the first element may be referred to as the second element. Similarly, the second element may be referred to as the first element. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0037] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0038] The flowcharts shown in the drawings are only exemplary illustrations and do not necessarily include all the contents and operations / steps, nor are they necessarily executed in the order described. For example, some operations / steps may be decomposed, while some operations / steps may be combined or partially combined, so the actual execution order may change according to the actual situation.

[0039] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more.

[0040] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are applicable to the following explanations.

[0041] Cloud computing: It refers to the delivery and usage model of IT infrastructure, which means obtaining the required resources through the network in a demand-driven and easily scalable manner; in a broad sense, cloud computing refers to the delivery and usage model of services, which means obtaining the required services through the network in a demand-driven and easily scalable manner. Such services can be related to IT and software, the Internet, or other services. Cloud computing is the product of the development and integration of traditional computer and network technologies such as grid computing, distributed computing, parallel computing, utility computing, network storage technologies, virtualization, and load balance.

[0042] Cloud storage: It is a new concept extended and developed on the basis of the cloud computing concept. A distributed cloud storage system (hereinafter referred to as the storage system) refers to a storage system that combines a large number of different types of storage devices (storage devices are also called storage nodes) in the network through cluster applications, grid technologies, and distributed file systems, and works together through application software or application interfaces to jointly provide data storage and business access functions to the outside world.

[0043] Non-relational database: It is a data storage mode used to store a large amount of non-relational data in the Web 2.0 era, meeting the requirements of concurrent access performance, storage capacity, scalability, and reliability.

[0044] MongoDB: It is an open-source, schema-free, document-oriented, high-performance distributed NoSQL database. A replica set in MongoDB contains multiple instances, and the instances may be in states such as primary and secondary. Data is written through the primary node, and then the operation log (oplog) is actively synchronized to other nodes based on the Raft-like replication set protocol, so as to ensure the consistency of data on all nodes in the replica set and form replicas with each other; read operations can be selectively distributed to all nodes, effectively improving the performance of database queries.

[0045] As a very flexible and widely used distributed NoSQL database product, MongoDB provides users with a very rich set of parameters and modes to control the read and write strategies of the database, namely read concern, write concern, and read preference. Users can configure the desired database read and write model through the combination of these parameters.

[0046] There are three parameters related to the read and write model control of MongoDB, namely: Read Concern, Write Concern, and Read Preference. Users can effectively control these three parameters to adjust the consistency and availability guarantee levels and establish corresponding read and write strategies, such as providing stronger consistency guarantees for the database or providing slightly weaker consistency guarantees but higher availability guarantees for the database. The following briefly introduces these three parameters:

[0047] I. Read Concern

[0048] The levels of read concern can be the following:

[0049] (1) Local level: It can read any data on the current node.

[0050] (2) Majority level: It can only read the data that has been successfully written to the majority of nodes.

[0051] (3) Linearizable level: Linear level, which can read all successful write operations that have been completed and confirmed by the majority of nodes before this operation.

[0052] (4) Available level: Available level, which provides greater partition tolerance in the sharded cluster scenario and is the same as the Local level in the replica set scenario.

[0053] (5) Snapshot level: Snapshot level, which is used in the multi-document transaction scenario.

[0054] The original intention of read concern is to solve the common "dirty read" problem in the database. For example, a user reads a piece of data from the primary node of MongoDB, but this data has not been synchronized to the majority of nodes, and then the primary node fails. After recovery, this primary node will roll back the data that has not been synchronized to the majority of nodes, resulting in the user reading "dirty data". When the read concern is specified as the Majority level, it can be ensured that the data read has been synchronized to the majority of nodes and will definitely not be rolled back, so there will be no "dirty read" problem. It should be noted that the Majority level does not guarantee that the data read is the latest.

[0055] II. Write Concern

[0056] The specific options for write concern include:

[0057] 1. w: How many nodes the data needs to be written to before returning to the client

[0058] (1) {w:0}: It means that no confirmation of the write operation is required for the client, which is applicable to scenarios where high write performance is required and data security requirements (data may be lost) are secondary.

[0059] (2) {w:1}: It means that the write operation is confirmed only on the current primary node. When w is specified as N, the confirmation is based on the number of members in the replica set. For example, when there are 3 nodes in the replica set, w:2 means that the write operation needs to be confirmed on two nodes before returning; w:3 means that the write operation needs to be confirmed by all nodes before returning.

[0060] (3) {w:majority}: It means that the write operation is confirmed on a majority of nodes in the replica set, which is applicable to scenarios where high data security requirements are required and write performance is secondary.

[0061] (4) Custom replication guarantee rule: It needs to be used in conjunction with "tag". For example, the tag of high-performance nodes is specified as Solid State Disk (SSD). The write operation is returned after these nodes confirm it.

[0062] 2. j: Whether to persist the journal before returning to the client when writing

[0063] (1) {j:true}: It means that the journal is persisted before returning to the client when writing.

[0064] (2) {j:false}: It means that the journal is not persisted before returning to the client when writing.

[0065] 3. wtimeout: It represents the write timeout.

[0066] When w greater than 1 or majority is specified, if a node fails, it may cause the write confirmation condition to never be met, thus blocking the write request. In this case, this problem can be avoided by setting the write timeout.

[0067] III. Read Preference

[0068] Read preference describes how a MongoDB client sends read requests to the primary node or secondary nodes in a replica set. Its specific options include:

[0069] (1) Primary: Read only from the primary node, which is the default mode.

[0070] (2) Primary Preferred: Prefer to read from the primary node and read from secondary nodes when the primary node is unavailable.

[0071] (3)Secondary: Read-only secondary node.

[0072] (4)Secondary Preferred: Prefer to read from secondary nodes. If secondary nodes are unavailable, read from the primary node.

[0073] (5)Nearest: Read from the node with the lowest network latency.

[0074] (6)Custom Rule: Customize the node to be read based on attributes such as the location or performance of nodes in the replica set. Nearest can be regarded as a general custom rule.

[0075] Due to various reasons (such as network congestion, low disk throughput, long-term high-load operation, etc.), the secondary nodes in the replica set may lag far behind the primary node. If users choose to read from secondary nodes, they may read overly old data. Therefore, in versions after 3.4, MongoDB introduced a new control parameter for read preference: maxStalenessSeconds. When users set the read preference to Secondary Preferred, they can specify this parameter to avoid reading overly old data.

[0076] IV. Causal Consistency

[0077] When an operation logically depends on a previous operation, there is a causal relationship between the two operations. For example, the insert operation of a document and the update operation on that document have a causal relationship. If the insert operation is not completed, the update operation cannot be executed. So-called causal consistency means that operations with a causal relationship meet their strict order requirements. In MongoDB, by specifying specific read concern and write concern, operations can be executed in an order that respects causal relationships, and users can also observe results that conform to causal consistency constraints.

[0078] The usage method is relatively simple and is roughly as follows:

[0079] a. The client starts a new causal consistency client session and specifies that both the read concern and write concern are "majority";

[0080] b. Execute a series of read and write operations, and MongoDB will return the global logical clock;

[0081] c. The client session will continuously track this clock to ensure causal consistency.

[0082] Causal consistency is implemented based on the global logical clock "ClusterTime". For two operations with a causal relationship, their timing is roughly as follows:

[0083] 1. The client sends a write operation to the primary node. {w:1} means that the primary node will not wait for the write operation to be replicated to other secondary nodes before returning, but will return immediately after committing it to its own operation log;

[0084] 2. The primary node calculates the global logical clock "ClusterTime" and increments it;

[0085] 3. The primary node returns the result containing the global logical clock (stored in the "OpTime" field) to the client;

[0086] 4. The client conditionally updates the view of "lastOperationTime" based on the received return value;

[0087] 5. The client sends a read operation to the secondary node. To ensure "reading the data it has written" (i.e., ensuring causal consistency), it specifies "afterClusterTime" to indicate the data after the read operation time;

[0088] 6. The secondary node checks if the data it expects to read is in its own operation log. If not, it will block here until the condition is met;

[0089] 7. The result is returned to the client;

[0090] To make MongoDB meet the requirements of causal consistency, the following conditions need to be met: 1) Specify the client session as the causal consistency category; 2) Specify both read concern and write concern as "Majority"; 3) All operations must be completed within the same client session. If not in the same session, the value of "afterClusterTime" needs to be set manually (this value can be obtained from other causally consistent sessions).

[0091] In MongoDB, the default read concern and write concern configurations are "local" and "{w:1}". Under such read and write strategies, the causal consistency of the database cannot be guaranteed. If the read preference is to read from secondary nodes first, old data, dirty data, etc. may be read.

[0092] MongoDB chooses to completely leave the trade-off issue among availability, consistency, and performance to the user for free choice. By choosing different read concerns, write concerns, and read preferences, the user can set their database into the desired read and write model. The common read and write models in business can be summarized as shown in Table 1 below:

[0093]

[0094] Table 1

[0095] As can be seen from Table 1, these traditional read-write models either fail to leverage the advantages of slave nodes, or there is a possibility of reading old data, or there are certain availability issues (node downtime may cause write operations to fail). The existing read-write models of MongoDB, including the newly introduced causal consistency model, still have the defect of being unable to meet the high availability, strong consistency, and high performance requirements of some users in complex scenarios.

[0096] In response, the embodiments of the present application provide a data processing method applied to a distributed system. The distributed system includes a master node and multiple slave nodes. The terminal device obtains the status identifiers of each slave node from the master node. The status identifier is used to indicate whether the slave node is in an available state. Among them, the slave nodes in the available state store copies of the data in the master node. Then, the terminal device can determine whether there is a target slave node in the available state in the distributed system according to the status identifiers of each slave node. If there is a target slave node, the data read request is sent to the target slave node, and the response data for the data read request returned by the target slave node is received. If there is no target slave node, the data read request is sent to the master node, and the response data for the data read request returned by the master node is received. The technical solution of the embodiments of the present application can better meet the high availability, strong consistency, and high performance requirements encountered by some services in the actual environment:

[0097] First of all, the technical solution of the embodiments of the present application uses the status identifier to indicate whether the slave node is in an available state. The slave nodes in the available state store copies of the data in the master node, indicating that the data write request only needs to be confirmed in the slave nodes in the available state. In other words, the data write request will not be blocked because there are slave nodes in the unavailable state in the distributed system, thus ensuring high availability;

[0098] In addition, when there is a target slave node in the available state, the data read request is sent to the target slave node. Because the slave nodes in the available state store copies of the data in the master node, therefore, sending the data read request to the target slave node can ensure that the data written previously is read, ensuring strong consistency of read after write;

[0099] Furthermore, even if there are nodes in the unavailable state in the distributed system, data can still be read from the slave nodes or the master node in the available state, thus leveraging the performance advantages of the remaining nodes and ensuring high performance.

[0100] Figure 1 FIG. shows a schematic diagram of an exemplary system architecture to which the technical solution of the embodiments of the present application can be applied. The system architecture 100 may include a terminal device 101 and a distributed system 102.

[0101] The terminal device 101 can be an actual hardware device, such as a general computer device, or a virtual software application, such as a client running on a computer device. The client can be a virtual machine, virtual machine software, database software, and so on.

[0102] When specifically implementing the distributed system 102 in this application, it can include a server, which can be used to provide storage, computing, and network resources for the distributed system. The server can be called a node. The distributed system 102 can select one of the servers as the master node and set one or several other servers to store the same data as the master node, that is, as a copy of the master node. The copy used as a backup can be called a slave node. For example Figure 1 the distributed system 102 includes a master node 102a and three slave nodes 102b, 102c, and 102d.

[0103] It should be noted that Figure 1 the three slave nodes 102b, 102c, and 102d shown do not constitute a limitation on the slave nodes in the distributed system. It can be understood that the distributed system 102 can include a larger number of slave nodes.

[0104] It should also be noted that the servers included in the distributed system 102 can be independent physical servers or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0105] The terminal device 101 can initiate service requests such as write requests, read requests, or data synchronization requests to the nodes in the distributed system 102. In an embodiment of this application, the terminal device 101 can obtain the status identifiers of each slave node in the distributed system 102 through a driver. The status identifier is used to indicate whether the slave node is in an available state. The slave node in the available state stores a copy of the master node data. After obtaining the status identifiers of each slave node, the terminal device 101 can determine whether there is a target slave node in the available state in the distributed system 102 according to the status identifiers of each slave node. If there is a target slave node, the terminal device 101 can send a data read request to the target slave node. On the contrary, if there is no target slave node, the terminal device 101 can send a data read request to the master node 102a.

[0106] In one embodiment of the present application, the terminal device 101 may further send a data write request to the master node 102a, so that the master node 102a writes the data specified in the data write request. The master node 102a writes the data and updates the operation log, so that the slave nodes can determine the data to be written by pulling the operation log. After determining the data to be written, the slave nodes can write a copy of the data locally.

[0107] In one embodiment of the present application, after a slave node in an available state writes a copy of the data locally, it needs to send a write copy completion message to the master node 102a. After receiving the write copy completion message returned by the slave node in the available state, the master node 102a may generate a response message for the data write request and send the response message for the data write request to the terminal device 101.

[0108] The implementation details of the technical solutions in the embodiments of the present application are elaborated in detail below:

[0109] Figure 2 The flowchart of a data processing method applied to a distributed system according to an embodiment of the present application is shown. This data processing method can be executed by a terminal device, and the terminal device can be Figure 1 the terminal device 101 shown in Figure 2 As shown, the method includes:

[0110] Step S210, obtain the status identifiers of each slave node from the master node, where the status identifier is used to indicate whether the slave node is in an available state. Among them, the slave node in the available state stores a copy of the data in the master node;

[0111] Step S220, determine whether there is a target slave node in the available state in the distributed system according to the status identifiers of each slave node;

[0112] Step S230, if there is the target slave node, send a data read request to the target slave node and receive the response data for the data read request returned by the target slave node;

[0113] Step S240, if there is no such target slave node, send the data read request to the master node and receive the response data for the data read request returned by the master node.

[0114] These steps are described in detail below.

[0115] In step S210, obtain the status identifiers of each slave node from the master node, where the status identifier is used to indicate whether the slave node is in an available state. Among them, the slave node in the available state stores a copy of the data in the master node.

[0116] Since the traditional database read-write model either has the possibility of reading old data or has certain availability issues, although the read mode with write concern as {w:majority} plus reading from replicas is already a compromise solution, it still lacks finer-grained control over consistency and high performance.

[0117] In addition, the introduction of the causal consistency model effectively strengthens the available read-write models and the level of consistency guarantee, but there are still many defects, such as: 1) Reading from replicas may cause blocking, resulting in unexpected degradation of read performance; 2) It is necessary to specify that both read / write concerns are "majority", and the write performance has not been improved compared to {w:majority}; 3) The application program must ensure that only one thread performs read and write operations in the causally consistent client session at a time. If causal consistency needs to be ensured in multiple client sessions, it is necessary to pass the global logical clock in the session, which is rather cumbersome; 4) It is necessary to use a driver that supports causally consistent client sessions.

[0118] It can be seen that the existing read-write models still cannot meet the high-availability, strong-consistency, and high-performance requirements of some customers in complex scenarios.

[0119] To solve this technical problem, in this embodiment, first, the master node in the distributed system can maintain the information of the status identifiers of each slave node. The status identifiers of each slave node are used to identify whether the slave node is in an available state. The status identifier can effectively distinguish between the slave nodes in the available state and those in the unavailable state.

[0120] In an embodiment of the present application, since the status identifiers of each slave node are maintained on the master node, if the master node fails, master-slave switchover can be performed to switch a slave node to a new master node. In this case, the status identifiers of each slave node will be maintained on the new master node.

[0121] It can be understood that the slave nodes in the available state can be used for data processing, while the slave nodes in the unavailable state cannot be used for data processing. For example, the user's data read request can be selectively sent to the nodes in the available state instead of those in the unavailable state; the user's data write request only needs to be confirmed on the slave nodes in the available state and does not need to be confirmed on the slave nodes in the unavailable state. That is to say, when the user's data write request is sent to the master node, after the master node writes the data, as long as the master node receives the message indicating that the copy has been written successfully on the slave nodes in the available state, it can return a notification message to the terminal device.

[0122] Because the user's data write requests only need to be confirmed by the slave nodes in the available state, rather than by the slave nodes in the unavailable state, in this embodiment, the slave nodes in the available state store copies of the data in the master node.

[0123] When the user has a need to read data, a data read request can be sent to the distributed system through the terminal device. The data read request is mainly used to request to read data from the nodes in the distributed system. Before the terminal device sends the data read request, the terminal device can obtain the status identifiers of each slave node from the master node of the distributed system through the driver.

[0124] In an embodiment of the present application, the master node can maintain the status identifiers of each slave node in the form of a node status view. A view is actually a virtual table. The form of the view is similar to that of a real data table, including a series of columns and row data with names. However, the view does not exist in the database in the form of a stored data value set. Therefore, maintaining the status identifiers of each slave node in the form of a node status view has a lower maintenance cost.

[0125] In this embodiment, the terminal device can obtain the node status view, which contains the status identifiers of each slave node. Then, according to the node status view, the terminal device can determine the status identifiers of each slave node, that is, determine whether each slave node is in the available state.

[0126] Step S220, determine whether there is a target slave node in the available state in the distributed system according to the status identifiers of each slave node.

[0127] After obtaining the status identifiers of each slave node from the master node, the terminal device can also determine whether there is a target slave node in the available state in the distributed system according to the status identifiers of each slave node.

[0128] Step S230, if there is the target slave node, send the data read request to the target slave node and receive the response data for the data read request returned by the target slave node.

[0129] In this embodiment, if the terminal device determines that there is a target slave node in the distributed system, the data read request can be sent to the target slave node. After submitting the data read request to the target slave node, receive the response data for the data read request returned from the target slave node.

[0130] Step S240, if there is no such target slave node, send the data read request to the master node and receive the response data for the data read request returned by the master node.

[0131] Conversely, if the terminal device determines that there is no target slave node, that is, there is no available slave node in the distributed system. In this case, the terminal device can send a data read request to the master node. After receiving the data read request, the master node returns the response data for the data read request to the terminal device.

[0132] It should be noted that in this embodiment, the master node is always in an available state. Therefore, when there is no target slave node in an available state, the master node can respond to the data read request, and then the master node returns the response data for the data read request to the terminal device.

[0133] Based on the technical solutions of the above embodiments, by using the status flag to indicate whether the slave node is in an available state, the slave node in an available state stores a copy of the data in the master node, which means that the data write request only needs to be confirmed in the available slave node. In other words, it means that the data write request will not be blocked due to the existence of an unavailable slave node in the distributed system, thus ensuring high availability. In addition, the data read request is sent to the target slave node in an available state, which can ensure that the previously written data can be read, guaranteeing the strong consistency of read-after-write. Moreover, even if there are nodes in an unavailable state, the data can still be read from the available slave node or the master node, thus giving play to the performance advantages of the remaining nodes and ensuring high performance.

[0134] In an embodiment of the present application, in addition to the data read request, the user may also have a need for data writing. In this embodiment, as Figure 3 shown, the data processing method further includes steps S310 - S320, which are described in detail as follows:

[0135] Step S310: Send a data write request to the master node to enable the master node to write the data specified in the data write request.

[0136] Specifically, the user can send a data write request to the master node through the terminal device. The data write request is mainly used to write the specified data into the master node of the distributed data system. Among them, the specified data to be written can be data such as video, audio, image, and text.

[0137] There are various ways for the master node to receive the data write request sent by the terminal device. In some embodiments, the terminal device may directly carry the data to be written in the data write request. The master node receives the data write request and writes the data to be written carried in the data write request. Of course, in other embodiments, the terminal device may also store the data to be written in a third-party database and send the storage address of the data to be written in the third-party database to the master node in the data write request. After receiving the data write request, the master node obtains the storage address of the data to be written stored in the third-party database according to the data write request. Then, the master node obtains the data in the third-party database according to the storage address. After obtaining the data, the master node may also send a prompt message to the terminal device to prompt the terminal device that the data has been successfully obtained.

[0138] Step S320: Receive the notification message sent by the master node, where the notification message is generated by the master node after receiving the write copy completion message sent by the available slave node.

[0139] After receiving the data write request, the master node can, in response to the data write request, write the data in the local of the master node.

[0140] In one embodiment, after writing the data, the master node may generate an operation log. A log refers to a record of the completed processing, that is, the master node can encapsulate the written data into an operation log. The available slave node can perform data synchronization by pulling the operation log, ultimately ensuring data consistency between the master node and the available slave node in the distributed system.

[0141] Among them, after writing a copy of the data locally, the available slave node may send a write copy completion message to the master node. After receiving the write copy completion message returned by the available slave node, the master node may generate a response message for the data write request and send the response message for the data write request to the terminal device.

[0142] In this embodiment, each data write request only needs to be confirmed on the available nodes, which is more flexible than {w:all} and {w:majority} in write concern because the number of nodes that need to be confirmed may change dynamically according to whether the nodes are in an available state. When there is only one available node, that is, after the master node receives the write copy completion message sent by this node, it can return a notification message to the terminal device. At this time, it is equivalent to {w:1}; when all slave nodes in the distributed system are in an available state, that is, the master node needs to receive the write copy completion messages sent by all slave nodes before it can return a notification message to the terminal device. At this time, it is equivalent to {w:all}.

[0143] Figure 4 The flowchart of a data processing method applied to a distributed system according to an embodiment of the present application is shown. The data processing method can be executed by a master node in the distributed system, and the distributed system can be Figure 1 the distributed system 102 shown in Figure 1 and the master node can be the master node 102a shown in Figure 4 As shown in

[0144] Step S410: Send the status identifiers of each slave node to the terminal device. The status identifiers are used to indicate whether the slave node is in an available state. Among them, the slave nodes in the available state store copies of the data in the master node;

[0145] Step S420: Receive the data read request sent by the terminal device. The data read request is sent by the terminal device when it determines that there is no target slave node in the available state in the distributed system according to the status identifiers of each slave node;

[0146] Step S430: Respond to the data read request and return the response data for the data read request to the terminal device.

[0147] The following is a detailed description of these steps:

[0148] In step S410, the status identifiers of each slave node are sent to the terminal device. The status identifiers are used to indicate whether the slave node is in an available state. Among them, the slave nodes in the available state store copies of the data in the master node.

[0149] Specifically, the master node can send the status identifiers of each slave node to the terminal device under the action of the terminal device driver. The status identifiers of each slave node are used to identify whether the slave node is in an available state, and the status identifiers of each slave node can effectively distinguish the slave nodes in the available state from the slave nodes in the unavailable state. Among them, the slave nodes in the available state store copies of the data in the master node.

[0150] In step S420, receive the data read request sent by the terminal device. The data read request is sent by the terminal device when it determines that there is no target slave node in the available state in the distributed system according to the status identifiers of each slave node.

[0151] When the terminal device learns from the status identifiers of each slave node that there is no target slave node in the available state in the distributed system, it can send the data read request to the master node, and the master node receives the data read request. The data read request is mainly used to request to read the data stored in the distributed system.

[0152] In step S430, respond to the data read request and return the response data for the data read request to the terminal device.

[0153] After receiving the data read request from the terminal device, the master node responds to the data read request and returns the response data for the data read request to the terminal device.

[0154] Reference Figure 5 , in an embodiment of the present application, when the terminal device initiates a data write request, it may first request from the master node. In this embodiment, the method specifically includes steps S510 - S520, which are described in detail as follows:

[0155] Step S510: Receive the data write request sent by the terminal device.

[0156] Among them, the data write request, that is, the request to write data, may be a write operation to request to change existing data, or a write operation to add new data.

[0157] When the user has a need to write data, the terminal device can send a data write request to the master node. It can be understood that the data write request sent by the terminal device to the master node may directly carry the data to be written, or may only carry the storage address of the data to be written in a third - party database, so that the master node can obtain the data to be written according to the storage address.

[0158] Step S520: Respond to the data write request, write the data specified to be written in the data write request, and generate an operation log, where the operation log is used to enable each slave node to determine the data to be written.

[0159] After receiving the data write request, the master node can respond to the data write request, write the data specified to be written in the data write request in the local of the master node, and can generate an operation log. Among them, the operation log contains the data written by the master node, so that each slave node in the distributed system can determine the data to be written by pulling the operation log.

[0160] In an embodiment of the present application, the master node can update the status identifiers of each slave node according to the latency time returned after each slave node pulls the operation log, and obtain new status identifiers of each slave node.

[0161] In one embodiment of the present application, the master node can update the status flags of each slave node according to the latency time returned after each slave node pulls the operation log. Specifically, it may include: If the master node does not receive the latency time, it can determine that the slave node that has not sent the latency time is in an unavailable state. If the slave node maintained on the master node before was in an available state, then at this time, the status flag of this slave node can be updated to be in an unavailable state; If the master node receives the latency time, and the received latency time is greater than the preset latency time, it indicates that the slave node lags far behind the master node. Therefore, the master node can also determine that the slave node that sent this latency time is in an unavailable state. If the slave node maintained on the master node before was in an available state, then at this time, the status flag of this slave node can be updated to be in an unavailable state.

[0162] In one embodiment of the present application, the master node can update the status flags of each slave node according to the latency time returned after the slave node pulls the operation log. After updating the status flags of each slave node, the master node can know which slave nodes are in an available state at this time. Therefore, if the master node receives a write replica completion message sent by a slave node in an available state, the master node can generate a notification message and return the notification message to the terminal device.

[0163] In this embodiment, each data write request only needs to be confirmed on the nodes in an available state, which is more flexible than {w:all} and {w:majority} in write concern because the number of nodes that need to be confirmed may change dynamically according to whether the nodes are in an available state. When there is only one node in an available state, that is, when the master node receives the write replica completion message sent by this node, it can return the notification message to the terminal device. At this time, it is equivalent to {w:1}; When all slave nodes in the distributed system are in an available state, that is, the master node needs to receive the write replica completion messages sent by all slave nodes before it can return the notification message to the terminal device. At this time, it is equivalent to {w:all}.

[0164] The following introduces the device embodiments of the present application, which can be used to execute the data processing method applied to the distributed system in the above embodiments of the present application. For the details not disclosed in the device embodiments of the present application, please refer to the embodiments of the data processing method applied to the distributed system in the above of the present application.

[0165] Figure 6 shows a block diagram of a data processing device applied to a distributed system according to an embodiment of the present application. Refer to Figure 6As shown, a data processing device 600 according to an embodiment of the present application is applied to a distributed system, which includes a master node and multiple slave nodes. The device includes: an acquisition unit 602, a determination unit 604, a first sending unit 606, and a second sending unit 608.

[0166] Among them, the acquisition unit 602 is configured to acquire the status identifiers of each slave node from the master node. The status identifier is used to indicate whether the slave node is in an available state. Among them, the slave node in the available state stores a copy of the data in the master node. The determination unit 604 is configured to determine whether there is a target slave node in the available state in the distributed system according to the status identifiers of each slave node. The first sending unit 606 is configured to, if there is the target slave node, send a data read request to the target slave node and receive the response data for the data read request returned by the target slave node. The second sending unit 608 is configured to, if there is no target slave node, send the data read request to the master node and receive the response data for the data read request returned by the master node.

[0167] In some embodiments of the present application, the device further includes: a write request sending unit configured to send a data write request to the master node so that the master node writes the data specified to be written in the data write request; a notification receiving unit configured to receive the notification information sent by the master node. The notification information is generated by the master node after receiving the write copy completion message sent by the slave node in the available state.

[0168] In some embodiments of the present application, the acquisition unit 602 is configured to: acquire a node status view from the master node, and the node status view includes the status identifiers of each slave node; determine the status identifiers of each slave node according to the node status view.

[0169] Figure 7 Shows a block diagram of a data processing device applied to a distributed system according to an embodiment of the present application.

[0170] See Figure 7 As shown, a data processing device 700 according to an embodiment of the present application is applied to a distributed system, which includes a master node and multiple slave nodes. The device includes: a third sending unit 702, a receiving unit 704, and a response unit 706.

[0171] Among them, the third sending unit 702 is configured to send the status identifiers of the slave nodes to the terminal device, where the status identifiers are used to indicate whether the slave nodes are in an available state. Among them, the slave nodes in the available state store copies of the data in the master node; the receiving unit 704 is configured to receive the data read request sent by the terminal device, and the data read request is sent by the terminal device when it determines that there is no target slave node in the available state in the distributed system according to the status identifiers of the slave nodes; the response unit 706 is configured to respond to the data read request and return the response data for the data read request to the terminal device.

[0172] In some embodiments of the present application, the device further includes: a write request receiving unit configured to receive the data write request sent by the terminal device; a writing unit configured to respond to the data write request, write the data specified to be written in the data write request, and generate an operation log, where the operation log is used to enable each slave node to determine the data to be written.

[0173] In some embodiments of the present application, the device further includes: an updating unit configured to update the status identifiers of the slave nodes according to the latency times returned after each slave node pulls the operation log, to obtain the new status identifiers of the slave nodes.

[0174] In some embodiments of the present application, the device further includes: a notification generating unit configured to generate a notification message if a write copy completion message sent by a slave node in the available state is received, where the slave node in the available state is determined according to the new status identifiers of the slave nodes; a notification sending unit configured to send the notification message to the terminal device.

[0175] In some embodiments of the present application, the updating unit is configured to: if the latency time is not received, determine that the slave node that has not sent the latency time is in an unavailable state; if the latency time is received and the latency time is greater than a preset latency time, determine that the slave node that has sent the latency time is in an unavailable state.

[0176] Figure 8 FIG. shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application.

[0177] It should be noted that Figure 8 The computer system 800 of the electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0178] Such as Figure 8As shown, computer system 800 includes a Central Processing Unit (CPU) 801, which can perform various appropriate actions and processes according to a program stored in a Read-Only Memory (ROM) 802 or a program loaded from a storage section 808 into a Random Access Memory (RAM) 803, such as executing the methods described in the above embodiments. In the RAM 803, various programs and data required for system operations are also stored. The CPU 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An Input / Output (I / O) interface 805 is also connected to the bus 804.

[0179] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including, for example, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc., and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read from it can be installed into the storage section 808 as needed.

[0180] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by a Central Processing Unit (CPU) 801, various functions defined in the system of the present application are executed.

[0181] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The computer program contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0182] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0183] The units involved in the embodiments of the present application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation to the units themselves.

[0184] As another aspect, the present application also provides a computer-readable medium, which can be included in the electronic device described in the above embodiments; or can exist alone without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the method described in the above embodiments.

[0185] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0186] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented in software or in the form of software combined with necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal device, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0187] After considering the specification and practicing the disclosed embodiments herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application.

[0188] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A data processing method applied to a distributed system, characterized in that, the distributed system includes a master node and multiple slave nodes, and the method includes: obtaining the status identifiers of the slave nodes from the master node, where the status identifiers are used to indicate whether the slave nodes are in an available state. Among them, the slave nodes in the available state store copies of the data in the master node; the status identifiers of the slave nodes are maintained on the master node, and when the master node fails, one of the available slave nodes is switched to a new master node; determining whether there is a target slave node in the available state in the distributed system according to the status identifiers of the slave nodes; if there is the target slave node, sending a data read request to the target slave node and receiving the response data for the data read request returned by the target slave node; if there is no such target slave node, sending the data read request to the master node and receiving the response data for the data read request returned by the master node.

2. The method according to claim 1, characterized in that, the method further includes: sending a data write request to the master node so that the master node writes the data specified in the data write request; receiving the notification information sent by the master node, where the notification information is generated by the master node after receiving the write copy completion message sent by the available slave node.

3. The method according to claim 1, characterized in that, obtaining the status identifiers of the slave nodes from the master node includes: obtaining a node status view from the master node, where the node status view contains the status identifiers of the slave nodes; determining the status identifiers of the slave nodes according to the node status view.

4. A data processing method applied to a distributed system, characterized in that, the distributed system includes a master node and multiple slave nodes, the method is applied to the master node, and the method includes: maintaining the status identifiers of the slave nodes and sending the status identifiers of the slave nodes to the terminal device, where the status identifiers are used to indicate whether the slave nodes are in an available state. Among them, the slave nodes in the available state store copies of the data in the master node, and when the master node fails, one of the available slave nodes is switched to a new master node; receiving the data read request sent by the terminal device, where the data read request is sent by the terminal device when it determines that there is no target slave node in the available state in the distributed system according to the status identifiers of the slave nodes; responding to the data read request and returning the response data for the data read request to the terminal device.

5. The method according to claim 4, characterized in that, the method further includes: receiving the data write request sent by the terminal device; responding to the data write request, writing the data specified in the data write request, and generating an operation log, where the operation log is used to enable the slave nodes to determine the data to be written.

6. The method according to claim 5, characterized in that, The method further includes: Updating the status identifiers of the slave nodes according to the latency times returned after pulling the operation logs from the respective slave nodes, to obtain new status identifiers of the respective slave nodes.

7. The method according to claim 6, wherein, The method further includes: If a write copy completion message sent by an available slave node is received, generating notification information, where the available slave node is determined according to the new status identifiers of the respective slave nodes; Sending the notification information to the terminal device.

8. The method according to claim 6, wherein, Updating the status identifiers of the respective slave nodes according to the latency times returned after pulling the operation logs from the respective slave nodes, to obtain new status identifiers of the respective slave nodes, includes: If the latency time is not received, determining that the slave node that did not send the latency time is in an unavailable state; If the latency time is received and the latency time is greater than a preset latency time, determining that the slave node that sent the latency time is in an unavailable state.

9. A data processing apparatus applied to a distributed system, wherein, The distributed system includes a master node and multiple slave nodes, and the apparatus includes: An obtaining unit, configured to obtain the status identifiers of the respective slave nodes from the master node, where the status identifier is used to indicate whether a slave node is in an available state, and among them, the slave node in an available state stores a copy of the data in the master node; the status identifiers of the respective slave nodes are maintained on the master node, and when the master node fails, one of the available slave nodes is switched to be the new master node; A determining unit, configured to determine whether there is a target slave node in an available state in the distributed system according to the status identifiers of the respective slave nodes; A first sending unit, configured to, if there is the target slave node, send a data read request to the target slave node, and receive response data for the data read request returned by the target slave node; A second sending unit, configured to, if there is no target slave node, send the data read request to the master node, and receive response data for the data read request returned by the master node.

10. A data processing apparatus applied to a distributed system, wherein, The distributed system includes a master node and multiple slave nodes, the apparatus is applied to the master node, and the apparatus includes: A third sending unit, configured to maintain the status identifiers of the respective slave nodes, and send the status identifiers of the respective slave nodes to a terminal device, where the status identifier is used to indicate whether a slave node is in an available state, and among them, the slave node in an available state stores a copy of the data in the master node, and when the master node fails, one of the available slave nodes is switched to be the new master node; A receiving unit, configured to receive a data read request sent by the terminal device, where the data read request is sent by the terminal device when it determines that there is no target slave node in an available state in the distributed system according to the status identifiers of the respective slave nodes; A response unit, configured to respond to the data read request and return response data for the data read request to the terminal device.

11. A computer-readable medium, having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, it implements the data processing method applied to a distributed system according to any one of claims 1-8.

12. An electronic device, characterized in that, comprising: a processor; and a memory, having computer-readable instructions stored thereon, and when the computer-readable instructions are executed by the processor, the data processing method applied to a distributed system according to any one of claims 1-8 is implemented.

13. A computer program product, characterized in that, the computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; a processor of a computer device reads and executes the computer instructions from the computer-readable storage medium, so that the computer device executes the data processing method applied to a distributed system according to any one of claims 1-8.

Citation Information

Patent Citations

  • Cluster data acquisition method, device and equipment

    CN110048896A