Distributed container cluster data push method and system based on Socket protocol
By introducing master and slave servers into distributed container clusters, using Redis to monitor the working status of slave servers, and automatically create alternative slave servers when abnormalities are abnormal, the problem of reduced efficiency in the containerized distributed environment is solved, and the independent recovery and rapid recovery of data push services are achieved.
Patent Information
- Application Number
- CN202310107936.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-09
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-02-09
AI Technical Summary
In a containerized distributed environment, when an exception occurs, automatically restarting the service or recreating the Socket connection may affect the data push of other target sources, resulting in reduced data push efficiency.
By introducing the master server and slave server in the distributed container cluster, the master server obtains the working status of the slave server from Redis, and automatically creates a substitute slave server when an exception occurs to realize the autonomous recovery of the data push service.
When the push service is abnormal from the server, this method can quickly restore the data push service, avoid affecting other services, and realize the rapid recovery of push service without affecting other services.
Smart Images

Figure CN116319783B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, in particular to the field of communication technology, and specifically to a distributed container cluster data push method and system based on the Socket protocol. Background Art
[0002] The Socket protocol is an agreement or a way for computers to communicate. Through the agreement of the Socket protocol, a computer can receive data from other computers and send data to other computers. After the connection is established between the server and the client based on the Socket protocol, both parties can continue to send data through the connection without having to create connections between the server and the client multiple times. Clustering refers to bringing together multiple servers to provide services together. Compared with a single server, a cluster can provide a higher service capacity; distribution refers to splitting jobs according to business and running them on different servers; distributed clusters refer to splitting different businesses and running them in collaboration by multiple servers according to a certain logic; distributed container clusters refer to containerizing all businesses based on distributed clusters.
[0003] Currently, in a containerized distributed environment, the data push service based on the Socket protocol needs to rely on the container's own fault detection mechanism or manual monitoring to determine whether the data push service is running normally. When a service anomaly is found, most will use the solution of automatically restarting the service or each node trying to re-create the Socket connection to restore the data push service. However, the above solution will cause the data push of other target sources to be affected by restarting the service after an anomaly occurs in the push service, thereby reducing the efficiency of data push. Summary of the invention
[0004] The present invention provides a distributed container cluster data push method and system based on Socket protocol.
[0005] According to a first aspect of the present disclosure, a distributed container cluster data push method based on the Socket protocol is provided. The method comprises:
[0006] The master server in the distributed container cluster obtains the working status of the first slave server in the distributed container cluster from redis and enters the first slave server workflow; the first slave server workflow includes:
[0007] When the working status of the first slave server is online, the master server queries the data point i of the last consumption of the first slave server from Redis, where i is a positive integer greater than or equal to 2;
[0008] The master server broadcasts a message to notify the first slave server to start the task;
[0009] After receiving the message of starting the task from the server, it will start consuming messages from data point i;
[0010] First, after the server obtains data from the data source, it will record the data point i;
[0011] First, the slave server pushes the data to the Socket target using the Socket channel;
[0012] First, the server determines whether the data push is normal based on the push result;
[0013] If the first slave server data push is normal, the data point i is recorded in redis, and new data is obtained from the data source for the next data push;
[0014] If the data push of the first slave server is abnormal, a message that the first slave server is working abnormally is sent to redis;
[0015] When the master server receives the message that the first slave server is working abnormally, it will continue to try to reconnect to the Socket target end;
[0016] When the master server successfully connects to the Socket target end, it will enter the next slave server workflow again.
[0017] According to the above aspects and any possible implementation manner, an implementation manner is further provided, wherein the method further includes:
[0018] When the working status of the first slave server is offline, the master server creates a second slave server to replace the first slave server;
[0019] When the second slave server is created successfully, the master server queries Redis for the data point i of the second slave server's last consumption;
[0020] The master server broadcasts a message to notify the second slave server to start the task;
[0021] After receiving the message of starting the task from the server, the second slave will start consuming messages from data point i. After obtaining the data from the data source, the second slave will record the data point i.
[0022] Second, the slave server pushes the data to the Socket target using the Socket channel;
[0023] Second, the slave server determines whether the data push is normal based on the push result;
[0024] If the data push from the second slave server is normal, the data point i is recorded in redis, and new data is obtained from the data source for the next data push;
[0025] If the data push of the second slave server is abnormal, a message about the abnormal operation of the second slave server is sent to redis;
[0026] When the master server receives the message that the second slave server is working abnormally, it will continue to try to reconnect to the Socket target end;
[0027] When the master server successfully connects to the Socket target end, it will enter the next slave server workflow again.
[0028] According to the above aspects and any possible implementation manner, an implementation manner is further provided, wherein the master server in the distributed container cluster obtains the working status of the first slave server in the distributed container cluster from redis, including:
[0029] The master server in the distributed container cluster monitors the status of data push tasks;
[0030] When the master server receives the start task instruction, the master server obtains the working status of the first slave server in the distributed container cluster from redis.
[0031] According to the above aspects and any possible implementation, an implementation is further provided, after the master server queries the data point position i of the last consumption by the first slave server from Redis, the method further includes:
[0032] The first slave server obtains data from the data source starting from data point i-1.
[0033] According to the above aspects and any possible implementation manner, an implementation manner is further provided, before the master server in the distributed container cluster obtains the working status of the first slave server in the distributed container cluster from redis, the method further includes:
[0034] All servers in the distributed container cluster trigger a workflow for determining the normal state of the primary server at the same time; the workflow for determining the normal state of the primary server includes:
[0035] All servers try to obtain the status information of the primary server in Redis;
[0036] If the status information of the master server is obtained in Redis, it is determined that the master server in the distributed container cluster is in a normal state.
[0037] According to the above aspects and any possible implementation manner, an implementation manner is further provided, before the master server in the distributed container cluster obtains the working status of the first slave server in the distributed container cluster from redis, the method further includes:
[0038] If the status information of the primary server is not obtained in Redis, the workflow of primary server election is triggered;
[0039] The workflow of the master server election includes:
[0040] All servers in the distributed container cluster try to store the same key in Redis at the same time, and determine the first server that stores the key in Redis as the master server, and the other servers as slave servers.
[0041] According to the above aspects and any possible implementation manner, an implementation manner is further provided, before the master server in the distributed container cluster obtains the working status of the first slave server in the distributed container cluster from redis, the method further includes:
[0042] The slave server reports the server working status to Redis so that Redis can record the online status of the slave server.
[0043] According to the above aspects and any possible implementation manner, an implementation manner is further provided, before the master server in the distributed container cluster obtains the working status of the first slave server in the distributed container cluster from redis, the method further includes:
[0044] The slave server obtains the status information of other servers from Redis.
[0045] As described above, and any possible implementation, a further implementation is provided, wherein the data source is a Kafka message queue.
[0046] According to the second aspect of the present disclosure, a distributed container cluster data push system based on the Socket protocol is provided. The system includes a master server, multiple slave servers and redis; wherein,
[0047] The master server in the distributed container cluster obtains the working status of the first slave server in the distributed container cluster from redis and enters the first slave server workflow; the first slave server workflow includes:
[0048] When the working status of the first slave server is normal, the master server queries the data point i of the last consumption of the first slave server from Redis, where i is a positive integer greater than or equal to 2;
[0049] The master server broadcasts a message to notify the first slave server to start the task;
[0050] After the first slave server receives the message to start the task, it will start consuming messages from data point i. After the first slave server obtains data from the data source, it will record data point i.
[0051] First, the slave server pushes the data to the Socket target using the Socket channel;
[0052] First, the server determines whether the data push is normal based on the push result;
[0053] If the first slave server data push is normal, the data point i is recorded in redis, and new data is obtained from the data source for the next data push;
[0054] If the data push of the first slave server is abnormal, a message that the first slave server is abnormal will be broadcast to redis;
[0055] When the master server receives the message that the first slave server is working abnormally, it will continue to try to reconnect to the Socket target end;
[0056] When the master server successfully connects to the Socket target end, it will enter the next slave server workflow again.
[0057] According to a third aspect of the present disclosure, an electronic device is provided, which includes a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the program, the method described above is implemented.
[0058] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.
[0059] The embodiment of the present application provides a distributed container cluster data push method and system based on the Socket protocol, which can obtain the working status of the first slave server in the distributed container cluster from redis through the master server in the distributed container cluster, and enter the first slave server workflow; wherein the first slave server workflow includes: when the working status of the first slave server is online, the master server queries the data point i of the last consumption of the first slave server from Redis, where i is a positive integer greater than or equal to 2; the master server broadcasts a message to notify the first slave server to start the task; after receiving the message to start the task, the first slave server will start consuming messages from data point i; after the first slave server obtains the data from the data source, it will record the data point i; the first slave server pushes the data to the Socket target end using the Socket channel; the first slave server sends the data to the target end according to ... The push result determines whether the data push is normal. If the data push of the first slave server is normal, the data point i is recorded in redis, and new data is obtained from the data source for the next data push. If the data push of the first slave server is abnormal, a message indicating that the first slave server is working abnormally is broadcast to redis. When the master server receives the message indicating that the first slave server is working abnormally, it will continue to try to reconnect to the Socket target end. When the master server successfully connects to the Socket target end, it will enter the next slave server workflow again. Based on this, when the push service of the first slave server is abnormal, the master server can take over the push service from the first slave server to restart it, and the push service abnormality can be autonomously recovered. After completing the push service, the master service enters the next slave server workflow, and the push service can be quickly restored without affecting other businesses.
[0060] It should be understood that the contents described in the summary of the invention are not intended to limit the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, among which:
[0062] Figure 1 A schematic diagram showing an exemplary operating environment in which embodiments of the present disclosure can be implemented;
[0063] Figure 2 A schematic diagram showing data flow according to an embodiment of the present disclosure is shown;
[0064] Figure 3Shows Figure 1 A schematic diagram of the interaction method between the sending end, the distributed container cluster data push system based on the Socket protocol, and the receiving end shown in ;
[0065] Figure 4 A schematic diagram is shown of triggering, ending, and abnormal recovery of a data push operation according to an embodiment of the present disclosure.
[0066] Figure 5 A schematic diagram showing a workflow for determining a normal state of a primary server according to an embodiment of the present disclosure is shown;
[0067] Figure 6 A schematic diagram showing a workflow of a master server election according to an embodiment of the present disclosure;
[0068] Figure 7 A block diagram of an exemplary electronic device capable of implementing embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0069] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0070] In addition, the term "and / or" in this article is only a description of the association relationship between the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0071] In the present disclosure, when an exception occurs in the push service of the first slave server, the main server can take over the first slave server to restart the push service, and can achieve autonomous recovery of the push service exception. After completing the push service, the main service enters the next slave server workflow, and can achieve rapid recovery of the push service without affecting other businesses.
[0072] Figure 1 1 is a schematic diagram of an exemplary operating environment 100 in which embodiments of the present disclosure can be implemented. Figure 1As shown, the operating environment 100 includes a sending end 101, a distributed container cluster data push system based on the Socket protocol 102, and a receiving end 103. Among them, the sending end 101 is the data source, the receiving end 103 is the Socket target end, and the distributed container cluster data push system based on the Socket protocol 102 includes a master server 1021, multiple slave servers, and redis 1023.
[0073] Among them, redis1023 includes a service registration center and a message broadcasting platform. The service registration center is used to provide service registration and service discovery for the Socket data push nodes in the distributed container cluster, and the message broadcasting platform is used to provide message broadcasting and subscription messages for the Socket data push nodes. The service registration center and the message broadcasting platform are created based on the single-threaded working mechanism and message broadcasting mechanism of the Remote Dictionary Server (Redis).
[0074] The socket data push node includes a master node and multiple slave nodes, that is, a master server 1021 and multiple slave servers. The multiple slave servers include a first slave server 1022 and other slave servers. The first slave server 1022 is any one of the multiple slave servers and can be represented as node A.
[0075] Figure 2 FIG. 1 shows a schematic diagram of data flow according to an embodiment of the present disclosure. Figure 2 As shown, the data flow process includes: obtaining data from a data source, and the Socket data push node pushes the obtained data to the Socket target end, that is, the data target address, through the Socket protocol.
[0076] Figure 3 Shows Figure 1 The schematic diagram of the interaction method 300 between the sending end 101, the distributed container cluster data push system 102 based on the Socket protocol, and the receiving end 103 shown in FIG. The specific steps in the distributed container cluster data push method based on the Socket protocol of the embodiment of the present disclosure can be found in Figure 3 .
[0077] In box 301, the master server 1021 in the distributed container cluster obtains the working status of the first slave server 1022 in the distributed container cluster from redis 1023, and enters the working process of the first slave server.
[0078] The first slave server workflow includes:
[0079] In box 302, when the working status of the first slave server 1022 is online, the master server 1021 queries the data point i of the last consumption of the first slave server 1022 from Redis 1023, where i is a positive integer greater than or equal to 2;
[0080] In block 303, the master server 1021 broadcasts a message to notify the first slave server 1022 to start the task;
[0081] In block 304, after receiving the message to start the task, the first slave server 1022 will start consuming messages from data point i;
[0082] In block 305 , the first slave server 1022 obtains data from the sender 101 .
[0083] In block 306 , after the first slave server 1022 obtains data from the transmitting end 101 , it records the data point i.
[0084] In block 307 , the first slave server 1022 pushes the data to the receiving end 103 using the Socket channel.
[0085] In block 308 , the receiving end 103 returns the data push result to the first slave server 1022 .
[0086] In block 309 , the first slave server 1022 determines whether the data push is normal based on the push result.
[0087] In box 310 , if the data push from the first slave server 1022 is normal, the data point i is recorded in redis 1023 .
[0088] In box 311, if the data push of the first slave server 1022 is abnormal, a message indicating that the first slave server 1022 is working abnormally is sent to redis 1023.
[0089] In box 312, redis 1023 broadcasts a message that the first slave server 1022 is operating abnormally.
[0090] In block 313 , after receiving the message that the first slave server 1022 is operating abnormally, the master server 1021 will continue to try to reconnect to the receiving end 103 .
[0091] In box 314, when the master server 1021 successfully connects to the receiving end 103, it will enter the next slave server workflow again.
[0092] In some embodiments, when the master server 1021 in the distributed container cluster enters the workflow of starting the slave server, the master server 1021 can obtain the working status of all slave servers from redis 1023. For ease of description, this disclosure takes the first slave server 1022 as an example.
[0093] When the working status of the first slave server 1022 is online, that is, when the first slave server 1022 is normal, the master server 1021 queries the data point i of the last consumption of the first slave server 1022 from redis1023, where i is a positive integer greater than or equal to 2. The master server 1021 broadcasts a message to notify the first slave server 1022 to start the task. After receiving the message to start the task, the first slave server 1022 will start consuming data from the specified point, that is, data point i; and after obtaining data from the data source each time, it will record this point, that is, data point i; and then push the data to the Socket target end using the Socket channel; and then judge whether this data push is normal based on the push result. If the data push of the first slave server 1022 is normal, the first slave server 1022 will record the point of this message consumption, that is, data point i, in redis1023, and pull new data from the data source to carry out the next data push process. If the data push of the first slave server 1022 is abnormal, the first slave server 1022 broadcasts a message of abnormal node operation to redis 1023. After receiving the message of abnormal operation of the first slave server 1022, the master server 1021 will continue to try to reconnect to the Socket target end. After the master server 1021 successfully connects to the Socket target end, it will enter the process of starting the slave server again.
[0094] In some embodiments, in order to achieve the effect of actively monitoring anomalies of large-scale cluster nodes, each slave server in multiple slave servers can be automatically traversed based on the workflow mode of the first slave server to achieve the effect of actively monitoring anomalies of large-scale cluster nodes in a distributed cluster environment.
[0095] According to the embodiments of the present disclosure, the following technical effects are achieved:
[0096] The master server in the distributed container cluster obtains the working status of the first slave server in the distributed container cluster from redis, and enters the working process of the first slave server; wherein, the working process of the first slave server includes: when the working status of the first slave server is online, the master server queries the data point i of the last consumption of the first slave server from Redis, where i is a positive integer greater than or equal to 2; the master server broadcasts a message to notify the first slave server to start the task; after receiving the message of starting the task, the first slave server will start consuming messages from data point i; after the first slave server obtains the data from the data source, it will record the data point i; the first slave server pushes the data to the Socket target end using the Socket channel; the first slave server determines whether the data push is normal according to the push result; if the first slave If the server data push is normal, the data point i will be recorded in redis, and new data will be obtained from the data source for the next data push; if the data push of the first slave server is abnormal, the message of the first slave server working abnormality will be broadcast to redis; when the master server receives the message of the first slave server working abnormality, it will continue to try to reconnect to the Socket target end; when the master server successfully connects to the Socket target end, it will enter the next slave server workflow again; based on this, when the push service of the first slave server is abnormal, the master server can take over the push service from the first slave server to restart, and the push service abnormality can be autonomously recovered. After completing the push service, the master service enters the next slave server workflow, and the push service can be quickly restored without affecting other businesses.
[0097] Figure 4 The schematic diagram of triggering, ending and abnormal recovery of data push work according to an embodiment of the present disclosure is shown. The specific steps in the above-mentioned distributed container cluster data push method based on Socket protocol are as follows: Figure 4 It is also shown in .
[0098] In some embodiments, in the above-mentioned distributed container cluster data push method based on Socket protocol, it also includes:
[0099] When the working status of the first slave server is offline, the master server creates a second slave server to replace the first slave server;
[0100] When the second slave server is created successfully, the master server queries Redis for the data point i of the second slave server's last consumption;
[0101] The master server broadcasts a message to notify the second slave server to start the task;
[0102] After receiving the message of starting the task from the server, the second slave will start consuming messages from data point i. After obtaining the data from the data source, the second slave will record the data point i.
[0103] Second, the slave server pushes the data to the Socket target using the Socket channel;
[0104] Second, the slave server determines whether the data push is normal based on the push result;
[0105] If the data push from the second slave server is normal, the data point i is recorded in redis, and new data is obtained from the data source for the next data push;
[0106] If the data push of the second slave server is abnormal, a message about the abnormal operation of the second slave server is sent to redis;
[0107] When the master server receives the message that the second slave server is working abnormally, it will continue to try to reconnect to the Socket target end;
[0108] When the master server successfully connects to the Socket target end, it will enter the next slave server workflow again.
[0109] like Figure 3 and Figure 4 As shown, when an exception occurs in the data push of the first slave server, that is, the working status of the first slave server is offline, it is necessary to recreate a new slave server - the second slave server to replace the first slave server, that is, node B replaces node A to quickly restore data push.
[0110] In some embodiments, the master server can create a second slave server to replace the first slave server by calling the API of the Kubernetes container cluster.
[0111] For the sake of brevity, the specific steps of the second slave server performing data push can refer to the specific steps of the first slave server performing data push, which will not be repeated here.
[0112] According to an embodiment of the present disclosure, when the first slave server works abnormally, a new slave server is created by the master server to replace the abnormal slave server for data push, thereby achieving the effect of quickly restoring data push.
[0113] It should be noted that in the case of pushing massive data through the Socket protocol connection, a new slave server is created through the master server to replace the abnormal slave server for data push. This can also achieve the effect of quickly recovering data without loss when the data push service sends an exception.
[0114] In some embodiments, the master server in the distributed container cluster obtains the working status of the first slave server in the distributed container cluster from redis, including:
[0115] The master server in the distributed container cluster monitors the status of data push tasks;
[0116] When the master server receives the start task instruction, the master server obtains the working status of the first slave server in the distributed container cluster from redis.
[0117] like Figure 4 As shown in the figure, the master server in the distributed container cluster continuously monitors the status of the data push task. When the master server receives the start task instruction, that is, it needs to start the task, then it enters the start slave server workflow, and the master server obtains the working status of the first slave server in the distributed container cluster from redis.
[0118] If the master server receives the start task instruction, that is, it does not need to start the task, the master server will start broadcasting the message instruction of the slave server to stop the data push task, so that redis broadcasts the message of the slave server stopping the data push task. After the slave server receives the message of the slave server stopping the data push task from redis, it subscribes to the stop task message and stops the task, and the process ends.
[0119] According to an embodiment of the present disclosure, starting a data push task based on a task start instruction can facilitate users to choose whether to push data, thereby providing users with more choices.
[0120] In some embodiments, after the master server queries Redis for the data point i last consumed by the first slave server, the method further includes:
[0121] The first slave server obtains data from the data source starting from data point i-1.
[0122] In the face of massive data migration, in order to prevent data push loss, the first slave server can be set to obtain data from the data source starting from data point i-1.
[0123] According to an embodiment of the present disclosure, by setting the first slave server to obtain data from the data source starting from data point i-1, the problem of data loss can be avoided, thereby further monitoring push service anomalies and improving the efficiency of quickly restoring push services.
[0124] In some embodiments, before the master server in the distributed container cluster obtains the working status of the first slave server in the distributed container cluster from redis, the method further includes:
[0125] All servers in the distributed container cluster trigger the workflow for determining the normal status of the primary server at the same time; the workflow for determining the normal status of the primary server includes:
[0126] All servers try to obtain the status information of the primary server in Redis;
[0127] If the status information of the master server is obtained in Redis, it is determined that the master server in the distributed container cluster is in a normal state.
[0128] like Figure 5 As shown in the figure, all servers in the distributed container cluster have the same scheduled task. All servers trigger the workflow of judging the normal status of the master server at the same time and enter the master node election trigger process. The server tries to obtain the status information of the master server in Redis.
[0129] If the status information of the primary server is not obtained, it means that there is no primary server, and the primary server election process, that is, the primary node election process, will be entered. Otherwise, the task of judging the health status of the primary server is completed.
[0130] According to the embodiments of the present disclosure, by adopting multiple nodes to perform a master node election mechanism, the effect of data push can be improved, and multi-node expansion can be supported, thereby further monitoring push service anomalies and improving the efficiency of quickly restoring push services.
[0131] In some embodiments, before the master server in the distributed container cluster obtains the working status of the first slave server in the distributed container cluster from redis, the method further includes:
[0132] If the status information of the primary server is not obtained in Redis, the workflow of primary server election is triggered;
[0133] The workflow of primary server election includes:
[0134] All servers in the distributed container cluster try to store the same key in Redis at the same time, and determine the first server that stores the key in Redis as the master server, and the other servers as slave servers.
[0135] like Figure 6 As shown in the figure, all servers enter the master node election process at the same time, and the servers try to store the same key in Redis. Since Redis is a single-threaded working mechanism, only one instruction can be executed at the same time, so the first node to store the key in Redis is the master node, that is, the master server, and other nodes that cannot store the key in Redis are slave nodes, that is, slave servers.
[0136] According to the embodiments of the present disclosure, based on the single-threaded working mechanism of redis, the election of the master server is performed so that under the master server, data push of multiple slave servers can start or stop working synchronously, thereby further monitoring the abnormalities of the push service and improving the efficiency of quickly recovering the push service.
[0137] In some embodiments, before the master server in the distributed container cluster obtains the working status of the first slave server in the distributed container cluster from redis, the method further includes:
[0138] The slave server reports the server working status to Redis so that Redis can record the online status of the slave server.
[0139] like Figure 4 and Figure 6 As shown, based on Redis recording the online status of the slave servers, the online status of all slave servers can be obtained from Redis to achieve the effect of actively monitoring abnormalities of large-scale cluster nodes, that is, offline nodes.
[0140] In some embodiments, before the master server in the distributed container cluster obtains the working status of the first slave server in the distributed container cluster from redis, the method further includes:
[0141] The slave server obtains the status information of other servers from Redis.
[0142] like Figure 4 and Figure 6 As shown, the servers store their respective election information in redis, and other servers can obtain all server role information and status information from redis.
[0143] According to an embodiment of the present disclosure, the slave server obtains the status information of other servers from Redis so that the slave server can determine the master server and the master server obtains the status information of other slave servers from Redis, thereby further monitoring the push service anomalies and improving the efficiency of quickly recovering the push service.
[0144] In some embodiments, the data source is a Kafka message queue.
[0145] like Figure 4 As shown, the data source used is the Kafka message queue, which can specify the consumption of data from a specified point, which is equivalent to each piece of data having a unique location sequence ID.
[0146] According to the embodiments of the present disclosure, by setting the data source as the Kafka message queue, it is convenient to locate the data point according to the data position sequence ID, so as to further monitor the push service anomalies and improve the efficiency of quickly restoring the push service.
[0147] As can be seen from the above, in the face of massive data migration, the multi-node Socket long connection transmission method is mostly used. When individual or a large number of nodes are abnormal, manual intervention is required to repair the data push task, and the problem of data loss cannot be quickly solved, the method of electing the master node, the master node controlling the slave node, and the slave node recording the message data point can be relied on to avoid the problem of node abnormality or node downtime and unable to recover quickly and data loss. And when an abnormality occurs on the Socket target end, all nodes will continue to reconnect, which may easily cause the Socket target end to be unable to quickly restore service. The solution of only the master node trying to reconnect and notifying other nodes after the reconnection is successful can well solve the impact of the above problems and save the hardware pressure of the local cluster environment.
[0148] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited by the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present disclosure.
[0149] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described module can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0150] In the technical solution disclosed herein, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0151] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0152] Figure 7 A schematic block diagram of an electronic device 700 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0153] The electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a ROM 702 or a computer program loaded from a storage unit 708 into a RAM 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An I / O interface 705 is also connected to the bus 704.
[0154] Multiple components in the electronic device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the electronic device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0155] The computing unit 701 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above, such as method 300. For example, in some embodiments, the method 300 may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the method 300 described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform the method 300 in any other appropriate manner (e.g., by means of firmware).
[0156] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0157] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0158] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0159] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0160] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0161] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0162] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0163] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A distributed container cluster data push method based on Socket protocol, It is characterized in that include: The master server in the distributed container cluster obtains the working status of the first slave server in the distributed container cluster from redis and enters the working process of the first slave server; The first slave server workflow includes: When the working status of the first slave server is online, the master server queries the data point i of the last consumption of the first slave server from Redis, where i is a positive integer greater than or equal to 2; The master server broadcasts a message to notify the first slave server to start the task; After receiving the message of starting the task from the server, it will start consuming messages from data point i; First, after the server obtains data from the data source, it will record the data point i; First, the slave server pushes the data to the Socket target using the Socket channel; First, the server determines whether the data push is normal based on the push result; If the first slave server data push is normal, the data point i is recorded in redis, and new data is obtained from the data source for the next data push; If the data push of the first slave server is abnormal, a message that the first slave server is working abnormally is sent to redis; When the master server receives the message that the first slave server is working abnormally, it will continue to try to reconnect to the Socket target end; When the master server successfully connects to the Socket target end, it will enter the next slave server workflow again.
2. The method according to claim 1, It is characterized in that The method further comprises: When the working status of the first slave server is offline, the master server creates a second slave server to replace the first slave server; When the second slave server is created successfully, the master server queries Redis for the data point i of the second slave server's last consumption; The master server broadcasts a message to notify the second slave server to start the task; After receiving the message of starting the task from the server, the second slave will start consuming messages from data point i. After obtaining the data from the data source, the second slave will record the data point i. Second, the slave server pushes the data to the Socket target using the Socket channel; Second, the slave server determines whether the data push is normal based on the push result; If the data push from the second slave server is normal, the data point i is recorded in redis, and new data is obtained from the data source for the next data push; If the data push of the second slave server is abnormal, a message about the abnormal operation of the second slave server is sent to redis; When the master server receives the message that the second slave server is working abnormally, it will continue to try to reconnect to the Socket target end; When the master server successfully connects to the Socket target end, it will enter the next slave server workflow again.
3. The method according to claim 1, It is characterized in that The master server in the distributed container cluster obtains the working status of the first slave server in the distributed container cluster from redis, including: The master server in the distributed container cluster monitors the status of data push tasks; When the master server receives the start task instruction, the master server obtains the working status of the first slave server in the distributed container cluster from redis.
4. The method according to claim 1, It is characterized in that After the master server queries Redis for the data point i last consumed by the first slave server, the method further includes: The first slave server obtains data from the data source starting from data point i-1.
5. The method according to claim 1, It is characterized in that Before the master server in the distributed container cluster obtains the working status of the first slave server in the distributed container cluster from redis, the method further includes: All servers in the distributed container cluster trigger a workflow for determining the normal state of the primary server at the same time; the workflow for determining the normal state of the primary server includes: All servers try to obtain the status information of the primary server in Redis; If the status information of the master server is obtained in Redis, it is determined that the master server in the distributed container cluster is in a normal state.
6. The method according to claim 5, It is characterized in that Before the master server in the distributed container cluster obtains the working status of the first slave server in the distributed container cluster from redis, the method further includes: If the status information of the primary server is not obtained in Redis, the workflow of primary server election is triggered; The workflow of the master server election includes: All servers in the distributed container cluster try to store the same key in Redis at the same time, and determine the first server that stores the key in Redis as the master server, and the other servers as slave servers.
7. The method according to claim 6, It is characterized in that Before the master server in the distributed container cluster obtains the working status of the first slave server in the distributed container cluster from redis, the method further includes: The slave server reports the server working status to Redis so that Redis can record the online status of the slave server.
8. The method according to claim 7, It is characterized in that Before the master server in the distributed container cluster obtains the working status of the first slave server in the distributed container cluster from redis, the method further includes: The slave server obtains the status information of other servers from Redis.
9. The method according to claim 1, It is characterized in that The data source is the Kafka message queue.
10. A distributed container cluster data push system based on Socket protocol, It is characterized in that It includes the master server, multiple slave servers and redis; among them, The master server in the distributed container cluster obtains the working status of the first slave server in the distributed container cluster from redis and enters the first slave server workflow; the first slave server workflow includes: When the working status of the first slave server is normal, the master server queries the data point i of the last consumption of the first slave server from Redis, where i is a positive integer greater than or equal to 2; The master server broadcasts a message to notify the first slave server to start the task; After the first slave server receives the message to start the task, it will start consuming messages from data point i. After the first slave server obtains data from the data source, it will record data point i. First, the slave server pushes the data to the Socket target using the Socket channel; First, the server determines whether the data push is normal based on the push result; If the first slave server data push is normal, the data point i is recorded in redis, and new data is obtained from the data source for the next data push; If the data push of the first slave server is abnormal, a message that the first slave server is abnormal will be broadcast to redis; When the master server receives the message that the first slave server is working abnormally, it will continue to try to reconnect to the Socket target end; When the master server successfully connects to the Socket target end, it will enter the next slave server workflow again.
Citation Information
Patent Citations
Method and system for achieving high availability of Redis cluster
CN105933407A
Distributed machine learning parameter adjusting system based on celery
CN110110863A