A testing method, device and equipment for a distributed timed message system
By introducing a fault injection mechanism in the distributed timing message system, simulating various fault conditions, testing the delivery process of timing messages and the stability of the system, the robustness problem of difficult to reliably test the distributed timing message system in the prior art is solved, and efficient system stability testing is achieved.
Patent Information
- Application Number
- CN202211038337.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-29
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-08-29
AI Technical Summary
The prior art is difficult to reliably test the robustness of a distributed timing message system, resulting in risks such as loss of timing messages, resending and delivery time not meeting expectations.
By introducing a fault injection mechanism between the message subscription client and the timed message server, various fault conditions are simulated and the delivery process of the timed message and the stability of the system are tested.
It realizes robustness testing of distributed timing message systems, and can conduct daily-level sustainable operation abnormality testing at low cost, reduce manual operation costs and improve system stability.
Smart Images

Figure CN115604164B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of testing technologies, and in particular, to a testing method, apparatus, and device for a distributed timed message system. Background Art
[0002] In a distributed timed message system, multiple nodes are responsible for delivering timed messages. They usually cooperate according to a certain strategy to improve efficiency and take disaster tolerance into account at the same time.
[0003] Since timed messages are not sent immediately but after a certain delay time, the uncertainty of the entire system is increased. During the delay time and in the process of delivering timed messages after the delay time expires, node state transitions may occur within the system, leading to various risks of distributed inconsistency. The intuitive manifestations of these risks include loss of timed messages, retransmission, and delivery time not meeting expectations, etc.
[0004] Based on this, a testing scheme that can reliably and low-costly test the robustness of a distributed timed message system is needed to guide the improvement of the system and reduce these risks. Summary of the Invention
[0005] One or more embodiments of this specification provide a testing method, apparatus, device, and storage medium for a distributed timed message system to solve the following technical problem: A testing scheme that can more reliably test the robustness of a distributed timed message system is needed to guide the improvement of the system and reduce these risks.
[0006] To solve the above technical problem, one or more embodiments of this specification are implemented as follows:
[0007] A testing method for a distributed timed message system provided by one or more embodiments of this specification, the system includes a message publishing client, a message subscribing client, and a timed message server. The method includes:
[0008] Batch initiate subscriptions to the timed messages that can be published by the message publishing client through the message subscribing client;
[0009] Construct a fault injection instruction according to the delay time corresponding to the timed message and send it to the timed message server to inject a specified type of fault into the timed message server;
[0010] Publish each of the timed messages subscribed by the message subscribing client to the timed message server through the message publishing client, so that the timed message server delivers the timed messages to the message subscribing client;
[0011] According to the subscription, verify the situation of the message subscription client receiving the timed message, and determine the test result of the system according to the verification result.
[0012] A distributed timed message system testing device provided by one or more embodiments of this specification. The system includes a message publishing client, a message subscription client, and a timed message server. The device includes:
[0013] A message subscription module, which batch initiates subscriptions to the timed messages that can be published by the message publishing client through the message subscription client;
[0014] A fault injection module, which constructs a fault injection instruction according to the delay time corresponding to the timed message and sends it to the timed message server to inject a specified type of fault into the timed message server;
[0015] A publishing and delivery module, which publishes each of the timed messages subscribed by the message subscription client to the timed message server through the message publishing client, so that the timed message server delivers the timed message to the message subscription client;
[0016] A result determination module, which verifies the situation of the message subscription client receiving the timed message according to the subscription, and determines the test result of the system according to the verification result.
[0017] A distributed timed message system testing device provided by one or more embodiments of this specification. The system includes a message publishing client, a message subscription client, and a timed message server. The device includes:
[0018] At least one processor; and,
[0019] A memory communicatively connected to the at least one processor; wherein,
[0020] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can:
[0021] Batch initiate subscriptions to the timed messages that can be published by the message publishing client through the message subscription client;
[0022] Construct a fault injection instruction according to the delay time corresponding to the timed message and send it to the timed message server to inject a specified type of fault into the timed message server;
[0023] Through the message publishing client, publish each of the timed messages subscribed by the message subscribing client to the timed message server, so that the timed message server delivers the timed messages to the message subscribing client;
[0024] According to the subscription, verify the situation of the message subscribing client receiving the timed messages, and determine the test result of the system according to the verification result.
[0025] A non-volatile computer storage medium provided by one or more embodiments of this specification, the medium stores computer-executable instructions, and the computer-executable instructions are set to:
[0026] Through the message subscribing client, batch initiate subscriptions to the timed messages that the message publishing client can publish;
[0027] According to the delay time corresponding to the timed message, construct a fault injection instruction and send it to the timed message server to inject a specified type of fault into the timed message server;
[0028] Through the message publishing client, publish each of the timed messages subscribed by the message subscribing client to the timed message server, so that the timed message server delivers the timed messages to the message subscribing client;
[0029] According to the subscription, verify the situation of the message subscribing client receiving the timed messages, and determine the test result of the system according to the verification result.
[0030] The above at least one technical solution adopted by one or more embodiments of this specification can achieve the following beneficial effects: For the use of timed messages based on a distributed cluster, an architecture with the cooperation of a message publishing client, a message subscribing client, and a timed message server is provided. The message publishing client publishes timed messages to the timed message server, and the timed message server delivers the timed messages to the message subscribing client. Active optional type of fault injection is performed on the timed message server, so that the delivery process of the timed messages is affected by the server fault. Furthermore, the reception situation (such as integrity, real-time, etc.) of a large number of timed messages by the message subscribing client under this influence is used as the basis for judging the system's steady state. It has good directivity and reliability, can realize daily-level sustainable operation anomaly testing and reduce manual operation costs, and better supports the test scenarios of the robustness of the distributed timed message system. Description of the Drawings
[0031] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0032] Figure 1 Schematic flowchart of a method for testing a distributed timed message system provided by one or more embodiments of this specification;
[0033] Figure 2 Schematic diagram of a partial architecture of a distributed timed message system provided by one or more embodiments of this specification;
[0034] Figure 3 Provided by one or more embodiments of this specification Figure 2 Schematic diagram of the scenario of partition takeover in the system in
[0035] Figure 4 Provided by one or more embodiments of this specification Figure 1 Schematic diagram of the principle of a specific implementation of the method in
[0036] Figure 5 Schematic diagram of the structure of a device for testing a distributed timed message system provided by one or more embodiments of this specification;
[0037] Figure 6 Schematic diagram of the structure of a device for testing a distributed timed message system provided by one or more embodiments of this specification. Detailed implementation
[0038] The embodiments of this specification provide a method, device, equipment, and storage medium for testing a distributed timed message system.
[0039] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0040] This application proposes a test solution for the robustness of a distributed timed message system with good applicability, and attempts to test and use it on product systems such as middleware timed message products. Among them, timed messages refer to messages that can only be consumed by consumers after a specified timestamp, and are used to solve some scenarios where there are time window requirements for message production and consumption, or scenarios where timed tasks are triggered by messages. Timed messages are often used to trigger the timed operation of specified business functions. For example, page refresh, transaction retry, payment reminder, running risk control strategies, collecting order data, reporting monitoring data, etc.; Robustness refers to whether software can still complete its work normally without crashing or freezing in the face of input errors, disk failures, network overloads, or intentional attacks.
[0041] On the premise of taking the integrity and real-time performance of timed messages as the test criteria, this solution constructs server-side fault injection by sending instructions, so as to realize the integration of daily-level stability testing and continuously discover problems related to product robustness. The following will be described in detail based on this idea.
[0042] Figure 1 It is a schematic flowchart of a test method for a distributed timed message system provided by one or more embodiments of this specification. This method can be applied to different business fields, such as: electronic payment business field, e-commerce business field, social business field, game business field, official business field, etc. This process can be executed on distributed timed message system devices in these fields and / or test devices connected to these systems. For example, a test program is pre-installed on these devices, and by running the test program, corresponding control instructions are sent to one end or multiple ends, and the instruction receiving end responds to the control instructions and specifically executes corresponding steps. Some input parameters or intermediate results in the process allow manual intervention and adjustment to help improve accuracy.
[0043] Figure 1 The process in
[0044] S102: Through the message subscription client, batch initiate subscriptions to the timed messages that the message publishing client can publish.
[0045] The distributed timed message system includes a message publishing client, a message subscription client, and a timed message server (which can be abbreviated as the server), and there can be multiple of each end, and they are pre-distributed.
[0046] In one or more embodiments of the present specification, the system at least includes a distributed cluster composed of multiple scheduled message service terminals, each scheduled message service terminal is a node in the distributed cluster, and the message publishing client and the message subscription client can also serve as nodes in the distributed cluster, or be outside the distributed cluster. Scheduled message services are provided to clients through the distributed cluster. The message subscription client subscribes to the scheduled message, and the message publishing client or the scheduled message service terminal itself publishes the scheduled message accordingly. The published scheduled message is delivered to the message subscription client by the scheduled message service terminal in a timely manner. During the test process, for example, the test script triggers the message subscription client to subscribe, triggers the message publishing client to publish, and then tests whether the delivery effect of the scheduled message service terminal is abnormal.
[0047] S104: constructing a fault injection instruction according to the delay time corresponding to the scheduled message and sending the instruction to the scheduled message server, so as to inject a specified type of fault into the scheduled message server.
[0048] In one or more embodiments of this specification, the background technology mentions that node state changes may occur in the system, which may lead to risks. This risk is mainly reflected in the impact on the delivery effect of the scheduled message server. In order to achieve daily-level fault reproduction and stability testing, the required faults are actively injected into the scheduled message server by constructing fault injection instructions. This is convenient for testing at any time on demand and can restore normal business in a timely manner.
[0049] The delay time corresponding to a scheduled message includes at least the scheduled duration. For example, if a scheduled duration of 30 seconds is set for a scheduled message, the scheduled message should be sent (for example, delivered, or published and delivered, etc.) or received 30 seconds after the scheduled actual time point. In this case, the delay time corresponding to the scheduled message can be 30 seconds. In addition, considering that sending and receiving messages also takes a small amount of time (for example, a duration of milliseconds), a little delay can be added to the scheduled duration, and the whole can be used as the delay time corresponding to the scheduled message.
[0050] This solution can also be used for testing in the scenario of instant messaging. However, since the message is sent instantly, and the stability and predictability of the server in this scenario are relatively clear, the test time window is narrower and the test effect is more limited, which is not conducive to fully and comprehensively simulating and mining various complex risk environments in real situations. Based on this, the solution of this application is mainly aimed at the scenario of scheduled messages. Through the coordination of a large number of scheduled messages with different delay times, as well as server-side fault injection of multiple optional types and optional strategies, various complex test scenarios can be constructed more realistically, which helps to achieve more efficient and effective testing results at a low cost.
[0051] In one or more embodiments of this specification, according to the delay time corresponding to the timing message, a fault injection instruction with a fault injection time not less than this delay time is constructed to perform server-side fault injection, where the fault injection time includes at least part of the effective duration of the injected fault. In this way, it is beneficial to make the server in the desired fault state at the key time points such as sending or receiving after the timing message is delayed. Therefore, the anomalies that may occur in the situation of receiving the timing message are more likely to be caused by this fault, which helps to clarify the correlation relationship and degree between the test influencing factors and the test results during the test process.
[0052] In one or more embodiments of this specification, during the test process, a large number of timing messages are involved at the same time, and the setting strategies for their delay times are diverse. For example, a consistent delay time is set, a stepped delay time is set, etc. Similarly, the servers participating in message delivery and message subscription clients can also be numerous, thereby constructing a more complex and realistic test environment and more accurately testing the robustness of the server.
[0053] S106: Through the message publishing client, publish each of the timing messages subscribed by the message subscription client to the timing message server, so that the timing message server delivers the timing message to the message subscription client.
[0054] In one or more embodiments of this specification, one or more message subscription clients subscribe to timing messages from a timing message server. The timing message is published by a message publishing client to the server, and the server delivers the published timing message to each message subscription client that has subscribed to this timing message. The injection of faults may affect the delivery action of the server, and thus affect the reception situation of the message subscription client.
[0055] In one or more embodiments of this specification, before or during the delivery of the timing message, the timing message server is in a state affected by the injected fault. Therefore, it is possible that this timing message server is taken over by another timing message server to continue the delivery, and a node state transition will occur during the process of taking over the service. This solution particularly focuses on the robustness performance of the server in this situation. It should be noted that this other timing message server can also be injected with faults. This solution not only focuses on the robustness of a single timing message server, but more on the robustness of the entire distributed cluster.
[0056] S108: According to the subscription, verify the situation of the message subscription client receiving the timing message, and determine the test result of the system according to the verification result.
[0057] In one or more embodiments of this specification, a subscription reflects the expectation for the timed messages to be received. However, due to the impact of faults, the situation where the message subscription client receives timed messages may not conform to this expectation. The degree of non - conformity with this expectation is determined through verification. For example, according to the subscription, the integrity and timeliness of the timed messages received by the message subscription client through the delivery of the server are verified to determine the impact of the specified type of injected faults on the delivery, thereby determining the robustness performance of the server. In addition, it is also possible to verify the interaction actions and state transition situations of one local server or multiple previous servers related to this timed message to see if they conform to the predetermined logic.
[0058] In one or more embodiments of this specification, the above - mentioned test process can be arranged as a daily - level test case and executed repeatedly to obtain more reliable global test results.
[0059] Through Figure 1 the method, for the use of timed messages based on a distributed cluster, an architecture with the cooperation of a message publishing client, a message subscription client, and a timed message server is provided. The message publishing client publishes timed messages to the timed message server, and the timed message server delivers the timed messages to the message subscription client. Active optional - type fault injection is performed on the timed message server so that the delivery process of the timed messages is affected by server faults. Furthermore, the reception situation of a large number of timed messages by the message subscription client under this influence (such as integrity, timeliness, etc.) is used as the basis for judging the system's steady state. It has good directivity and reliability, can achieve daily - level sustainable operation abnormal testing and reduce manual operation costs, and better supports the test scenario of the robustness of the distributed timed message system.
[0060] Based on Figure 1 the method, this specification also provides some specific implementation schemes and extended schemes of this method, which will be further described below.
[0061] More intuitively, in combination with an exemplary distributed timed message system, it will be further described. Refer to Figure 2 . Figure 2 It is a partial architecture schematic diagram of a distributed timed message system provided in one or more embodiments of this specification.
[0062] In Figure 2Among them, the producer is the above-mentioned message publishing client, the consumer is the above-mentioned message subscribing client, there are multiple timed message servers, such as server A, server B, and server C shown in the partition table, and there is also a partition coordinator responsible for coordinating each timed message server. This system can logically divide all messages (including the timed messages to be delivered) with the timed message server as the dimension, obtain multiple partitions, and allocate them to the corresponding timed message servers. The corresponding timed message server provides services such as delivery for the messages belonging to this partition. For example, the partitions corresponding to the current server A are P1 and P2, the partitions corresponding to server B are P3 and P4, and the partitions corresponding to server B are P5 and P6. The partition coordinator synchronizes the partition configuration information among each timed message server and allocates partitions.
[0063] When working properly, the producer produces messages and inputs them into the corresponding timed message server. It can also send timed instructions to the server to make the message a timed message, or send instructions to cancel the timing. On the timed message server, filters can be set (when there are multiple filters, they can form a filter chain) to filter the input messages, and filter out the messages that meet the rules and store them in the router. In this system, a wheel-shaped data structure called a time wheel is exemplarily used to store timed messages. The time wheel is divided into multiple grids. According to the timed duration of the timed message, it is determined which grid it should be stored in. When the timed duration is relatively long (longer than the duration represented by one week of the time wheel), it is necessary to determine which grid the timed message should be stored in by adding the number of rounds and the number of grids, and wait until the number of rounds is reached and the grid is reached to send out the timed message. The storage area can be divided into a short-term storage area and a long-term storage area, and can be selected for use according to actual needs. The timed message server will also time through the time wheel trigger, and when the timed duration expires, it will trigger a query to the storage for the corresponding timed message, and through the delivery router, deliver this message as an output message to the corresponding consumer.
[0064] In practical applications, when the number of timed message servers in the distributed cluster changes (such as machine scaling, replacement, downtime, etc.), in order to ensure that the newly launched server can start providing services with the partition configuration information, or the partitions of the offline server are taken over by other servers, the partition reassignment of the cluster will be triggered. See Figure 3 , Figure 3 provided by one or more embodiments of this specification Figure 2 for the scenario schematic diagram of partition takeover in the system in
[0065] In Figure 3In the example, initially, partition P3 is managed by server A, and server B is the backup responsible for P3. When server A is temporarily unable to be responsible for P3 due to reasons such as failure or active shutdown, server B will take over P3. Server A shuts down P3 by shutting down the timer of P3 and updating the detection point accordingly. Server B starts the invalid message compensation task by establishing the timer of P3 and reading the partition detection point to restore P3. For example, before shutting down P3, the index reaches the time point 1007. When taking over, for P3, messages corresponding to the time points from 1007 can be obtained. For example, messages corresponding to the time points from 1007 to 1009 are obtained, and then the current index is updated to 1010, and server B continues to serve P3. In this process, there will be a partition state transition stage, which will lead to the risk of distributed inconsistency.
[0066] Similarly, after injecting a specified type of fault into the scheduled message server, the injection of the specified type of fault can trigger in the distributed cluster: at least one scheduled message server newly launches a service responsible for delivering scheduled messages, or at least one scheduled message server takes over the service of delivering the scheduled messages from another scheduled message server. For the entire system, in response to the specified type of fault injected into the server, it may enter the partition state transition stage. In the partition state transition stage, at least part of the partitions are reallocated or switched to disaster recovery state (for example, different servers take over the service to achieve multi-backup disaster recovery).
[0067] In one or more embodiments of the present specification, when the test process is actually executed, the fault injection time period may be relatively long, and during this period, the more critical partition state transition stage may account for a small part of it. The delay time corresponding to the scheduled message is pre-set, and it may not hit the partition state transition stage. This may reduce the probability of abnormal situations, which is not conducive to discovering problems through testing. To address this problem, this solution adopts an adaptively adjusted delay time, so that the scheduled message essentially becomes an adaptive dynamically timed scheduled message. For example, when the scheduled message server waits and executes the step of delivering the scheduled message, it can determine whether the delay time corresponding to the scheduled message matches the partition state transition stage. If not, the delay time is adjusted accordingly to force the scheduled message to be delivered in the partition state transition stage, thereby helping to cause anomalies with a higher probability and improve testing benefits.
[0068] According to the above description, one or more embodiments of this specification also provide Figure 1 A schematic diagram of a specific implementation scheme of the method is shown in FIG. Figure 4 As shown. Combined Figure 4 , the specific implementation scheme is described, and the exemplary steps include the following:
[0069] Environment setup: In a stable test environment, set up one or more fixed timed message publishing clients and timed message subscribing clients. The timed message server starts a specified instruction receiver to expose an HTTP service externally for receiving fault injection.
[0070] Data preparation: By cleaning up residual data, ensure that there are no residual messages in each client, and batch initiate subscriptions and corresponding publications of timed messages. For example, initiate 100 subscriptions of timed messages with a 30 - second delay (the magnitude can be adjusted arbitrarily). After triggering, check that the initialization of these 100 subscriptions is successful and that the subscribed messages have not been received yet.
[0071] Send fault injection instructions: Asynchronously initiate fault injection after data preparation for a time not less than the delay time of the timed message. The injected faults include at least one of the following types of atomic faults: system crash, CPU overheating, disk full, IO exception, memory overheating, network packet delay, duplication and loss, JVM method - level exception, etc. At the same time, the timed message server asynchronously delivers timed messages based on the data published by the message publishing client. After receiving the messages delivered by the server, the message subscribing client consumes them. Through this step, a scenario for the distributed timed message system to handle timed messages under abnormal conditions is constructed.
[0072] Result verification: At the expected time, for example, after 30 seconds, verify that the message subscribing end has received the above 100 batch timed messages. Based on this, judge that the integrity and real - time performance of the timed messages have not been affected by server exceptions. It is possible to check whether the timed messages are consumed successfully as expected, thereby verifying the robustness of the system. After this test, the residual data can be cleaned up in a timely manner for subsequent tests.
[0073] Generate a baseline: Orchestrate the test process and scenarios corresponding to the test results into daily test tasks, trigger and run them multiple times, and generate a robustness test baseline based on the results of one or multiple runs. If the test results are abnormal, adjust them accordingly and then generate a baseline.
[0074] In one or more embodiments of this specification, multiple timed message servers in the distributed cluster participate in the test process. Therefore, it is possible for the timed message servers to switch and take over services. In order to expose as many abnormalities as possible in a single test, a chained fault - tracking injection scheme is adopted, so that the fault injection can be transferred in a timely manner accordingly as the service is transferred. In this way, it is also possible to avoid wasting resources by injecting faults into timed message servers on a large scale in advance.
[0075] For example, after injecting a specified type of fault into a scheduled message server, the scheduled message server into which the fault is injected is used as the first server, and it is determined that in response to the fault, a second server among multiple scheduled message servers will take over the service of delivering scheduled messages from the first server. The specified type of fault is injected into the second server through the first server, and can be injected during the takeover switching process, which helps to avoid increasing the instruction interaction overhead.
[0076] Furthermore, based on this idea, a gradually weakening fault injection is designed to extend the service takeover chain (composed of multiple servers that take over the service in turn). Using the above example to illustrate, the service interference effect corresponding to the fault injected into the second server can be made lower than the service interference effect corresponding to the fault injected into the first server (for example, the service interference effect corresponding to a CPU surge is usually lower than the service interference effect corresponding to a crash), because the higher the service interference effect of a fault, the more likely it is that the server will switch to take over, which may lead to multiple relay takeovers for the same scheduled message. Similarly, if the second server is unable to deliver the message due to a fault, the third server may continue to take over the service, and the second server will inject a fault that further reduces the service interference effect into the third server. In addition to the relay fault injection method, a unified end can also be responsible for the injection of all faults, and the corresponding accuracy and flexibility may be affected. In addition to gradually reducing the service interference effect, the faults injected into the service takeover chain can also be at least partially differentiated in type (for example, the first injection is a downtime failure, the second injection is a CPU surge, etc.), thereby increasing the complexity of the scenario and helping to more efficiently discover system anomalies.
[0077] In one or more embodiments of the present specification, some types of faults have instant volatility in actual applications, such as high CPU and high memory. Most of the time, since the service interference effect does not reach the critical value, it may not cause the server to switch to take over, and the fault state will continue for a period of time. In this case, it is considered to simulate not only the peak state of the fault but also the trough state of the fault during the test, because the trough state may be the key to the system to maintain robustness, that is, although the system may be abnormal in the peak state of the fault, it may also survive the peak state and successfully complete the business in the trough state. Of course, in actual applications, the trough state may be fleeting. This solution considers simulating this trough state by inserting a safe time slot during the test process to give the same server a chance to turn around, rather than steadily suppressing it with the fault, so as to more objectively evaluate the robustness of the system under the state of drastic fluctuations in actual business.
[0078] Specifically, for example, after injecting a specified type of fault into the timed message server, it can be determined whether, by injecting the specified type of fault, at least one timed message server has been triggered to take over the service of delivering timed messages from another timed message server. If not, a safe time slot is inserted during the fault injection effective period. During the safe time slot, the injected fault does not take effect, and the length of the safe time slot is less than or even much less than the injection effective period, which helps to accurately evaluate the sensitivity and response speed of the system robustness. If the server can seize the safe time slot to successfully deliver the timed message, the sensitivity and response speed of the robustness are relatively high, and the entire system is relatively more reliable.
[0079] Based on the same idea, one or more embodiments of this specification also provide a corresponding apparatus and device for the above method, such as Figure 5 、 Figure 6 shown.
[0080] Figure 5 It is a schematic structural diagram of a distributed timed message system testing apparatus provided by one or more embodiments of this specification. The system includes a message publishing client, a message subscribing client, and a timed message server. The apparatus includes:
[0081] A message subscribing module 502, through the message subscribing client, batch initiates subscriptions to the timed messages that can be published by the message publishing client;
[0082] A fault injection module 504, according to the delay time corresponding to the timed message, constructs a fault injection instruction and sends it to the timed message server to inject a specified type of fault into the timed message server;
[0083] A publishing and delivery module 506, through the message publishing client, publishes each of the timed messages subscribed by the message subscribing client to the timed message server, so that the timed message server delivers the timed message to the message subscribing client;
[0084] A result determination module 508, according to the subscription, verifies the situation of the message subscribing client receiving the timed message, and determines the test result of the system according to the verification result.
[0085] Optionally, the fault injection module 504 constructs a fault injection instruction whose corresponding fault injection time is not less than the delay time according to the delay time corresponding to the timed message.
[0086] Optionally, the system includes a distributed cluster composed of multiple timed message servers;
[0087] After injecting a specified type of fault into the timing message server, the fault injection module 504 triggers, through the injection of the specified type of fault, in the distributed cluster: at least one new timing message server goes online to be responsible for delivering the timing message service, or at least one timing message server takes over the service of delivering the timing message from another timing message server.
[0088] Optionally, it further includes:
[0089] The partition management module 510 logically divides all the timing messages to be delivered with the timing message server as the dimension, obtains multiple partitions, and assigns them to the corresponding timing message servers;
[0090] After injecting a specified type of fault into the timing message server, the partition management module 510 enters the partition state transition phase in response to the injected specified type of fault;
[0091] In the partition state transition phase, at least some of the partitions are reallocated or switched to the disaster tolerance state.
[0092] Optionally, the fault injection module 504 determines whether the delay time corresponding to the timing message matches the partition state transition phase;
[0093] If not, the delay time is adjusted accordingly to force an attempt to deliver the timing message during the partition state transition phase.
[0094] Optionally, after injecting a specified type of fault into the timing message server, the fault injection module 504 uses the timing message server injected with the fault as the first server;
[0095] It is determined that in response to the fault, a second server among the multiple timing message servers will take over the service of delivering the timing message from the first server;
[0096] Through the first server, a specified type of fault is injected into the second server.
[0097] Optionally, the service disruption effect corresponding to the fault injected into the second server is lower than the service disruption effect corresponding to the fault injected into the first server.
[0098] Optionally, after injecting a specified type of fault into the timing message server, the fault injection module 504 determines whether, through the injection of the specified type of fault, at least one timing message server has taken over the service of delivering the timing message from another timing message server;
[0099] If not, a safe time slot is inserted during the injection effective time period of the fault, during which the fault does not take effect, and the length of the safe time slot is less than the injection effective time period.
[0100] Optionally, the fault injection module 504 injects at least one of the following types of atomic faults into the timing message server:
[0101] System crash, CPU overheating, disk full, IO exception, memory overheating, network packet delay, repetition and loss, JVM device-level exception.
[0102] Optionally, the result determination module 508 verifies the integrity and real-time performance of the timing messages received by the timing message server through the delivery according to the subscription, so as to determine that the delivery is affected by the specified type of fault injected.
[0103] Optionally, after determining the test result of the system according to the verification result, the result determination module 508 arranges the test process corresponding to the test result into a daily test task and runs it multiple times;
[0104] Generate a test baseline according to the results of the multiple runs.
[0105] Figure 6 A structural schematic diagram of a distributed timing message system test device provided for one or more embodiments of this specification. The system includes a message publishing client, a message subscribing client, and a timing message server. The device includes:
[0106] At least one processor; and,
[0107] A memory communicatively connected to the at least one processor; wherein,
[0108] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to:
[0109] Batch initiate subscriptions to the timing messages that can be published by the message publishing client through the message subscribing client;
[0110] Construct a fault injection instruction according to the delay time corresponding to the timing message and send it to the timing message server to inject a specified type of fault into the timing message server;
[0111] Publish each of the timing messages subscribed by the message subscribing client to the timing message server through the message publishing client, so that the timing message server delivers the timing messages to the message subscribing client;
[0112] According to the subscription, verify the situation of the message subscription client receiving the timed message, and determine the test result of the system according to the verification result.
[0113] Communication may be carried out between the processor and the memory through a bus, and the device may further include an input / output interface for communicating with other devices.
[0114] Based on the same idea, one or more embodiments of this specification also provide a non-volatile computer storage medium corresponding to the Figure 1 method in, storing computer-executable instructions, and the computer-executable instructions are set as:
[0115] Through the message subscription client, batch initiate subscriptions to the timed messages that the message publishing client can publish;
[0116] According to the delay time corresponding to the timed message, construct a fault injection instruction and send it to the timed message server to inject a specified type of fault into the timed message server;
[0117] Through the message publishing client, publish each of the timed messages subscribed by the message subscription client to the timed message server, so that the timed message server delivers the timed message to the message subscription client;
[0118] According to the subscription, verify the situation of the message subscription client receiving the timed message, and determine the test result of the system according to the verification result.
[0119] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not just one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0120] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to make the controller implement the same function in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.
[0121] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0122] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0123] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, the embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0124] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0125] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0126] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0127] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0128] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0129] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0130] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0131] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0132] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment, and non-volatile computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0133] The specific embodiments of this specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0134] The above description is only for one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various changes and modifications can be made to one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope of the claims of this specification.
Claims
1. A testing method for a distributed timed message system, the system including a message publishing client, a message subscribing client, and a timed message server, the method comprises: Batch initiating subscriptions to the timed messages that can be published by the message publishing client through the message subscribing client; Constructing a fault injection instruction according to the delay time corresponding to the timed message and sending it to the timed message server to inject a specified type of fault into the timed message server, specifically including: constructing a fault injection instruction with a corresponding fault injection time not less than the delay time according to the delay time corresponding to the timed message; Publishing each of the timed messages subscribed by the message subscribing client to the timed message server through the message publishing client, so that the timed message server delivers the timed message to the message subscribing client; Verifying the situation of the message subscribing client receiving the timed message according to the subscription, and determining the test result of the system according to the verification result; The system includes a distributed cluster composed of multiple timed message servers; After injecting the specified type of fault into the timed message server, the method further includes: Triggering in the distributed cluster through the injection of the specified type of fault: at least one timed message server newly goes online to be responsible for delivering the service of the timed message, or at least one timed message server takes over the service of delivering the timed message from another timed message server.
2. The method according to claim 1, further comprises: Logically dividing all the timed messages to be delivered with the timed message server as the dimension to obtain multiple partitions, and assigning them to the corresponding timed message servers; After injecting the specified type of fault into the timed message server, the method further includes: Responding to the injected specified type of fault, entering the partition state transition stage; In the partition state transition stage, at least some of the partitions are reallocated or switched to the disaster recovery state.
3. The method according to claim 2, when the timed message server delivers the timed message to the message subscribing client, further comprises: Judging whether the delay time corresponding to the timed message matches the partition state transition stage; If not, then correspondingly adjusting the delay time to force an attempt to deliver the timed message in the partition state transition stage.
4. The method according to claim 1, after injecting the specified type of fault into the timed message server, the method further comprises: Regarding the timed message server injected with the fault as the first server; Determining that in response to the fault, a second server among the multiple timed message servers is to take over the service of delivering the timed message from the first server; Injecting a specified type of fault into the second server through the first server.
5. The method according to claim 4, the service disruption effect corresponding to the fault injected into the second server is lower than the service disruption effect corresponding to the fault injected into the first server.
6. The method according to claim 1, after injecting a fault of a specified type into the timing message server, the method further includes: determining whether at least one timing message server is triggered to take over the service of delivering the timing message from another timing message server by injecting the fault of the specified type; if not, inserting a safe time slot during the effective period of the injection of the fault, during which the fault is ineffective, and the length of the safe time slot is less than the effective period of the injection.
7. The method according to claim 1, the injecting a fault of a specified type into the timing message server specifically includes: injecting at least one of the following types of atomic faults into the timing message server: shutdown, CPU overheating, disk full, IO exception, memory overheating, network packet delay, repetition and loss, JVM method-level exception.
8. The method according to claim 1, the verifying the situation of the message subscription client receiving the timing message according to the subscription specifically includes: verifying the integrity and timeliness of the timing message received by the message subscription client through the delivery according to the subscription, so as to determine that the delivery is affected by the injected fault of the specified type.
9. The method according to claim 1, after determining the test result of the system according to the verification result, the method further includes: orchestrating the test process corresponding to the test result into a daily test task and running it multiple times; generating a test baseline according to the results of the multiple runs.
10. A distributed timing message system test device, the system includes a message publishing client, a message subscription client, and a timing message server, the device includes: a message subscription module, which batch initiates subscriptions to the timing messages that can be published by the message publishing client through the message subscription client; a fault injection module, which constructs a fault injection instruction according to the delay time corresponding to the timing message and sends it to the timing message server to inject a fault of a specified type into the timing message server, specifically including: constructing a fault injection instruction with a corresponding fault injection time not less than the delay time according to the delay time corresponding to the timing message; a publishing and delivery module, which publishes each of the timing messages subscribed by the message subscription client to the timing message server through the message publishing client, so that the timing message server delivers the timing message to the message subscription client; a result determination module, which verifies the situation of the message subscription client receiving the timing message according to the subscription and determines the test result of the system according to the verification result; the system includes a distributed cluster composed of multiple timing message servers; After injecting a specified type of fault into the timed message server, the fault injection module triggers, through the injection of the specified type of fault, in the distributed cluster: at least one new timed message server goes online to be responsible for delivering the timed message service, or at least one timed message server takes over the service of delivering the timed message from another timed message server.
11. The apparatus according to claim 10, further comprises: A partition management module logically divides all the timed messages to be delivered by taking the timed message server as a dimension, obtains a plurality of partitions, and assigns them to the corresponding timed message servers; After injecting a specified type of fault into the timed message server, the partition management module enters a partition state transition phase in response to the injected specified type of fault; In the partition state transition phase, at least some of the partitions are reallocated or switched to a disaster recovery state.
12. The apparatus according to claim 11, wherein the fault injection module determines whether the delay time corresponding to the timed message matches the partition state transition phase; If not, the delay time is adjusted accordingly to force an attempt to deliver the timed message during the partition state transition phase.
13. The apparatus according to claim 10, after injecting a specified type of fault into the timed message server, the fault injection module designates the timed message server injected with the fault as the first server; Determine that in response to the fault, a second server among the plurality of timed message servers is to take over the service of delivering the timed message from the first server; Inject a specified type of fault into the second server through the first server.
14. The apparatus according to claim 13, wherein the service disruption effect corresponding to the fault injected into the second server is lower than the service disruption effect corresponding to the fault injected into the first server.
15. The apparatus according to claim 10, after injecting a specified type of fault into the timed message server, the fault injection module determines whether at least one timed message server takes over the service of delivering the timed message from another timed message server through the injection of the specified type of fault; If not, a safe time slot is inserted during the effective period of the injection of the fault, and during the safe time slot, the fault does not take effect, and the length of the safe time slot is less than the effective period of the injection.
16. The apparatus according to claim 10, wherein the fault injection module injects at least one of the following types of atomic faults into the timed message server: Shutdown, CPU spikes, disk full, IO exception, memory spikes, network packet delay, duplication and loss, JVM device-level exception.
17. The apparatus according to claim 10, wherein the result determination module checks the integrity and timeliness of the timed messages received by the timed message server through the delivery according to the subscription, so as to determine the impact of the injected specified type of fault on the delivery.
18. For the device according to claim 10, after the result determination module determines the test result of the system based on the verification result, it schedules the test process corresponding to the test result as a daily test task and runs it multiple times; Generate a test baseline based on the results of the multiple runs.
19. A distributed timed message system test device, the system includes a message publishing client, a message subscribing client, and a timed message server, and the device includes: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: Batch initiate subscriptions to the timed messages that can be published by the message publishing client through the message subscribing client; Construct a fault injection instruction according to the delay time corresponding to the timed message and send it to the timed message server to inject a specified type of fault into the timed message server, specifically including: constructing a fault injection instruction whose corresponding fault injection time is not less than the delay time according to the delay time corresponding to the timed message; Publish each of the timed messages subscribed by the message subscribing client to the timed message server through the message publishing client, so that the timed message server delivers the timed messages to the message subscribing client; Verify the situation of the message subscribing client receiving the timed messages according to the subscription, and determine the test result of the system according to the verification result; The system includes a distributed cluster composed of multiple timed message servers; After injecting the specified type of fault into the timed message server, it further includes: Trigger in the distributed cluster through the injection of the specified type of fault: at least one timed message server newly goes online to be responsible for delivering the service of the timed message, or at least one timed message server takes over the service of delivering the timed message from another timed message server.
Citation Information
Patent Citations
System testing method and apparatus thereof
CN105868097A
Test method, device and system and computer readable storage medium
CN109981406A