Decentralized domain event publishing and subscribing method and system

By built-in publishing and subscription services in the microservice instance cluster, using the Gossip protocol and PAFD algorithm to achieve state synchronization and fault detection, the problems of large resource consumption and low business isolation capabilities in the microservice architecture are solved, and the stability and scalability of the system are improved.

CN120263845APending Publication Date: 2025-07-04BEIJING BAIJU YIXING TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510327120.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the existing microservice architecture, the deployment resources of message queue services are consumed largely and the business isolation capabilities are low, which poses the availability risk of full hangs.

Method used

The publish subscription service is built into the microservice instance cluster, and the Gossip protocol and PAFD algorithm are used to realize state synchronization and fault detection between nodes, combined with push-pull event delivery strategy to ensure high availability and stability.

Benefits of technology

It reduces the deployment resource consumption of small and medium-sized microservice clusters, improves the isolation capability of the business line and the stability of the system, supports rapid horizontal scaling, and provides intelligent fault detection through D-S evidence theory to improve decision efficiency and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263845A_ABST
    Figure CN120263845A_ABST
Patent Text Reader

Abstract

The invention discloses a decentralized domain event publishing and subscribing method and system. The invention relates to the technical field of decentralized services. Based on a Gossip protocol and a PAFD algorithm, service instance cluster nodes are kept in state synchronization; an event storage structure for holding event data and consumption status is initialized. And after the service instance receives the business operation request, generating a field event and publishing the field event to a local event storage. According to the invention, the publishing and subscribing service is built in the micro-service instance cluster, so that close integration of business logic and message service is realized, and extra framework service deployment resources are not needed. Meanwhile, the D-S evidence theory not only provides a method for processing uncertain information, but also provides intelligent support for the decision-making process. In fault detection, the system can automatically make decisions according to fused information, such as isolating fault nodes, triggering a recovery mechanism and the like, so that the decision efficiency and reliability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of decentralized services, specifically to the field of domain event publishing and subscribing, and particularly to a decentralized domain event publishing and subscribing method and system. Background Art

[0002] The microservices architecture splits large applications into a series of small, autonomous services. Each service focuses on solving specific business problems and interacts through lightweight communication mechanisms. This architectural pattern brings significant advantages. To effectively solve the business synchronization problem under the microservices architecture, the event-driven synchronization mechanism has emerged. As the result and impact of business operations, domain events are passed between microservices through event messages, achieving decoupling and parallel processing between services.

[0003] In the event-driven business synchronization mechanism, message queue services (such as traditional Kafka, ActiveMQ, etc., for specific logic, refer to Figure 2 ) provide the functions of publishing and subscribing event messages, ensuring the reliable transmission of event messages between microservices. However, as a separately deployed public service, message queue services also face some limitations:

[0004] (1) High deployment resource consumption: Message queue services require independent hardware resources and operation and maintenance support, increasing the deployment cost of the system.

[0005] (2) Business isolation problem: Large and small businesses both rely on a single public service, and may affect each other due to resource contention or improper configuration, increasing the complexity and risk of the system.

[0006] Therefore, the present invention proposes a decentralized domain event publishing and subscribing method and system. Summary of the Invention

[0007] In view of this, the present invention hopes to provide a decentralized domain event publishing and subscribing method and system to solve or alleviate the technical problems existing in the prior art, that is:

[0008] (1) Existing solutions require providing deployment resources for framework services in addition to business deployment resources, which consumes a large amount of deployment resources for various small and medium-sized microservice clusters.

[0009] (2) The business line isolation ability of the centrally dependent public service is relatively low, and there is an availability risk of all services failing when one fails.

[0010] The technical solution of the present invention is implemented as follows:

[0011] In a first aspect, a decentralized domain event publishing and subscribing method:

[0012] (1) Overview:

[0013] The present invention aims to build a publish - subscribe service into a microservice instance cluster to reduce the consumption of additional framework service deployment resources. By using the Gossip protocol and the PAFD algorithm (Phi Accrual Failure Detector), it realizes state synchronization and fault detection between nodes, ensuring the high availability and stability of the service. Each microservice instance is allowed to participate in event publishing and subscribing while undertaking business logic processing, thus achieving the maximization of resource utilization. In addition, a push - pull combined event delivery strategy is adopted to address the backpressure problem of event writing and consumption at the consumer end in a distributed environment, avoiding service downtime due to unbalanced publish - subscribe rates. Generally, it reduces the implementation difficulty of microservice instance deployment, improves the resource utilization efficiency, and provides strong support for the rapid horizontal expansion of the microservice cluster.

[0014] (2) Technical Solution:

[0015] After configuring a microservice instance cluster and each instance has the ability to provide a publish - subscribe service, the following steps are executed.

[0016] 2.1 Step S1, Initialization:

[0017] Based on the Gossip protocol and the PAFD algorithm, keep the nodes in the service instance cluster in state synchronization; initialize an event storage structure for storing event data and consumption status.

[0018] 2.1.1 Step S100, Keep the nodes in the service instance cluster in state synchronization based on the Gossip protocol:

[0019] Each node randomly selects other nodes at regular intervals for state exchange. Through multiple iterations, state synchronization of all nodes in the cluster is achieved. For example:

[0020]

[0021]

[0022] 2.1.2 Step S101, Detect the node state based on the PAFD algorithm:

[0023] The PAFD algorithm determines whether a node is alive by accumulating observed data (such as heartbeat messages). Each node maintains a "suspicion list" about other nodes and updates this list according to the received messages: Let S(t) be the degree of suspicion of node A about node B at time t, and H(t) be whether node A receives a heartbeat message from node B at time t (1 for received, 0 for not received), then the update rule is:

[0024] S(t + 1) = α·S(t) + (1 - α)·(1 - H(t));

[0025] Where α is a decay factor used to adjust the influence of historical suspicion degree on the current suspicion degree, and its value is assigned through the D-S evidence theory algorithm. The method includes the following steps S1010 to S1013.

[0026] 2.1.2.1 Step S1010, collecting evidence, including:

[0027] (1) Communication delay a: Regularly measure the communication delay between node A and node B and record it.

[0028] (2) Packet loss rate b: Statistically calculate the packet loss rate during the communication between node A and node B.

[0029] (3) Historical interaction record c: Record the historical interaction situation between node A and node B, including the number of successful and failed interactions.

[0030] Then form a data set D: D = [(a1, b1, c1), (a2, b2, c2),..., (a n , b n , c n )];

[0031] Where (a i , b i , c i ) is the i-th evidence and n is the number of evidences.

[0032] 2.1.2.2 Step S1011, constructing the basic probability assignment BPA:

[0033] For each evidence (a i , b i , c i ), divide multiple focal elements o (i.e., the states or intervals of ) according to its value range. Then assign a basic probability e representing the possibility of the state or interval corresponding to the focal element o occurring to each focal element o. For example:

[0034] For each evidence (a i , b i , c i ), divide the focal elements according to the value ranges of the communication delay, packet loss rate, and historical interaction record. Let O be the set of focal elements, and assign a basic probability e(o) to each focal element o ∈ O, representing the possibility of the state or interval corresponding to the focal element o occurring.

[0035] Specifically, focal elements can be divided according to the actual measured values of communication delay, packet loss rate, and historical interaction records, and a basic probability is assigned to each focal element. For example, for communication delay, it can be divided into three focal elements: "low delay", "medium delay", and "high delay", and basic probabilities are assigned to each focal element according to historical data. Similarly, for packet loss rate and historical interaction records, they are divided into "low packet loss", "medium packet loss", "high packet loss", "low interaction", "medium interaction", and "high interaction". These value ranges can be either average value ranges assigned based on historical probabilities or value ranges assigned subjectively.

[0036] 2.1.2.3 Step S1012, combining evidence:

[0037] Using the combination rule of D-S evidence theory, the basic probability assignments e of multiple evidences (a i , b i , c i ) are combined into a joint basic probability assignment m. Let m(B alive ) and m(B dead ) represent the joint basic probability assignments of the survival and death states of node B respectively. Calculate the belief degree Bel(B alive ) and plausibility Pl(B alive ) of the survival state of node B according to the joint basic probability assignment:

[0038]

[0039] 2.1.2.4 Step S1013, dynamically adjusting the attenuation factor α:

[0040] Define a boolean function isReliable to judge whether the survival state of node B is reliable according to the thresholds of belief degree and plausibility. If isReliable is true, then decrease the value of the attenuation factor α; if it is false, then increase the value of α:

[0041] isReliable = (Bel(B alive ) ≥ θ Bel ) ∧ (Pl(B alive ) ≤ θ Pl )

[0042] where θ Bel and θ Pl are the thresholds of belief degree and plausibility respectively. ∧ is a logical AND operation; if both logical expressions are true, then the result of the whole expression is true; otherwise it is false.

[0043] If isReliable is true, then decrease the value of the attenuation factor α; if it is false, then increase the value of α. Specifically, two adjustment factors δ can be setdecrease and δ increase , used to control the adjustment amplitude of the attenuation factor α:

[0044]

[0045] where α old is the current attenuation factor value, and α new is the adjusted attenuation factor value. By dynamically adjusting the attenuation factor α, it can better adapt to the actual communication status and the change of trust degree among nodes.

[0046] 2.1.3 Step S102, initialize the event storage structure:

[0047] Initialize an event storage structure for each microservice instance to save event data and consumption status. The data structure is as follows:

[0048] EventStore:

[0049] - events: list of events

[0050] - consumptionStates: map of eventID to consumptionState

[0051] 2.2 Step S2, event publishing:

[0052] When the service instance receives a business operation request, it generates a domain event and publishes the domain event to the local event storage. The Gossip protocol synchronizes the event to other service instances, and after other service instances receive the synchronized event, they update the local event storage.

[0053] 2.2.1 Step S200, receive a business operation request:

[0054] The service instance receives a business operation request from the client or other services, and generates one or more domain events. Among them, the domain event is a fact describing the impact of the business operation on the system state; the service instance publishes the generated domain event to its local event storage.

[0055] 2.2.2 Step S201, synchronize events using the Gossip protocol:

[0056] The service instance uses the Gossip protocol to synchronize the events in the local event storage to other service instances; after other service instances receive the synchronized events from the Gossip protocol, they update their local event storage according to the received synchronized events.

[0057] 2.3 Step S3, event subscription and consumption:

[0058] When a service instance registers for corresponding event types and there are new events in the local event store, the events are pushed to the corresponding service instances according to the registration information; after receiving the events, the service instances process them and update the consumption status in the local event store, and then synchronize the consumption status to other service instances through the Gossip protocol.

[0059] 2.3.1 Step S300, service instance registers event types:

[0060] The service instance registers event types with the event management system or message broker, including the event type, the identifier of the service instance, and the callback interface or queue for receiving events.

[0061] The event management system or the service instance itself monitors the local event store to detect whether new events have arrived. New events refer to those events that have not been pushed to the registered service instances yet.

[0062] 2.3.2 Step S301, push events according to registration information:

[0063] When new events are detected, the events are pushed to the corresponding service instances according to the registration information of the service instances through callback interface calls, message queue sending, or asynchronous communication mechanisms.

[0064] 2.3.3 Step S302, perform business logic processing:

[0065] The service instance executes corresponding business logic according to the content of the event, including updating the local database, calling other services, and generating new events. After processing the event, the service instance updates the consumption status of the event in the local event store, including processed, processing, and / or processing failed.

[0066] 2.3.4 Step S303, synchronize consumption status through the Gossip protocol:

[0067] The service instance uses the Gossip protocol to synchronize the updated consumption status to other service instances. This ensures that all service instances are aware of the current processing status of the events and can avoid duplicate processing or missed processing.

[0068] 2.4 Step S4, fault detection:

[0069] The PAFD algorithm is used to detect the status of service instances; when a fault of a certain service instance is detected, a fault recovery mechanism is triggered for repair, and the events and data on the faulty instance are migrated to other surviving instances, and at the same time, the node status information in the cluster is updated.

[0070] 2.4.1 Step S400, collect monitoring data:

[0071] Including response time, error rate, and CPU usage.

[0072] 2.4.2 Step S401, applying the PAFD algorithm:

[0073] Use the PAFD algorithm to analyze the monitoring data to detect whether a service instance has failed; Let X t be the monitoring data vector at time t, T be the preset threshold vector, and R t be the probe signal response received at time t:

[0074]

[0075] where FaultDetected t is a boolean value indicating whether a failure is detected at time t. When the monitoring data X t exceeds the threshold T or the probe signal response R t is null (i.e., there is no response), a failure is considered to be detected. null is any of the following states:

[0076] 1) No response to the probe signal: When a probe signal is actively sent to the service instance to check its status, if the service instance does not respond, the response value is set to null. This indicates that the probe attempt has failed because the service instance has crashed, the network is down, or for other reasons.

[0077] 2) Data missing: During the monitoring or log collection process, if some expected data points or records are missing, the relevant fields are marked as null. This indicates that the data is incomplete or there is a problem with the collection process.

[0078] 3) Uninitialized or unassigned variable: In programming or algorithm implementation, if a variable is declared but not initialized or assigned, its default value is null. In the fault detection algorithm, such a variable represents a state that has not been set or detected yet.

[0079] 4) Invalid or incorrect state: null is a pointer or reference that does not point to any valid object or instance.

[0080] 2.4.3 Step S402, fault detection decision:

[0081] Based on the output of the PAFD algorithm, determine whether the service instance is in a fault state. If a failure is detected, trigger the fault recovery mechanism for repair and enter S403; otherwise, continue monitoring or enter S5.

[0082] 2.4.4 Step S403, execute the fault recovery mechanism:

[0083] Migrate the events and data on the failed instance to other surviving service instances. Update the node status information in the cluster to reflect and resolve the current status of the failed instance.

[0084] 2.5 Step S5, Expansion and Maintenance Process:

[0085] When expanding the microservice cluster is required, add new service instances and automatically synchronize the existing events and data:

[0086] S500, Add new service instances: Be compatible with the existing instances and handle requests;

[0087] S501, Automatically synchronize the existing events and data: Use the data replication or event synchronization mechanism to automatically synchronize the existing events and data to the newly added service instances;

[0088] S502, Update service registration and discovery: Update the service registration and discovery mechanism so that the new instances can be correctly discovered and called by the clients and other services.

[0089] (3) Mechanism for Solving Technical Problems:

[0090] (1) For the problem that the existing solution needs to provide additional deployment resources for the framework service outside the business deployment resources, the present invention realizes the tight integration of the business logic and the publish-subscribe service by building the publish-subscribe service into the microservice instance cluster. In this way, there is no need to provide separate deployment resources for the framework service, thus significantly reducing the deployment resource consumption of small and medium-sized microservice clusters. Each microservice instance is responsible for both the business logic processing and the publishing and subscribing of events, achieving the maximization of resource sharing and utilization.

[0091] (2) For the problem that the business line isolation ability of the centrally dependent common service is relatively low and there is a risk of all services failing when one fails, the present invention eliminates the need for the centrally dependent common service through a decentralized design. Each microservice instance has the ability to provide the publish-subscribe service, and the nodes synchronize their states through the Gossip protocol, achieving high availability and stability of the service. Even if a certain node fails, other nodes can quickly take over its tasks to ensure the continuity and stability of the business. This decentralized design not only improves the business line isolation ability but also effectively reduces the risk of the overall service unavailability caused by a single point of failure.

[0092] In the second aspect, a decentralized domain event publish-subscribe system:

[0093] The system includes a processor and a memory connected to the processor. Program instructions are stored in the memory. When the program instructions are executed by the processor, the processor executes the domain event publish-subscribe method as described above. And the processor is connected to,

[0094] (1) The event publishing module responsible for event publishing and synchronization.

[0095] (2) The event storage module used to save event data and consumption status;

[0096] (3) The event subscription and consumption module used to push new events to corresponding service instances according to registration information.

[0097] (4) The status synchronization and fault detection module that implements status synchronization and fault detection between nodes based on the Gossip protocol and the PAFD algorithm.

[0098] (5) The cluster management module responsible for the management of the entire microservice instance cluster.

[0099] (6) The business logic processing module responsible for processing business requests and generating domain events.

[0100] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0101] I. Efficient resource utilization: By integrating the publish-subscribe service into the microservice instance cluster, the present invention realizes the tight integration of business logic and message service without the need for additional framework service deployment resources. It significantly reduces the deployment cost of small and medium-sized microservice clusters, enabling resources to be utilized more efficiently. At the same time, the D-S evidence theory of the present invention not only provides a method for processing uncertain information but also provides intelligent support for the decision-making process. In fault detection, the system can automatically make decisions based on the fused information, such as isolating faulty nodes, triggering recovery mechanisms, etc., thereby improving the decision-making efficiency and reliability.

[0102] II. Comprehensive active and passive fault detection strategies and multi-source information fusion technology: The D-S evidence theory can process uncertain information from multiple different sources and fuse this information into a more reliable and comprehensive judgment through the combination rule. In fault detection, this is equivalent to being able to utilize the information provided by multiple monitoring metrics and detection means to more accurately judge the node status. It allows the system to consider information from multiple sources simultaneously and perform intelligent fusion on this information. This enables the system to still make accurate judgments even when there are false alarms or missed alarms in a single monitoring metric. This ability enhances the robustness of the system, enabling it to better cope with complex and changing operating environments. And by detecting faulty nodes in a timely and accurate manner, the system can quickly isolate these nodes to prevent the spread of faults. At the same time, combined with the fault recovery mechanism, the system can automatically or manually migrate the events and data on the faulty nodes to other surviving instances to ensure the continuity and availability of the service.

[0103] III. Improvement in High Availability and Stability: The present invention adopts the Gossip protocol and the Phi Accrual Failure Detector algorithm to ensure state synchronization and fault detection among nodes. The decentralized design eliminates the need for a centralized and dependent common service. Even if a node fails, other nodes can quickly take over its tasks, thereby improving the high availability and stability of the service.

[0104] IV. Enhancement of Business Line Isolation Ability: Since each microservice instance has the ability to publish and subscribe to services, the present invention effectively enhances the isolation ability of business lines. Even if a problem occurs in a certain business line, it will not affect the normal operation of other business lines, reducing the risk of overall service unavailability caused by a single point of failure.

[0105] V. Flexible Expansion and Maintenance: The present invention supports the rapid horizontal expansion of microservice clusters. When the cluster needs to be expanded, only new service instances need to be added to automatically synchronize existing events and data. This reduces the complexity of expansion and improves the maintainability of the system. In a microservice architecture or a distributed system, as the business develops, the system needs to be dynamically expanded to handle higher loads. The fault detection method based on the PAFD algorithm and D-S evidence theory can support this dynamic expansion process, ensuring that newly added service instances can be quickly monitored and included in the fault detection scope after joining the cluster.

[0106] VI. Avoidance of the Problem of Imbalanced Publishing and Subscribing Rates: The present invention adopts a push-pull combined event delivery strategy to effectively address the backlog problem of event writing and consumption at the consumer end in a distributed environment. This avoids service downtime caused by imbalanced publishing and subscribing rates and improves the stability and processing capacity of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0107] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0108] Figure 1 It is a schematic diagram of the method flow of the present invention;

[0109] Figure 2 It is a relationship diagram between Kafka and business lines;

[0110] Figure 3 It is a schematic diagram of the cluster network adopted in the embodiment of the present invention;

[0111] Figure 4 It is a schematic diagram of the system composition of the present invention. Detailed implementation manners

[0112] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following will describe in detail the specific implementation manners of the present invention with reference to the accompanying drawings. Many specific details are set forth in the following description to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below;

[0113] It should be noted that the embodiments in this specification are described in a progressive manner, and the key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0114] Explanation of related terms:

[0115] (1) Decentralization: A network structure and management method aiming to improve the robustness, transparency, and fairness of the system;

[0116] (2) Domain event: A state change event published by the business domain model;

[0117] (3) Publish-subscribe: A message passing pattern, where the publisher creates a message and sends it to a shared message channel (such as a message queue or topic), and subscribers can subscribe to these messages to receive notifications when the messages are published;

[0118] (4) Broker: The management node that delivers events to the consumer side;

[0119] (5) Gossip protocol: A decentralized distributed protocol that achieves data synchronization by randomly exchanging information between nodes, similar to the spread of rumors or epidemics.

[0120] (6) PAFD algorithm (Phi Accrual Failure Detector): An algorithm for detecting node failures in a distributed system, which determines whether a node is alive by accumulating observed data to improve the reliability and fault tolerance of the system.

[0121] (7) Event data and consumption status: Event data refers to the information generated in a microservice that needs to be processed by other service instances; the consumption status refers to the progress or result of the subscription and processing of these event data.

[0122] (8) Business Logic: The code logic in a microservice instance for handling business requests, generating domain events, etc., which is the core function and a key part for implementing business functions.

[0123] (9) Node Status Information: In a microservice cluster, the current status information of each service instance (node), including whether it is alive, data synchronization status, load status, etc.

[0124] (10) D-S Evidence Theory (Dempster-Shafer Theory): A mathematical tool for dealing with uncertain information, which provides a set of formal methods to combine evidence from different sources and calculate the probability of hypotheses. In this embodiment, the D-S evidence theory is used to dynamically adjust the value of the attenuation factor α to better reflect the actual communication status and trust degree between nodes.

[0125] Embodiment 1: As Figure 1 、 3 shown, this embodiment provides an application of a decentralized domain event publishing and subscribing method in the information processing scenario of an online car-hailing platform, including the following steps S1 to S5.

[0126] In this embodiment, regarding step S1, Initialization: Based on the Gossip protocol and the PAFD algorithm, keep the state synchronization of the service instance cluster nodes; Initialize the event storage structure for saving event data and consumption status.

[0127] Specifically, in step S100, keep the state synchronization of the service instance cluster nodes based on the Gossip protocol: In the online car-hailing platform, the service instance cluster nodes include user service, order service, driver service, etc. These service instances need to keep state synchronization to ensure that the platform can handle user requests normally. Through the Gossip protocol, each service instance node randomly selects other nodes regularly for state exchange. For example, the user service node selects to exchange states with the order service node or the driver service node. Through multiple iterations, the states of all nodes in the cluster will reach synchronization.

[0128] Specifically, in step S101, detect the node status based on the PAFD algorithm: In the online car-hailing platform, the normality of the node status is directly related to the stability and availability of the platform. Therefore, it is crucial to use the PAFD algorithm to detect the node status. The PAFD algorithm determines whether a node is alive by accumulating observed data (such as heartbeat messages). Each node maintains a "suspicion list" about other nodes and updates this list according to the received messages: Let S(t) be the degree of suspicion of node A about node B at time t, and H(t) be whether node A receives the heartbeat message of node B at time t (received is 1, not received is 0), then the update rule is:

[0129] S(t + 1) = α·S(t) + (1 - α)·(1 - H(t));

[0130] Where α is a decay factor used to adjust the influence of historical doubt degree on the current doubt degree;

[0131] Preferably, the decay factor α is assigned by the D - S evidence theory algorithm, and the method includes the following steps S1010 to S1013.

[0132] Step S1010, collect evidence, including:

[0133] (1) Communication delay a: Regularly measure the communication delay between the user service node and the order service node.

[0134] (2) Packet loss rate b: Statistically calculate the packet loss rate during the communication between the user service node and the order service node.

[0135] (3) Historical interaction record c: Record the historical interaction situation between the user service node and the order service node.

[0136] Then form a data set D: D = [(a1, b1, c1), (a2, b2, c2),..., (a n , b n , c n )];

[0137] Where (a i , b i , c i ) is the i - th evidence, and n is the number of evidences.

[0138] Step S1011, construct the basic probability assignment BPA: Divide the focal elements according to the value ranges of the communication delay, packet loss rate, and historical interaction record, such as "low delay", "medium delay", "high delay", "low packet loss", "medium packet loss", "high packet loss", "low interaction", "medium interaction", "high interaction". Assign basic probabilities to each focal element to represent the possibility of the corresponding state or interval occurring.

[0139] It can be understood that these value ranges can be either the average value ranges assigned based on historical probabilities or the subjectively assigned value ranges.

[0140] Step S1012, combine evidences: Use the combination rule of the D - S evidence theory to combine the respective basic probability assignments e of multiple evidences (a i , b i , c i ) into a combined basic probability assignment m. Let m(B alive ) and m(B dead) represent the combined basic probability assignments for the alive and dead states of node B respectively. Calculate the belief degree Bel(B alive ) and the plausibility degree Pl(B alive ) based on the combined basic probability assignment:

[0141]

[0142] Step S1013, dynamically adjust the attenuation factor α: Based on the Boolean function isReliable, determine whether the alive state of node B is reliable according to the thresholds of the belief degree and the plausibility degree. If isReliable is true, then decrease the value of the attenuation factor α; if it is false, then increase the value of α:

[0143] isReliable = (Bel(B alive ) ≥ θ Bel ) ∧ (Pl(B alive ) ≤ θ Pl )

[0144] where θ Bel and θ Pl are the thresholds of the belief degree and the plausibility degree respectively. ∧ is the logical AND operation; if both logical expressions are true, then the result of the entire expression is true; otherwise it is false.

[0145] If isReliable is true, then decrease the value of the attenuation factor α; if it is false, then increase the value of α. Specifically, two adjustment factors δ decrease and δ increase can be set to control the adjustment amplitude of the attenuation factor α:

[0146]

[0147] where α old is the current value of the attenuation factor, and α new is the adjusted value of the attenuation factor. By dynamically adjusting the attenuation factor α, it can better adapt to the actual communication situation between nodes and the changes in the degree of trust.

[0148] Specifically, in step S102, initialize the event storage structure: In the online car-hailing platform, each microservice instance (such as the user service, order service) needs to initialize an event storage structure to save event data and consumption status. This event storage structure includes an event list and a consumption status mapping table. The event list is used to store the events that occur, while the consumption status mapping table records the consumption status of each event (such as processed, unprocessed). By initializing such an event storage structure, the online car-hailing platform can ensure the orderly processing of events and the correct tracking of status, thereby improving the reliability and availability of the platform.

[0149] It should be noted that in step S1, regarding the Gossip protocol: by randomly selecting nodes periodically for status exchange, the status synchronization of all nodes in the cluster is achieved. This mechanism can ensure that even if some nodes fail, other nodes can still maintain the latest status information.

[0150] It should be noted that in step S1, regarding the PAFD algorithm: by accumulating observed data (such as heartbeat messages) to determine whether a node is alive. Combining with the D-S evidence theory, the PAFD algorithm can more accurately evaluate the survival status of nodes and dynamically adjust the attenuation factor to adapt to the actual communication situation and the change of trust degree among nodes.

[0151] It should be noted that in step S1, regarding the event storage structure: a structured event storage mechanism is provided for each microservice instance to ensure the orderly processing of events and the correct tracking of status. This helps to improve the reliability and availability of the platform, because even in case of failure or restart, event data will not be lost and can be restored to the correct processing state.

[0152] Furthermore, the Python execution program of step S1 is as follows:

[0153]

[0154]

[0155]

[0156]

[0157] In the above program, the Gossip protocol exchanges status by randomly selecting other nodes as peers through multiple iterations. The PAFD algorithm collects evidence (random values in this example) about peer nodes. Basic probability assignment (BPA) is performed based on the evidence. The D-S evidence theory combination rule is used to combine the evidence and calculate the degree of trust and likelihood. The attenuation factor α is dynamically adjusted according to the degree of trust and likelihood. The event storage structure initializes an instance of the EventStore class, which has an event list and a consumption status mapping table. This structure is used to store and process event data and track the consumption status of each event.

[0158] In this embodiment, regarding step S2: Event Publishing:

[0159] Specifically, in step S200, receive a business operation request: A ride-hailing user initiates a ride request through the platform. The service instance of the platform receives this request and generates a domain event in the "ride request" domain. A domain event is a fact that describes the impact of a business operation (such as a ride request) on the system state (such as order status, vehicle allocation). The service instance publishes the event to its local event store for subsequent processing and synchronization.

[0160] Specifically, in step S201, use the Gossip protocol to synchronize the event: Other service instances of the platform (such as order processing, vehicle scheduling) need to synchronize this "ride request" event through the Gossip protocol. The Gossip protocol is a decentralized communication protocol used to synchronize data in a distributed system. Here, it ensures that all relevant service instances can receive the "ride request" event and update their local event stores to maintain data consistency.

[0161] Furthermore, the Python execution program for the above step S2 is as follows:

[0162]

[0163]

[0164] Among them, the ServiceInstance class represents a service instance, which has a local event store local_event_store. When the service instance receives a business operation request, it generates a domain event and stores it in the local event store. In actual applications, the Gossip protocol can be used to synchronize the event to other service instances.

[0165] In this embodiment, regarding step S3: Event subscription and consumption:

[0166] Specifically, in step S300, the service instance registers the event type: The order processing service instance registers the "ride request" event type with the event management system and provides a callback interface for receiving events. By registering the event type, the service instances indicate that they are interested in specific types of events and provide a mechanism for receiving events.

[0167] Specifically, in step S301, push the event according to the registration information: When the "ride request" event arrives, the event management system pushes the event to the order processing service instance through the callback interface. The event management system monitors the local event store, and once a new event arrives, it pushes the event to the service instances according to their registration information.

[0168] Specifically, in step S302, perform business logic processing: After the order processing service instance receives the "ride request" event, it executes the corresponding business logic, such as creating an order, allocating a vehicle, etc. The service instance executes the business logic according to the content of the event and updates the local database and / or calls other services. After processing the event, they update the consumption status in the local event store.

[0169] Specifically, in step S303, synchronize the consumption status through the Gossip protocol: The order processing service instance uses the Gossip protocol to synchronize the consumption status (such as processed) of the "ride request" event to other service instances. Synchronizing the consumption status through the Gossip protocol ensures that all service instances are aware of the current processing status of the event, thus avoiding duplicate processing or missed processing.

[0170] Furthermore, the Python execution program for step S3 above is as follows:

[0171]

[0172]

[0173] Among them, the EventManagementSystem class represents the event management system, which maintains a dictionary of event subscribers event_subscribers. Service instances can register the event types they are interested in with the event management system. When the event management system decides to push an event, it will push the event to the corresponding service instance according to the registration information.

[0174] In this embodiment, regarding step S4: Fault detection:

[0175] Specifically, in step S400, collect monitoring data: The platform collects monitoring data such as the response time, error rate, and CPU usage rate of service instances.

[0176] Specifically, in step S401, apply the PAFD algorithm: The platform uses the PAFD algorithm to analyze the monitoring data to detect whether a service instance has a fault. Let X t be the monitoring data vector at time t, T be the preset threshold vector, and R t be the detection signal response received at time t:

[0177]

[0178] Among them, FaultDetected t is a boolean value indicating whether a fault is detected at time t. When the monitoring data X t exceeds the threshold T or the detection signal response R tWhen it is empty (i.e., there is no response), a fault is considered detected. null is the fault state;

[0179] It can be understood that the PAFD algorithm is a fault detection algorithm that analyzes monitoring data to detect whether a service instance has a fault. If the monitoring data exceeds a preset threshold or the probe signal has no response, a fault is considered detected.

[0180] Specifically, in step S402, fault detection decision: If the PAFD algorithm detects that a certain service instance has a fault, the platform will trigger the fault recovery mechanism. According to the output of the PAFD algorithm, the platform determines whether the service instance is in a fault state and decides whether to trigger the fault recovery mechanism accordingly.

[0181] Specifically, in step S403, execute the fault recovery mechanism: The platform migrates the events and data on the faulty service instance to other surviving service instances and updates the node status information in the cluster. The fault recovery mechanism ensures the continuity of the system and the integrity of the data when a service instance has a fault. By migrating events and data and updating the node status information, the system can continue to run normally and process user requests.

[0182] Furthermore, the Python execution program for the above step S4 is as follows:

[0183]

[0184]

[0185] Among them, the PAFD class represents the PAFD algorithm, which has a threshold dictionary threshold for defining the thresholds of response time, error rate, and CPU usage. The detect_fault method receives the monitoring data as input and determines whether the service instance has a fault according to the thresholds. If any item in the monitoring data exceeds its corresponding threshold, a fault is considered detected.

[0186] In this embodiment, regarding step S5, the extension and maintenance process: As the number of users of the online car-hailing platform grows and the service demand increases, the platform needs to expand its microservice cluster to handle more requests. For example, the order processing service needs to be expanded due to high concurrent requests. Then the following steps need to be executed:

[0187] (1) The newly added order processing service instance needs to obtain the existing order events and data so that it can immediately start processing new order requests:

[0188] P1. When the new instance starts, it will automatically connect to the event bus or message queue.

[0189] P2. The new instance receives existing order events and data from the event bus or message queue.

[0190] P3. The new instance stores these events and data in its local database for subsequent processing.

[0191] P4. The new instance starts listening for new order requests and processes them using synchronized data.

[0192] It can be understood that the data replication or event synchronization mechanism is crucial for ensuring that all instances in the microservice cluster have the same data. By synchronizing these events and data to the new instance, it can be ensured that the new instance can immediately start processing requests without waiting for data synchronization to complete.

[0193] (2) Update service registration and discovery: To ensure that the client and other services of the online car-hailing platform can discover and call the newly added order processing service instance, the service registration and discovery mechanism needs to be updated:

[0194] P1. When the new instance starts, it automatically registers itself with the service registry.

[0195] P2. The service registry updates its internal service list, including the address and port information of the new instance.

[0196] P3. The client and other services query the service registry through the service discovery mechanism to obtain the latest service list.

[0197] P4. The client and other services start sending order requests to the new instance for processing.

[0198] It can be understood that the service registration and discovery mechanism is a key component in the microservice architecture. It allows service instances to automatically register themselves when starting and unregister automatically when shutting down. The client and other services can discover available service instances by querying the service registry and send requests to these instances for processing. By updating the service registration and discovery mechanism, it can be ensured that the new instance can be correctly discovered and called.

[0199] Furthermore, the Python execution program for step S5 above is as follows:

[0200]

[0201]

[0202] In the above program, when it is necessary to expand the microservice cluster, a new service instance can be started. After the new instance starts, it will automatically register with the service registry so that other services and clients can discover and call it.

[0203] The new instance needs to consume existing events and data from a message queue or event bus. The consumed events and data will be saved in the local database or memory storage of the new instance for subsequent processing.

[0204] Through the registration operation of the service registry, the information of the new instance will be added to the list of the service registry. Other services and clients can query the service registry to obtain the latest service list and discover the existence of the new instance.

[0205] Embodiment 2: As Figure 4 shown, based on Embodiment 1, this embodiment further provides a decentralized domain event publishing and subscribing system:

[0206] The system includes a processor and a memory connected to the processor. Program instructions are stored in the memory. When the program instructions are executed by the processor, the processor executes the domain event publishing and subscribing method as described above. And the processor is connected to

[0207] (1) An event publishing module responsible for event publishing and synchronization: responsible for publishing these events to the local event storage after the microservice instance receives a business operation request and generates a domain event. At the same time, this module is also responsible for synchronizing the events to other service instances through the Gossip protocol to ensure the consistency and reliability of the events.

[0208] (2) An event storage module for saving event data and consumption status;

[0209] (3) An event subscription and consumption module for pushing new events to corresponding service instances according to registration information: allowing service instances to register interested event types and pushing new events to corresponding service instances according to the registration information. After the service instance receives an event, it triggers business logic processing and updates the consumption status in the event storage. At the same time, the consumption status is synchronized to other service instances through the Gossip protocol to maintain the consistency of the state within the cluster.

[0210] (4) A state synchronization and fault detection module based on the Gossip protocol and the PAFD algorithm to achieve state synchronization and fault detection between nodes: regularly detects the state of service instances. When a fault is detected, it triggers a fault recovery mechanism, such as migrating the events and data on the faulty instance to other surviving instances.

[0211] (5) A cluster management module responsible for the management of the entire microservice instance cluster: including node joining, exiting, status monitoring, etc. Provides interfaces for viewing and modifying cluster configurations, supporting the dynamic expansion and maintenance of the cluster.

[0212] All of the above embodiments merely represent the implementation manners of the relevant actual applications of the present invention. The descriptions thereof are relatively specific and detailed, but should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.

Claims

1. A decentralized domain event publishing and subscribing method, including configuring a microservice instance cluster and each instance has the ability to publish and subscribe services, characterized in that, After the service instance receives a business operation request from the client, the following steps are implemented: S1. Based on the Gossip protocol and the PAFD algorithm, keep the state synchronization of the service instance cluster nodes; initialize an event storage structure for saving event data and consumption status; S2. When the service instance receives a business operation request, generate a domain event and publish the domain event to the local event storage; The Gossip protocol synchronizes the event to other service instances, and after other service instances receive the synchronized event, update the local event storage; S3. When the service instance registers the corresponding event type and there are new events in the local event storage, push them to the corresponding service instance; update the consumption status in the local event storage, and synchronize the consumption status to other service instances through the Gossip protocol; S4. When it is detected that a certain service instance fails, trigger a fault recovery mechanism for repair, migrate the events and data on the faulty instance to other surviving instances, and at the same time update the node status information in the cluster.

2. The domain event publishing and subscribing method according to claim 1, characterized in that: The implementation method of the above S1 includes: S100. Each node randomly selects other nodes at regular intervals for status exchange, and realizes the status synchronization of all nodes in the cluster through iteration; S101. The PAFD algorithm accumulates observation data:: Let S(t) be the degree of suspicion of node A about node B at time t, and H(t) be whether node A receives the heartbeat message of node B at time t, then the update rule is: S(t + 1) = α·S(t) + (1 - α)·(1 - H(t)); Among them, α is a decay factor used to adjust the influence of historical suspicion degree on the current suspicion degree; S102. Initialize an event storage structure for saving event data and consumption status for each microservice instance.

3. The domain event publishing and subscribing method according to claim 2, wherein: In the above S101, the update method of the decay factor α is: S1010, Collect evidence, including the communication delay a, packet loss rate b, and historical interaction record c between node A and node B, to form a data set D: D = [(a1, b1, c1), (a2, b2, c2),..., (a n , b n , c n )]; where (a i , b i , c i ) is the i-th piece of evidence, and n is the number of pieces of evidence; S1011, for each piece of evidence (a i , b i , c i ), divide multiple focal elements o according to its value range and assign a basic probability e representing the nature of the state or interval corresponding to the focal element o; S1012, using the combination rule of D-S evidence theory, combine the basic probability assignments e of multiple evidences (a i , b i , c i ) into a combined basic probability assignment m; S1013. Based on the boolean function isReliable, judge whether the survival status of node B is reliable according to the thresholds of trust degree and likelihood degree; if isReliable is true, reduce the value of the decay factor α; if it is false, increase the value of α.

4. The field event publishing and subscribing method according to claim 3, characterized in that: In the step S1012, m(B alive ) and m(B dead ) respectively represent the combined basic probability assignments for the alive and dead states of node B; calculate the belief degree Bel(B alive ) and the plausibility degree Pl(B alive ) for the alive state of node B according to the combined basic probability assignments: In the above S1013, the boolean function isReliable: isReliable = (Bel(B alive ) ≥ θ Bel ) ∧ (Pl(B alive ) ≤ θ Pl ) where θ Bel and θ Pl are the thresholds of the degree of belief and the likelihood, respectively; ∧ is the logical AND operation; if both logical expressions are true, the result of the whole expression is true; otherwise it is false.

5. The domain event publishing and subscribing method according to any one of claims 1 to 4, characterized in that: The implementation method of the above S2 includes: S200. The service instance receives a business operation request from the client or other services, generates one or more domain events; the service instance publishes the generated domain events to its local event storage; S201. The service instance uses the Gossip protocol to synchronize the events in the local event storage to other service instances; after other service instances receive the synchronized events from the Gossip protocol, update their local event storage according to the received synchronized events.

6. The domain event publishing and subscribing method according to any one of claims 1 to 4, characterized in that: The implementation method of the above S3 includes: S300. The service instance registers the event type with the event management system or message broker, including the event type, the identifier of the service instance, and the callback interface or queue for receiving events; S301. When a new event is detected, according to the registration information of the service instance, push the event to the corresponding service instance through callback interface call, message queue sending or asynchronous communication mechanism In S302, execute corresponding business logic, including updating the local database, invoking other services, and generating new events; In S303, the service instance uses the Gossip protocol to synchronize the updated consumption status to other service instances.

7. The domain event publishing and subscribing method according to any one of claims 1 to 4, characterized in that: The implementation method of S4 includes: In S400, collect monitoring data, including response time, error rate, and CPU usage rate; In S401, use the PAFD algorithm to analyze the monitoring data to detect whether a service instance fails; In S402, determine whether the service instance is in a failure state. If a failure is detected, trigger a failure recovery mechanism for repair and enter S403. Otherwise, continue monitoring; In S403, migrate the events and data on the failed instance to other surviving service instances; update the node status information in the cluster to reflect and resolve the current state of the failed instance.

8. The domain event publishing and subscribing method according to any one of claims 1 to 4, characterized in that: It also includes S5: When it is necessary to expand the microservice cluster, add new service instances and automatically synchronize the existing events and data.

9. A decentralized domain event publishing and subscribing system, characterized in that: The system includes a processor and a memory connected to the processor. Program instructions are stored in the memory. When the program instructions are executed by the processor, the processor executes the domain event publishing and subscribing method according to any one of claims 1-8.

10. The domain event publishing and subscribing system according to claim 9, characterized in that: The processor is connected to an event publishing module responsible for event publishing and synchronization; an event storage module for saving event data and consumption status; an event subscribing and consuming module for pushing new events to corresponding service instances according to registration information; a status synchronization and failure detection module for implementing status synchronization and failure detection between nodes based on the Gossip protocol and the PAFD algorithm; a cluster management module responsible for managing the entire microservice instance cluster; a business logic processing module responsible for processing business requests and generating domain events.

Citation Information

Cited By

  • Bullet screen live broadcast interaction method, server, program product and storage medium

    CN120915971A