Risk control engine system, self-healing method and electronic equipment
Through the dual monitoring mechanism of the central control node and the working node, the shortcomings of the traditional risk control engine in anomaly detection and self-healing are solved, timely and accurate self-healing processing is achieved, and the stability and reliability of the system are improved.
Patent Information
- Application Number
- CN202511004710.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-09-30
AI Technical Summary
When faced with complex and ever-changing business scenarios and high-concurrency requests, traditional risk control engines lack the node's own active monitoring and self-healing capabilities, resulting in untimely anomaly detection and inaccurate processing. They are unable to adopt the optimal recovery strategy based on specific circumstances, and the self-healing effect is poor.
A dual-track monitoring mechanism of central control nodes and working nodes is adopted to promptly detect node anomalies through the dual monitoring mode, and the optimal self-healing strategy is determined according to the anomaly level and type to avoid operation conflicts and improve self-healing efficiency.
It improves the system's abnormal perception ability and response efficiency, avoids resource waste, enhances the system's stability and recovery ability in complex abnormal scenarios, and has higher reliability and adaptability.
Smart Images

Figure CN120729700A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wind control technology, and in particular to a wind control engine system, a self-healing method, and an electronic device. Background Art
[0002] With the rapid development of internet technology, the scale of business in areas such as e-commerce, live streaming, and fintech has exploded, and the volume of user activity data has increased exponentially. In the areas of risk management and compliance, real-time monitoring and processing of massive amounts of user behavior data has become a core requirement for ensuring the safe operation of platforms. Therefore, risk control engine systems play a key role.
[0003] However, traditional risk control engines, faced with complex and ever-changing business scenarios and high-concurrency requests, have a limited monitoring mechanism, often relying solely on external monitoring and lacking the node's own proactive monitoring and self-healing capabilities. This results in delayed anomaly detection and inaccurate handling. Furthermore, a uniform approach is applied to anomalies of varying types and severity, failing to implement optimal recovery strategies based on specific circumstances, resulting in poor node self-healing effectiveness. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a risk control engine system, a self-healing method, and an electronic device that can promptly detect node anomalies through a dual monitoring mode and adopt the optimal recovery strategy to avoid operation conflicts and improve self-healing efficiency.
[0005] In order to achieve the above object, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides a self-healing method for a risk control engine system, wherein the risk control engine system includes a central control node and a working node; the method includes: using the central control node and the working node to monitor the health status of the working node; when an abnormality is detected, determining a self-healing strategy according to the abnormality level; performing a self-healing operation on the working node based on the self-healing strategy; when the same abnormality is detected simultaneously by the working node and the central control node, determining whether the central control node or the working node performs the self-healing operation according to the priority of the processing method corresponding to the abnormality type.
[0006] In a second aspect, the present invention provides a risk control engine system, which includes a central control node and a working node; the risk control engine system is used to execute the risk control engine system self-healing method as described in any of the aforementioned implementation methods.
[0007] In a third aspect, the present invention provides an electronic device, which is equipped with the risk control engine system described in the above embodiment.
[0008] The risk control engine system, self-healing method, and electronic device provided by the embodiments of the present invention, in the self-healing method of the risk control engine system provided by the embodiments of the present invention, the central control node and the working node are used to monitor the health status of the working node at the same time. This dual-track monitoring mechanism improves the overall abnormality perception capability of the system; when an abnormality is detected, it will be divided into levels according to the severity of the abnormality, and the corresponding self-healing operation will be determined according to the preset graded abnormality self-healing strategy. Such hierarchical processing not only improves the response efficiency, but also avoids resource waste; and when the same abnormality is detected by the working node and the central control node at the same time, the preset abnormality type and processing method priority mechanism is used to determine which party should perform the self-healing operation, effectively avoiding conflicts or repeated processing caused by the simultaneous operation of the two, and also improving the stability and recovery capability of the system in complex abnormal scenarios. Compared with the traditional single monitoring and processing mode, it has higher reliability and adaptability.
[0009] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0011] Figure 1 An architectural diagram of the risk control engine system provided by an embodiment of the present invention; Figure 2 A functional module diagram of a central control node provided by an embodiment of the present invention; Figure 3 A schematic diagram of the internal structure of a working node provided in an embodiment of the present invention; Figure 4 A schematic flow chart of a self-healing method for a risk control engine system according to an embodiment of the present invention; Figure 5 An interaction diagram between a working node and a central control node provided by an embodiment of the present invention; Figure 6 A schematic diagram of rule execution within a risk control engine system provided by an embodiment of the present invention; Figure 7 This is a structural block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0012] The following will be combined with the accompanying drawings to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.
[0013] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but is merely intended to represent selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative work are within the scope of protection of the present invention.
[0014] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
[0015] See Figure 1 , Figure 1 This is an architectural diagram of the risk control engine system provided by an embodiment of the present invention. Figure 1 The risk control engine system 10 includes a central control node 101 and multiple working nodes 102, wherein the working nodes 102 adopt a distributed deployment mode, which can realize a high reliability, high availability and high scalability risk control rule execution environment.
[0016] In the embodiment of the present invention, the central control node 101 serves as the coordination center of the entire distributed system and is responsible for managing the complete life cycle of the working node 102. For ease of understanding, please refer to Figure 2 , Figure 2 This is a functional module diagram of the central control node provided by an embodiment of the present invention.
[0017] exist Figure 2 In [1], the central control node 101 includes a rule updater, a distributed node manager, a health monitor, and a graceful shutdown processor, which work together to implement the above functions of the central control node 101. Specifically: The Rule Update Manager is responsible for detecting updates to business rules within the system and distributing them to the Distributed Node Manager. The Distributed Node Manager is responsible for creating and destroying worker nodes 102 and distributing business rules to each worker node 102 for execution. The Health Monitor is responsible for detecting anomalies in each worker node 102 and helping abnormal nodes recover. The Graceful Shutdown Processor coordinates an orderly shutdown process, ensuring data consistency and complete resource release even during system termination. This will be discussed in detail later.
[0018] based on Figure 2 The structure of the central control node 101 is as follows: calling the distributed node manager's initialization node method, which can create and initialize a specified number of working nodes 102 in a distributed environment based on the system configuration. Each working node 102 is started through a preset startup method, which calls the initialization node method to create the actual processing process and returns the node instance and its initial status information to the distributed node manager. After the node initialization is completed, the central control node 101 starts the health monitor and regularly performs health checks on all working nodes 102. If any node is found to be abnormal or unresponsive, it will immediately notify the distributed node manager, which will then adopt a corresponding recovery strategy based on the severity of the abnormality.
[0019] Next, we will introduce the working node 102. Figure 3 , Figure 3 Schematic diagram of the internal structure of a working node provided by an embodiment of the present invention. Each working node 102 is an independent execution entity.
[0020] exist Figure 3 In the example, worker node 102 includes a complete server, application layer, symmetric layer, and runtime environment. During operation, worker node 102 is responsible for starting the server, which then initializes the application layer. The application layer then loads and configures the symmetric layer. The symmetric layer completes preparations for data encryption and decryption. Finally, the runtime environment is activated and the required resources and configuration are loaded. This series of operations forms a complete risk control rule execution chain.
[0021] In this embodiment of the present invention, each worker node 102 can also perform self-monitoring and self-healing. For example, it can monitor internal process status, automatically perform local self-healing operations when an anomaly is detected, and report the self-healing event to the health monitor; for another example, it can perform optimization adjustments when resource usage approaches a limit.
[0022] In the embodiment of the present invention, the working nodes 102 may be deployed in a resource isolation manner, which can ensure that a problem of one working node 102 does not spread to other working nodes, thereby greatly improving the stability and reliability of the system.
[0023] Combine Figures 1 to 3 The risk control engine system 10 provided in the embodiment of the present invention can also coordinate the interaction between the central control node 101 and the working node 102 and control the operating status. For example, as an example, the system can use two global state managers SYSTEM_STATE and ACTIVE_NODES. SYSTEM_STATE is used to control the operating status and graceful shutdown process of the system. When the system needs to be shut down, the central control node 101 will update the SYSTEM_STATE flag. After each working node 102 detects the state change, it will perform a cleanup operation and exit safely. ACTIVE_NODES stores all currently active working node 101 instances and their health status information. The central control node 101 uses this collection to achieve unified management and health monitoring of the working nodes 102.
[0024] In this embodiment of the present invention, the risk control engine system 10 also supports dynamic system scaling and intelligent load balancing. Specifically, the central control node 101 can dynamically adjust the number and distribution of working nodes 102 based on system load and health status, automatically adding working nodes 102 during peak periods to improve processing capacity and reducing working nodes 102 during off-peak periods to optimize resource utilization. This flexible resource management mechanism enables the system to efficiently respond to changing business needs and provide continuous and stable risk control services.
[0025] Based on the above Figures 1 to 3 The system architecture shown in the figure, an embodiment of the present invention also provides a risk control engine system self-healing method, which is essentially a dual-track self-healing mechanism with a central control node and a working node in parallel. This mechanism can handle abnormalities of different degrees in a graded manner and realizes the coordination of the dual-track self-healing process to avoid duplication or conflict of self-healing operations.
[0026] See Figure 4 , Figure 4 This is a schematic flow chart of a self-healing method for a risk control engine system provided by an embodiment of the present invention. The method includes steps S401 to S404, which are described as follows: S401: Using the central control node and the working nodes to monitor the health status of the working nodes; S402: When an abnormality is detected, a self-healing strategy is determined based on the abnormality level; S403: Performing a self-healing operation on the working node based on the self-healing strategy; S404: When the same exception is detected by the working node and the central control node at the same time, it is determined whether the central control node or the working node performs the self-healing operation according to the priority of the processing method corresponding to the exception type.
[0027] In the self-healing method of the risk control engine system provided in an embodiment of the present invention, the central control node and the working node are used to monitor the health status of the working node at the same time. This dual-track monitoring mechanism improves the overall abnormality perception capability of the system; when an abnormality is detected, it is divided into levels according to the severity of the abnormality, and the corresponding self-healing operation is determined according to the preset graded abnormality self-healing strategy. Such hierarchical processing not only improves the response efficiency but also avoids waste of resources; and when the same abnormality is detected by the working node and the central control node at the same time, the preset abnormality type and processing method priority mechanism is used to determine which party should perform the self-healing operation, effectively avoiding conflicts or repeated processing caused by the simultaneous operation of the two, and also improving the stability and recovery capability of the system in complex abnormal scenarios. Compared with the traditional single monitoring and processing mode, it has higher reliability and adaptability.
[0028] It should be noted that Figure 1 There's no clear order of execution between step S404, S402, and S403. Specifically, after S403 completes, step S404 isn't executed immediately. Similarly, after S404 completes and the self-healing target is determined, the system can return to S403 to execute the corresponding self-healing operation based on that target. This flexible process design ensures the system can dynamically adjust the self-healing strategy based on actual conditions, thereby improving the flexibility and adaptability of the self-healing process.
[0029] Next, the embodiment of the present invention will introduce the process of achieving the above technical effects through the above self-healing process in conjunction with relevant drawings.
[0030] In step S401, the central control node and the working nodes are used to monitor the health status of the working nodes. This is essentially a dual-track parallel monitoring mode of central monitoring and node self-monitoring. The specific implementation method is: For the central control node, the central control node can regularly monitor the health status of all working nodes at a fixed frequency to promptly detect whether there are any abnormal conditions in the working nodes.
[0031] As an optional implementation, a central control node can send health check requests to each working node to verify its connection status, response time and basic functions. This inspection method can detect obvious anomalies such as node offline and response timeout.
[0032] For worker nodes, each can continuously monitor its local health status. Optionally, a separate monitoring agent can be deployed within the worker node to continuously collect detailed operational metrics, including but not limited to: memory usage, CPU load, garbage collection frequency, request response time, rule execution success rate, and other multi-dimensional health data.
[0033] Optionally, in this embodiment of the present invention, the working nodes can also send monitored indicators to the central control node through both regular reporting and exception-triggered reporting. The central control node integrates the indicators in the active reports received from the working nodes with its own health monitoring results for analysis, forming a more comprehensive and accurate node health assessment. This multi-dimensional health profile enables the system to identify potential problems that are difficult to detect with a single monitoring method, such as gradual anomalies such as slow memory leaks and gradually declining performance.
[0034] Optionally, in this embodiment of the present invention, when a worker node detects an abnormality (e.g., memory usage approaching a threshold), it can immediately send an early warning notification to the central control node, rather than waiting for the next scheduled report. This proactive early warning mechanism significantly reduces the latency of anomaly detection, enabling the system to take preventative measures before problems escalate.
[0035] Through the above implementation, the embodiments of the present invention have the following advantages over the single central monitoring mode of the system: Under the central monitoring mechanism, the system will perform a comprehensive health check on each working node to detect its connection status, resource usage, and processing capacity; under the node self-monitoring mechanism, each working node can continuously monitor its own health indicators, including memory allocation, garbage collection frequency, thread status, request queue length, etc., and actively take self-healing measures when an anomaly is detected. By establishing a dual monitoring mechanism of working nodes and central control nodes, the system can promptly and accurately identify various anomalies and trigger the corresponding level of self-healing strategy based on the type and severity of the anomaly, ensuring that the continuity of the overall service is not affected.
[0036] The following will describe the self-healing strategies after the working nodes and the central control node respectively detect an abnormality, referring to step S402.
[0037] In step S402, the self-healing strategy specifies whether the working node or the central control node will perform the self-healing operation. The working node or central control node determines the tasks assigned to it and the other node during the self-healing operation. Different abnormality levels correspond to different self-healing strategies. Next, we will first describe the abnormality level classification method used in this embodiment of the present invention.
[0038] In this embodiment of the present invention, anomaly levels can be divided into four levels: mild, moderate, severe, and extremely severe. Different levels of anomaly correspond to different self-healing strategies. The basis for the classification of anomaly levels is based on the consideration of the recovery capabilities of the working node itself.
[0039] For example, minor anomalies may include brief performance fluctuations, temporary connection interruptions, or slight memory usage. These types of anomalies are tolerable and self-recoverable for worker nodes within a short period of time, allowing them to recover on their own. Moderate anomalies, on the other hand, may involve prolonged resource overload, partial service response timeouts, or localized functional failures. These anomalies may exceed the self-healing capabilities of a single worker node. The worker node can attempt to recover on its own, and if self-healing is unsuccessful, the central control node intervenes. Severe anomalies, on the other hand, typically involve critical component crashes, network partitions, and other major issues that could lead to service interruptions or data loss. These anomalies typically require direct intervention from the central control node. Therefore, categorizing anomaly levels not only helps accurately match appropriate self-healing strategies but also effectively improves system robustness and operational efficiency in complex operating environments.
[0040] It should be noted that the abnormality level classification method provided in the embodiment of the present invention is a relative classification method, not an absolute classification method. For the above classification method, the embodiment of the present invention is based on the recovery capability of the working node itself. Of course, the classification can also be combined with other factors, which are not limited here.
[0041] In this embodiment of the present invention, the target for minor and moderate anomalies is the worker node itself, while the target for severe and extremely severe anomalies is the central control node. Therefore, when a worker node or central control node detects an anomaly, a corresponding self-healing strategy can be developed based on the following circumstances.
[0042] Case 1: Anomalies are detected by the working node itself.
[0043] In the embodiment of the present invention, if a working node detects an abnormality in itself, the self-healing strategy is as shown in steps a1 to a3, as follows: Step a1: Determine whether the operation target corresponding to the abnormal level is a working node or a central control node; If yes, go to step a2, otherwise go to step a3; Step a2: Determine that the working node performs local self-healing operations and reports the self-healing event to the central control node; Step a3: Determine that the working node sends an abnormality report to the central control node to instruct the central control node to perform a self-healing operation on the working node.
[0044] Taking minor anomalies as an example, if a working node detects a minor anomaly (such as a brief performance fluctuation, temporary connection interruption, or slight memory growth), the working node will perform local self-healing operations, such as resource optimization and connection retry, and report the self-healing event to the central control node without the need for central control node intervention. This approach reduces the burden on the central control node and improves the overall efficiency of the system.
[0045] Considering that moderate anomalies (such as persistent memory leaks and persistent CPU overloads) may require a period of time for the working node to recover, in order to ensure that no intervention occurs during the recovery period and that operation conflicts do not occur, the present invention provides the following solutions: Step b1: The working node applies to the central control node for a self-healing time window; In this embodiment of the present invention, a self-healing time window is used to suppress other self-healing operations on working nodes. Because the self-healing time window is requested from the central control node only after a working node self-detects an anomaly, the central control node may have already detected the anomaly and issued an intervention instruction before that. This could lead to operational conflicts between the central control node and the working nodes. Therefore, by setting a self-healing time window, only the working nodes are allowed to perform self-healing operations, while blocking intervention instructions from the central control node.
[0046] Step b2: After the working node successfully self-heals within the self-healing time window, it reports the self-healing success result to the central control node.
[0047] In the above implementation, the working nodes can also report abnormal conditions to the central control node. The central control node prepares the resources required for the self-healing of the working nodes and monitors the self-healing process of the working nodes throughout the process without directly intervening.
[0048] If a worker node self-healing fails or times out (i.e., it doesn't complete within the self-healing window), indicating that the worker node is incapable of recovering from the specified level of anomaly or recovering too slowly, the central control node will take over self-healing operations for the worker node and escalate the anomaly level. In this case, the central control node itself takes over self-healing operations, ensuring successful self-healing. This also allows the central control node to directly execute self-healing operations the next time such an anomaly occurs, improving the worker node's self-healing speed.
[0049] In this embodiment of the present invention, if a working node cannot recover from a serious anomaly (such as frequent crashes or persistent response timeouts), it can send an emergency report to the central control node while attempting to save critical status data. Upon receiving the report, the central control node immediately initiates high-level recovery procedures, such as restarting the node or preparing a replacement node to take over. For extremely serious anomalies (such as hardware failures or network partitions), the central control node can immediately create a replacement node, migrate the workload, and isolate the failed working node.
[0050] In one embodiment of the present invention, when a working node detects an anomaly through self-monitoring, it can also immediately report the anomaly type, severity, and measures taken to the central control node. The central control node maintains a global self-healing status table to record the anomaly status and self-healing progress of each node.
[0051] The above describes the process of using the working nodes to detect anomalies and formulate a self-healing strategy. Next, we will introduce the self-healing strategy after the central control node detects anomalies.
[0052] Case 2: The central control node detects an abnormality in the working node.
[0053] In this embodiment of the present invention, if the central control node detects that an abnormality has occurred in a working node, the self-healing strategy is as shown in steps c1 to c3, as follows: Step c1: Determine whether the operation target corresponding to the abnormal level is a working node or a central control node; If yes, go to step c2, otherwise go to step c3; Step c2: Determine that the central control node notifies the working node to perform a local self-healing operation; Step c3: Determine that the central control node performs a self-healing operation on the working node.
[0054] In this embodiment of the present invention, similar to the self-healing strategy for working nodes, when a central control node detects a minor or moderate anomaly, it can notify the working node to perform local self-healing operations. If a severe or extremely severe anomaly is detected, the central control node itself performs self-healing operations on the working node. The self-healing operations performed by the central control node in this case are consistent with those described above and will not be repeated here.
[0055] In an embodiment of the present invention, for serious abnormalities, the central control node can also determine whether the working nodes need to cooperate; if necessary, the central control node will notify the working nodes to perform collaborative operations, otherwise no processing will be performed. This can improve the self-healing efficiency and avoid excessive pressure on the central control node.
[0056] To facilitate understanding of the above implementation, the following uses three main scenarios: node offline, resource anomaly, or normal operation as examples to introduce various abnormality-specific self-healing operations: 1. Node offline situation: Node offline is a serious extreme abnormality. The central control node will first try to reestablish a connection with the abnormal working node. If the reconnection is successful, the node service is restored and the node registry is updated. If the reconnection fails, the central control node will automatically create a new replacement node and migrate the workload of the abnormal working node to the new replacement node. During the recovery / replacement process, the central control node will ensure that the request is routed correctly to maintain business continuity.
[0057] 2. Abnormal resource situation: In the embodiment of the present invention, resource abnormalities may include: memory abnormalities, CPU overload, disk I / O abnormalities and other abnormalities, and the working node may perform a local self-healing operation.
[0058] When the exception is a memory exception, an embodiment of the present invention provides a hierarchical processing strategy based on different abnormality levels, namely: determining the abnormality level; and performing self-healing operations on the working node based on the self-healing measures corresponding to the abnormality level.
[0059] For example, in the case of mild anomalies, garbage collection is triggered by the worker node to release unnecessary caches. In the case of moderate anomalies, memory defragmentation is performed by the worker node, and non-critical data is selectively unloaded. In the case of severe anomalies, the node process is restarted by the worker node while saving critical status data.
[0060] Of course, the above classification of abnormality levels is also a relative classification method, not an absolute one. Relevant technicians can flexibly define the classification method, which is not limited here.
[0061] When the exception is CPU overload, the worker node adjusts the processing thread priority to ensure that critical tasks are executed first; or the worker node temporarily lowers the processing priority of non-critical requests; the worker node can also be used to divert some requests to other nodes.
[0062] When the exception is a disk I / O exception, use the worker nodes to optimize the data access pattern and reduce disk operations; or move frequently accessed data to the memory cache; the worker nodes can also limit non-critical write operations.
[0063] 3. Nodes that are operating normally: their health status records can be updated in a timely manner, performance indicators can be collected for load balancing and capacity planning, and they will wait for the next monitoring.
[0064] The above situations illustrate the self-healing strategies after the working node and the central control node each detect an anomaly. According to the self-healing strategy, the working node or the central control node may perform a self-healing operation on the working node, ie, execute step S403.
[0065] In step S403, after each self-healing operation, the system performs a self-healing evaluation, comparing health indicators before and after the self-healing operation to determine whether the self-healing operation was successful. This evaluation mechanism may include comparing changes in key health indicators before and after the self-healing operation, such as memory usage, CPU load, and response time; monitoring the system's stable operating time and performance fluctuations after the self-healing operation; and analyzing the impact of the self-healing process on business continuity and user experience. If the self-healing operation is successful, the system updates the node's health status and logs the self-healing event. If the self-healing operation fails, the system will attempt a higher-level self-healing strategy or, after multiple failures, mark the node as requiring manual intervention.
[0066] The self-healing coordination mechanism described above avoids duplicate or conflicting self-healing operations. For example, when node self-monitoring detects that memory usage exceeds a threshold, it first attempts local optimization and simultaneously reports the abnormal status to the central control node. The central control node records this event and allows the node sufficient time to recover. If the node fails to recover, the processing level is escalated, triggering a higher-level self-healing strategy. This dual-track self-healing mechanism ensures the system can quickly detect and respond to various failure scenarios, significantly improving system stability and reliability in distributed environments. Furthermore, the entire process is essentially a progressive intervention strategy, starting with minimal intervention (on the working node) and escalating to higher-level intervention (on the central control node) only when necessary. This approach minimizes the impact on normal business operations while ensuring that problems are effectively resolved. It also avoids unnecessary central intervention and ensures that serious problems are addressed promptly, significantly improving the system's self-healing efficiency and reliability.
[0067] Next, the embodiment of the present invention will be described in detail how to avoid operation conflicts or duplications between the working nodes and the central control node, see step S404.
[0068] In step S404, considering that the self-healing mechanism of the working node and the intervention measures of the central control node may be triggered simultaneously, if the two operations are executed simultaneously, they may interfere with each other or cause resource waste. For example, the memory usage of a working node quickly rises to 85%, triggering the abnormality threshold. The node self-monitoring detects the memory abnormality and prepares to perform local self-healing operations (garbage collection, cache cleaning). At the same time, the central control node also detects the memory abnormality of the node during regular inspection and prepares to initiate central intervention (perhaps to replace the node or restart the node). At this time, the conflict is: if the working node is currently performing garbage collection and cache cleaning, this process may take several seconds, but the system has determined that the central control node should restart the node, which will interrupt the ongoing self-healing operation. This not only wastes the self-healing resources already invested, but may also cause the request being processed to fail.
[0069] To address the above issues, the present invention first sets different handling methods for different types of exceptions. Each handling method has a corresponding operation target (working node or central control node). Each handling method also has a preset priority, with a higher priority indicating a better handling method. Therefore, step S404 can be executed as follows: Step d1: Obtain the processing method with the highest priority preset for the exception type; Step d2: Determine whether the operation target is a central control node or a working node based on the processing mode; Step d3: The determined operation target performs a self-healing operation on the working node according to the processing method.
[0070] In this embodiment of the present invention, a single exception type can be handled in multiple ways, each of which can be executed by either a worker node or a central control node. Therefore, this embodiment of the present invention pre-defines a priority for each handling method. For example, for a memory leak, local optimization of a worker node typically takes priority; whereas, for an offline worker node, replacement by the central control node takes priority. Based on this priority, the optimal handling method can be determined, and the target action can be determined.
[0071] The above implementation method can avoid the conflict between the self-healing operation of the working node and the intervention of the central control node, ensure the success of self-healing, and improve the self-healing effect.
[0072] For a comprehensive understanding of the above implementations, see Figure 5 , Figure 5 An interaction diagram of a working node and a central control node provided by an embodiment of the present invention, from Figure 5 It can be seen that the embodiments of the present invention can ensure that the system can promptly detect and resolve various abnormal situations.
[0073] For example, in the risk control engine system of an e-commerce platform, a working node encountered the following problem when processing high-risk transaction assessments: When the working node executes a complex machine learning model, the node self-monitoring detects that the memory usage continues to rise, reaching the warning threshold of 85%; the node self-healing executor immediately triggers memory optimization: cleaning non-critical caches and performing garbage collection; at the same time, the node reports the memory abnormality warning status to the central control node; after optimization, the memory usage still rises to 90%, and the node upgrades the self-healing measures: performing memory defragmentation and uninstalling non-essential components; the central control node receives continuous abnormality reports and prepares a backup node in case takeover is needed; after the second level of optimization, the memory usage drops to 75%, and the node reports the self-healing success to the central control node; the central control node verifies that the node status has returned to normal, but continues to closely monitor the node for a period of time; the system records the entire self-healing event for subsequent analysis and optimization.
[0074] This multi-layered, coordinated self-healing mechanism not only significantly improves system reliability but also significantly reduces the need for manual intervention. Through intelligent fault detection, hierarchical processing, and automatic recovery, the system maintains stable operation in a variety of complex scenarios. Even in the face of challenges such as node failures, resource exhaustion, or network fluctuations, it can quickly self-heal and restore normal service status, avoiding operational conflicts and providing continuous and reliable support for risk control operations.
[0075] Next, the embodiment of the present invention can also demonstrate in detail the interactive process of the self-healing risk control engine system 10 in actual operation through the rule execution process, especially the automatic processing mechanism in key scenarios such as working node abnormalities and rule updates. Figure 6 , Figure 6 A schematic diagram of rule execution within the risk control engine system provided by an embodiment of the present invention.
[0076] In the initial state, the client's risk control request is routed to the first working node through the load balancer. The first working node processes request 1 and returns a risk control decision response. This is the basic process for the normal operation of the system.
[0077] When the first working node encounters an exception due to resource exhaustion while processing risk control request 2, the dual-track self-healing mechanism of this embodiment of the present invention is triggered. For example, the working node's self-monitoring mechanism first detects the abnormal state and attempts to repair it locally, while simultaneously reporting the abnormality to the central control node. Upon receiving the abnormality report, the central control node immediately assesses the severity of the abnormality and triggers the corresponding level of self-healing process.
[0078] At the same time, the load balancer receives the risk control request 3 from the client and intelligently routes the request to the healthy second working node based on the latest node health status information to ensure that business processing is not affected.
[0079] After the first working node completes its self-healing process, the central control node confirms its return to normal and updates the node availability information in the load balancer. The load balancer can then reroute requests to the first working node, restoring the system to a fully normal load distribution state. This automated health monitoring and self-healing mechanism significantly improves system availability and stability.
[0080] For rule updates, the central control node regularly checks the update status of the rule registry. When a new version of a rule is discovered, a detailed rolling update plan is developed. The update process uses a parallelized rolling strategy. First, the new rules are prepared and loaded on the second worker node. Simultaneously, traffic is temporarily redirected to the first worker node via a load balancer to ensure that the second worker node does not process requests while the new rules are being applied. After the second worker node completes the update, its traffic is restored, and the same update process is then repeated on the first worker node. This rolling update strategy ensures service continuity during the rule update process, and clients do not perceive any interruptions.
[0081] The entire diagram illustrates the high availability, self-healing capabilities, and smooth upgrades of the self-healing distributed risk control engine in complex scenarios. Through mechanisms such as health monitoring, intelligent routing, automatic recovery, and zero-downtime updates, the system ensures business continuity while enabling timely updates and application of risk control rules, providing efficient and reliable technical support for business risk control scenarios.
[0082] In summary, the risk control engine self-healing method provided by the embodiments of the present invention has the following advantages: It provides a dual-track self-healing architecture with a central control node and worker nodes running in parallel: the central control node regularly monitors the health of all worker nodes, while each worker node continuously monitors its local health status. When an anomaly occurs on a worker node, a corresponding self-healing strategy is triggered based on the anomaly level (minor, moderate, severe, or extremely severe), implementing progressive intervention, starting with minimal intervention (local self-healing by the worker node) and escalating to a higher level of intervention (self-healing by the central control node) only when necessary. This approach minimizes the impact on normal operations, avoids unnecessary intervention by the central control node, and improves self-healing efficiency. Furthermore, self-healing operations are coordinated between the central control node and worker nodes to avoid duplication or conflicts. This multi-level, collaborative self-healing mechanism not only significantly improves system reliability but also significantly reduces the need for manual intervention, providing continuous and reliable support for risk control operations.
[0083] See Figure 7 , Figure 7 The electronic device provided in an embodiment of the present invention includes a memory 701, a processor 702, and a communication interface 703. The memory 701, processor 702, and communication interface 703 are electrically connected to each other, directly or indirectly, to enable data transmission or interaction. For example, these components may be electrically connected to each other via one or more communication buses or signal lines.
[0084] Optionally, the bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0085] In the embodiment of the present invention, the processor 702 can be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiment of the present invention. A general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the embodiment of the present invention can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The software module can be located in the memory 701, and the processor 702 reads the program instructions in the memory 701 and completes the steps of the above method in combination with its hardware.
[0086] In an embodiment of the present invention, the memory 701 may be a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), or a volatile memory (Volatile Memory), such as RAM. The memory may also be any other medium that can be used to carry or store the desired program executable code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory in an embodiment of the present invention may also be a circuit or any other device that can implement a storage function, for storing instructions and / or data.
[0087] The memory 701 can be used to store software programs and modules, such as the instructions / modules of the risk control engine system 10 provided in the embodiment of the present invention. These can be stored in the memory 701 in the form of software or firmware or embedded in the operating system (OS) of the electronic device 70. The processor 702 executes the software programs and modules stored in the memory 701 to perform various functional applications and data processing. The communication interface 703 can be used for signaling or data communication with other node devices.
[0088] I understand. Figure 7 The structure shown is for illustration only. The electronic device 7 may also include Figure 7 More or fewer components than shown, or with Figure 4 Different configurations shown. Figure 7 The components shown may be implemented in hardware, software, or a combination thereof.
[0089] Based on the above embodiments, the present invention also provides a readable storage medium, which stores a computer program. When the computer program is executed by a computer, the computer executes the risk control engine system self-healing method provided in the above embodiments. For specific implementation, please refer to the method embodiment and will not be repeated here.
[0090] An embodiment of the present invention can also provide a computer program product for executing a self-healing method for a risk control engine system, including a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method in the previous method embodiment. For specific implementation, please refer to the method embodiment and will not be repeated here.
[0091] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, the indirect coupling or communication connection of the device or unit can be electrical, mechanical or other forms.
[0092] In addition, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the objectives of the embodiments of the present invention.
[0093] Furthermore, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0094] It should be noted that if the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, electronic device, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.
[0095] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A self-healing method for a risk control engine system, characterized in that: The risk control engine system includes a central control node and a working node; the method includes: Using the central control node and the working node to monitor the health status of the working node; When an abnormality is detected, a self-healing strategy is determined based on the abnormality level; Performing a self-healing operation on the working node based on the self-healing strategy; When the same exception is detected simultaneously by the working node and the central control node, it is determined whether the central control node or the working node performs the self-healing operation according to the priority of the processing method corresponding to the exception type.
2. The self-healing method of the risk control engine system according to claim 1, characterized in that: When an anomaly is detected, a self-healing strategy is determined based on the anomaly level, including: When an anomaly is detected by the working node, determining whether the operation target corresponding to the anomaly level is the working node or the central control node; If it is a working node, determining that the working node performs a local self-healing operation and reports a self-healing event to the central control node; If it is a central control node, it is determined that the working node sends an exception report to the central control node to instruct the central control node to perform a self-healing operation on the working node.
3. The self-healing method of the risk control engine system according to claim 2, characterized in that: When an anomaly is detected, a self-healing strategy is determined based on the anomaly level, including: The working node applies to the central control node for a self-healing time window; wherein the self-healing time window is used to suppress other self-healing operations on the working node; After the working node successfully self-heals within the self-healing time window, the working node reports a self-healing success result to the central control node.
4. The self-healing method of the risk control engine system according to claim 3, characterized in that: When an anomaly is detected, a self-healing strategy is determined based on the anomaly level, including: If the self-healing operation of the working node fails or times out, it is determined that the central control node will take over the self-healing operation of the working node and upgrade the abnormality level.
5. The self-healing method of the risk control engine system according to claim 1, characterized in that: When an anomaly is detected, a self-healing strategy is determined based on the anomaly level, including: When an anomaly is detected by the central control node, determining whether the operation target corresponding to the anomaly level is a working node or a central control node; If it is a working node, determining that the central control node notifies the working node to perform a local self-healing operation; If it is a central control node, it is determined that the central control node performs a self-healing operation on the working node.
6. The self-healing method of the risk control engine system according to claim 5, characterized in that: When an anomaly is detected, a self-healing strategy is determined based on the anomaly level, including: Determining, by the central control node according to the abnormality level, whether the working nodes need to cooperate; If so, it is determined that the central control node notifies the working node to perform the collaborative operation, otherwise no processing is performed.
7. The self-healing method of the risk control engine system according to claim 1, characterized in that: Determining whether the central control node or the working node performs the self-healing operation according to the priority of the processing method corresponding to the exception type includes: Obtaining the processing method with the highest priority preset for the exception type; Determining whether the operation target is the central control node or the working node according to the processing mode; The self-healing operation is performed on the working node according to the determined operation target and the processing method.
8. The self-healing method for a risk control engine system according to any one of claims 1 to 7, characterized in that: The method further comprises: When the abnormality is a memory abnormality, determining the degree of the abnormality; A self-healing operation is performed on the working node according to the self-healing measure corresponding to the abnormality degree.
9. A risk control engine system, characterized in that: The risk control engine system includes a central control node and a working node; the risk control engine system is used to execute the risk control engine system self-healing method according to any one of claims 1 to 8.
10. An electronic device, characterized in that: A risk control engine system as described in claim 9 is deployed.