Enterprise-level micro-service retry treatment method

By defining and managing RPC links in the microservice system, initializing and storing link information, dismantling and monitoring links, and dynamically adjusting retry strategies, the problem of unstable success rate of RPC calls among microservices is solved, and system stability and resource utilization efficiency are improved.

CN120128620APending Publication Date: 2025-06-10SHANDONG INSPUR DIGITAL BUSINESS TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510384380.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

RPC calls between microservices are unstable due to network jitter, delay, etc., resulting in unstable success rate, frequent timeouts or failures, triggering retry mechanisms, increasing system resource consumption, and may lead to request storms and service avalanche effects.

Method used

By defining and loading the RPC link of the service module, initializing and storing link information to the Redis database, disassembly links for management and monitoring, maintaining the success and failure counts of each call path, and determining whether to perform retry based on the predefined retry success rate threshold.

Benefits of technology

It improves the success rate of RPC calls, ensures the stability of the system link, reduces the consumption of system resources, and avoids request storms and service avalanche effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128620A_ABST
    Figure CN120128620A_ABST
Patent Text Reader

Abstract

The invention provides an enterprise-level micro-service retry treatment method, and belongs to the technical field of enterprise service data processing. The method comprises the following steps: defining RPC links of all modules of a service, and loading the modules adopting the RPC links in a service system according to a preset rule; initializing RPC link information of each module and pre-storing the RPC link information in a Redis database; a value corresponding to the RPC link information is set to be serialized JSON (JavaScript Object Notation) data; disassembling the RPC link to perform service management and monitoring; converting each disassembled calling path into an independent key value pair, and maintaining request success count and failure count in a time period; and monitoring a request path of a user requesting to obtain the target service module, and determining a retry request according to a predefined retry success rate threshold. By limiting the success rate of the single-node service retry request and reasonably distributing the time-out time of the whole link, the retry success rate is improved, and the stability of the system link is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of enterprise service data processing, and particularly relates to an enterprise-level microservice retry governance method. Background Art

[0002] In the current technical environment, with the rapid development of mobile Internet technology, traditional centralized applications have gradually been replaced by distributed microservice architectures. The microservice architecture, with its characteristics of independent deployment and independent operation, has greatly improved the flexibility and scalability of the system.

[0003] However, this architecture also brings new challenges, especially in the communication between microservices. Usually, the interaction between microservices relies on remote procedure calls (RPCs). However, due to network jitter, latency, and complex service links, the success rate of RPC calls is unstable, and timeouts or failures frequently occur, triggering the retry mechanism of upstream services. This not only increases the consumption of system resources but also may lead to request storms and service avalanche effects. Summary of the Invention

[0004] To solve at least one aspect of the technical problems in the background art, this application provides an enterprise-level microservice retry governance method, which can improve the retry success rate and ensure the stability of the system link by restricting the success rate of retry requests for single-node services and reasonably allocating the overall link timeout time.

[0005] The technical solution adopted by this application is as follows:

[0006] An embodiment of this application provides an enterprise-level microservice retry governance method, including:

[0007] Define the RPC links for each module of the business, and load the modules using RPC links in the business system according to preset rules;

[0008] Initialize the RPC link information for each module and store it in the Redis database in advance, where the cache Key corresponding to the link information is concatenated according to the rule of business scenario + "RPC" + business module interface path + "_C";

[0009] Set the corresponding value of the RPC link information as serialized JSON data, and the JSON data includes the expected total time consumption, call hierarchy, and call path;

[0010] Decompose the RPC link for business management and monitoring;

[0011] Convert each of the decomposed call paths into independent key-value pairs, and maintain the request success count and failure count within a time period;

[0012] Monitor the request path for the target business module in the user request, and determine the retry request according to the predefined retry success rate threshold.

[0013] According to the enterprise-level microservice retry governance method provided by the embodiments of the present application, first, define the RPC links of each module of the service, and load these modules using preset rules to ensure that each service call has a clear and unique identifier. Then, initialize the RPC link information of each module and store it in the Redis database, and splice the cache key using specific rules to improve data retrieval efficiency and management convenience. Use JSON data containing key information such as the expected total time consumption, call hierarchy, and path as the value of the link information, making this information structured and easy to process. Further, refine service management and monitoring by disassembling the RPC link, decomposing the complex link into multiple independent service nodes, facilitating more accurate performance monitoring and problem troubleshooting. Then, convert each call path into an independent key-value pair, and maintain the success count and failure count of the request within a certain period of time, thereby dynamically adjusting the retry policy, reducing ineffective retries, and saving system resources. Finally, monitor the request path of the target business module in the user request, and decide whether to perform a retry based on the predefined retry success rate threshold, effectively avoiding system overload and service avalanche effects caused by improper retries. In summary, the enterprise-level microservice retry governance method provided by the embodiments of the present application not only improves the success rate of RPC calls, but also ensures the stability of the system link and reduces the consumption of system resources. This method significantly enhances the overall stability and response efficiency of the system through fine control and intelligent monitoring of microservice communication, providing reliable technical support for enterprise-level applications.

[0014] According to an embodiment of the present application, the defining of the RPC links of each module of the service and the loading of the modules using RPC links in the service system according to preset rules are specifically as follows:

[0015] Set the interface path of the service module as the RPC request entry, and use the interface path as the unique code;

[0016] Initialize the relevant RPC link information to the Redis database based on the unique code.

[0017] According to an embodiment of the present application, the initializing of the RPC link information of each module and pre-storing it in the Redis database, where the cache key corresponding to the link information is spliced according to the rule of business scenario + "RPC" + service module interface path + "_C", specifically as follows:

[0018] When the system is started for the first time, generate a unique cache key according to the business scenario and interface path, and associate the cache key with the RPC link information of the corresponding module;

[0019] Store the serialized JSON data containing the expected total elapsed time, call hierarchy, and call path as the value of the cache key in the Redis database.

[0020] According to an embodiment of the present application, setting the corresponding value of the RPC link information as serialized JSON data, the JSON data including the expected total elapsed time, call hierarchy, and call path, specifically:

[0021] Construct a JSON object containing the expected total elapsed time, call hierarchy, and call path;

[0022] In the call path part of the JSON object, record the information of each layer of the caller and the service being called, and set the retry request success rate threshold for each service node;

[0023] Serialize the constructed JSON object into a string format and store it as the value under the corresponding cache key in the Redis database.

[0024] According to an embodiment of the present application, disassembling the RPC link for business management and monitoring, specifically:

[0025] Decompose the entire RPC link into multiple independent service nodes according to the call hierarchy and call path of the RPC link;

[0026] Generate a unique key-value pair for each service node, where the Key is composed of the business scenario, business module interface path, and service node identifier;

[0027] Store the request success and failure counts of each service node in the Redis queue and update the statistical data according to a preset period.

[0028] According to an embodiment of the present application, converting each of the disassembled call paths into independent key-value pairs and maintaining the request success count and failure count within a time period, specifically:

[0029] Create a queue in Redis based on each generated key, where each element in the queue is a serialized JSON string;

[0030] Update the latest statistical data in the queue at a preset time interval, add the latest count value to the head of the queue, and remove the data at the tail of the queue.

[0031] According to an embodiment of the present application, monitoring the request path for the user to request the target business module and determining the retry request according to the predefined retry success rate threshold, specifically:

[0032] Obtain the request path of the current user, and convert it into the actual RPC link key in combination with the business scenario type, and obtain the defined RPC link information from Redis through this key;

[0033] Determine whether the current request is a retry request. If it is a retry request, calculate the ratio of the number of failed requests to the number of successful requests of the corresponding service node within the historical time.

[0034] According to an embodiment of the present application, the determining whether the current request is a retry request. If it is a retry request, calculating the ratio of the number of failed requests to the number of successful requests of the corresponding service node within the historical time is specifically as follows:

[0035] If the calculated success rate is lower than the preset retry success rate threshold, reject this retry request to avoid invalid calls.

[0036] According to an embodiment of the present application, the method further includes:

[0037] Regularly generate and send a monitoring report, and the report content includes the success rate, failure rate, average response time and the number of retry requests of the service node;

[0038] When it is detected that the error rate of a specific service node increases, trigger an early warning mechanism, and switch to an alternative service node or increase resource allocation.

[0039] According to an embodiment of the present application, the method further includes:

[0040] When it is detected that a specific service node has a persistent high error rate and cannot be solved by switching to an alternative service node or increasing resource allocation, start a deep diagnosis process, and the deep diagnosis process includes collecting detailed log information and performing performance tests. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0042] Figure 1 It is a schematic flowchart of the enterprise-level microservice retry governance method provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] In order to more clearly illustrate the overall concept of the present application, the following will be described in detail by way of examples in combination with the drawings of the specification.

[0044] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, the present application may be practiced in other ways than those described herein. Therefore, the scope of protection of the present application is not limited by the specific embodiments disclosed below. It should be noted that, without conflict, the embodiments of the present application and the features in each embodiment may be combined with each other.

[0045] In the present application, unless otherwise clearly defined and limited, the first feature being "on" or "under" the second feature may be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0046] As Figure 1 shown, an enterprise-level microservice retry governance method provided by an embodiment of the present application includes:

[0047] Step 100: Define the RPC links of each module of the service, and load the modules using RPC links in the service system according to a preset rule.

[0048] Step 200: Initialize the RPC link information of each module and store it in the Redis database in advance, where the cache Key corresponding to the link information is spliced according to the rule of business scenario + "RPC" + business module interface path + "_C".

[0049] Step 300: Set the corresponding value of the RPC link information as serialized JSON data, and the JSON data includes the expected total time consumption, call hierarchy, and call path.

[0050] Step 400: Decompose the RPC link for service management and monitoring.

[0051] Step 500: Convert each decomposed call path into an independent key-value pair, and maintain the request success count and failure count within a time period.

[0052] Step 600: Monitor the request path for the user to request the target service module, and determine the retry request according to the predefined retry success rate threshold.

[0053] In step 100, it involves defining the RPC links for each module in the business system and loading them according to preset rules. By using the interface path as the unique identifier, it ensures that each service call has a clear entry point. This approach not only simplifies subsequent configuration and maintenance work but also improves the manageability and scalability of the system. In addition, adopting a unified rule to define and load these modules helps to establish a standardized service interaction process, thus enhancing development and operation efficiency.

[0054] This structured definition method makes the communication between microservices more transparent, facilitating monitoring and troubleshooting. It provides a clear way to manage and identify different RPC links, reducing the confusion caused by unclear or duplicate definitions. At the same time, it also lays a foundation for subsequent steps such as link information initialization, disassembly, and monitoring.

[0055] In step 200, the RPC link information of all modules is initialized and stored in the Redis database. A cache key (business scenario + "RPC" + business module interface path + "_C") is concatenated using specific rules. This can quickly access and update link information, greatly improving data retrieval efficiency. By leveraging a high-performance in-memory database like Redis, the latency caused by looking up link information can be significantly reduced, enhancing the system's response speed.

[0056] It improves the reliability and flexibility of the system. The fast data access capability supports the requirements of real-time link management and dynamic adjustment strategies, while ensuring stable operation in high-concurrency scenarios. In addition, based on the distributed characteristics of Redis, cross-node data sharing and synchronization can be easily achieved, further enhancing the system's scalability and fault tolerance.

[0057] In step 300, JSON data containing key information such as the estimated total duration, call hierarchy, and call path is serialized and stored as the value of the link information. This structured representation method is not only convenient for parsing and processing but also supports complex query requirements. In this way, the retry strategy can be flexibly adjusted, and the detailed description requirements for service links in different business scenarios can be met.

[0058] It provides strong data analysis support, making it possible to evaluate and optimize the performance of specific service nodes. This not only helps to improve the success rate of individual services but also enables the discovery of potential bottlenecks and the adoption of corresponding measures through in-depth analysis of the entire link, thereby enhancing the overall service quality.

[0059] In step 400, the entire RPC link is decomposed into multiple independent service nodes to facilitate refined monitoring and troubleshooting. By disassembling according to the call hierarchy and call path, the problem can be more accurately located, optimizing service governance. This method allows for independent management and monitoring of each service node, greatly enhancing the accuracy and efficiency of system operation and maintenance.

[0060] It can significantly improve the understanding and control capabilities of complex service links. The refined management means enable the system to quickly locate and solve problems without affecting other services, reducing maintenance costs while enhancing the stability and availability of the system.

[0061] In step 500, each call path is converted into independent key-value pairs, and queue records of the success and failure counts of requests are created in Redis. The purpose of this is to dynamically adjust the retry policy, reduce ineffective retries, and save system resources. By regularly updating the statistical data, more reasonable retry decisions can be made based on the latest information.

[0062] It can effectively avoid system overload and service avalanche effects caused by improper retries. The real-time monitoring and dynamic adjustment mechanism ensure the effective utilization of resources while improving the overall stability of the system. In addition, this method also supports personalized customization of retry policies to better adapt to different business requirements.

[0063] In step 600, the request paths of the target business modules of user requests are monitored, and whether to perform a retry is determined based on a predefined success rate threshold. This step avoids unnecessary retry operations through intelligent monitoring and judgment, protecting downstream nodes from the impact of invalid requests. Decisions made based on real-time data contribute to maintaining the efficient operation of the system.

[0064] By precisely controlling the retry behavior, not only can system overload be prevented, but also the user experience can be improved, reducing the waiting time caused by frequent retries. This method effectively balances the relationship between the retry mechanism and system load, providing a solid foundation for building robust enterprise-level applications.

[0065] According to the enterprise-level microservice retry governance method provided by the embodiments of the present application, first, define the RPC links of each module of the business, and load these modules according to preset rules to ensure that each service call has a clear and unique identifier. Then, initialize the RPC link information of each module and store it in the Redis database, and splice the cache keys using specific rules to improve data retrieval efficiency and management convenience. Use JSON data containing key information such as the expected total duration, call hierarchy, and path as the value of the link information, making this information structured and easy to process. Further, refine business management and monitoring by disassembling the RPC link, decomposing the complex link into multiple independent service nodes, facilitating more accurate performance monitoring and problem troubleshooting. Then, convert each call path into an independent key-value pair, and maintain the success count and failure count of requests within a certain period of time, thereby dynamically adjusting the retry policy, reducing ineffective retries, and saving system resources. Finally, monitor the request path of the target business module of the user request, and decide whether to perform a retry based on a predefined retry success rate threshold, effectively avoiding system overload and service avalanche effects caused by improper retries. In summary, the enterprise-level microservice retry governance method provided by the embodiments of the present application not only improves the success rate of RPC calls, but also ensures the stability of the system link and reduces the consumption of system resources. This method significantly enhances the overall stability and response efficiency of the system through fine control and intelligent monitoring of microservice communication, providing reliable technical support for enterprise-level applications.

[0066] In some embodiments of the present application, define the RPC links of each module of the business, and load the modules using RPC links in the business system according to preset rules, specifically:

[0067] Set the interface path of the business module as the RPC request entry, and use the interface path as the unique code;

[0068] Initialize the associated RPC link information to the Redis database based on the unique code.

[0069] Specifically, first, set the interface path of the business module as the RPC request entry and use it as the unique code. This approach provides a clear and unique identifier for each service call. In this way, the system can clearly identify and manage different RPC links, ensuring that each call has an accurate entry point. This method not only simplifies the configuration process but also improves the maintainability and scalability of the system, making it more convenient and efficient to add or modify services. In addition, initializing based on the unique code helps to establish a standardized service interaction process and reduce confusion caused by unclear or repeated definitions.

[0070] This structured definition method greatly improves the transparency and controllability of the system. Since each RPC link has a unique identifier, management and monitoring become more intuitive and easy to operate, facilitating quick location and problem-solving. At the same time, initializing the associated RPC link information into the Redis database enables efficient access and update, reduces the time delay in searching for link information, and enhances the system's response speed and overall performance. This method lays a solid foundation for subsequent steps such as link decomposition, monitoring, and dynamic adjustment strategies, further enhancing the flexibility and reliability of the system.

[0071] In some embodiments of the present application, the RPC link information of each module is initialized and pre-stored in the Redis database, where the cache Key corresponding to the link information is concatenated according to the rule of business scenario + "RPC" + business module interface path + "_C", specifically as follows:

[0072] When the system is started for the first time, a unique cache Key is generated according to the business scenario and interface path, and the cache Key is associated with the RPC link information of the corresponding module;

[0073] The serialized JSON data containing the expected total time consumption, call hierarchy, and call path is stored in the Redis database as the value of the cache Key.

[0074] Specifically, when the system is started for the first time, a unique cache Key is generated according to the business scenario and interface path, and the cache Key is associated with the RPC link information of the corresponding module. This process ensures that each RPC link information has a clear storage location in the Redis database. By adopting the concatenation rule of "business scenario + 'RPC' + business module interface path + '_C'", not only can cache Key conflicts be avoided, but also the ownership and usage of the link can be intuitively reflected. This design method makes the management and access of link information more efficient, and is also convenient for subsequent expansion of new business scenarios or service modules. In addition, storing the serialized JSON data containing the expected total time consumption, call hierarchy, and call path in the Redis as the value of the cache Key provides a structured description of the link information, facilitating quick parsing and dynamic adjustment.

[0075] This method significantly improves the response speed and scalability of the system. With the high-performance read and write capabilities of Redis, the system can quickly retrieve and update link information, thereby reducing the latency caused by searching for link information and enhancing the overall performance. At the same time, the design of serializing JSON data makes the link information highly flexible and can easily adapt to complex business requirement changes. In addition, through the strong association between the unique cache Key and the link information, the system can achieve precise link management, providing reliable data support for subsequent monitoring, retry strategy adjustment, and fault troubleshooting, further enhancing the stability and operation and maintenance efficiency of the system.

[0076] In some embodiments of the present application, the corresponding value of the RPC link information is set to serialized JSON data, and the JSON data includes the expected total elapsed time, call hierarchy, and call path. Specifically:

[0077] Construct a JSON object containing the expected total elapsed time, call hierarchy, and call path;

[0078] In the call path part of the JSON object, record the information of the calling party and the service being called at each layer, and set the retry request success rate threshold for each service node;

[0079] Serialize the constructed JSON object into a string format and store it as a value under the corresponding cache Key in the Redis database.

[0080] Specifically, constructing a JSON object containing the expected total elapsed time, call hierarchy, and call path provides a detailed structured description for each RPC link. During the construction process, special attention is paid to recording the information of the calling party and the service being called at each layer and setting the retry request success rate threshold for each service node. This method not only helps to clearly display the hierarchical structure of the entire link and the relationship between each node but also provides accurate data support for subsequent monitoring and fault troubleshooting. In this way, complex business scenarios can be flexibly handled to ensure that each service node can dynamically adjust its behavior according to the actual situation.

[0081] These details are stored as a serialized JSON string under the corresponding cache key in the Redis database, making data retrieval and update extremely efficient. This design allows the system to quickly obtain the required link information, reducing the search time and improving the response speed. In addition, due to the good readability and scalability of the JSON format, it is convenient for later maintenance and upgrade. More importantly, by recording the service call information at each layer and setting the retry request success rate threshold, the system can make more reasonable retry decisions based on real-time data, reduce the number of ineffective retries, optimize the resource utilization efficiency, and effectively avoid system overload and service avalanche effects caused by improper retries. This not only improves the stability and reliability of the system but also provides a consistent and efficient experience for users.

[0082] In some embodiments of the present application, the RPC link is disassembled for business management and monitoring, specifically as follows:

[0083] According to the call hierarchy and call path of the RPC link, the entire RPC link is decomposed into multiple independent service nodes;

[0084] A unique key-value pair is generated for each service node, where the Key is composed of the business scenario, business module interface path, and service node identifier;

[0085] The request success and failure counts of each service node are stored in the Redis queue, and the statistical data is updated according to a preset period.

[0086] Specifically, according to the call hierarchy and call path of the RPC link, the entire RPC link is decomposed into multiple independent service nodes. This process makes the complex link structure clear and manageable by refining the roles and responsibilities of each service node. Each service node represents a specific service call point. This method helps to accurately locate the performance bottleneck or fault location and take optimization measures for specific nodes. In addition, it also supports independent management and monitoring of each service node, improving the operation and maintenance efficiency and response speed of the system.

[0087] Generate a unique key-value pair for each service node (the key consists of the business scenario, the business module interface path, and the service node identifier), and use the Redis queue to store the success and failure counts of requests. This not only achieves efficient data management and access but also provides data support for dynamically adjusting the retry strategy. Update the statistical data according to a preset period to ensure that the system can make accurate judgments and decisions based on the latest information. This mechanism can effectively reduce ineffective retries, optimize the resource usage efficiency, and thus improve the overall service quality. At the same time, this method enhances the transparency and controllability of the system, enabling managers to monitor the service status in real time, discover and solve problems in a timely manner, and further ensure the stability and reliability of the system. Through such a design, enterprise-level applications can still maintain an efficient and stable running state in the face of high concurrency and complex business logics.

[0088] In some embodiments of the present application, each disassembled call path is converted into an independent key-value pair, and the request success count and failure count within a time period are maintained. Specifically:

[0089] Based on each generated key, create a queue in Redis, where each element in the queue is a serialized JSON string;

[0090] Update the latest statistical data in the queue at a preset time interval, add the latest count value to the head of the queue, and remove the data at the tail of the queue.

[0091] Specifically, based on each generated key, create a queue in Redis, and design each element in the queue as a serialized JSON string. This design method enables the success count and failure count of each request to be stored in a structured form, facilitating quick parsing and use. By storing the statistical data in the Redis queue, the system can efficiently record the running status of each service node, and at the same time utilize the high-performance read and write capabilities of Redis to ensure low latency for data update and retrieval. In addition, storing data in the form of serialized JSON strings not only enhances the readability and extensibility of the data but also supports future possible function expansions, such as adding new statistical dimensions or metrics.

[0092] Update the latest statistical data in the queue at preset time intervals, add the latest count value to the head of the queue, and remove the data at the tail of the queue. This sliding window mechanism can effectively reflect the service performance in the recent period. This method avoids the storage pressure caused by accumulating a large amount of data for a long time, and also ensures the timeliness and accuracy of statistical data. By dynamically maintaining the request success and failure counts within a time period, the system can make more reasonable retry decisions based on the latest data, reduce the occurrence of ineffective retries, and optimize the resource utilization efficiency. In addition, this mechanism also provides reliable data support for monitoring and warning, enabling managers to detect abnormal situations in a timely manner and take measures, thereby further improving the stability and reliability of the system.

[0093] In some embodiments of the present application, monitor the request path for the user to request the target business module, and determine the retry request according to the predefined retry success rate threshold. Specifically:

[0094] Obtain the request path of the current user, and combine it with the business scenario type to convert it into an actual RPC link Key. Through this Key, obtain the predefined RPC link information from Redis;

[0095] Determine whether the current request is a retry request. If it is a retry request, calculate the ratio of the number of failed requests to the number of successful requests of the corresponding service node in the historical time.

[0096] Specifically, obtain the request path of the current user, and combine it with the business scenario type to convert it into an actual RPC link Key. Through this Key, obtain the predefined RPC link information from Redis. This process realizes the accurate positioning of user requests and the rapid retrieval of link information. By combining the request path with the business scenario to generate a unique link Key, the system can accurately identify the target business module and its corresponding RPC link information. This method not only improves the efficiency of data retrieval, but also ensures the accuracy and consistency of link information, providing a reliable basis for subsequent retry decisions.

[0097] When it is determined that the current request is a retry request, the system will calculate the ratio of the number of failed requests to the number of successful requests of the corresponding service node in the historical time to evaluate whether it meets the predefined retry success rate threshold. This dynamic judgment mechanism based on historical statistical data can effectively avoid resource waste and service overload problems caused by blind retries. At the same time, through independent monitoring and analysis of each service node, the system can more accurately control the retry behavior, reduce the occurrence of ineffective retries, and thus improve the overall service quality. In addition, this method enhances the intelligence level of the system, making the retry strategy more flexible and adaptable, and capable of providing an efficient and stable user experience in different business scenarios.

[0098] In some embodiments of the present application, it is determined whether the current request is a retry request. If it is a retry request, the ratio of the number of failed requests to the number of successful requests of the corresponding service node within a historical time is calculated. Specifically:

[0099] If the calculated success rate is lower than the preset retry success rate threshold, this retry request is rejected to avoid invalid calls.

[0100] Specifically, the system determines whether the current retry request meets the preset retry success rate threshold by calculating the ratio of the number of failed requests to the number of successful requests of the corresponding service node within a historical time. This process is based on real-time statistical data for dynamic analysis, ensuring the accuracy and rationality of the decision-making. If the calculated success rate is lower than the preset threshold, this retry request is directly rejected, thus avoiding resource waste and increased service pressure caused by invalid calls. This method effectively prevents request storms or service avalanche effects caused by frequent retries through strict control of the success rate, improving the stability and reliability of the system.

[0101] This intelligent retry control mechanism significantly optimizes the utilization efficiency of system resources. By rejecting retry requests with low success rates, the system can preferentially allocate limited resources to requests that are more likely to succeed, thereby improving the overall service quality. In addition, this method also reduces the pressure on downstream services and the additional load caused by invalid calls, further ensuring the efficient operation of the system. At the same time, the dynamic evaluation based on historical data makes the retry strategy more scientific and flexible, capable of adapting to the requirements of different business scenarios and providing users with a more stable and smooth service experience.

[0102] In some embodiments of the present application, the method further includes:

[0103] Regularly generate and send monitoring reports, and the report content includes the success rate, failure rate, average response time, and number of retry requests of the service node;

[0104] When the error rate of a specific service node is detected to increase, trigger an early warning mechanism and switch to an alternative service node or increase resource allocation.

[0105] Specifically, the system comprehensively records the key performance indicators of each service node, including the success rate, failure rate, average response time, and number of retry requests, by regularly generating and sending monitoring reports. These data provide an operation and maintenance personnel with a clear view of the service running status, facilitating the analysis of the overall performance and potential problems of the system. By long-term tracking and comparison of these indicators, abnormal trends or performance bottlenecks can be detected in a timely manner, so as to take targeted optimization measures. In addition, the content of the monitoring report can also serve as an important basis for subsequent capacity planning and service governance, helping the system find a balance between scalability and stability.

[0106] When the error rate of a specific service node is detected to increase, the system can automatically trigger an early warning mechanism and switch to a standby service node or increase resource allocation according to a preset strategy. This proactive risk management method significantly improves the fault tolerance and availability of the system. Through rapid response and dynamic adjustment, the system can solve problems before or at the initial stage of a fault, avoiding affecting the user experience. At the same time, this method also enhances the elasticity of the system, enabling it to operate stably in scenarios of high concurrency or sudden traffic. Generally speaking, this design combining monitoring reports and early warning mechanisms not only improves the reliability of the system but also provides strong support for the efficient operation and continuous optimization of enterprises.

[0107] In some embodiments of the present application, the method further includes:

[0108] When it is detected that a specific service node has a persistent high error rate and cannot be solved by switching to a standby service node or increasing resource allocation, start a deep diagnosis process, which includes collecting detailed log information and performing performance tests.

[0109] Specifically, when it is detected that a specific service node has a persistent high error rate and the problem cannot be solved by conventional means (such as switching to a standby service node or increasing resource allocation), the system will start a deep diagnosis process. This process includes collecting detailed log information and performing performance tests to comprehensively analyze the root cause of the problem. By deeply mining the log data, potential code defects, configuration errors, or external dependency problems can be located; while performance tests can evaluate the performance of the service node under different load conditions, helping to identify performance bottlenecks or resource contention situations. This method not only provides a scientific basis for troubleshooting but also lays a foundation for subsequent optimization and repair work.

[0110] The deep diagnosis process significantly improves the system's problem-solving ability and stability. Through proactive intervention and comprehensive analysis, complex or hidden problems can be quickly found and solved, avoiding their long-term impact on the overall system. In addition, this mechanism also enhances the observability of the system, enabling the operation and maintenance team to better understand the operating status and potential risks of the service node. At the same time, the results of the deep diagnosis can be used as experience accumulation for improving system design and optimizing operation and maintenance strategies, thereby further enhancing the reliability and robustness of the system and providing more solid technical support for enterprise-level applications. What is not described in this application can be implemented by adopting or referring to existing technologies.

[0111] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments.

[0112] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and variations can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. An enterprise-level microservice retry management method, characterized in that: include: Define the RPC links of each business module, and load the modules using RPC links in the business system according to preset rules; Initialize the RPC link information of each module and store it in the Redis database in advance, where the cache key corresponding to the link information is spliced ​​according to the rule of business scenario + "RPC" + business module interface path + "_C"; Set the corresponding value of the RPC link information to serialized JSON data, where the JSON data includes the estimated total time consumption, call level and call path; Disassembling the RPC link for business management and monitoring; Convert each of the decomposed call paths into an independent key-value pair, and maintain the request success count and failure count within the time period; Monitor user requests to obtain the request path of the target business module, and determine the retry request based on the predefined retry success rate threshold.

2. The enterprise-level microservice retry management method according to claim 1 is characterized in that: The RPC links of each business module are defined, and the modules using the RPC links in the business system are loaded according to preset rules, specifically: Set the interface path of the business module as the RPC request entry, and use the interface path as the unique code; The associated RPC link information is initialized into the Redis database based on the unique code.

3. The enterprise-level microservice retry management method according to claim 1 is characterized in that: The RPC link information of each module is initialized and pre-stored in the Redis database, wherein the cache key corresponding to the link information is spliced ​​according to the rule of business scenario + "RPC" + business module interface path + "_C", specifically: When the system is started for the first time, a unique cache key is generated according to the business scenario and interface path, and the cache key is associated with the RPC link information of the corresponding module; The serialized JSON data including the expected total time consumption, call level and call path is used as the value of the cache Key and stored in the Redis database.

4. The enterprise-level microservice retry management method according to claim 1, characterized in that: The value corresponding to the RPC link information is set as serialized JSON data, and the JSON data includes the estimated total time consumption, the call level and the call path, specifically: Construct a JSON object containing the expected total time, call hierarchy, and call path; In the call path section of the JSON object, record the information of each layer of callers and service callees, and set the retry request success rate threshold for each service node; Serialize the constructed JSON object into a string format and store it as a value under the corresponding cache key in the Redis database.

5. The enterprise-level microservice retry management method according to claim 4 is characterized in that: The disassembling of the RPC link to perform service management and monitoring is specifically as follows: According to the call level and call path of the RPC link, the entire RPC link is decomposed into multiple independent service nodes; Generate a unique key-value pair for each service node, where the key consists of the business scenario, business module interface path, and service node identifier; The request success and failure counts of each service node are stored in the Redis queue, and the statistics are updated according to the preset period.

6. The enterprise-level microservice retry management method according to claim 1, characterized in that: The decomposed call paths are converted into independent key-value pairs, and the request success count and failure count within the time period are maintained, specifically: Based on each generated key, create a queue in Redis, where each element in the queue is a serialized JSON string; Update the latest statistics in the queue at a preset time interval, add the latest count value to the head of the queue, and remove the data at the tail of the queue.

7. The enterprise-level microservice retry management method according to claim 1, characterized in that: The monitoring user requests to obtain the request path of the target business module, and determine the retry request according to the predefined retry success rate threshold, specifically: Get the current user's request path, and convert it into the actual RPC link key based on the business scenario type. Use the key to get the defined RPC link information from Redis. Determine whether the current request is a retry request. If it is a retry request, calculate the ratio of the number of failed requests to the number of successful requests of the corresponding service node in the historical time.

8. The enterprise-level microservice retry management method according to claim 7, characterized in that: The determination of whether the current request is a retry request, if it is a retry request, then the ratio of the number of failed requests to the number of successful requests of the corresponding service node in the historical time is calculated, specifically: If the calculated success rate is lower than the preset retry success rate threshold, the retry request will be rejected to avoid invalid calls.

9. The enterprise-level microservice retry management method according to any one of claims 1 to 8, characterized in that: The method also includes: Generate and send monitoring reports regularly, including the success rate, failure rate, average response time and number of retry requests of service nodes; When an increase in the error rate of a specific service node is detected, an early warning mechanism is triggered, and a backup service node is switched or resource allocation is increased.

10. The enterprise-level microservice retry management method according to claim 9, characterized in that: The method also includes: When a persistent high error rate is detected in a specific service node and cannot be resolved by switching to a backup service node or increasing resource allocation, a deep diagnosis process is initiated, which includes collecting detailed log information and performing performance testing.

Citation Information

Cited By

  • Multi-protocol unified arrangement-oriented asynchronous notification and retry system and method

    CN121418490A