Operation and maintenance management and control method, system, storage medium, electronic device, and computer program product

By collecting multi-source monitoring data from the cloud service platform and conducting acceptance tests, target traffic control rules are obtained, enabling operation and maintenance management of cloud gateways and regional data centers. This solves the security and stability problems caused by unreasonable traffic control in the cloud service platform, and improves the security and stability of the system.

WO2026152878A1PCT designated stage Publication Date: 2026-07-23CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD
Filing Date
2025-11-21
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

The lack of a reasonable traffic control mechanism in cloud service platforms leads to service failures and stability issues, such as cloud service resource exhaustion, service layer cache penetration, service layer interface timeout response, and service circuit breakers.

Method used

Collect multi-source monitoring data, and obtain target traffic control rules by conducting various types of acceptance tests on predefined rules. Based on these rules, perform traffic analysis and operation and maintenance management on cloud gateways and regional data centers, including automatic expansion and contraction, adjustment of rate limiting thresholds, fault recovery and degradation control, etc.

Benefits of technology

It improves the security and stability of the cloud service platform, ensures the rationality and effectiveness of traffic control, and reduces system failures and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025136872_23072026_PF_FP_ABST
    Figure CN2025136872_23072026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of cloud services, and provides an operation and maintenance management and control method, a system, a storage medium, an electronic device, and a computer program product. The method is applied to a management and control service system. The management and control service system comprises: a cloud gateway and a plurality of regional data centers corresponding to the cloud gateway. The method comprises: acquiring multi-source monitoring data, wherein the multi-source monitoring data comprises gateway monitoring data of the cloud gateway and regional monitoring data of the plurality of regional data centers; performing traffic analysis on the multi-source monitoring data on the basis of a target traffic control rule to obtain an analysis result, wherein the target traffic control rule is obtained by performing a plurality of types of acceptance tests on a predefined rule; and performing operation and maintenance management and control on the cloud gateway and / or the plurality of regional data centers on the basis of the analysis result. The present disclosure solves the technical problem in the related art of low security and stability of the management and control service system caused by a lack of a traffic control mechanism or an unreasonable traffic control mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Operation and maintenance management methods, systems, storage media, electronic devices and computer program products Technical Field

[0001] This disclosure relates to the field of cloud service technology, and more specifically, to an operation and maintenance management method, system, storage medium, electronic device, and computer program product. Background Technology

[0002] In cloud service platform operation and maintenance (O&M) scenarios, a significant portion of service failures are caused by a lack of traffic control mechanisms or by inadequate traffic control mechanisms. For example, while single-user rate limiting may be implemented for some cloud firewalls during O&M, it may fail to restrict security token API calls. A surge in security token API requests could exhaust cloud service resources or even cause the cloud service platform to crash. Another example is the lack of rate limiting for access layer gateway interfaces on some cloud service consoles, or the ineffective rate limiting mechanisms implemented for access layer gateways. This could lead to service layer cache penetration, resulting in database unresponsiveness or service layer API timeouts. Furthermore, third-party service anomalies can cause service circuit breakers, impacting the secure and stable operation of the cloud service platform. Therefore, more effective O&M strategies are needed to address these situations.

[0003] Therefore, how to establish a reasonable and effective traffic control mechanism in the operation and maintenance management of cloud service platforms to improve their security and stability has become one of the important technical issues in related fields. Currently, no effective solution has been proposed to address these issues.

[0004] There is currently no effective solution to the above problems. Summary of the Invention

[0005] This disclosure provides an operation and maintenance management method, system, storage medium, electronic device, and computer program product to at least solve the technical problem of low security and stability of management service systems caused by the lack of traffic control mechanisms or unreasonable traffic control mechanisms in related technologies.

[0006] According to one aspect of the embodiments of this disclosure, an operation and maintenance management method is provided, applied to a management and control service system. The management and control service system includes a cloud gateway and multiple regional data centers corresponding to the cloud gateway. The operation and maintenance management method includes: collecting multi-source monitoring data, wherein the multi-source monitoring data includes gateway monitoring data of the cloud gateway and regional monitoring data of the multiple regional data centers; performing traffic analysis on the multi-source monitoring data based on target traffic control rules to obtain analysis results, wherein the target traffic control rules are obtained by performing various types of acceptance tests on predefined rules; and performing operation and maintenance management on the cloud gateway and / or multiple regional data centers based on the analysis results.

[0007] According to another aspect of the embodiments of this disclosure, an operation and maintenance management method is provided, comprising: obtaining an operation and maintenance management request through a first application programming interface, wherein the operation and maintenance management request is used to obtain multi-source monitoring data in a management and control service system; and returning an operation and maintenance management response through a second application programming interface, wherein the response data carried in the operation and maintenance management response includes: an operation and maintenance management report, wherein the operation and maintenance management report is obtained by performing operation and maintenance management on the cloud gateway and / or multiple regional data centers of the management and control service system through analysis results, and the analysis results are obtained according to the operation and maintenance management method described above.

[0008] According to another aspect of the present disclosure, an operation and maintenance management system is provided, including: an operation and maintenance management cloud component, a cloud gateway, and multiple regional data centers corresponding to the cloud gateway, wherein the operation and maintenance management cloud component is configured to provide operation and maintenance management cloud services to implement the operation and maintenance management method of any of the above.

[0009] According to another aspect of the present disclosure, an electronic device is provided, including: a memory storing an executable program; and a processor configured to run the program, wherein the program executes the operation and maintenance management method of any one of the above-mentioned methods when it runs.

[0010] According to another aspect of the present disclosure, a computer-readable storage medium is provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform any of the above-mentioned operation and maintenance management methods.

[0011] According to another aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the operation and maintenance management method described above.

[0012] In this embodiment of the disclosure, the management and control service system collects multi-source monitoring data, which includes: gateway monitoring data of the cloud gateway and regional monitoring data of multiple regional data centers; traffic analysis is performed on the multi-source monitoring data based on target traffic control rules to obtain analysis results, wherein the target traffic control rules are obtained by performing various types of acceptance tests on predefined rules; and operation and maintenance management of the cloud gateway and / or multiple regional data centers are performed based on the analysis results.

[0013] It is noteworthy that the target traffic control rules in this embodiment are obtained by conducting various types of acceptance tests on predefined rules. On the one hand, using predefined rules as the basis for target traffic control rules allows for the definition of specific traffic control rules for different application scenarios or different operation and maintenance management needs, resulting in strong flexibility and applicability. On the other hand, the various types of acceptance tests ensure the rationality and effectiveness of the target traffic control rules in managing the traffic of the management service system. Based on this, using target traffic control rules to perform traffic analysis on multi-source monitoring data corresponding to cloud gateways and multiple regional data centers in the management service system, and then performing operation and maintenance management based on the analysis results, can improve the security and stability of the management service system.

[0014] In other words, the embodiments disclosed herein achieve the purpose of using the target traffic control rules obtained from acceptance testing to perform operation and maintenance management of the management service system, thereby realizing the technical effect of enhancing the rationality of traffic control rule settings to improve the security and stability of the management service system (especially the cloud management service platform), and thus solving the technical problem of low security and stability of the management service system caused by the lack of traffic control mechanism or unreasonable traffic control mechanism in related technologies.

[0015] It is worth noting that the above general description and the following detailed description are merely for illustrative and explanatory purposes and do not constitute a limitation thereof. Attached Figure Description

[0016] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this disclosure, illustrate exemplary embodiments of the present disclosure and are used to explain the disclosure, but do not constitute an undue limitation of the disclosure. In the drawings:

[0017] Figure 1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing operation and maintenance management methods;

[0018] Figure 2 is a flowchart of an operation and maintenance management method according to an embodiment of the present disclosure;

[0019] Figure 3 is a schematic diagram of an optional rule-based acceptance test scheme according to an embodiment of the present disclosure;

[0020] Figure 4 is a schematic diagram of an optional development and production environment for a managed cloud service according to an embodiment of the present disclosure;

[0021] Figure 5 is a schematic diagram of the architecture of an optional management and control service system according to an embodiment of the present disclosure;

[0022] Figure 6 is a schematic diagram of an optional global flow control scheme according to an embodiment of the present disclosure;

[0023] Figure 7 is a schematic diagram of a node structure in an optional stand-alone flow control component according to an embodiment of the present disclosure;

[0024] Figure 8 is a schematic diagram of an optional problem discovery scheme according to an embodiment of the present disclosure;

[0025] Figure 9 is a schematic diagram of an optional problem location and problem-solving process according to an embodiment of the present disclosure;

[0026] Figure 10 is a schematic diagram of an optional flexible capacity operation and maintenance management scheme according to an embodiment of the present disclosure;

[0027] Figure 11 is a flowchart of another operation and maintenance management method according to an embodiment of the present disclosure;

[0028] Figure 12 is a structural block diagram of an operation and maintenance management device according to an embodiment of the present disclosure;

[0029] Figure 13 is a structural block diagram of another operation and maintenance management device according to an embodiment of the present disclosure;

[0030] Figure 14 is a structural block diagram of an electronic device according to an embodiment of the present disclosure. Detailed Implementation

[0031] To enable those skilled in the art to better understand the present disclosure, the technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present disclosure, and not all embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present disclosure.

[0032] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0033] First, some of the nouns or terms that appear in the description of the embodiments of this disclosure are to be interpreted as follows.

[0034] Flow control (also known as "flow regulation" or "flow control") refers to the technology used in a cloud computing environment to manage and regulate network traffic, requests, and resource usage of cloud services (or cloud applications). Flow control is used to balance the load on cloud service platforms, prevent resource overload, and ensure service stability.

[0035] Traffic control component: Enables traffic control for distributed, multi-language, and heterogeneous service architectures. It provides multi-dimensional traffic control capabilities within cloud service platforms, including traffic routing, traffic limiting, traffic shaping, circuit breaking and degradation, adaptive overload protection, and hotspot traffic protection.

[0036] A cloud gateway is a gateway located at the access layer of a cloud service platform. As the entry point for cloud services, the cloud gateway manages traffic, performs security verification, and routes requests, ensuring efficient and secure distribution of external requests to the backend services of the cloud service platform. The cloud gateway can also implement traffic control and parameter validation to maintain the system stability of the cloud service platform.

[0037] Cloud monitoring services refer to cloud services within a cloud service platform used for real-time collection, analysis, and visualization of cloud resource and application data. These services help operations and maintenance personnel quickly identify anomalies in the cloud service platform, provide early warnings and alerts, ensure the healthy operation of the cloud service platform, and optimize resource utilization.

[0038] Log cloud service refers to a cloud service platform that provides functions such as data collection, processing, analysis, querying, visualization, and alerting for log data (such as log files, trace files, metric files, etc.). Log services can enhance the digital capabilities of applications connected to the cloud platform.

[0039] According to the embodiments of this disclosure, an embodiment of an operation and maintenance management method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0040] The method embodiment provided in Embodiment 1 of this disclosure can be executed in a mobile terminal, computer terminal, or similar computing device. Figure 1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing an operation and maintenance management method. As shown in Figure 1, the computer terminal 10 (or mobile device) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) (processor 102 may include, but is not limited to, a microprocessor (MCU) or a field programmable gate array (FPGA), etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, the computer terminal 10 may also include: a display, an input / output interface, a Universal Serial Bus (USB) port (which may be included as one of the ports of a computer bus), a network interface, a cursor control device (such as a mouse, touchpad, etc.), a keyboard, a power supply, and / or a camera.

[0041] Those skilled in the art will understand that the structure shown in FIG1 is merely illustrative and does not limit the structure of the electronic device described above. For example, the computer terminal 10 may also include more or fewer components than shown in FIG1, or have a different configuration than that shown in FIG1.

[0042] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuitry are generally referred to herein as "data processing circuitry". This data processing circuitry may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in embodiments of this disclosure, the data processing circuitry serves as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0043] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the operation and maintenance management method in this embodiment of the present disclosure. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the above-mentioned operation and maintenance management method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0044] The transmission device 106 is used to connect to a network via a network interface to receive or send data. Specific examples of the network described above may include wired and / or wireless networks provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0045] The display shown in Figure 1 can be, for example, a touchscreen liquid crystal display (LCD), which allows the user to interact with the user interface of the computer terminal 10 (or mobile device).

[0046] Under the aforementioned operating environment, this disclosure provides an operation and maintenance management method as shown in Figure 2. Figure 2 is a flowchart of the operation and maintenance management method according to an embodiment of this disclosure. The above-mentioned operation and maintenance management method is applied to a management and control service system, which includes a cloud gateway and multiple regional data centers corresponding to the cloud gateway. As shown in Figure 2, the method may include the following steps S202 to S206.

[0047] Step S202: Collect multi-source monitoring data, which includes: gateway monitoring data of the cloud gateway and regional monitoring data of multiple regional data centers.

[0048] The aforementioned management and control service system can be one of the core components of a cloud service platform, and the management and control service system can be a cloud service within the cloud service platform.

[0049] The management and control service system can manage access to and use of resources on the cloud service platform, such as traffic management, permission verification, and resource scheduling. The management and control service system works in conjunction with other components of the cloud service platform (such as database services and middleware components) to ensure the stable operation of the cloud service platform.

[0050] In other words, the management and control service system is a component of the cloud service platform. Its main responsibility is to manage various services on the cloud platform (including but not limited to storage, computing, and networking), providing functions such as service creation, configuration, monitoring, and fault handling. The aforementioned operation and maintenance management methods are applied to the management and control service system to provide managed cloud services for the cloud service platform.

[0051] In particular, in terms of traffic control and elastic scaling, the management and control service system can work closely with cloud gateways and container services to achieve comprehensive protection and intelligent resource scheduling from the access layer to the service layer.

[0052] The aforementioned cloud gateway can serve as the gateway entry point for the access layer of a cloud service platform. The cloud gateway is used to process and manage all Hypertext Transfer Protocol (HTTP) traffic entering and leaving the cloud service platform. In addition to providing basic traffic routing, user authentication, and parameter validation functions, the cloud gateway also possesses traffic control and early warning capabilities, enabling it to monitor and manage global traffic and prevent traffic surges from impacting backend services.

[0053] Regional data centers correspond to the deployment locations of cloud service platforms in different geographical regions. Each regional data center can be a geographically distributed collection of computing, storage, and network resources established by a cloud service provider to provide services. Each regional data center includes independent service clusters (including multiple service units) and service resources, which are used to provide regional service support and data storage. By managing the setup of multiple regional data centers in the service system, considering the regionality of data, compliance, and reducing network requirements, efficient and stable operation and maintenance management services (such as traffic control services, performance monitoring services, fault detection services, fault recovery services, data processing services, etc.) can be provided to users in different regions.

[0054] Among the aforementioned multi-source monitoring data, the cloud gateway monitoring data reflects the real-time status and performance metrics of the cloud gateway. Gateway monitoring data can include: HTTP traffic metrics processed by the cloud gateway (e.g., total number of requests, request rate, response time, error rate, etc.), global interface traffic of the cloud gateway, and user traffic accessed by the cloud gateway (specifically, traffic data for each user). Through gateway monitoring data, the load on the cloud gateway can be monitored in real time, enabling traffic anomaly warnings, rate limiting control, and capacity adjustments to ensure the stability of the cloud gateway in processing access requests.

[0055] The aforementioned multi-source monitoring data includes regional monitoring data from multiple regional data centers. This data can encompass the usage of various resources within these data centers, such as processor resources, memory resources, and network bandwidth resources. Furthermore, regional monitoring data can also include service operation status information within the corresponding regional data centers, such as service response time, request success rate, and log information. By utilizing regional monitoring data, anomalies in resource usage and service operation can be identified in each regional data center, enabling targeted operation and maintenance management across multiple regional data centers. This ensures high availability and performance optimization of the management service system across geographical regions.

[0056] In application scenarios, the operation and maintenance management process can comprehensively and in real time collect (i.e. monitor) monitoring data of all levels of components within the management service system (including cloud gateways and multiple regional data centers, and may also include service units and service layer functional components in multiple regional data centers), providing a data foundation for subsequent traffic analysis.

[0057] Step S204: Perform traffic analysis on multi-source monitoring data based on the target traffic control rules to obtain analysis results. The target traffic control rules are obtained by performing various types of acceptance tests on predefined rules.

[0058] The aforementioned target traffic control rules characterize the traffic control mechanism of the management and control service system, preventing service failures caused by overload. These rules can be used to determine traffic limiting methods implemented across multiple dimensions (e.g., global, interface, user, and individual service dimensions). Traffic limiting methods can include threshold-based rate limiting and algorithmic rate limiting (e.g., fixed-time window algorithms, token bucket algorithms). By setting target traffic control rules, the management and control service system can quickly identify and respond to traffic surges (e.g., implementing system protection measures such as degradation and circuit breaking) to ensure system security and stability.

[0059] The aforementioned predefined rules can be determined by predefined rule parameters. These rule parameters are pre-set by operations and maintenance personnel or threads based on service characteristics, historical traffic data, service requirements, and performance test results within the management and control service system. Based on this, various types of acceptance tests are performed on the predefined rules to obtain the aforementioned target traffic control rules. These various types of acceptance tests may include performance acceptance testing, fault simulation testing, interface rate limiting testing, user rate limiting testing, and system performance stress testing.

[0060] In some embodiments, the predefined rules may include predefined rate limiting rules, degradation strategies, and circuit breaker conditions.

[0061] It should be noted that the various types of acceptance tests mentioned above can be integrated into the Continuous Integration and Continuous Deployment (CICD) pipeline of the management and control service system. During the development and deployment phases of the management and control service system, acceptance tests are performed on predefined rules or currently used traffic control rules to determine the target traffic control rules to be used. These target traffic control rules may include rate limiting thresholds, degradation strategies, circuit breaker conditions, etc., set for the management and control service system.

[0062] By using target traffic control rules obtained through various types of acceptance tests to perform traffic analysis on multi-source monitoring data, it is possible to ensure that the target traffic control rules remain available and effective under various possible conditions, thereby ensuring that the analysis results obtained from the traffic analysis have high reference value for operation and maintenance management.

[0063] Based on target traffic control rules, multi-source monitoring data is analyzed to assess the real-time traffic status of the management and control service system. Specifically, target traffic control rules can include multiple rate-limiting thresholds corresponding to multiple sets of traffic data from multi-source monitoring data (such as global traffic data of the cloud gateway, interface traffic data of the cloud gateway, and individual machine traffic data of multiple service machines in the regional data center). When analyzing multi-source monitoring data, it is determined whether multiple sets of traffic data are close to or exceed the corresponding rate-limiting thresholds to obtain the analysis results. For example, when the interface traffic of the cloud gateway in the multi-source monitoring data reaches the preset rate-limiting threshold in the target traffic control rules, the analysis results will include the traffic alarm data corresponding to the interface traffic of the cloud gateway.

[0064] It should be noted that during the analysis of multi-source monitoring data, it is also possible to locate traffic alarms appearing in the management and control service system. In other words, the analysis results can include the traffic alarm location results corresponding to the traffic alarm data.

[0065] Based on step S204 above, by analyzing multi-source monitoring data and evaluating the real-time traffic status of the management and control service system according to the target traffic control rules, the analysis results are obtained, which can provide a basis for decision-making in subsequent operation and maintenance management.

[0066] Step S206: Based on the analysis results, perform operation and maintenance management on the cloud gateway and / or multiple regional data centers.

[0067] Based on the analysis results, control targets can be identified from cloud gateways and multiple regional data centers. For example, a control target could be a cloud gateway, a specific regional data center, or certain individual service machines within a regional data center. Furthermore, the analysis results can be used to perform operational and maintenance control on these control targets. Control strategies during operation and maintenance control can include, but are not limited to: automatic scaling up, automatic scaling down, adjusting rate limiting thresholds, fault recovery, degradation control, and circuit breaker control.

[0068] For example, when analysis results show that service instances are insufficient to handle current or future traffic demands, the system will automatically add service instances (i.e., automatic scaling up) to increase processing capacity. When analysis results show that traffic has decreased and the current resource level is no longer needed, the system will automatically reduce service instances (i.e., automatic scaling down) to optimize costs. When analysis results indicate that the current rate limiting rules are no longer applicable, the rate limiting thresholds in the traffic control rules can be adjusted to adapt to the actual traffic situation. When analysis results indicate that abnormal traffic or service failures have been detected in the management service system, degradation strategies or fault recovery procedures can be initiated to maintain service availability and reduce the scope of the failure's impact.

[0069] Through steps S202 to S206 above, multi-source monitoring data is collected in the management and control service system. The multi-source monitoring data includes: gateway monitoring data of the cloud gateway and regional monitoring data of multiple regional data centers. Traffic analysis is performed on the multi-source monitoring data based on target traffic control rules to obtain analysis results. The target traffic control rules are obtained by conducting various types of acceptance tests on predefined rules. Based on the analysis results, operation and maintenance management of the cloud gateway and / or multiple regional data centers is carried out.

[0070] It is noteworthy that the target traffic control rules in this embodiment are obtained by conducting various types of acceptance tests on predefined rules. On the one hand, using predefined rules as the basis for target traffic control rules allows for the definition of specific traffic control rules for different application scenarios or different operation and maintenance management needs, resulting in strong flexibility and applicability. On the other hand, the various types of acceptance tests ensure the rationality and effectiveness of the target traffic control rules in managing the traffic of the management service system. Based on this, using target traffic control rules to perform traffic analysis on multi-source monitoring data corresponding to cloud gateways and multiple regional data centers in the management service system, and then performing operation and maintenance management based on the analysis results, can improve the security and stability of the management service system.

[0071] In other words, the embodiments disclosed herein achieve the purpose of using the target traffic control rules obtained from acceptance testing to perform operation and maintenance management of the management service system, thereby realizing the technical effect of enhancing the rationality of traffic control rule settings to improve the security and stability of the management service system (especially the cloud management service platform), and thus solving the technical problem of low security and stability of the management service system caused by the lack of traffic control mechanism or unreasonable traffic control mechanism in related technologies.

[0072] The following provides further explanation of other optional implementation steps included in the operation and maintenance management method provided in the embodiments of this disclosure.

[0073] In one optional embodiment, the operation and maintenance management method further includes the following steps:

[0074] Step S211: In response to the continuous integration deployment event corresponding to the management and control service system, use various types of test cases to conduct acceptance testing on the predefined rules and obtain the target test results;

[0075] Step S212: Based on the target test results, configure and publish the predefined rules to obtain the target traffic control rules.

[0076] The aforementioned continuous integration deployment events can serve as trigger events for the CICD pipeline. These continuous integration deployment events may include, but are not limited to: the management service system completing the development of new features or modifying code, predefined rules being adjusted or updated (such as adjusting rate limiting thresholds, updating degradation and circuit breaker policies, etc.), system health checks, and production environment anomalies.

[0077] In the application scenario, when the aforementioned continuous integration deployment event is detected, the CICD pipeline is triggered to deploy and release the management and control service system. In particular, in this embodiment of the disclosure, the process of acceptance testing of predefined rules is integrated into the CICD pipeline; that is, when the aforementioned continuous integration deployment event occurs, acceptance testing of predefined rules is also triggered.

[0078] Specifically, various types of test cases are used to automatically conduct acceptance tests on the predefined rules to obtain the aforementioned target test results. These target test results characterize whether the predefined rules can reasonably and effectively manage and control the current management service system, especially whether they can reasonably and effectively implement rate limiting. Furthermore, based on the target test results, the predefined rules are configured and published to obtain the aforementioned target traffic control rules.

[0079] It should be noted that the various types of test cases mentioned above are test cases used to verify the effectiveness of predefined rules. These test cases may include, but are not limited to, performance test cases, fault simulation test cases, high concurrency test cases, etc. Each test case is used to simulate specific use cases or abnormal situations to verify the performance and adaptability of predefined rules under different conditions.

[0080] It should be noted that the above-mentioned predefined rules can be operation and maintenance management rules predefined by the developers, traffic management rules used by the management service system before the start of this acceptance test, or rules updated and determined based on the rule parameters updated by the developers and the traffic management rules used by the management service system.

[0081] During the configuration release process described above, predefined rules that meet the testing conditions are adjusted and then released to the runtime environment of the management and control service system. In other words, after the effectiveness of the predefined rules is proven in the acceptance testing, these predefined rules are formally configured into the management and control service system to obtain the target traffic control rules.

[0082] The process of using multiple types of test cases to perform acceptance testing on predefined rules can also be achieved through manual testing and human-computer interaction verification synchronized with the CICD pipeline, or by using preset performance automation testing tools. This disclosure does not limit the specific implementation method described above.

[0083] Through steps S211 to S212 above, in this embodiment of the disclosure, the acceptance testing process is triggered by a continuous integration deployment event. After integrating the acceptance test into the CICD pipeline, the rule acceptance testing can be normalized and automated. Furthermore, the predefined rules that pass the acceptance test can be automatically configured and published as target traffic control rules, which can promote the collaborative work between the development and deployment phase and the operation and maintenance phase of the management and control service system, and improve the stability and response speed of the management and control service system.

[0084] In an optional embodiment, in step S211, acceptance testing of predefined rules is performed using multiple types of test cases to obtain target test results, including the following method steps:

[0085] Step S213: Perform test orchestration based on predefined rules and multiple types of test cases to obtain test instructions to be executed. Among them, the multiple types of test cases include: performance recovery test cases and fault simulation test cases.

[0086] Step S214: Execute test commands using the virtual test container of the management and control service system to obtain execution results;

[0087] Step S215: Generate the target test results based on the execution results and predefined rules.

[0088] The performance recovery test cases described above are used to verify the performance of the management and control service system under specific loads, ensuring that the system can maintain good response times and success rates even under high traffic. For example, performance recovery test cases can simulate 1.5 times the normal traffic as stress traffic to observe whether the various performance indicators (such as response time and throughput) of the management and control service system meet the expected conditions under stress traffic.

[0089] The above fault simulation test cases are used to simulate possible faults or abnormal situations (such as network interruption, database crash, unavailability of third-party services, etc.) in the management and control service system to verify whether the degradation strategy or fault recovery mechanism under the predefined rules can be executed correctly, and to ensure that the management and control service system can still provide basic services or recover quickly in the event of a fault.

[0090] In application scenarios, during the test orchestration process based on the aforementioned predefined rules and various types of test cases, the test cases and predefined rules are organized into an orderly test process, which is characterized by test instructions to be executed.

[0091] Specifically, the aforementioned acceptance testing process is integrated into the CICD pipeline, and the aforementioned test orchestration is implemented through a test framework (such as Cucumber). This test framework converts test cases written in natural language into machine-executable instructions (such as load test instructions for k6, fault injection instructions for ChaosBlade, etc.) and ensures that these instructions can be correctly executed according to predefined rules in a specific environment, thereby obtaining test instructions to be executed.

[0092] The virtual test container mentioned above can be a test container image. For example, the management and control service system mentioned above can be an Elastic Block Storage (EBS) cloud service. Correspondingly, the virtual test container mentioned above can be the "ebs-ctrl-test" container image in EBS.

[0093] Virtual test containers are used to execute test instructions generated from test orchestration in an isolated environment. They provide all the software environment and tools needed to run test cases, ensuring that test instructions are executed under conditions as consistent as possible with the production environment, thus avoiding test errors caused by environmental differences.

[0094] The execution results mentioned above can include: response time, success rate, and traffic data from performance regression tests, as well as service degradation and recovery during fault drill tests. These results can serve as an important basis for evaluating the effectiveness of predefined rules.

[0095] Based on the execution results of the test instructions, the predefined rules are evaluated to determine whether they meet expectations. For example, if the performance regression test results show that request latency and success rate meet expectations, it is considered that the rate limiting rules and system protection mechanisms in the predefined rules can operate effectively under high load. If the fault simulation test results show that the degradation circuit breaker strategy in the predefined rules can be correctly triggered and protect the system under preset fault scenarios, it is considered that the effectiveness of the degradation circuit breaker strategy has been verified. The above target test results are derived from the conclusion transformation of the above evaluation results of the predefined rules. That is, the target test results are used to determine which rules or strategies in the predefined rules have passed the test and which rules or strategies need further optimization or adjustment.

[0096] In addition, the results of the aforementioned target tests will be used to generate test reports, providing detailed feedback to the development team and operations personnel so that traffic control rules or system architecture can be adjusted as needed to ensure the stability and responsiveness of cloud management services.

[0097] In an exemplary application scenario, a rule-based acceptance testing scheme is provided as shown in Figure 3. As shown in Figure 3, the Continuous Integration Deployment (CICD) pipeline is integrated with performance regression testing and fault simulation testing. When the CICD pipeline is triggered, performance regression test cases and fault simulation test cases are input into the elastic control test container. In the elastic control test container, the orchestration component orchestrates predefined rules (input from the CICD pipeline), performance regression test cases, and fault simulation test cases to obtain test instructions; the execution component executes these test instructions to obtain execution results; based on the execution results and predefined rules, a report is generated, resulting in a test report corresponding to the target test results. This test report can be used for visualization.

[0098] It should be noted that during the development phase, reasonable performance metrics can be set for the interface, and the corresponding rate limiting rules can be verified through preset load test cases, providing a basis for further determining reasonable rate limiting thresholds. In some cases, backup links can be set up for core links in the management service system, and the effectiveness of the backup links can be verified through fault drill tests; circuit breaking strategies can also be designed for non-core links, and the effectiveness of the circuit breaking strategies can be verified through fault drill tests.

[0099] Through steps S213 to S215, this embodiment of the disclosure ensures comprehensive testing of the system under different loads and fault scenarios by orchestrating and executing various types of test cases, thereby improving the system's stability and robustness. Utilizing virtual test containers isolates the production environment, preventing testing activities from impacting actual users and protecting the purity of test data and processes. Test orchestration, execution, and result analysis are all automated, which not only improves testing efficiency and reduces human error but also allows for large-scale testing and frequent test iterations. The target test results can be fed back instantly to assist in rapid adjustments to predefined rules, improving the service quality and user experience of the management service system. Therefore, this embodiment of the disclosure enables comprehensive, efficient, and automated testing of the management service system, ensuring the effectiveness of target traffic control rules, thereby improving the stability and response speed of the management service system and reducing operational costs and system fault recovery time.

[0100] In an optional embodiment, the execution result includes interface performance data of the management service system during the execution of test instructions; in step S215, the target test result is generated based on the execution result and predefined rules, including the following method steps:

[0101] Step S2151: Obtain the interface performance metrics corresponding to the predefined rules;

[0102] Step S2152: Generate target test results based on interface performance data and interface performance indicators. The target test results are used to characterize whether the interface performance data meets the acceptance conditions corresponding to the interface performance indicators.

[0103] When test commands are executed in the virtual test container, the system collects and records a series of data metrics related to interface performance, resulting in the aforementioned interface performance data. Interface performance data is a key basis for evaluating the system's performance under specific load or fault scenarios.

[0104] Interface performance data may include, but is not limited to: response time, success rate, throughput, and resource utilization. Response time refers to the system's response time to a request, i.e., the time from receiving a request to returning a result. Success rate refers to the proportion of requests received within a specific time period that are successfully processed and return a normal response. Throughput refers to the number of requests the system can process per unit of time. Resource utilization can include CPU utilization, memory utilization, network bandwidth utilization, etc., reflecting the system's resource consumption when processing requests.

[0105] The aforementioned interface performance metrics can be performance standards related to predefined rules, used to measure whether the interface's performance meets expectations under specific load or failure scenarios. These performance metrics can serve as the basis for test case acceptance. Interface performance metrics typically include upper limits for response time, lower limits for success rate, upper limits for throughput, and upper limits for resource utilization. Correspondingly, the acceptance criteria for these performance metrics can include: interface response time not exceeding the upper limit, interface success rate not lower than the lower limit, interface throughput not exceeding the upper limit, and the corresponding resource utilization not exceeding the upper limit.

[0106] If the interface performance data meets the acceptance criteria, it means that the predefined rule acceptance test has passed; otherwise, it means that the predefined rule needs to be further analyzed and optimized.

[0107] The interface performance data is compared and analyzed with the interface performance metrics to generate target test results. Specifically, if the interface performance data meets the acceptance criteria corresponding to the interface performance metrics for multiple types of test cases, the target test result will indicate that the predefined rule has passed the acceptance test, and the predefined rule is considered reasonable and effective. If the interface performance data fails to meet the acceptance criteria for a certain type of test case, the target test result will indicate that the predefined rule has failed the acceptance test. In this case, the target test result may also include specific problem information from the management service system that caused the interface performance data to fail to meet the acceptance criteria (e.g., problem type, problem location, and suggestions for eliminating the problem).

[0108] Through steps S2151 to S2152, this embodiment of the disclosure compares the test results with the interface performance indicators corresponding to the predefined rules to ensure that the system can meet the predetermined performance requirements in actual operation, such as response time, success rate, and throughput. If the test results show that the system performance does not meet the acceptance criteria, the above-mentioned target test results can help maintenance personnel or developers quickly locate the problem, adjust the predefined rules, or optimize the system architecture to improve system performance. The target test results can reflect the performance of the predefined rules in the actual test environment, which helps to improve the accuracy of rule configuration and avoid the system overload caused by overly lenient rule settings, or the user experience being affected by overly strict rules. By monitoring whether the interface performance data meets the indicators, timely measures can be taken to avoid system failures, enhance system stability, and provide more reliable cloud management services. Therefore, this embodiment of the disclosure not only verifies the effectiveness of the predefined rules, but also ensures that the interface performance of the management service system meets the design standards, thereby improving the overall performance, stability, and reliability of the system and providing users with a better service experience.

[0109] In an optional embodiment, in step S212, the predefined rules are configured and published based on the target test results to obtain the target traffic control rules, including the following method steps:

[0110] Step S2121: In response to determining that the interface performance data meets the acceptance criteria based on the target test results, the predefined rules are configured and published to obtain the target traffic control rules.

[0111] In the deployment phase of the management and control service system, predefined rules are configured and published. Based on the actual operating environment of the management and control service system, the information to be configured is determined. According to the information to be configured, the predefined rule configuration is converted into a rule configuration file. The rule configuration file is then pushed and published to the corresponding service instance in the management and control service system to complete the configuration publication.

[0112] After the configuration is published, the predefined rules are converted into target traffic control rules, which are then applied in the actual operating environment of the management and control service system. These target traffic control rules are used to monitor and regulate system traffic in real time, ensuring stable system operation and preventing system failures caused by traffic surges or anomalies.

[0113] Furthermore, step S2121 above can be used as part of the training iteration process, allowing the target traffic control rule to be continuously optimized and iterated based on feedback information from the actual operating environment after the target traffic control rule is released, so as to adapt to the ever-changing service operation requirements and traffic patterns.

[0114] Through the aforementioned step S2121, in this embodiment of the present disclosure, if the target test results show that the interface performance data meets the acceptance criteria, the predefined rules will be configured and published. This means that the application of the rules in the actual environment will be more effective and applicable, avoiding system performance problems or failures caused by improper rule configuration. By conducting rigorous testing and verification before rule publication, the workload of manual intervention and fault recovery required by operation and maintenance personnel in the production environment is reduced, improving the efficiency of fault handling.

[0115] In one optional embodiment, the predefined rules include an initial rate limiting rule and an initial degradation rule, wherein the initial rate limiting rule is used to perform rate limiting control on the management service system during the execution of various types of test cases, and the initial degradation rule is used to perform single-machine degradation control on the management service system during the execution of various types of test cases.

[0116] The initial rate limiting and initial degradation rules described above are for acceptance testing. The initial rate limiting rule is determined based on a predefined rate limiting threshold. The initial degradation rule is determined based on predefined degradation parameters.

[0117] In application scenarios, a set of traffic control rules pre-defined during the development or rule design phase serves as the initial rate limiting rule. This rule limits the number of requests or resource usage within a specific time window to prevent system overload. The aforementioned initial degradation rule corresponds to the rate limiting rule. The initial degradation rule is used to proactively reduce certain functions or performance of a service when the system detects that resources are approaching a critical point or that a service is malfunctioning, in order to protect core services from being affected.

[0118] During the execution of various types of test cases, the management and control service system will implement rate limiting control on the cloud gateway and regional data center (including multiple service machines) within the system according to the initial rate limiting rules, and the management and control service system will implement single-machine degradation control on multiple service machines in the regional data center according to the initial degradation rules.

[0119] In an exemplary application scenario, the management and control service system is a management and control cloud service within a cloud service platform. The development and production environment of this management and control cloud service is shown in Figure 4. During the development phase, developers can define rules for the management and control cloud service. Specifically, based on the performance expectations of the management and control cloud service, interface performance metrics can be predefined, and initial rate limiting rules and initial degradation rules can be defined in conjunction with the management and control requirements of the management and control cloud service.

[0120] As shown in Figure 4, the predefined interface performance metrics, initial rate limiting rules, and initial degradation rules will undergo rule acceptance testing within the Continuous Integration and Deployment (CICD) pipeline. Specifically, during rule acceptance, performance regression testing and fault simulation testing will be performed on the initial rate limiting and initial degradation rules based on the interface performance metrics. It should be noted that the performance regression test cases and fault simulation test cases used during rule acceptance can be pre-built by the developers. If the predefined initial rate limiting and initial degradation rules pass the acceptance test, the deployment phase begins, and the rules are published.

[0121] As shown in Figure 4, during the deployment phase, rules are published based on the initial rate limiting rules and initial degradation rules that have passed the acceptance test. The resulting target traffic control rules may include: global rate limiting rules, interface rate limiting rules, user rate limiting rules, single-machine degradation rules, and single-machine protection rules.

[0122] Based on the above optional embodiments, this disclosure introduces predefined initial rate limiting rules and initial degradation rules, which not only improves the ability of the management and control service system to cope with high traffic and abnormal situations, but also provides support for system optimization and operation and maintenance, thereby improving the overall system stability and user experience.

[0123] In one optional embodiment, the target flow control rules include global control rules and single-machine control rules, and the analysis results include global monitoring results and single-machine monitoring results; in step S204, flow analysis is performed on multi-source monitoring data based on the target flow control rules to obtain analysis results, including the following method steps:

[0124] Step S241: Perform global traffic analysis on the gateway monitoring data based on global control rules to obtain global monitoring results;

[0125] Step S242: Perform single-machine traffic analysis on the area monitoring data based on the single-machine control rules to obtain the single-machine monitoring results.

[0126] The aforementioned global control rules can be traffic control rules applied to the entire management and control service system or the access layer cloud gateway. Global control rules are used to ensure, from a macroscopic perspective, that the inbound traffic level of the management and control service system remains within an acceptable and safe range. These global control rules can include time-window-based global rate limiting rules (including global rate limiting thresholds), user-level rate limiting rules (including user rate limiting thresholds), and interface-level rate limiting rules (including gateway interface rate limiting thresholds). Global control rules can be executed by the cloud gateway of the management and control service system. The cloud gateway is responsible for processing and managing all HTTP traffic entering the management and control service system, ensuring that the distribution and processing of traffic do not overload backend services.

[0127] The aforementioned single-machine control rules are used to constrain and manage the traffic control and resource allocation of individual service machines (such as a single service node or a single service instance) within the service system. These single-machine control rules can be executed by the corresponding single-machine flow control component (such as the Sentinel component) to prevent the service machine from crashing due to sudden traffic surges, while ensuring that high-priority interface calls can still be processed even under resource constraints. These single-machine control rules may include interface-level single-machine rate limiting rules (including single-machine interface rate limiting thresholds), overall single-machine rate limiting rules (including overall single-machine rate limiting thresholds), and system protection policies.

[0128] The aforementioned global monitoring results are assessments of global traffic conditions obtained through real-time analysis of global control rules and gateway monitoring data. These results indicate whether the management and control service system is currently experiencing traffic peaks, whether there are abnormal requests exceeding preset global control rules, and whether rate limiting or degradation policies have been triggered. Through these global monitoring results, operations and maintenance personnel can promptly understand the overall traffic situation of the system and determine whether adjustments to global control rules or other operational measures are necessary.

[0129] The aforementioned single-machine monitoring results are assessments of the traffic status of a service machine obtained through real-time analysis of regional monitoring data based on single-machine control rules. Single-machine monitoring results characterize the performance of a specific service machine (such as a server node or service instance) when processing requests, and whether single-machine rate limiting rules or system protection policies have been triggered. Single-machine monitoring results help to clearly define the load distribution and resource utilization efficiency within the management service system, and aid in identifying bottleneck nodes and optimizing resource allocation.

[0130] Through steps S241 to S242, this embodiment of the present disclosure achieves real-time traffic monitoring and analysis at both global and single-machine levels. The global and single-machine monitoring results provide real-time traffic status feedback, helping maintenance personnel to promptly detect traffic anomalies and resource shortages, thereby enabling rapid response and preventing system overload or failure. The implementation of target traffic control rules in the management service system ensures that traffic is reasonably allocated and controlled within safe limits. In particular, the combined use of global and single-machine control rules considers both the overall traffic balance of the management service system and the resource protection of individual service machines, effectively preventing service interruptions caused by traffic surges. Through the linkage between real-time analysis results and target traffic control rules, the management service system can also automatically adjust traffic processing strategies (such as automatic scaling up and down, dynamic adjustment of rate limiting thresholds, etc.) to cope with constantly changing traffic demands, demonstrating strong adaptability and stability. As described above, by monitoring and analyzing the effect of traffic control in real time and combining it with the execution of target traffic control rules, this embodiment of the present disclosure not only enhances the management and control service system's capabilities in traffic control and resource protection, but also improves the system's adaptability and operational efficiency, ensuring service stability and efficient resource utilization.

[0131] In an optional embodiment, the global monitoring results include global early warning results; in step S241, global traffic analysis is performed on the gateway monitoring data based on global control rules to obtain global monitoring results, including the following method steps:

[0132] Step S2411: Determine the global rate limiting threshold, gateway interface rate limiting threshold, and user rate limiting threshold according to the global control rules;

[0133] Step S2412: Based on the global rate limiting threshold, the gateway interface rate limiting threshold, and the user rate limiting threshold, perform rate limiting early warning analysis on the gateway monitoring data to obtain the global early warning result. The global early warning result is used to identify the location of the rate limiting control in the management and control service system.

[0134] The aforementioned global early warning results are derived from rate limiting early warning analysis based on global control rules and real-time monitoring data. These results can be used to identify which parts of the management and control service system or at which points in time may require rate limiting to cope with upcoming traffic peaks or abnormal situations. Global early warning results can include warning information determined based on global rate limiting thresholds, gateway interface rate limiting thresholds, and user rate limiting thresholds, helping operations personnel or the management and control service system automatically identify urgent traffic control needs and providing a basis for taking rate limiting measures.

[0135] In application scenarios, operations and maintenance personnel or management and control service systems will calculate a global rate limiting threshold based on global control rules, combined with historical traffic data, business needs, and system performance. This global rate limiting threshold can be regarded as the upper limit of the total traffic that the management and control service system can safely handle.

[0136] Based on the aforementioned global control rules, gateway interface rate limiting thresholds can also be determined. These thresholds can include multiple rate limiting values ​​set for specific interfaces processed by the cloud gateway. The gateway interface rate limiting thresholds take into account the interface characteristics, historical usage, and resource consumption of the Application Programming Interface (API), ensuring that abnormal traffic from a single API does not lead to a degradation in overall system performance.

[0137] Based on the aforementioned global control rules, user rate limiting thresholds can also be determined. User rate limiting thresholds can be traffic limits set for specific users or user groups. User rate limiting thresholds can be used to prevent abnormal traffic from a single user or user group from affecting the service experience of other users, helping to balance resource usage among different users.

[0138] In application scenarios, gateway monitoring data is compared with the aforementioned global rate limiting thresholds, gateway interface rate limiting thresholds, and user rate limiting thresholds to detect whether an upcoming traffic peak may exceed these thresholds. For example, the system obtains the global predicted value, gateway interface predicted value, and user traffic predicted value for the peak traffic period, determines whether the global predicted value is close to or exceeds the aforementioned global rate limiting threshold, determines whether the gateway interface predicted value is close to or exceeds the aforementioned gateway interface rate limiting threshold, and determines whether the user traffic predicted value is close to or exceeds the aforementioned user rate limiting threshold. Specifically, the above rate limiting early warning analysis not only considers the volume of traffic but also assesses traffic trends and patterns.

[0139] After analyzing the rate limiting warnings, the global warning results will identify some APIs, users or user groups, and system components or parts in the management service system that may require rate limiting measures, thus pinpointing the locations to be subject to rate limiting control. These global warning results provide clear operational instructions for maintenance personnel or tools, helping them to adjust rate limiting configurations or activate rate limiting contingency plans to protect the system from traffic surges.

[0140] In an exemplary application scenario, the architecture of the management and control service system is shown in Figure 5. The access layer of the management and control service system includes a cloud gateway, for which global rate limiting, interface rate limiting, and user rate limiting are set. Applications connected to the cloud gateway are referred to as unified management and control applications, and the gateway monitoring data corresponding to the cloud gateway is real-time application traffic data for these unified management and control applications. During the operation and maintenance phase, the cloud gateway performs rate limiting early warning analysis on the gateway monitoring data based on the global rate limiting thresholds, gateway interface rate limiting thresholds, and user rate limiting thresholds determined by the global control rules in the target traffic control rules. This yields global early warning results, thereby determining whether there are rate limiting early warnings at the cloud gateway at the global, interface, or user level.

[0141] Specifically, the global flow control scheme corresponding to the aforementioned cloud gateway-based global flow control mechanism is shown in Figure 6. The user (or the user through an application) initiates a request to the cloud gateway, which receives and processes the request. Further, the cloud gateway performs a fixed-window traffic check to ensure the current traffic is within acceptable limits. If the traffic exceeds a threshold (determined by global flow control rules), corresponding flow control measures may be triggered. If the traffic check passes, the request is sent to the local database for processing. Further, based on the processing result from the local database, a synchronous flow control value is determined. When the current flow control value is less than 70%, the asynchronous aggregation process is initiated. Asynchronous aggregation can handle lower traffic data, improving system response speed. When the current flow control value is greater than or equal to 70% and less than 85%, the asynchronous storage process is initiated. Asynchronous storage can handle medium traffic data without affecting system performance. When the current flow control value is greater than or equal to 85%, the synchronous storage process is initiated. Synchronous storage ensures that high traffic data can be processed and stored immediately, guaranteeing system stability and reliability. Ultimately, all data involved in the global flow control process will be stored in a key-value storage system (such as Redis) to ensure data consistency and durability.

[0142] Applying the method provided in this disclosure to the global flow control process shown in Figure 6, API rate limiting implemented through global control rules can prevent a large amount of interface traffic from overwhelming backend services. Similarly, user rate limiting implemented through global control rules can prevent a large amount of user request traffic from overwhelming backend services. In other words, this disclosure provides stability to cloud gateway traffic control, thereby improving the security and stability of the entire system.

[0143] Through steps S2411 to S2412 above, this embodiment of the disclosure enables proactive traffic management at the global level of the management and control service system. That is, the management and control service system identifies areas requiring rate limiting before traffic peaks arrive, ensuring the timeliness and effectiveness of traffic control measures. Precise traffic warnings prevent over- or under-allocation of resources, ensuring that the management and control service system can allocate resources reasonably during traffic surges, prioritizing requests from core businesses or high-priority APIs. Global warning results help maintain the stable operation of the entire system during periods of abnormal or peak traffic, avoiding service interruptions or performance degradation due to improper traffic control, and enhancing the overall stability of the system. By pre-identifying global rate limiting needs, measures can be taken to prevent or mitigate the impact of traffic surges on the entire system, thereby reducing the occurrence of failures and improving service availability and user satisfaction. Furthermore, global warning results can be identified and responded to by automated systems, helping to achieve automatic rate limiting or automatic capacity expansion during the operation and maintenance phase, reducing the operational burden and improving operational efficiency.

[0144] In an optional embodiment, the single-machine monitoring results include single-machine flow limiting warning results and single-machine degradation warning results; in step S242, single-machine traffic analysis is performed on the area monitoring data based on single-machine control rules to obtain single-machine monitoring results, including the following method steps:

[0145] Step S2421: Determine the single-machine current limiting threshold, single-machine interface current limiting threshold, and single-machine degradation threshold according to the single-machine control rules;

[0146] Step S2422: Based on the single-machine rate limiting threshold and the single-machine interface rate limiting threshold, perform rate limiting early warning analysis on the regional monitoring data to obtain the single-machine rate limiting early warning result. The single-machine rate limiting early warning result is used to identify the service single machine to be rate-limited and controlled in multiple regional data centers.

[0147] Step S2423: Based on the single-machine degradation threshold, perform degradation early warning analysis on the regional monitoring data to obtain the single-machine degradation early warning result. The single-machine degradation early warning result is used to identify the service single machines in multiple regional data centers that need to be degraded and managed.

[0148] The above-mentioned single-machine rate limiting early warning results are obtained by analyzing regional monitoring data based on single-machine control rules. These results identify which service machines in multiple regional data centers may require rate limiting control; in other words, they determine the service machines to be subject to rate limiting. The results may also include real-time traffic statistics for the service machine and a comparative analysis with the single-machine rate limiting threshold. This helps operations personnel or automated operations tools identify service machines whose traffic is approaching or exceeding the threshold, allowing for proactive rate limiting measures to prevent overload.

[0149] The aforementioned single-machine degradation warning results focus more on the resource utilization and impact of abnormal services on individual service machines within a regional data center. Degradation warning analysis identifies which individual service machines in a regional data center have resource utilization rates approaching or reaching preset degradation thresholds, or pose potential failure risks, resulting in the aforementioned single-machine degradation warning results. Single-machine rate limiting warning results are used to identify which individual service machines in multiple regional data centers may require degradation control; that is, to determine the individual service machines to be degraded.

[0150] In application scenarios, operations and maintenance personnel or automated operations and maintenance tools will calculate the single-machine rate limiting threshold, single-machine interface rate limiting threshold, and single-machine degradation threshold based on single-machine control rules, combined with historical data, business requirements, and system performance.

[0151] The aforementioned single-machine rate limiting threshold can be a traffic limit set for a single service machine. The single-machine rate limiting threshold is used to ensure that the traffic level of a single service machine (such as a single server node or service instance) does not exceed the processing capacity of that single service machine, thereby avoiding failures in the management service system caused by local overload.

[0152] The aforementioned single-machine interface rate limiting threshold can be a rate limiting threshold set for a specific API on a single service machine. The single-machine interface rate limiting threshold is used to prevent a surge in traffic to a specific API from affecting the stable operation of the entire single service machine or the entire management and control service system.

[0153] The aforementioned single-machine degradation threshold can be a resource usage or traffic threshold set for a single service machine under abnormal conditions. When the threshold is approached or reached, the management and control service system will automatically trigger degradation measures for that single service machine.

[0154] Based on this, by using single-machine rate limiting thresholds and single-machine interface rate limiting thresholds, rate limiting early warning analysis is performed on regional monitoring data to identify service machines whose traffic is about to exceed the thresholds, and single-machine rate limiting early warning results are obtained. By using single-machine degradation thresholds, degradation early warning analysis is performed on regional monitoring data to identify service machines whose resource utilization is close to or has reached the threshold, or which have a risk of failure, and single-machine degradation early warning results are obtained.

[0155] As shown in Figure 5, the service layer of the management and control service system (showing a regional data center in Figure 5) includes a unified access component and a service discovery component. The unified access component processes traffic data forwarded from the cloud gateway, while the service discovery component receives and processes application traffic data from independent applications (i.e., applications that directly access the service layer without going through the cloud gateway). The service layer can perform load balancing control based on the unified access component and the service discovery component, forwarding and distributing the traffic data to be processed to multiple service machines. Each service machine corresponds to a single-machine flow control component. Figure 5 exemplarily shows two single-machine flow control components: single-machine flow control component 1 is associated with the service machine corresponding to the unified access component, and single-machine flow control component 2 is associated with the service machine corresponding to the service discovery component.

[0156] As shown in Figure 5, for each standalone flow control component, interface rate limiting, standalone rate limiting (also known as standalone total rate limiting), degradation circuit breaking, and system protection can be implemented. The regional monitoring data corresponding to the multiple regional data centers mentioned above can include traffic data related to the unified access component, service discovery component, and each standalone flow control group within each regional data center. Based on the standalone rate limiting threshold, standalone interface rate limiting threshold, and standalone degradation threshold determined by the standalone control rules in the target flow control rules, rate limiting early warning analysis is performed on the regional monitoring data to obtain the aforementioned standalone monitoring results. During the operation and maintenance phase, based on the standalone rate limiting early warning results in the standalone monitoring results, the service standalone to be rate-limited and managed can be identified. This service standalone to be rate-limited and managed can be the service standalone associated with the standalone flow control component that has issued a rate limiting early warning. Based on the standalone degradation early warning results in the standalone monitoring results, the service standalone to be degraded and managed can be identified.

[0157] Based on the architecture of the management and control service system shown in Figure 5, the operation and maintenance management method provided in this embodiment involves two levels of traffic control: global flow control based on the cloud gateway and single-machine flow control based on the regional data center. Global flow control typically prioritizes user rate limiting and API rate limiting. To protect services from being overwhelmed by traffic, single-machine flow control prioritizes API rate limiting and total single-machine rate limiting; to prevent single-machine services from being dragged down by third-party applications or anomalies, single-machine flow control also needs to have the capabilities of degradation circuit breaking and system protection.

[0158] In application scenarios, traditional standalone flow control components may include a Statistic Slot, a System Slot, a Flow Slot, and a Degrade Slot. These traditional standalone flow control components lack rate limiting / degrade warning functionality. Therefore, this embodiment extends the functionality of the standalone flow control component. Figure 7 is a schematic diagram of the node structure in an extended standalone flow control component. As shown in Figure 7, a rate limiting / degrade warning node is constructed using a custom node (My Custom Slot), and this node is positioned between the System Slot and the Flow Slot. Based on this, in the management service system shown in Figure 5, at least some of the multiple standalone flow control components can adopt the aforementioned extended standalone flow control component, enabling the service layer to have rate limiting / degrade warning functionality for individual service machines.

[0159] Through steps S2421 to S2423 described above, this embodiment of the present disclosure achieves precise traffic control at the single-machine level. Specifically, by generating single-machine rate limiting and degradation warning results, the management service system allows for targeted rate limiting or degradation control of individual service machines, avoiding unnecessary resource waste and service interruptions. Through single-machine rate limiting warning analysis, the management strategy can be dynamically adjusted based on real-time regional monitoring data, enabling the management service system to quickly respond to changes in traffic and resource usage, enhancing its adaptability. Automated warning analysis helps maintenance tools or personnel quickly locate problematic single machines and take corresponding measures, reducing troubleshooting time and improving maintenance response speed and efficiency. Therefore, this embodiment of the present disclosure, through real-time monitoring and warning analysis at the single-machine level, achieves precise traffic and resource control for each service machine within the management service system. This not only enhances the system's stability and adaptability in the face of localized overload or resource shortages but also optimizes the service experience, improves maintenance efficiency, and provides data support for fault prevention and rapid recovery of the management service system.

[0160] In an optional embodiment, in step S206, the operation and maintenance management of the cloud gateway and / or multiple regional data centers is performed based on the analysis results, including the following method steps:

[0161] Step S261: Use the elastic management cloud service and global monitoring results to perform elastic capacity operation and maintenance management on the cloud gateway. The elastic management cloud service is implemented based on the target service layer architecture. The main service area in the target service layer architecture is deployed as a physical machine cluster, and the extended service area in the target service layer architecture is deployed as a container cluster.

[0162] Step S262: Utilize the elastic management cloud service and single-machine monitoring results to perform elastic capacity operation and maintenance management on some or all of the data centers in multiple regions.

[0163] The aforementioned elastic cloud management service can be an operation and maintenance management service designed based on the target service layer architecture. Elastic cloud management services can automatically or manually adjust the capacity of cloud services to cope with fluctuations in system traffic. Elastic cloud management services typically include functions such as automatic scaling, resource scheduling, and load balancing to ensure high availability and high performance of the service.

[0164] The aforementioned target service layer architecture can be a Serverless Operations and Maintenance (ASO) platform. The ASO platform includes a main service area and extended service areas. The main service area is deployed using a physical machine cluster, meaning it uses high-performance, stable hardware resources to support the system's backbone or critical services, ensuring stable service even under high traffic or failure conditions. The extended service area is deployed using a container cluster, leveraging the lightweight, portable, and resource isolation characteristics of containers to achieve rapid scaling and efficient resource utilization, adapting to the processing needs of non-critical services or traffic peaks.

[0165] In application scenarios, when performing elastic capacity operation and maintenance management on cloud gateways, the elastic management cloud service will automatically adjust the capacity of the cloud gateway based on the early warning information in the global monitoring results. For example, it will increase or decrease the number of cloud gateway instances to cope with upcoming traffic peaks or abnormal situations. This elastic capacity operation and maintenance management takes into account both the global distribution of traffic and makes full use of the stability of physical machine clusters and the flexibility of container clusters, ensuring that the cloud gateway can operate stably under various traffic scenarios while optimizing resource utilization.

[0166] Based on the above single-machine monitoring results, data centers to be elastically managed are identified from multiple regional data centers. These data centers can be some or all of the multiple regional data centers. When performing elastic capacity operation and maintenance management on some or all of the multiple regional data centers, the elastic management cloud service will intelligently adjust the capacity of the service machines in the data centers to be elastically managed based on the early warning information in the single-machine monitoring results. For example, it may start new container instances to increase service processing capacity, or shut down redundant instances to release resources. This elastic capacity operation and maintenance management for multiple regional data centers is applicable not only to cloud gateways but also to backend processing services or database services, ensuring that resources can be dynamically optimized based on real-time monitoring data in data centers in different regions, thereby improving the overall performance and stability of the system.

[0167] Taking the development and production environment of the managed cloud service shown in Figure 4 as an example, during the operation and maintenance phase, based on the target traffic control rules (including global rate limiting rules, interface rate limiting rules, user rate limiting rules, single-machine degradation rules, and single-machine protection rules) released during the deployment phase, real-time traffic data corresponding to the managed cloud service can be monitored and analyzed, and the following three stages can be completed: problem discovery, problem localization, and problem resolution. The aforementioned real-time traffic data can include global traffic data corresponding to the managed cloud service (corresponding to the cloud gateway) and single-machine cluster traffic data (corresponding to the single-machine clusters set up in the regional data center).

[0168] As shown in Figure 4, in the problem discovery phase, traffic alarm / early warning analysis is performed on real-time traffic data based on the target traffic control rules to discover the following possible problems: rate limiting alarm / early warning, degradation alarm / early warning, single-machine self-protection alarm / early warning, and interface / user traffic monitoring.

[0169] As shown in Figure 4, after a problem is discovered, the location tools of the cloud service platform can be used in the problem localization phase to pinpoint the location of the alarm / warning corresponding to the discovered alarm / warning problem. Specifically, the problem localization phase can utilize end-to-end logs and application monitoring link tracing to determine the alarm / warning location. Furthermore, based on the problem discovery data and problem localization data, a visual interface for a single-machine monitoring dashboard and a gateway availability dashboard can be generated and displayed.

[0170] As shown in Figure 4, after locating the problem, in the problem-solving stage, the elastic management cloud service can be used to achieve automatic scaling up and down, and other components configured in the management service system can be used to perform management measures such as modifying rate limits, prioritizing interface access, and degrading / circuiting services to resolve the discovered problem.

[0171] Specifically, this disclosure provides a problem discovery scheme as shown in Figure 8. First, during data collection, multi-source data is collected from both the service unit and the cloud gateway. For example, multiple data sources for multi-source monitoring data include: the cloud gateway's metric delivery endpoint, the operation and maintenance monitoring platform's collection endpoint, and the log service's collection endpoint. Second, during data aggregation, the operation and maintenance monitoring platform and the log service aggregate and analyze the multi-source monitoring data to obtain analysis results. These analysis results may include problem discovery data. Further, multiple channels of alarms / warnings are issued for the problem discovery data through an integrated alarm channel. For example, the alarm methods corresponding to the integrated alarm channel may include: SMS, telephone, email, application messages, etc.

[0172] In an exemplary application scenario, a problem localization and resolution process is provided as shown in Figure 9. After monitoring multi-source data related to a single-machine cluster of services, the multi-source monitoring platform detects a problem and generates an alarm. This alarm is sent to the development team (which can be developers or automated development tools) through an integrated alarm channel. Upon receiving the alarm, the development team displays the system's multi-source monitoring data and alarm information in a visual monitoring system. They can also perform full-link tracing based on the request ID and trace ID corresponding to the alarm to obtain the problem localization information. Further, the development team begins a status check based on the alarm and problem localization information to determine if the current system is experiencing any of the following: insufficient capacity, rate limiting issues, or dependency anomalies. If insufficient capacity is detected, an automatic scaling mechanism is triggered. If rate limiting is in effect, the current rate limiting is checked to ensure it meets expectations. If it does, the status quo is maintained; otherwise, the rate limiting threshold is adjusted. If dependency anomalies are present, the anomaly is checked to determine if it is located on a core link. If it is, a circuit breaker mechanism is triggered; otherwise, degradation processing is implemented. Based on the entire process from alarm notification to problem location and handling, as shown in Figure 9, it is possible to ensure that appropriate operation and maintenance control measures can be taken when different problems are discovered, thereby guaranteeing the stability and reliability of the system.

[0173] In an exemplary application scenario, an elastic capacity operation and maintenance management solution is provided, as shown in Figure 10. As shown in Figure 10, in the management service system, application traffic data is accessed to the service layer through a unified access component and a service discovery component, and then forwarded to the service cluster through load balancing. Specifically, the service cluster includes a physical machine cluster in the main service area. During load balancing, multiple physical machine interfaces in the physical machine cluster are accessed through the main service cluster interface gateway. The service cluster also includes a container cluster in the extended service area. During load balancing, multiple container interfaces in the container cluster are accessed through the extended service cluster interface gateway.

[0174] Furthermore, after processing application traffic data (such as user requests), the service cluster continues to allocate elastic capacity for operation and maintenance management through the service discovery component. Through the interfaces, tasks, and workflows of the main service cluster, it connects to the corresponding elastic components of the main service area (denoted as elastic component Z1… in Figure 10) to achieve elastic operation and maintenance management of the physical machine cluster. Through the interfaces of the extended service cluster, it connects to the corresponding elastic components of the extended service area (denoted as elastic component K1… in Figure 10) to achieve elastic operation and maintenance management of the container cluster.

[0175] Specifically, when performing elastic operation and maintenance management, the elastic components corresponding to the main service area and the elastic components corresponding to the extended service area execute corresponding elastic strategies (such as timed elastic strategies, system indicator elastic strategies, and custom indicator elastic strategies) to achieve automatic scaling up or automatic scaling down.

[0176] Through steps S261 to S262 above, the elastic capacity management cloud service introduced in this embodiment can dynamically adjust resource allocation based on real-time monitoring data. This ensures that the system can automatically expand when traffic surges or resources are scarce, and automatically shrink when traffic decreases, avoiding resource waste. By combining global monitoring results and single-machine monitoring results in elastic capacity operation and maintenance management, the system can quickly identify and respond to traffic anomalies, preventing service interruptions or performance degradation caused by insufficient resources or overload. Elastic capacity operation and maintenance management not only prevents faults but also accelerates the fault recovery process by rapidly adjusting the capacity of service machines in cloud gateways and regional data centers, reducing the impact of faults on business and user experience. Automated elastic capacity operation and maintenance management reduces the need for manual intervention by operation and maintenance personnel, improving operation and maintenance efficiency. Elastic capacity operation and maintenance management in multi-regional data centers supports the global deployment and operation of cloud services, ensuring that users in different regions can obtain a high-performance and stable service experience. As described above, this embodiment of the disclosure achieves efficient resource management of cloud gateways and regional data centers in the management service system through automated and intelligent capacity operation and maintenance control, combined with a flexible deployment architecture of physical machine clusters and container clusters. This enhances the stability and adaptability of the management service system and improves operation and maintenance efficiency and user experience.

[0177] In the aforementioned operating environment, this disclosure also provides an operation and maintenance management method as shown in Figure 11. Figure 11 is a flowchart of another operation and maintenance management method according to an embodiment of this disclosure. As shown in Figure 11, the operation and maintenance management method includes:

[0178] Step S1101: Obtain an operation and maintenance management request through the first application programming interface, wherein the operation and maintenance management request is used to obtain multi-source monitoring data in the management and control service system;

[0179] Step S1102: Return the operation and maintenance management response through the second application programming interface, wherein the response data carried in the operation and maintenance management response includes: operation and maintenance management report.

[0180] The aforementioned operation and maintenance management report is obtained by performing operation and maintenance management on the cloud gateway and / or multiple regional data centers of the management and maintenance service system through analysis results. The analysis results are obtained according to any of the aforementioned operation and maintenance management methods.

[0181] The operation and maintenance management method described in this embodiment can run on a cloud server to provide cloud operation and maintenance management services to clients. The client sends an operation and maintenance management request by calling a first application programming interface (API). After obtaining the request through the first API, the cloud server generates an operation and maintenance management report according to the method and further returns the report to the client through a second API.

[0182] The first and second application programming interfaces (APIs) mentioned above can be the same or different APIs. In one optional embodiment, the interface parameters in the first and second APIs may include, but are not limited to: a global interface identifier, an interface signing key, an interface timestamp, an interface request identifier, and a system call credential identifier. The first API can use a GET request or a POST request as the interface request method to obtain a file processing request. The second API can use a lightweight data exchange format (such as JavaScript Object Notation, or JSON for short) to return a file processing response.

[0183] In this embodiment, an operation and maintenance (O&M) control request is obtained through a first application programming interface (API), wherein the O&M control request is used to obtain multi-source monitoring data from the management and control service system. An O&M control response is returned through a second API, wherein the response data carried in the O&M control response includes an O&M control report. The O&M control report is obtained by performing O&M control on the cloud gateway and / or multiple regional data centers of the management and control service system based on analysis results. The analysis results are obtained according to any of the aforementioned O&M control methods. In this embodiment, the target traffic control rule is obtained by conducting various types of acceptance tests on predefined rules. On the one hand, using predefined rules as the basis for the target traffic control rule supports defining specific traffic control rules for different application scenarios or different O&M control needs, resulting in a flexible and applicable solution. On the other hand, various types of acceptance tests ensure the rationality and effectiveness of the target traffic control rule in managing the traffic of the management and control service system. Based on this, using the target traffic control rule to perform traffic analysis on the multi-source monitoring data corresponding to the cloud gateway and multiple regional data centers in the management and control service system, and then performing O&M control based on the analysis results, can improve the security and stability of the management and control service system.

[0184] In other words, the embodiments disclosed herein achieve the purpose of using the target traffic control rules obtained from acceptance testing to perform operation and maintenance management of the management service system, thereby realizing the technical effect of enhancing the rationality of traffic control rule settings to improve the security and stability of the management service system (especially the cloud management service platform), and thus solving the technical problem of low security and stability of the management service system caused by the lack of traffic control mechanism or unreasonable traffic control mechanism in related technologies.

[0185] It should be noted that the preferred embodiments of steps S1101 to S1102 described above can be found in the foregoing description, and will not be repeated here.

[0186] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0187] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this disclosure is not limited to the described order of actions, because according to this disclosure, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this disclosure.

[0188] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory (ROM), random access memory (RAM), magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this disclosure.

[0189] According to an embodiment of this disclosure, an apparatus embodiment for implementing the above-described operation and maintenance management method is also provided. Figure 12 is a schematic diagram of the structure of an operation and maintenance management apparatus according to an embodiment of this disclosure. The apparatus is deployed in a management and control service system, which includes a cloud gateway and multiple regional data centers corresponding to the cloud gateway. As shown in Figure 12, the apparatus includes: a collection module 1201, configured to collect multi-source monitoring data, wherein the multi-source monitoring data includes: gateway monitoring data of the cloud gateway and regional monitoring data of multiple regional data centers; an analysis module 1202, configured to perform traffic analysis on the multi-source monitoring data based on target traffic control rules to obtain analysis results, wherein the target traffic control rules are obtained by performing various types of acceptance tests on predefined rules; and a management and control module 1203, configured to perform operation and maintenance management on the cloud gateway and / or multiple regional data centers based on the analysis results.

[0190] Optionally, in addition to all the modules mentioned above, the above-mentioned operation and maintenance management device also includes: a testing module (not shown in the figure), which is configured to respond to the continuous integration deployment event corresponding to the management and control service system, use various types of test cases to perform acceptance testing on predefined rules, and obtain target test results; and configure and publish the predefined rules according to the target test results to obtain target traffic control rules.

[0191] Optionally, the above testing module is also configured to: orchestrate tests based on predefined rules and multiple types of test cases to obtain test instructions to be executed, wherein the multiple types of test cases include: performance recovery test cases and fault simulation test cases; execute test instructions using the virtual test container of the management and control service system to obtain execution results; and generate target test results based on the execution results and predefined rules.

[0192] Optionally, the execution result includes the interface performance data of the management and control service system during the execution of the test instructions; the above-mentioned test module is also configured to: obtain the interface performance indicators corresponding to the predefined rules; generate the target test result based on the interface performance data and the interface performance indicators, wherein the target test result is used to characterize whether the interface performance data meets the acceptance conditions corresponding to the interface performance indicators.

[0193] Optionally, the above testing module is also configured to: in response to determining that the interface performance data meets the acceptance criteria based on the target test results, configure and publish predefined rules to obtain the target traffic control rules.

[0194] Optionally, in the above-mentioned operation and maintenance management device, the predefined rules include initial rate limiting rules and initial degradation rules. The initial rate limiting rules are used to perform rate limiting control on the management service system during the execution of various types of test cases, and the initial degradation rules are used to perform single-machine degradation control on the management service system during the execution of various types of test cases.

[0195] Optionally, the target traffic control rules include global control rules and single-machine control rules, and the analysis results include global monitoring results and single-machine monitoring results; the above analysis module 1202 is also configured to: perform global traffic analysis on gateway monitoring data based on global control rules to obtain global monitoring results; and perform single-machine traffic analysis on regional monitoring data based on single-machine control rules to obtain single-machine monitoring results.

[0196] Optionally, the global monitoring results include global early warning results; the analysis module 1202 is further configured to: determine the global rate limiting threshold, the gateway interface rate limiting threshold, and the user rate limiting threshold according to the global control rules; and perform rate limiting early warning analysis on the gateway monitoring data based on the global rate limiting threshold, the gateway interface rate limiting threshold, and the user rate limiting threshold to obtain global early warning results, wherein the global early warning results are used to identify the location to be rate-limited in the management and control service system.

[0197] Optionally, the single-machine monitoring results include single-machine rate limiting warning results and single-machine degradation warning results; the analysis module 1202 is further configured to: determine single-machine rate limiting thresholds, single-machine interface rate limiting thresholds, and single-machine degradation thresholds according to single-machine control rules; perform rate limiting warning analysis on regional monitoring data based on single-machine rate limiting thresholds and single-machine interface rate limiting thresholds to obtain single-machine rate limiting warning results, wherein the single-machine rate limiting warning results are used to identify service machines in multiple regional data centers that need to be rate-limited and managed; and perform degradation warning analysis on regional monitoring data based on single-machine degradation thresholds to obtain single-machine degradation warning results, wherein the single-machine degradation warning results are used to identify service machines in multiple regional data centers that need to be degraded and managed.

[0198] Optionally, the aforementioned management module 1203 is further configured to: utilize the elastic management cloud service and global monitoring results to perform elastic capacity operation and maintenance management on the cloud gateway, wherein the elastic management cloud service is implemented based on the target service layer architecture, the main service area in the target service layer architecture is deployed as a physical machine cluster, and the extended service area in the target service layer architecture is deployed as a container cluster; utilize the elastic management cloud service and single machine monitoring results to perform elastic capacity operation and maintenance management on some or all of the data centers in multiple regions.

[0199] It should be noted that the above-mentioned acquisition module 1201, analysis module 1202 and control module 1203 correspond to steps S202 to S206 in the embodiment. The three modules and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in the foregoing embodiment.

[0200] According to an embodiment of this disclosure, an apparatus embodiment for implementing the above-described operation and maintenance management method is also provided. Figure 13 is a schematic diagram of another operation and maintenance management apparatus according to an embodiment of this disclosure. As shown in Figure 13, the apparatus includes: an acquisition module 1301, configured to acquire an operation and maintenance management request through a first application programming interface, wherein the operation and maintenance management request is used to acquire multi-source monitoring data in the management and control service system; and a response module 1302, configured to return an operation and maintenance management response through a second application programming interface, wherein the response data carried in the operation and maintenance management response includes: an operation and maintenance management report, which is obtained by performing operation and maintenance management on the cloud gateway and / or multiple regional data centers of the management and control service system through analysis results, and the analysis results are obtained according to any of the aforementioned operation and maintenance management methods.

[0201] It should be noted that the above-mentioned acquisition module 1301 and response module 1302 correspond to steps S1101 to S1102 in the embodiment. The two modules and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in the foregoing embodiment.

[0202] It should be noted that the above-mentioned modules or units may be hardware or software components stored in memory and processed by one or more processors. The above-mentioned modules may also be part of a device and run in a computer terminal.

[0203] It should be noted that the preferred implementation of this embodiment can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.

[0204] It should be noted that the preferred embodiments involved in the above embodiments of this disclosure are the same as the solutions, application scenarios and implementation processes provided in the above embodiments, but are not limited to the solutions provided in the above embodiments.

[0205] The embodiments of this disclosure can provide an operation and maintenance management system, including: an operation and maintenance management cloud component, a cloud gateway, and multiple regional data centers corresponding to the cloud gateway, wherein the operation and maintenance management cloud component is configured to provide operation and maintenance management cloud services to realize any of the aforementioned operation and maintenance management methods.

[0206] Embodiments of this disclosure may provide an electronic device, including: a memory storing an executable program; and a processor configured to run the program, wherein the program executes the operation and maintenance management method of any of the foregoing embodiments during runtime.

[0207] Figure 14 is a structural block diagram of an electronic device according to an embodiment of the present disclosure. As shown in Figure 14, the electronic device 140 may include: one or more (only one is shown in the figure) processors 142, memory 144, memory controller, and peripheral interfaces.

[0208] The aforementioned electronic device can be understood as an integrated smart terminal, including but not limited to servers, desktop computers, personal computers (PCs), all-in-one model machines, etc., and the electronic device may have the model described in the above embodiments of this disclosure pre-installed.

[0209] Specifically, this electronic device can pre-install various types of models, including but not limited to models in natural language processing, visual processing, speech processing, code processing, and multimodal task processing, thus providing diverse model selection. In different product forms, this electronic device can support one or more model usage methods, including but not limited to model training, model invocation, model fine-tuning, model deployment, model inference, and application. In some product forms, this electronic device also supports model management, including but not limited to multi-type model management (supporting the management of discriminative, generative, and other model types), model version control (supporting the control of different model versions), and model evaluation (evaluating model performance and effectiveness based on model evaluation tools). In other product forms, this electronic device can also create applications based on models, providing Application Programming Interface (API) invocation capabilities. Models can be invoked into the created applications through the API interface, and application management tools are provided to manage and monitor the applications.

[0210] Furthermore, the electronic device may also include data management (supporting the creation and management of model tuning datasets), a training center (providing rich training resources to help users learn and master artificial intelligence (AI) technology), and basic control capabilities (providing enterprise-level basic control capabilities to ensure the security and efficient operation of the system). Through the above functions, it provides a comprehensive and integrated device for AI development, training, deployment, and application.

[0211] The memory can be configured to store software programs and modules, such as the program instructions / modules corresponding to the operation and maintenance management method and device in the embodiments of this disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the operation and maintenance management method in the above embodiments. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of the above networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0212] The processor can invoke the executable program stored in the memory through the transmission device to execute any of the above-described operation and maintenance management methods in the above embodiments.

[0213] It will be understood by those skilled in the art that the structure shown in FIG14 is merely illustrative, and the electronic device may also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, and a mobile internet device (MID), etc. FIG14 does not limit the structure of the above-described electronic device. For example, the electronic device 140 may also include more or fewer components (such as a network interface, a display device, etc.) than shown in FIG14, or have a different configuration than shown in FIG14.

[0214] Those skilled in the art will understand that all or part of the steps in the various operation and maintenance management methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0215] Embodiments of this disclosure also provide a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium includes a stored executable program, wherein, when the executable program runs, it controls the device where the computer-readable storage medium is located to execute any of the aforementioned operation and maintenance management methods.

[0216] Optionally, in this embodiment, the storage medium may be located in an electronic device.

[0217] Optionally, in this embodiment, the computer-readable storage medium is configured to store an executable program, and when the executable program runs, it controls the device where the computer-readable storage medium is located to execute any of the above-described operation and maintenance management methods in the above embodiments.

[0218] Embodiments of this disclosure also provide a computer program product. Optionally, in this embodiment, the computer program product may include a computer program that, when executed by a processor, implements the operation and maintenance management method provided in the above embodiments.

[0219] Embodiments of this disclosure also provide a computer program product. Optionally, the computer program product may include a non-volatile computer-readable storage medium, which may be configured to store a computer program that, when executed by a processor, implements the operation and maintenance management method provided in the above embodiments.

[0220] Embodiments of this disclosure also provide a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, it implements the operation and maintenance management method provided in the above embodiments.

[0221] In the above embodiments of this disclosure, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0222] In the several embodiments provided in this disclosure, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between units or modules, and may be electrical or other forms.

[0223] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0224] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0225] If the integrated units described above are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the operation and maintenance management methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, ROM, RAM, portable hard drives, magnetic disks, or optical disks.

[0226] The above description is only a preferred embodiment of this disclosure. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles of this disclosure, and these improvements and modifications should also be considered within the scope of protection of this disclosure.

Claims

1. An operation and maintenance management method, applied to a management service system, the management service system comprising: The operation and maintenance management method includes: (The cloud gateway and the multiple regional data centers corresponding to the cloud gateway are also mentioned.) Collect multi-source monitoring data, wherein the multi-source monitoring data includes: gateway monitoring data of the cloud gateway, and regional monitoring data of the multiple regional data centers; Traffic analysis is performed on the multi-source monitoring data based on the target traffic control rules to obtain analysis results. The target traffic control rules are obtained by performing various types of acceptance tests on predefined rules. Based on the analysis results, the cloud gateway and / or the multiple regional data centers are operated and managed.

2. The operation and maintenance management method according to claim 1, wherein The operation and maintenance management method also includes: In response to the continuous integration deployment event corresponding to the management and control service system, the predefined rules are subjected to acceptance testing using various types of test cases to obtain the target test results; Based on the target test results, the predefined rules are configured and published to obtain the target traffic control rules.

3. The operation and maintenance management method according to claim 2, wherein, The predefined rules are subjected to acceptance testing using the various types of test cases to obtain the target test results, including: Test orchestration is performed based on the predefined rules and the various types of test cases to obtain test instructions to be executed. The various types of test cases include: performance recovery test cases and fault simulation test cases. The test instructions are executed using the virtual test container of the management and control service system, and the execution results are obtained. The target test result is generated based on the execution result and the predefined rules.

4. The operation and maintenance management method according to claim 3, wherein, The execution result includes the interface performance data of the management and control service system during the execution of the test command; Based on the execution result and the predefined rules, the target test result is generated as follows: Obtain the interface performance metrics corresponding to the predefined rules; Based on the interface performance data and the interface performance indicators, the target test result is generated, wherein the target test result is used to characterize whether the interface performance data meets the acceptance conditions corresponding to the interface performance indicators.

5. The operation and maintenance management method according to claim 4, wherein, Based on the target test results, the predefined rules are configured and published to obtain the target traffic control rules, including: In response to determining that the interface performance data meets the acceptance criteria based on the target test results, the predefined rules are configured and published to obtain the target traffic control rules.

6. The operation and maintenance management method according to claim 2, wherein, The predefined rules include initial rate limiting rules and initial degradation rules. The initial rate limiting rules are used to perform rate limiting control on the management and control service system during the execution of the various types of test cases, and the initial degradation rules are used to perform single-machine degradation control on the management and control service system during the execution of the various types of test cases.

7. The operation and maintenance management method according to claim 1, wherein, The target flow control rules include global control rules and single-machine control rules, and the analysis results include global monitoring results and single-machine monitoring results. Based on the target flow control rules, flow analysis is performed on the multi-source monitoring data, and the analysis results include: Based on the global control rules, global traffic analysis is performed on the gateway monitoring data to obtain the global monitoring results; Based on the single-machine control rules, single-machine traffic analysis is performed on the regional monitoring data to obtain the single-machine monitoring results.

8. The operation and maintenance management method according to claim 7, wherein, The global monitoring results include global early warning results; Based on the global control rules, global traffic analysis is performed on the gateway monitoring data to obtain the following global monitoring results: Based on the global control rules, determine the global rate limiting threshold, the gateway interface rate limiting threshold, and the user rate limiting threshold; Based on the global rate limiting threshold, the gateway interface rate limiting threshold, and the user rate limiting threshold, rate limiting early warning analysis is performed on the gateway monitoring data to obtain the global early warning result, wherein the global early warning result is used to identify the location to be rate-limited in the management and control service system.

9. The operation and maintenance management method according to claim 7, wherein, The single-machine monitoring results include single-machine rate limiting warning results and single-machine degradation warning results; Based on the aforementioned single-machine control rules, single-machine traffic analysis is performed on the regional monitoring data to obtain the single-machine monitoring results, including: Based on the single-machine control rules, determine the single-machine current limiting threshold, the single-machine interface current limiting threshold, and the single-machine degradation threshold; Based on the single-machine rate limiting threshold and the single-machine interface rate limiting threshold, rate limiting early warning analysis is performed on the regional monitoring data to obtain the single-machine rate limiting early warning result, wherein the single-machine rate limiting early warning result is used to identify the service single machine to be rate-limited and controlled in the multiple regional data centers; Based on the single-machine degradation threshold, degradation early warning analysis is performed on the regional monitoring data to obtain the single-machine degradation early warning result, wherein the single-machine degradation early warning result is used to identify the service single machines in the multiple regional data centers that need to be degraded and managed.

10. The operation and maintenance management method according to claim 7, wherein, Based on the analysis results, the operation and maintenance management of the cloud gateway and / or the multiple regional data centers includes: The elastic capacity operation and maintenance management of the cloud gateway is performed using the elastic management cloud service and the global monitoring results. The elastic management cloud service is implemented based on the target service layer architecture. The main service area in the target service layer architecture is deployed as a physical machine cluster, and the extended service area in the target service layer architecture is deployed as a container cluster. Using the elastic management cloud service and the single-machine monitoring results, elastic capacity operation and maintenance management can be carried out on some or all of the multiple regional data centers.

11. The operation and maintenance management method according to claim 2, wherein, The continuous integration deployment event includes at least one of the following: the management and control service system completes the development of new functions or modifies the code, the predefined rules are adjusted or updated, system health checks are performed, and the production environment is abnormal.

12. The operation and maintenance management method according to claim 3, wherein, Based on the predefined rules and the various types of test cases, test orchestration is performed to obtain the test instructions to be executed, including: The testing framework converts the various types of test cases written in natural language into machine-executable instructions, which include load testing instructions and fault injection instructions.

13. The operation and maintenance management method according to claim 3, wherein, The fault simulation test cases are used to verify the effectiveness of the backup links of the core links in the management and control service system, and / or to verify the effectiveness of the circuit breaker strategy of the non-core links in the management and control service system.

14. The operation and maintenance management method according to claim 9, wherein, The single-machine control rules are executed by the single-machine flow control component corresponding to the service single machine. The single-machine flow control component includes a system protection node, a flow control node, and a rate limiting / degradation early warning node.

15. The operation and maintenance management method according to claim 10, wherein, The elastic capacity operation and maintenance management is implemented through elastic policies executed by elastic components. These elastic policies include: timed elastic policies, system indicator elastic policies, and custom indicator elastic policies.

16. An operation and maintenance management method, comprising: The operation and maintenance management request is obtained through the first application programming interface, wherein the operation and maintenance management request is used to obtain multi-source monitoring data in the management and control service system; The operation and maintenance management response is returned through the second application programming interface, wherein the response data carried in the operation and maintenance management response includes: an operation and maintenance management report, which is obtained by performing operation and maintenance management on the cloud gateway and / or multiple regional data centers of the management and control service system through analysis results, and the analysis results are obtained in accordance with the operation and maintenance management method described in any one of claims 1 to 15.

17. An operation and maintenance management system, comprising: The system includes an operation and maintenance management cloud component, a cloud gateway, and multiple regional data centers corresponding to the cloud gateway, wherein the operation and maintenance management cloud component is configured to provide operation and maintenance management cloud services to implement the operation and maintenance management method according to any one of claims 1 to 16.

18. An electronic device comprising: Memory, which stores executable programs; A processor is configured to run the program, wherein the program executes the operation and maintenance management method according to any one of claims 1 to 16 when it runs.

19. A computer-readable storage medium comprising a stored executable program, wherein, When the executable program is executed, it controls the device containing the computer-readable storage medium to perform the operation and maintenance management method according to any one of claims 1 to 16.

20. A computer program product, comprising a computer program that, when executed by a processor, implements the operation and maintenance management method according to any one of claims 1 to 16.