Service health state monitoring system and electronic equipment
By introducing multiple monitoring modules into the service monitoring system, collecting and storing key information and determining whether the warning mechanism is triggered, the problem that the existing system cannot effectively cover key monitoring elements is solved, and comprehensive monitoring and early warning of the service health status is achieved, improving the stability and reliability of the system.
Patent Information
- Application Number
- CN202510063660.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-09
AI Technical Summary
Existing service monitoring systems cannot effectively cover key monitoring elements, such as memory usage, number of function calls and execution time, making it difficult to understand the operation efficiency and resource consumption of functions, and fail to detect potential performance bottlenecks in a timely manner.
It provides a monitoring system for serving health status, including an associated system monitoring module, a function usage monitoring module, a business indicator monitoring module, an exit request monitoring module, an entry request module, a hardware occupation module and a log status module. It collects and stores relevant information through dynamic agents or interfaces, and determines whether the warning mechanism is triggered.
It realizes comprehensive monitoring of the health status of the service, can promptly detect abnormalities and issue early warnings, helping operation and maintenance personnel to quickly locate and solve problems, and improve the stability and reliability of the system.
Smart Images

Figure CN119961093A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a monitoring system and electronic equipment for a service health status. Background Art
[0002] In the field of computer technology, especially when it comes to service monitoring, existing monitoring systems have significant limitations. Most current monitoring modes focus only on the basic conditions of the machine, such as simple monitoring of hardware information such as CPU and memory, as well as the general situation of network requests in the ingress direction and limited log records. However, for many key monitoring elements, the existing system cannot effectively cover them. At the function level, important data such as memory usage, number of calls, and execution time are missing, which makes it difficult for developers to gain an in-depth understanding of the function's operating efficiency and resource consumption, and unable to promptly discover potential performance bottlenecks. In the network request dimension, the actual call details in the egress direction, including the specific parameters of the request and detailed analysis of the response results, are unclear, resulting in an inability to accurately determine the root cause of the problem when there is a problem in the interaction between systems.
[0003] For the associated systems that the services rely on, such as databases, RabbitMQ, Kafka, ES, Redis, MongoDB, etc., the monitoring of their operating status and key indicators is seriously insufficient. In particular, the queue system (such as RabbitMQ / Kafka) that plays an important role in asynchronous processing lacks an effective monitoring mechanism, which often causes system failures when the queue is backlogged, and even system abnormalities are not known until users report the problem. In addition, the existing monitoring system cannot provide accurate information on whether the processing results of key businesses are successful, which brings great uncertainty to the normal development of the business and user experience. This incomplete and in-depth monitoring status has seriously restricted the stability of the system, performance optimization, and timely resolution of problems. A more complete and powerful monitoring system is urgently needed to fill these gaps. Summary of the invention
[0004] In view of this, an embodiment of the present application provides a service health status monitoring system and electronic device for solving at least one technical problem.
[0005] The embodiment of the present application provides a service health status monitoring system, including: an associated system monitoring module, which is used to collect indicator information of different associated systems based on preset collection indicators and through existing links of associated systems used by servers, and mark and save the indicator information based on the collection target and collection time, and judge whether the collected indicator information triggers the associated system early warning mechanism; a function usage monitoring module, which captures the target function through a dynamic proxy or a function data interface, collects usage information of the target function, the usage information includes at least one of the number of calls, function signature information, execution results, hardware occupancy, and execution time, stores the usage information, and judges whether the collected usage information triggers the function early warning mechanism; a business indicator monitoring module, which is used to collect business information of different businesses through interfaces of different businesses based on preset business indicators, and monitor the business information storage, and judge whether the collected business information triggers the function warning mechanism; the export request monitoring module is used to realize the network request through the dynamic proxy or export data interface, collect the issued network request information, record the number of requests, request initiation time, request end time and response result, detect whether the request is successful, and judge whether the collected network request triggers the export request warning mechanism; the input request module is used to realize the network request through the dynamic proxy or export data interface, collect the received network request information, record the number of requests, request initiation time, request end time and response result, detect whether the request is successful, and judge whether the collected network request triggers the export request warning mechanism; the hardware occupancy module is used to obtain the hardware occupancy data by accessing the hardware data interface, and the hardware includes CPU, memory or hard disk; the log status module is used to collect server log information.
[0006] Optionally, when the associated system is a database, the collection indicators include at least one of the SQL execution time, prepared statements, transaction execution time and database connection pool status of the database system; when the associated system is a queue system, the collection indicators include at least one of the number of messages, processing time and queue backlog of the queue system; when the associated system is a cache system, the collection indicators include at least one of the cache value, empty value rate and hit rate of the cache system.
[0007] Optionally, the HTTP request monitoring module in the export direction is used to: define an interface of a network acquirer, wherein the interface includes a fetch method to send an HTTP request; through Spring AOP, create a WebRequesterProxy proxy class, which inserts additional logic before and after calling the fetch method; before calling the fetch method, collect the target URL configuration from the nacos configuration center, increase the request count value by one, and record the network request information of the request, wherein the network request information includes at least one of the URL address, request header information, parameters and request method; when initiating a network request, record the request initiation time; after the Fetch method is executed, capture the response result, record the response code, message type, message length, response header information and response result; check the response code, if the response code is between the response thresholds, mark the request as successful, and increase the successful request value by one; otherwise, mark it as failed, and increase the failed request value by one; record the end time of the request and calculate the request duration; determine whether the collected network request triggers the export request early warning mechanism, wherein the export request early warning mechanism includes at least one of the response time and the response success rate.
[0008] Optionally, when the associated system is a database, the associated system monitoring module is used to: collect database connections, connect to the database through a driver, record the connection start time and the connection success time, calculate the connection time, record the query time of the proxy object of the query interface, collect the proxy object of the query interface through Spring AOP, record the query start time and end time, calculate the query time, and determine whether the collected usage information triggers a function warning mechanism; and monitor the status of the database connection pool, including at least one of the number of idle connections, the number of active connections and the maximum number of connections in the connection pool.
[0009] Optionally, the function usage monitoring module is used to: increase the execution count before calling the target function, record the method name, input parameter type, input parameter value, current time and memory usage size when the method is executed, collect the method execution results, record the execution results and exception information, calculate the time consumed by the method execution and the memory usage, and determine whether the collected business information triggers the function early warning mechanism; monitoring of asynchronous functions, including the submission time, start execution time and completion time of asynchronous tasks.
[0010] Optionally, the business information includes at least one of: login status, order creation, payment completion and order delivery. After collecting the business information, a detailed business indicator analysis report is provided.
[0011] The present invention further provides an electronic device, comprising the system as described in any one of the above items.
[0012] The present invention further provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the system as described in any one of the above items is run.
[0013] According to the embodiment of the present application, through the associated system monitoring module, information of different associated systems (such as database, queue system, cache system) can be collected based on preset indicators, including the sql execution time of the database, the number of messages in the queue system, etc., which can timely detect anomalies and issue warnings based on configuration, effectively prevent system failures caused by associated system problems, and improve stability. The function usage monitoring module uses dynamic proxy to capture target function information, covering the number of calls, execution results, etc., which is convenient for optimizing the code in advance and avoiding risks from entering the production environment. Function problems can also be discovered and handled in time according to the warning configuration. The business indicator monitoring module collects key business information based on preset business indicators, provides analysis reports, supports flexible customization of warning rules, ensures the normal flow of business, and improves user experience. The export and import request monitoring module collects network request information, detects request status, and configures warning mechanisms, which can help track call links and solve network failures. The hardware occupancy module obtains hardware occupancy data, and the log status module collects log information, which together provide rich data for system performance evaluation and troubleshooting, and comprehensively improve the reliability, stability and optimization capabilities of system operation.
[0014] The service health status monitoring system and electronic device of the present invention cover the risk points that may cause system abnormalities in the project, and have the following significant technical effects: through the associated system monitoring module, the use of the associated system can be monitored in real time to prevent the associated system from being unavailable without knowledge or the system abnormalities caused by perception lag. At the same time, the module can be expanded on demand, has no limitations, has no strong coupling relationship with the system, and can be configured with different early warning trigger mechanisms and notification methods on demand. The network request monitoring module in the export direction can clearly obtain the actual request situation of the third party called internally by the system, and can be used to analyze the actual availability of third-party resources and other situations, so as to make corresponding adjustments to make the system more stable. The function usage monitoring module can record information such as the number of function calls, memory usage, execution time, etc., to help operation and maintenance personnel complete analysis and optimization before going online, and avoid bringing risks to the production environment. The business indicator monitoring module can flexibly define key business indicators and early warning rules to meet highly customized business needs. When the business is abnormal, it can be discovered and issued in time to help the operation and maintenance personnel quickly locate and solve the problem. In summary, the service health status monitoring system and electronic equipment of the present invention can promptly detect abnormal status of the system and issue warnings, providing clear results and rapid positioning means for operation and maintenance personnel, thereby ensuring the stability and reliability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solution of the embodiments of the present application, the following briefly introduces the drawings in the embodiments of the present application.
[0016] Figure 1 It is a schematic diagram of the system architecture of an embodiment of the present application.
[0017] Figure 2 It is a module diagram of a service health status monitoring system according to an embodiment of the present application.
[0018] Figure 3 This is a workflow diagram of the HTTP request monitoring module in the export direction of the embodiment of the present application.
[0019] Figure 4 It is a workflow diagram of the associated system monitoring module of an embodiment of the present application.
[0020] Figure 5 It is a workflow diagram of the function usage monitoring module of the embodiment of the present application.
[0021] Figure 6 It is a schematic diagram of an electronic device of a service health status monitoring system used to implement an embodiment of the present application. DETAILED DESCRIPTION
[0022] The principles and spirit of the present application will be described below with reference to several exemplary embodiments. It should be understood that the purpose of providing these embodiments is to make the principles and spirit of the present application clearer and more thorough, so that those skilled in the art can better understand and implement the principles and spirit of the present application. The exemplary embodiments provided herein are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments herein, all other embodiments obtained by ordinary technicians of the art without creative work are within the scope of protection of this application.
[0023] Embodiments of the present application relate to terminal devices and / or servers. Those skilled in the art will appreciate that the embodiments of the present application may be implemented as a system, apparatus, device, method, computer-readable storage medium, or computer program product. Therefore, the present disclosure may be specifically implemented in at least one of the following forms: complete hardware, complete software, or a combination of hardware and software. According to the embodiments of the present application, the present application claims protection for a monitoring system, electronic device, computer-readable storage medium, and computer program product for the health status of a service. Figure 1 A schematic diagram of a system architecture of an embodiment of the present application is shown. Figure 1As shown, the system includes a terminal device 102 and a server 104. Among them, the terminal device 102 may include at least one of the following: a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart TV, various wearable devices, an augmented reality AR device, a virtual reality VR device, etc. A client can be installed on the terminal device 102. For example, the client can be a client that specifically performs a specific function (such as an application app), or a client with multiple application applets (different functions) embedded in it, or a client logged in through a browser. The user can operate on the terminal device 102. For example, the user can open the client installed on the terminal device 102 and input instructions through the client operation, or the user can open the browser installed on the terminal device 102 and input instructions through the browser operation. After the terminal device 102 receives the instruction input by the user, the request information containing the instruction is sent to the server 104. After receiving the request information, the server 104 performs corresponding processing and then returns the processing result information to the terminal device 102. The user instruction is completed through a series of data processing and information interaction.
[0024] In this document, terms such as first, second, third, etc. are only used to distinguish one entity (or operation) from another entity (or operation), but not to require or imply any order or relationship between these entities (or operations).
[0025] The present application embodiment provides a system for monitoring the health status of a service, such as Figure 2 As shown, the service health status monitoring system includes:
[0026] The associated system monitoring module 201 is used to collect indicator information of different associated systems based on preset collection indicators through the existing links of the associated systems used by the server, mark and save the indicator information based on the collection target and collection time, and determine whether the collected indicator information triggers the associated system early warning mechanism.
[0027] The function usage monitoring module 202 captures the target function through a dynamic proxy or a function data interface, collects usage information of the target function, the usage information includes at least one of the number of calls, function signature information, execution result, hardware occupancy, and execution time, stores the usage information, and determines whether the collected usage information triggers a function early warning mechanism;
[0028] The business indicator monitoring module 203 is used to collect business information of different businesses through interfaces of different businesses based on preset business indicators, store the business information, and determine whether the collected business information triggers a function warning mechanism.
[0029] The egress request monitoring module 204 is used to implement network requests through dynamic proxy or egress data interface, collect the issued network request information, record the number of requests, request initiation time, request end time and response result, detect whether the request is successful, and determine whether the collected network request triggers the egress request early warning mechanism.
[0030] The entry request module 205 is used to implement network requests through dynamic proxy or exit data interface, collect received network request information, record the number of requests, request initiation time, request end time and response results, detect whether the request is successful, and determine whether the collected network request triggers the exit request early warning mechanism.
[0031] The hardware occupancy module 206 is used to obtain hardware occupancy data by accessing the hardware data interface, and the hardware includes CPU, memory or hard disk.
[0032] The log status module 207 is used to collect server log information.
[0033] Embodiments of the invention The service health status monitoring system and electronic device of the invention cover the risk points that may cause system abnormalities in the project. Through the associated system monitoring module, the use of the associated system can be monitored in real time to prevent the system abnormalities caused by the unavailability of the associated system without knowledge or perception lag. The network request monitoring module in the export direction can clearly obtain the actual request situation of calling the third party inside the system, and can analyze the actual availability of the third-party resources and other situations, so as to make corresponding adjustments to make the system more stable. The function usage monitoring module can record the number of function calls, memory usage, execution time and other information, helping the operation and maintenance personnel to complete the analysis and optimization before going online, and avoid bringing risks to the production environment. The business indicator monitoring module can flexibly define key business indicators and warning rules to meet highly customized business needs. When the business is abnormal, it can be discovered and issued in time to help the operation and maintenance personnel quickly locate and solve the problem. The service health status monitoring system and electronic device of the invention can detect the abnormal state of the system in time and issue a warning, providing clear results and fast positioning means for the operation and maintenance personnel, thereby ensuring the stability and reliability of the system.
[0034] like Figure 3 In some embodiments, the present application can check the network request status in the export direction, solve the problem of unclear actual call results between services, and provide the actual results of each link in tracing the call link of complex systems. In some embodiments, the HTTP request monitoring module in the export direction is used to:
[0035] Step 301: define a network acquirer interface, wherein the interface includes a fetch method to send an HTTP request.
[0036] Step 302: Create a WebRequesterProxy proxy class through Spring AOP, which inserts additional logic before and after calling the fetch method.
[0037] Step 303: Before calling the fetch method, collect the target URL configuration from the nacos configuration center, increase the request count value by one, and record the requested network request information, which includes at least one of the URL address, request header information, parameters and request method.
[0038] Step 304: When initiating a network request, record the time when the request is initiated.
[0039] Step 305: After the Fetch method is executed, the response result is captured, and the response code, message type, message length, response header information, and response result are recorded.
[0040] Step 306: Check the response code, if the response code is within the response threshold, mark the request as successful, and increase the successful request value by 1. Otherwise, mark it as failed, and increase the failed request value by 1.
[0041] Step 307: Record the end time of the request and calculate the request duration.
[0042] Step 308: Determine whether the collected network request triggers an egress request warning mechanism, wherein the egress request warning mechanism includes at least one of a response time and a response success rate.
[0043] The embodiment of the present invention implements a network request tool in the form of a dynamic proxy (Proxy or CGLIB) or an interface, collects detailed information of the request issued (including but not limited to URL, HEADER, request parameters, etc.), and supports configuration of the current request and collection indicators to be collected. Specifically, the network request client can be implemented in the form of a dynamic proxy (Proxy or CGLIB) or an interface. Get the configuration of the target to be collected. If there is no configuration, it is defaulted not to collect or collect all targets. Record the number of network requests, collect detailed information of the request issued (including but not limited to URL, HEADER, request parameters, etc.), and detect whether there is a request initiation time in the request. If not, use the current time as the request initiation time. Initiate a network request, collect the response result (response code, Header, etc.), and detect whether the request carries the request end time in the network request. If not, use the current time as the request end time. Process the response code and other information to record whether the request is successful. According to the warning rule configuration, detect whether the warning trigger condition is met. If it is met, issue a warning.
[0044] In a specific embodiment, the system defines an interface named WebRequester, which includes a fetch() method to send HTTP requests. Then, through Spring AOP, a WebRequesterProxy proxy class is created, which inserts additional logic before and after calling the fetch() method.
[0045] Before calling the fetch method, get the target URL configuration from the nacos configuration center, first increase the request count value +1, and record the detailed information of the request, including URL address, request header information, parameters and request method (GET / POST / HEADER), etc.
[0046] When a network request is initiated, the request initiation time startTime is recorded, and the fetch method execution uses a library such as OkHttp to actually send the HTTP request.
[0047] After the Fetch method is executed, the response result is captured, and the response code, Content-Type, Content-Length and other response header information, as well as the response result are recorded. At the same time, the system will check the response code. If the response code is between 200-299, the request is marked as successful, and the number of successful requests successCount value +1; otherwise, it is marked as failed, and the number of successful requests errorCount value +1. At the same time, the system will record the end time of the request endTime and calculate the request duration: responseTime = endTime-startTime.
[0048] At the same time, the warning task will start the warning rule check for this URL. For example, if the maximum response time is configured to be 500ms, then when responseTime>500ms, the warning of this rule will be triggered; if the failure rate is configured to be 5%, then when the number of requests count=100 and the number of failures errorCount>5, the warning of this rule will be triggered.
[0049] At the same time, after receiving the warning task, the early warning system sends a warning notification through text messages, emails, etc.
[0050] In some embodiments, when the associated system is a database, the collected indicators include at least one of the sql execution time, prepared statements, transaction execution time, and database connection pool status of the database system. When the associated system is a queue system, the collected indicators include at least one of the number of messages, processing time, and queue backlog of the queue system; when the associated system is a cache system, the collected indicators include at least one of the cache value, null value rate, and hit rate of the cache system.
[0051] like Figure 4 In some embodiments, when the associated system is a database, the associated system monitoring module is used to:
[0052] Step 401: Collect database connections, connect to the database through a driver, record the connection start time and connection success time, calculate the connection time, record the query time of the proxy object of the query interface, collect the proxy object of the query interface through Spring AOP, record the query start time and end time, calculate the query time, and determine whether the collected usage information triggers the function warning mechanism.
[0053] Step 402: Monitoring the database connection pool status, including at least one of the number of idle connections, the number of active connections, and the maximum number of connections in the connection pool.
[0054] In some embodiments, the present application can use the existing connections of the associated systems used by the project, collect key information of different associated systems according to the indicators that need to be collected in the configuration, and issue warnings according to the configurable early warning mechanism. Specifically, obtain the existing connections of the associated systems used by the project. Get the configuration of the associated system. When there is no configuration, it is defaulted not to collect or collect all indicators. Collect indicators according to the configuration, such as: database system, collect SQL execution time, prepared statements, transaction execution time, etc.; queue system (RabbitMQ or Kafka, etc.), collect the number of messages in the queue and the number of messages being processed, and record the start time and end time of this message processing when processing messages; cache system (Redis or Memcache, etc.), when the cache value of a key reaches the threshold, it is considered that there is a cache risk, when the null value rate of a key reaches a certain proportion, it is considered that there is a cache penetration risk, and when the cache does not set the validity period, it is considered that the cache system is at risk. The preservation of indicator data takes the target object and the current time as necessary tags to store the indicator data collected this time. Process the collected indicators and detect whether the early warning rules are triggered. For example, in a queue system, the first warning rule is that when the number of unprocessed messages in the queue reaches a certain threshold, an alarm is triggered; the second warning rule is that when the average response time of message processing exceeds the set threshold within a certain period of time, an alarm is triggered. And so on.
[0055] In a specific embodiment, the database connection DataSource object can be obtained, the connection jdbcUrl can be obtained, the database can be connected through the driver, the connection start time StartTime and the time when the connection is successful ConnectedTime can be recorded, and the time taken for this connection ConnectTime can be obtained through ConnectedTime-StartTime. At the same time, the number of connections ConnectedCount is increased by 1.
[0056] Get the proxy object SelectMapperProxy of the query interface through Spring Aop, where the select() method is used to query the database. Record the query start time queryStartTime before querying, execute the SQL query, and record the query completion time queryEndTime after obtaining the query results. The query time can be calculated as: queryTime = queryEndTime - queryStartTime.
[0057] At the same time, the early warning detection task will obtain the database early warning configuration from the nacos configuration center. If the slow query sql is defined as 300ms, then when queryTime>300ms, this query is a slow query. Record the query sql statement and parameters, and trigger the early warning rule. The early warning touch task is the same as other modules.
[0058] like Figure 5 In some embodiments, the function usage monitoring module is used to:
[0059] Step 501: Before calling the target function, increase the execution count, record the method name, input parameter type, input parameter value, current time and memory usage during method execution, collect method execution results, record execution results and exception information, calculate the time consumed by method execution and memory usage, and determine whether the collected business information triggers the function early warning mechanism.
[0060] Step 502: Monitoring of asynchronous functions, including submission time, start execution time, and completion time of asynchronous tasks.
[0061] This application can use dynamic proxy (Proxy or CGLIB) or interface (interface) to collect function memory usage, execution times, execution time and other information according to the configured functions that need to be collected. Specifically, the target function is captured in the form of dynamic proxy (Proxy or CGLIB) or interface (Interface). Get the configuration of the target that needs to be collected. If there is no configuration, it will not collect or collect all targets by default. Record the number of function calls, function signature information, execution results, space occupied, execution time and other information. According to the collected indicator data, the results are stored in a storage medium. According to the warning configuration, the warning indicators are calculated. And the results are stored in a storage medium. According to the warning rule configuration, check whether the warning trigger conditions are met. If met, a warning is issued.
[0062] In a specific embodiment, a class named Executor is defined in the system, which has a method named execute(). An ExecutorProxy proxy class is created through Spring AOP.
[0063] Before calling the execute method, first increase the execution count executeCount value +1, and record the method name, input parameter type, input parameter value, etc. when the method is executed, and record the current time executeStartTime and the current memory usage size memoryBeforeExecute. Then, execute the execute method, obtain the method execution result, record the current time executeEndTime, the current memory usage size memoryAfterExecute, record the method execution result, capture possible exceptions, and record the executeErrorCount value +1, and record detailed stack information; at the same time, you can calculate the time consumed by this method execution executeTime = executeEndTime-executeStartTime, usedMemory = memoryAfterExecute-meoryBeforeExecute.
[0064] At the same time, the warning task will enable the warning rule check for this method. If the maximum execution time is configured to be 500ms, then when executeTime>500ms, the warning for this rule will be triggered; if the exception rate is configured to be 5%, then when the number of requests executeCount=100 and the number of failures executeErrorCount>5, the warning for this rule will be triggered.
[0065] At the same time, after receiving the warning task, the early warning system sends a warning notification through text messages, emails, etc.
[0066] In some embodiments, the business information includes at least one of: login status, order creation, payment completion and order shipment. After collecting the business information, a detailed business indicator analysis report is provided.
[0067] The embodiment of the present application can provide a method for collecting and reporting business indicator data, and provide an indicator data processing interface for implementation. During the implementation process, the business party collects data in the key business through the provided data collection and reporting method, and realizes the processing of indicator data. Then the system will trigger an early warning when the business indicator meets the early warning condition according to the definition of the pre-defined key indicators and the early warning triggering rules. Specifically, the key business indicators can be defined, and the business goals that need to be collected, the indicator parameters that can be collected, and the collection time or collection cycle of the indicators can be clearly defined. The business indicator configuration that needs to be collected is obtained, and no collection is performed when there is no configuration. In the business logic, the collected indicator parameters are set by the collection method provided by this system. The indicator processing interface provided by this system is implemented to process the collected business indicator data and the early warning indicators required in the early warning configuration. This system will store the collected and processed results in the storage medium, and detect whether the early warning indicators are met according to the early warning configuration. If the early warning triggering conditions are met, the early warning information is sent to the early warning target by configuring the early warning path.
[0068] In a specific embodiment, first, the system defines a business indicator: loginFailCount, the number of login failures. When logging in, if the login is unsuccessful, the loginFailCount value is +1, and if the login is successful, the loginSuccessCount value is +1. At the same time, the warning configuration of this business indicator is obtained from the nacos configuration center. For example, if the loginFailCount failure rate is 3%, when the warning task detects loginFailCount / loginFailCount+loginSuccessCount>0.03, the warning is triggered.
[0069] By adopting the embodiments of the present application, risk points in the project that may cause system abnormalities are covered, abnormal states of the system can be discovered in a timely manner and early warnings can be issued, and clear results can be provided when tracking problems, so as to quickly locate and resolve abnormal situations in the system.
[0070] Among them, the monitoring of related systems can be expanded on demand, has no limitations, has no strong coupling relationship with the system, and different early warning trigger mechanisms and notification methods can be configured on demand.
[0071] Network requests in the export direction can clearly obtain the actual request status of the third party calls within the system, and can be used to analyze the actual availability of third-party resources and make adjustments to stabilize the system.
[0072] You can use information such as the number of function calls and memory usage to complete analysis and optimization before going online, without having to bring risks to the production environment.
[0073] The definition and early warning of key business indicators can be flexibly defined to support highly customized business needs.
[0074] The electronic device in the embodiment of the present application may be a user terminal device, a server, other computing devices, or a cloud server. Figure 6 A schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application is shown. The electronic device may include a processor 601 and a memory 602 storing computer program instructions. When the processor 601 executes the computer program instructions, the process or function of any of the above-mentioned embodiments of the method is implemented.
[0075] Specifically, the processor 601 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application. The memory 602 may include a large-capacity memory for data or instructions. For example, the memory 602 may be at least one of the following: a hard disk drive (HDD), a read-only memory (ROM), a random access memory (RAM), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a tape, a universal serial bus (USB) drive or other physical / tangible memory storage device. For another example, the memory 602 may include a removable or non-removable (or fixed) medium. For another example, the memory 602 may be inside or outside the integrated gateway disaster recovery device. The memory 602 may be a non-volatile solid-state memory. In other words, the memory 602 usually includes a tangible (non-transitory) computer-readable storage medium (such as a memory device) encoded with computer-executable instructions, and when the software is executed (such as executed by one or more processors), the operation described in the method of the embodiment of the present application can be performed. The processor 601 implements the process or function of any method in the above embodiments by reading and executing the computer program instructions stored in the memory 602.
[0076] In one example, Figure 6The electronic device shown may also include a communication interface 603 and a bus 610. Among them, the processor 601, the memory 602, and the communication interface 603 are connected through the bus 610 and complete the communication between each other. The communication interface 603 is mainly used to realize the communication between each module, device, unit and / or device in the embodiment of the present application. The bus 610 includes hardware, software or both, and can couple the components of the online data traffic billing device to each other. For example, the bus may include at least one of the following: an accelerated graphics port (AGP) or other graphics bus, an enhanced industrial standard architecture (EISA) bus, a front-end bus (FSB), a hypertransport (HT) interconnect, an industrial standard architecture (ISA) bus, an infinite bandwidth interconnect, a low pin count (LPC) bus, a memory bus, a micro channel architecture (MCA) bus, a peripheral component interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a serial advanced technology attachment (SATA) bus, a video electronics standard association local (VLB) bus or other suitable buses. The bus 610 may include one or more buses. Although the embodiments of the present application describe or illustrate a specific bus, the embodiments of the present application may consider any suitable bus or interconnection method.
[0077] In combination with the method in the above embodiments, an embodiment of the present application also provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the process or function of any method in the above embodiments is implemented.
[0078] In addition, an embodiment of the present application further provides a computer program product, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the process or function of any one of the methods in the above embodiments is implemented.
[0079] The above exemplarily describes the flowcharts and / or block diagrams of the methods, devices, systems and computer program products of the embodiments of the present application, and describes the relevant various aspects. It should be understood that each box or combination thereof in the flowchart and / or block diagram can be implemented by computer program instructions, or by dedicated hardware that performs a specified function or action, or by a combination of dedicated hardware and computer instructions. For example, these computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to form a machine that enables these instructions executed by such a processor to enable the implementation of the functions / actions specified in each box or combination thereof in the flowchart and / or block diagram. Such a processor can be a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit.
[0080] The functional block shown in the structured block diagram of the embodiment of the present application can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), suitable firmware, a plug-in, a function card, etc. When implemented in software, it is a program or code segment that is used to perform the required task. A program or code segment can be stored in a memory, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0081] It should be noted that the present application is not limited to the specific configurations and processes described above or shown in the figures. The above is only a specific implementation mode of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the described system, device, module or unit can refer to the corresponding process in the method embodiment without further description. It should be understood that the scope of protection of the present application is not limited to this. Any technician familiar with the technical field can think of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and these modifications or substitutions should be included in the scope of protection of the present application.
Claims
1. A system for monitoring the health status of a service, comprising: The associated system monitoring module is used to collect indicator information of different associated systems based on preset collection indicators through the existing links of the associated systems used by the server, mark and save the indicator information based on the collection target and collection time, and determine whether the collected indicator information triggers the associated system early warning mechanism; The function usage monitoring module captures the target function through a dynamic proxy or a function data interface, collects usage information of the target function, the usage information includes at least one of the number of calls, function signature information, execution result, hardware occupancy, and execution time, stores the usage information, and determines whether the collected usage information triggers a function early warning mechanism; The business indicator monitoring module is used to collect business information of different businesses through interfaces of different businesses based on preset business indicators, store the business information, and determine whether the collected business information triggers the function early warning mechanism; The egress request monitoring module is used to implement network requests through dynamic proxy or egress data interface, collect information about the network requests sent, record the number of requests, request initiation time, request end time and response results, detect whether the request is successful, and determine whether the collected network request triggers the egress request early warning mechanism; The entry request module is used to implement network requests through dynamic proxy or egress data interface, collect received network request information, record the number of requests, request initiation time, request end time and response result, detect whether the request is successful, and determine whether the collected network request triggers the egress request early warning mechanism; A hardware occupancy module is used to obtain hardware occupancy data by accessing a hardware data interface, wherein the hardware includes a CPU, a memory or a hard disk; The log status module is used to collect server log information.
2. The server monitoring system according to claim 1, wherein when the associated system is a database, the collection indicators include at least one of the SQL execution time, prepared statements, transaction execution time and database connection pool status of the database system; when the associated system is a queue system, the collection indicators include at least one of the number of messages, processing time and queue backlog of the queue system; when the associated system is a cache system, the collection indicators include at least one of the cache value, empty value rate and hit rate of the cache system.
3. The server monitoring system according to claim 1, wherein the egress request monitoring module is further configured to: Define a network retrieval interface, which contains a fetch method to send HTTP requests; Through Spring AOP, create a WebRequesterProxy proxy class that inserts additional logic before and after calling the fetch method; Before calling the fetch method, collect the target URL configuration from the nacos configuration center, increase the request count value by one, and record the network request information of the request, which includes at least one of the URL address, request header information, parameters and request method; When initiating a network request, record the time when the request is initiated; After the Fetch method is executed, the response result is captured, and the response code, message type, message length, response header information, and response result are recorded; Check the response code. If the response code is within the response threshold, mark the request as successful and increase the success request value by 1. Otherwise, mark it as failed and increase the failure request value by 1. Record the end time of the request and calculate the request duration; The collected network request determines whether it triggers an egress request early warning mechanism, wherein the egress request early warning mechanism includes at least one of a response time and a response success rate.
4. The server monitoring system according to claim 1, wherein when the associated system is a database, the associated system monitoring module is further used to: Collect database connections, connect to the database through the driver, record the connection start time and connection success time, calculate the connection time, record the query time of the proxy object of the query interface, collect the proxy object of the query interface through Spring AOP, record the query start time and end time, calculate the query time, and determine whether the collected usage information triggers the function warning mechanism; and The monitoring of the database connection pool status includes at least one of the number of idle connections, the number of active connections, and the maximum number of connections in the connection pool.
5. The server monitoring system according to claim 1, wherein the function usage monitoring module is further used for: Before calling the target function, increase the execution count, record the method name, input parameter type, input parameter value, current time and memory usage during method execution, collect method execution results, record execution results and exception information, calculate the time and memory usage of method execution, and determine whether the collected business information triggers the function warning mechanism; Monitoring of asynchronous functions, including submission time, start execution time, and completion time of asynchronous tasks.
6. The server monitoring system according to claim 1, wherein the service information includes: At least one of the following: login status, order creation, payment completion, and order delivery. After collecting business information, a detailed business indicator analysis report is provided.
7. An electronic device, comprising the system according to any one of claims 1 to 6.