Gray scale monitoring method and device, electronic equipment and computer program product

By adopting the index classification and grading strategy in the grayscale monitoring system, the monitoring indicators are classified and graded, and by analyzing and comparing the monitoring indicator values ​​of the online version and the grayscale version, the root cause of faults of abnormal events is determined, and the problem of systematic index grading and classification monitoring and positioning of the root cause of abnormal events is solved in the existing technology, and efficient monitoring and investigation efficiency is achieved.

CN120086095APending Publication Date: 2025-06-03HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510252940.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing basic grayscale monitoring system cannot realize systematic index grading and classification monitoring, and cannot locate the root cause of abnormal events during grayscale release.

Method used

The pre-configured indicator classification and grading strategy is adopted to classify and rank monitoring indicators to improve the correlation coverage and accuracy of monitoring indicators, and the root cause of abnormal events is determined by analyzing and comparing the monitoring indicator values ​​of the online version and the grayscale version.

Benefits of technology

The systematic classification and hierarchical monitoring of monitoring indicators has been realized, which improves the coverage and accuracy of monitoring indicators, and effectively improves the efficiency of checking abnormal events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086095A_ABST
    Figure CN120086095A_ABST
Patent Text Reader

Abstract

The invention provides a gray level monitoring method and device, electronic equipment and a computer program product, and relates to the technical field of software testing and operation and maintenance. The method comprises the following steps: determining a monitoring index which needs to be monitored when an application program performs gray release, wherein the monitoring index is determined based on a pre-configured index classification and grading strategy; obtaining a gray level monitoring index value and an online monitoring index value corresponding to the gray level version and the online version of the application program respectively; and analyzing and comparing the gray level monitoring index value and the online monitoring index value, and determining a fault root cause of an abnormal event generated in the gray level monitoring process of the application program. According to the method and the device, the monitoring indexes can be classified and graded in the gray level monitoring process, and the correlation coverage degree and the accuracy of the monitoring indexes are improved; and meanwhile, monitoring abnormity identification can be realized, and the troubleshooting efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] This section aims to provide background or context for the embodiments of the present disclosure stated in the claims. The description herein is not admitted to be prior art merely because it is included in this section.

[0003] Gray release means allowing a change to be first implemented in a small scope in the production environment and gradually expanded to all users. This method can achieve a smooth transition, enabling some users to continue using the old features while some users start using the new features, and gradually expanding the scope of use of the new features based on feedback. If a software is about to launch a brand-new function or undergo a relatively major overhaul in the near future, a small-scale trial can be carried out first, and then the scope can be gradually expanded until the brand-new function covers all system users. Summary of the Invention

[0004] In related solutions, the basic gray-scale monitoring system can identify the differences between the change-phase indicators and the online version through technical means for basic monitoring. However, the above basic gray-scale monitoring system can only perform basic monitoring, unable to achieve systematic indicator grading and classification monitoring, and unable to locate the root cause of abnormal events during the gray release process.

[0005] Therefore, the present disclosure proposes an improved gray-scale monitoring method to classify and grade monitoring indicators using an indicator classification and grading strategy during the gray-scale monitoring process, improving the correlation coverage and accuracy of the monitoring indicators; and at the same time, being able to identify monitoring anomalies and improve the troubleshooting efficiency.

[0006] In this context, embodiments of the present disclosure are expected to provide a gray-scale monitoring method, a gray-scale monitoring device, a computer-readable storage medium, an electronic device, and a computer program product.

[0007] In the first aspect of the embodiments of the present disclosure, a gray-scale monitoring method is provided, including: determining monitoring indicators required for gray release of an application program, where the monitoring indicators are determined based on a pre-configured indicator classification and grading strategy; obtaining the gray-scale monitoring indicator values and online monitoring indicator values corresponding to the gray-scale version and the online version of the application program respectively; analyzing and comparing the gray-scale monitoring indicator values and the online monitoring indicator values to determine the root cause of the abnormal event generated during the gray-scale monitoring of the application program.

[0008] In one embodiment of the present disclosure, determining the monitoring metrics required for the gray release of the application program includes: obtaining a pre-configured metric classification and grading strategy generated based on the business domain of the application program; determining the monitoring metrics based on the metric classification and grading strategy, where the monitoring metrics include one or more of application metrics, associated business metrics, metrics associated with the current change, and public opinion metrics.

[0009] In one embodiment of the present disclosure, determining the monitoring metrics based on the metric classification and grading strategy includes: determining the basic metrics associated with the application program and using the basic metrics as the application metrics; determining the change services corresponding to the gray release and the associated businesses of the change services, and determining the associated business metrics based on the associated businesses and the metric classification and grading strategy; determining the change-associated metrics corresponding to the gray release based on the metric classification and grading strategy as the metrics associated with the current change; and generating the public opinion metrics according to the public opinion keywords that need to be concerned about during the gray release of the application program.

[0010] In one embodiment of the present disclosure, the metrics associated with the current change include one or more of processor parameter metrics, database change parameter metrics, data persistence layer change parameter metrics, message queue parameter metrics, and custom monitoring parameter metrics.

[0011] In one embodiment of the present disclosure, obtaining the gray monitoring metric values and the online monitoring metric values corresponding to the gray version and the online version of the application program respectively includes: determining the monitoring scope of the gray monitoring, where the monitoring scope includes one or more of the monitored user scope, the monitored business scope, and the specified monitoring time period; determining the gray monitoring metric values corresponding to the gray version based on the monitoring scope and the monitoring metrics; and determining the online monitoring metric values corresponding to the online version based on the monitoring scope and the monitoring metrics.

[0012] In one embodiment of the present disclosure, analyzing and comparing the gray monitoring metric values and the online monitoring metric values to determine the root cause of the abnormal event generated during the gray monitoring of the application program includes: analyzing and comparing the gray monitoring metric values and the online monitoring metric values to determine the initial failure type corresponding to the abnormal event; determining the change event corresponding to the gray release and obtaining the event-related data of the change event; and combining the initial failure type and the event-related data to locate the root cause of the failure.

[0013] In one embodiment of the present disclosure, the analysis and comparison of the gray-scale monitoring metric value and the online monitoring metric value to determine the initial fault type corresponding to the abnormal event includes: generating time-series data to be analyzed according to the gray-scale monitoring metric value and the online monitoring metric value; performing data preprocessing on the time-series data to be analyzed to obtain the metric data to be analyzed; performing a bottom-up pre-clipping process on the metric data to be analyzed to obtain target associated metric data; performing a top-down anomaly localization process on the target associated metric data to obtain an initial anomaly localization result; and performing a root cause localization and voting decision process on the initial anomaly localization result to obtain the initial fault type.

[0014] In one embodiment of the present disclosure, the determination of the change event corresponding to the gray-scale release and the acquisition of the event-related data of the change event include: determining the event center associated with the monitoring platform, and obtaining the change content within a specified time interval before and after the change of the current gray-scale version based on the event center, where the change content includes server-side change data and client-side change data; determining the change activity data within a specified time interval before and after the change of the current gray-scale version through the activity guarantee platform associated with the monitoring platform; determining the upstream and downstream applications of the application program, and obtaining the application change data of the upstream and downstream applications within the specified time period of the change of the current gray-scale version; and determining the abnormal metric data corresponding to the online version within the specified time period of the change of the current gray-scale version.

[0015] In one embodiment of the present disclosure, the method further includes: determining the analysis dimension corresponding to the event-related data, where the analysis dimension includes one or more of an associated field, an associated application, and an associated resource; determining the associated field metric value corresponding to the event-related data based on the associated field; determining the associated application metric value corresponding to the event-related data based on the associated application; determining the associated resource metric value corresponding to the event-related data based on the associated resource; and monitoring the associated field metric value, the associated application metric value, and the associated resource metric value through a monitoring list.

[0016] In a second aspect of the embodiments of the present disclosure, a gray-scale monitoring device is provided, including: a monitoring metric determination module, configured to determine the monitoring metrics required for the gray-scale release of an application program, where the monitoring metrics are determined based on a pre-configured metric classification and grading strategy; a metric value acquisition module, configured to acquire the gray-scale monitoring metric values and the online monitoring metric values corresponding to the gray-scale version and the online version of the application program respectively; and a fault root cause determination module, configured to analyze and compare the gray-scale monitoring metric values and the online monitoring metric values to determine the fault root cause of the abnormal event generated during the gray-scale monitoring of the application program.

[0017] In one embodiment of the present disclosure, the monitoring metric determination module includes a monitoring metric determination unit, configured to: obtain a pre-configured metric classification and grading strategy generated based on the business domain of the application program; determine the monitoring metrics based on the metric classification and grading strategy, where the monitoring metrics include one or more of application metrics, associated business metrics, metrics associated with the current change, and public opinion metrics.

[0018] In one embodiment of the present disclosure, the monitoring metric determination unit includes a monitoring metric determination subunit, configured to: determine the basic metrics associated with the application program and use the basic metrics as the application metrics; determine the change services corresponding to the gray release and the associated businesses of the change services, and determine the associated business metrics based on the associated businesses and the metric classification and grading strategy; determine the change-associated metrics corresponding to the gray release based on the metric classification and grading strategy as the metrics associated with the current change; generate the public opinion metrics according to the public opinion keywords that need to be concerned about for the gray release of the application program.

[0019] In one embodiment of the present disclosure, the metric value acquisition module includes a metric value acquisition unit, configured to: determine the monitoring scope of the gray monitoring, where the monitoring scope includes one or more of a monitoring user scope, a monitoring business scope, and a specified monitoring time period; determine the gray monitoring metric values corresponding to the gray version based on the monitoring scope and the monitoring metrics; determine the online monitoring metric values corresponding to the online version based on the monitoring scope and the monitoring metrics.

[0020] In one embodiment of the present disclosure, the root cause determination module includes a root cause determination unit, configured to: analyze and compare the gray monitoring metric values and the online monitoring metric values to determine the initial fault type corresponding to the abnormal event; determine the change event corresponding to the gray release and obtain the event-related data of the change event; combine the initial fault type and the event-related data to locate the root cause of the fault.

[0021] In one embodiment of the present disclosure, the root cause determination unit includes a fault type determination subunit, configured to: generate time series data to be analyzed according to the gray monitoring metric values and the online monitoring metric values; perform data preprocessing on the time series data to be analyzed to obtain the metric data to be analyzed; perform a bottom-up pre-clipping process on the metric data to be analyzed to obtain the target associated metric data; perform a top-down anomaly localization process on the target associated metric data to obtain an initial anomaly localization result; perform root cause localization and voting decision processing on the initial anomaly localization result to obtain the initial fault type.

[0022] In one embodiment of the present disclosure, the fault root cause determination unit includes an event data acquisition subunit, configured to: determine the event center associated with the monitoring platform, and acquire the change content within a specified time period before and after the current gray-scale version change based on the event center, where the change content includes server-side change data and client-side change data; determine the change activity data within a specified time period before and after the current gray-scale version change through the activity guarantee platform associated with the monitoring platform; determine the upstream and downstream applications of the application program, and acquire the application change data of the upstream and downstream applications within a specified time period of the current gray-scale version change; determine the abnormal index data corresponding to the online version within a specified time period of the current gray-scale version change.

[0023] In one embodiment of the present disclosure, the gray-scale monitoring device further includes an index monitoring module, configured to: determine the analysis dimensions corresponding to the event-related data, where the analysis dimensions include one or more of an associated field, an associated application, and an associated resource; determine the associated field index value corresponding to the event-related data based on the associated field; determine the associated application index value corresponding to the event-related data based on the associated application; determine the associated resource index value corresponding to the event-related data based on the associated resource; monitor the associated field index value, the associated application index value, and the associated resource index value through a monitoring list.

[0024] In the third aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the gray-scale monitoring method as described above.

[0025] In the fourth aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; and a memory, on which a computer-readable instruction is stored, and when the computer-readable instruction is executed by the processor, it implements the gray-scale monitoring method as described above.

[0026] According to the fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the gray-scale monitoring method as described in any one of the above.

[0027] According to the technical solution of the embodiments of the present disclosure, on the one hand, a solution for classifying and grading the monitoring indicators is implemented by using a pre-configured index classification and grading strategy, which improves the coverage and accuracy of the monitoring indicators. On the other hand, by comparing the monitoring indicator values of the online version and the gray-scale version, the fault root cause of the abnormal events generated during the gray-scale monitoring can be located, which can effectively improve the troubleshooting efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown by way of example and not limitation, wherein:

[0029] Figure 1 A schematic block diagram of a system architecture of an exemplary application scenario according to some embodiments of the present disclosure is schematically shown;

[0030] Figure 2 A flowchart of a grayscale monitoring method according to some embodiments of the present disclosure is schematically shown;

[0031] Figure 3 A page diagram showing the change trends of grayscale versions and online version metrics according to some embodiments of the present disclosure is schematically shown;

[0032] Figure 4 A flowchart of root cause location detection for abnormal events according to some embodiments of the present disclosure is schematically shown;

[0033] Figure 5 A flowchart of analyzing change events before and after the current version change according to some embodiments of the present disclosure is schematically shown;

[0034] Figure 6 A schematic diagram of fault location according to abnormal metrics according to some embodiments of the present disclosure is schematically shown;

[0035] Figure 7 A schematic block diagram of a grayscale monitoring device according to some embodiments of the present disclosure is schematically shown;

[0036] Figure 8 A schematic diagram of a storage medium according to an exemplary embodiment of the present disclosure is schematically shown;

[0037] Figure 9 A block diagram of an electronic device according to an exemplary embodiment of the invention is schematically shown.

[0038] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts. Detailed Embodiments

[0039] The principles and spirit of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and thus implement the present disclosure, and not to limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to be able to fully convey the scope of the present disclosure to those skilled in the art.

[0040] Those skilled in the art know that the embodiments of the present disclosure can be implemented as a system, device, equipment, method, or computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0041] According to the embodiments of the present disclosure, a grayscale monitoring method, a grayscale monitoring device, a computer-readable storage medium, an electronic device, and a computer program product are provided.

[0042] In this article, it should be understood that the terms involved, for example, grayscale monitoring realizes the capabilities of grayscale alarm and grayscale real-time monitoring through real-time log processing and data storage, combined with subdivision dimensions such as version number, environment, and whether the grayscale dimension information is hit. The goal of the change control system is to prevent and manage constraints in a systematic manner throughout the life cycle of the change; the change in the change control system refers to adding, modifying, or deleting any content that may have a direct or indirect impact on the production environment.

[0043] In addition, the number of any element in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.

[0044] Next, referring to several representative embodiments of the present disclosure, the principles and spirits of the present disclosure will be elaborated in detail. Summary of the Invention

[0046] When software is released in grayscale, a basic grayscale monitoring system is usually used for grayscale monitoring. By technical means, the differences between the change stage indicators and the online version are identified for a basic monitoring to intuitively display various indicators of the entire system operation. For example, the basic grayscale monitoring system can perform real-time monitoring on performance indicators, service status, business indicators, and errors and exceptions.

[0047] However, the above basic grayscale monitoring system cannot support systematic index grading and index classification, and cannot locate the root cause of anomalies based on the monitored abnormal indicators.

[0048] Based on the above, in one embodiment of the present disclosure, the monitoring indicators required for the grayscale release of the application program are determined. The monitoring indicators are determined based on a pre-configured index classification and grading strategy; the grayscale monitoring index values and the online monitoring index values corresponding to the grayscale version and the online version of the application program are obtained; the grayscale monitoring index values and the online monitoring index values are analyzed and compared to determine the root cause of the abnormal event generated during the grayscale monitoring of the application program. During the grayscale monitoring process, the monitoring indicators can be classified and graded to improve the correlation coverage and accuracy of the monitoring indicators; at the same time, it can realize the identification of monitoring anomalies and improve the troubleshooting efficiency.

[0049] The following specifically introduces various non-limiting embodiments of the present disclosure.

[0050] Overview of Application Scenarios

[0051] First, refer to Figure 1 , Figure 1 which shows a schematic block diagram of a system architecture of an exemplary application scenario of a grayscale monitoring method and apparatus to which embodiments of the present disclosure can be applied.

[0052] As Figure 1 shown, the system architecture 100 may include one or more of the terminal devices 101, 102, 103, the network 104, and the server 105. The network 104 is used to provide a medium for a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. The terminal devices 101, 102, 103 may be various electronic devices with a display screen, including but not limited to desktop computers, portable computers, smart phones, and tablet computers, etc. It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0053] are merely illustrative. According to the implementation requirements, there may be any number of terminal devices, networks, and servers. For example, the server 105 may be a server cluster composed of multiple servers, etc.

[0054] It should be understood that Figure 1 the application scenario shown is only an example in which the embodiments of the present disclosure can be implemented. The scope of application of the embodiments of the present disclosure is not limited by any aspect of this application scenario.

[0055] Exemplary Method

[0056] The following combines Figure 1 the application scenarios of Figure 2 to describe the grayscale monitoring method according to the exemplary embodiments of the present disclosure. It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principle of the present disclosure, and the embodiments of the present disclosure are not limited in this regard. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.

[0057] The present disclosure first provides a grayscale monitoring method. The execution subject of this method can be a terminal device or a server, and the present disclosure does not make special limitations in this regard. In this exemplary embodiment, the case where the server executes this method is taken as an example for illustration.

[0058] Referring to Figure 2 as shown, the grayscale monitoring method may include the following steps S210 to S230:

[0059] Step S210, determining the monitoring metrics required for grayscale release of the application program, where the monitoring metrics are determined based on a pre-configured metric classification and grading strategy;

[0060] Step S220, obtaining the grayscale monitoring metric values and online monitoring metric values corresponding to the grayscale version and the online version of the application program respectively;

[0061] Step S230, analyzing and comparing the grayscale monitoring metric values and the online monitoring metric values to determine the root cause of the abnormal event generated during the grayscale monitoring of the application program.

[0062] In the grayscale monitoring method provided by this exemplary embodiment, on the one hand, a pre-configured metric classification and grading strategy is adopted to implement a scheme for classifying and grading the monitoring metrics, improving the coverage and accuracy of the monitoring metrics. On the other hand, by comparing the monitoring metric values of the online version and the grayscale version, the root cause of the abnormal event generated during the grayscale monitoring can be located, effectively improving the troubleshooting efficiency.

[0063] Next, the above steps of this exemplary embodiment will be described in more detail.

[0064] In step S210, the monitoring metrics required for grayscale release of the application program are determined, and the monitoring metrics are determined based on a pre-configured metric classification and grading strategy.

[0065] In some example embodiments, the monitoring metrics may be relevant metrics used to measure and reflect the running status and performance of an application. The metric classification and grading strategy may be a specific strategy for classifying multiple monitoring metrics by type and level according to the business domain.

[0066] When performing a gray release on an application, the monitoring of various metrics of the application can be achieved through a gray monitoring tool, and the monitoring metrics to be monitored during the gray release of the application are determined, such as response time, error rate, etc. To solve the problem that the basic gray monitoring system can only perform basic monitoring and cannot classify and grade the monitoring metrics in a systematic manner, the present disclosure uses a pre-configured metric classification and grading strategy to determine the monitoring metrics required during the gray release process.

[0067] In step S220, the gray monitoring metric values and the online monitoring metric values corresponding to the gray version and the online version of the application are obtained.

[0068] In some example embodiments, the gray monitoring metric values may be specific metrics or status values reflecting the running status and performance of the gray version of the application. The online monitoring metric values may be specific metrics or status values reflecting the running status and performance of the currently running online version of the application.

[0069] After determining the monitoring metrics using the metric classification and grading strategy, the gray monitoring tool can respectively obtain the gray monitoring metric values corresponding to the gray users using the gray version, and the online monitoring metric values corresponding to the users of the old version currently running online using the online version, so as to evaluate the advantages and disadvantages of the new version to be released after comparing and analyzing the two types of monitoring metric values.

[0070] In step S230, the gray monitoring metric values and the online monitoring metric values are analyzed and compared to determine the root cause of the failure of the abnormal event generated during the gray monitoring of the application.

[0071] In some example embodiments, the abnormal event may be an abnormal event generated during the running of the application. The root cause of the failure may refer to the fundamental reason for the failure of the system or application.

[0072] After obtaining the monitoring metric values of the online version and the gray-scale version, a comparative analysis is performed on the gray-scale monitoring metric values and the online monitoring metric values. By analyzing and comparing the above-mentioned monitoring metric values, the advantages and disadvantages of the gray-scale version compared with the current online version can be determined. Additionally, during the gray-scale monitoring process, if an abnormal event occurs in the application program, by analyzing and comparing the monitored monitoring metric values, the root cause of the fault corresponding to the abnormal event can be determined. The abnormal events that may occur in the application program can be classified into two categories. One is serious problems that cannot be handled in the program, such as memory overflow, etc.; the other is various unexpected events that may be encountered in the program, including runtime exceptions and compilation exceptions, etc.

[0073] In an embodiment of the present disclosure, for step S210, determining the monitoring metrics required for the gray-scale release of the application program includes: obtaining the pre-configured metric classification and grading strategy, where the metric classification and grading strategy is generated based on the business domain of the application program; determining the monitoring metrics based on the metric classification and grading strategy, and the monitoring metrics include one or more of application metrics, associated business metrics, metrics associated with the current change, and public opinion metrics.

[0074] Among them, the business domain may be the domain involved in the specific business associated with the application program. The application metrics may be the basic metrics corresponding to the application program. The associated business metrics may be the monitoring metrics corresponding to the business associated with the application program. The metrics associated with the current change may be the operation metrics affected by the current version change of the application program. The public opinion metrics may be the metrics used to measure the attention and discussion degree of users on a certain event of the application program.

[0075] When determining the monitoring metrics during the gray-scale monitoring process, the pre-configured metric classification and grading strategy can be obtained first. The present disclosure can configure the corresponding metric classification and grading strategy based on the business domain of the application program. For example, configuring the metric classification and grading strategy according to the business domain, specifically, it can be classified according to classification methods such as business type and system level.

[0076] The monitoring metrics determined based on the metric classification and grading strategy may include, but are not limited to, application metrics, associated business metrics, metrics associated with the current change, and public opinion metrics, etc. For the above four types of monitoring metrics, they can also be classified according to the importance level, such as dividing the monitoring metrics into first-level metrics (key metrics), second-level metrics (important metrics), and third-level metrics (general metrics).

[0077] Specifically, the first-level indicators (key indicators) can be indicators that have a significant impact on the business and system. Once an anomaly occurs, it may lead to business interruption, significant economic losses, or seriously affect the user experience. For example, the transaction success rate of an e-commerce platform, the availability of a payment system, etc. For the first-level indicators, real-time monitoring is required, strict thresholds and alarm rules are set, and once the indicator exceeds the normal range, an alarm is immediately sent to notify relevant personnel.

[0078] The second-level indicators (important indicators) can be indicators that have a certain impact on the business and system but do not immediately lead to serious consequences. Abnormal situations may affect some functions or performance of the business. For example, page load time, CPU usage rate of the server, etc. For the second-level monitoring indicators, high-frequency monitoring can be adopted, reasonable thresholds and alarm rules are set, and when the indicator shows abnormal fluctuations, a warning is sent in a timely manner.

[0079] The third-level indicators (general indicators) can be indicators that have a relatively small impact on the business and system, mainly used to provide reference information and conduct trend analysis. For example, the number of log records, the space occupied by temporary files of the system, etc. For the third-level indicators, regular monitoring can be adopted, the thresholds and alarm rules can be appropriately relaxed, and the long-term change trend of the indicators is concerned. By classifying and grading the indicators, the coverage and accuracy of the monitoring indicators can be effectively improved.

[0080] In some other exemplary embodiments of the present disclosure, other classification and grading rules can also be adopted according to specific data analysis requirements to configure the indicator classification and grading strategy. For example, the monitoring indicator categories can be divided according to the system level, and for another example, they can be divided into four important levels, five important levels, etc. according to the importance level of the indicators. The present disclosure does not make any special limitations on the specific basis for configuring the indicator classification and grading strategy.

[0081] In an embodiment of the present disclosure, determining the monitoring indicators based on the indicator classification and grading strategy includes: determining the basic indicators associated with the application program and using the basic indicators as application indicators; determining the changed services corresponding to the gray release and the associated businesses of the changed services, and determining the associated business indicators based on the associated businesses and the indicator classification and grading strategy; determining the change-related indicators corresponding to the gray release based on the indicator classification and grading strategy as the change-related indicators for this time; generating public opinion indicators according to the public opinion keywords that need to be concerned about for the gray release of the application program.

[0082] Among them, the changed service can be an application service that changes due to the version change of the application program. The associated business can be the specific business associated with the changed service. The change-related indicator can be an indicator that changes due to the release of this gray version. The public opinion keyword can be the keyword used to determine the public opinion indicator.

[0083] When determining monitoring metrics, basic metrics related to the application can be determined as application metrics. Application metrics can be system performance metrics of the application. For example, application metrics can include, but are not limited to, the availability rate of the Application Programming Interface (API), the availability rate of the Response Time (RT) of the server, throughput, CPU usage, memory usage, disk I / O, network bandwidth, exception logs, etc. These metrics reflect the basic environment status of the system operation, as well as the operation status and performance of the application.

[0084] When determining associated business metrics, the changed services resulting from the current version change can be determined first, and then the associated businesses related to the changed services can be determined. Thus, based on the determined associated businesses and in combination with the metric classification and grading strategy, the associated business metrics can be determined. In this embodiment, the process of determining monitoring metrics according to business service changes when a music platform undergoes a version change will be used as an example for illustration.

[0085] For example, when the underlying service of the music library changes due to a version upgrade of the music platform, the upstream businesses corresponding to the music library service, such as the play service and the playlist service, etc., all need to be concerned and monitored. Business personnel can also set monitoring metrics for specific business processes, such as the order processing duration and payment success rate of the purchase service connected to the music platform. These metrics are used to monitor the efficiency and quality of the business process. Taking the music platform as an example, when the changed business is a music library change, the associated business metrics can include, but are not limited to, the playlist play success rate, the song collection rate, and the core metrics of the music library business.

[0086] The associated business metrics can be classified at the business level and can be divided into core business metrics and auxiliary business metrics. The core business metrics can include the number of registered users, such as very important person (VIP) users, etc.; the auxiliary business metrics can be the display times of specified content and the click volume of users on the specified content, etc.

[0087] The associated business metrics can be classified at the user experience level and can be divided into interaction experience metrics and content experience metrics, etc. The interface operation response time can be the response time of the system when the user performs operations such as play, pause, and song switching. An overly long response time will seriously affect the user experience, and it is necessary to focus on whether the new version causes the operation response to slow down in gray-box monitoring. The function availability can refer to checking whether the new function can be used normally, for example, whether functions such as new song recommendation and personalized playlist have faults or anomalies.

[0088] Content experience metrics may include song playback smoothness and audio quality satisfaction. Song playback smoothness refers to monitoring whether there are issues such as stuttering and buffering during song playback, especially in different network environments. Audio quality satisfaction can be evaluated through user feedback or technical metrics (such as audio bitrate, channels, etc.) to assess the impact of the new version on audio quality.

[0089] When performing a gray release, corresponding change-related metrics can also be determined based on the metric classification and grading strategy as the change-related metrics for this time. The change-related metrics for this time may include processor parameter metrics related to intensive computing, database change parameter metrics related to database performance, data persistence layer change parameter metrics reflecting changes in the data persistence layer, message queue parameter metrics focusing on message queue performance, and custom monitoring parameter metrics determined according to business requirements, etc.

[0090] In addition, sentiment metrics can be generated based on the sentiment keywords that need to be concerned about during the gray release of the application. For example, when a music platform performs a gray release, the sentiment keywords that need to be concerned about for this application change can be "playback", "playlist", "login", etc. Therefore, corresponding sentiment metrics can be generated based on sentiment keywords such as "playback" and "login". Through the above metric classification and grading strategy, a general systematic metric monitoring solution is provided, facilitating the effective monitoring of various metrics during the gray release process.

[0091] In an embodiment of the present disclosure, the change-related metrics for this time include one or more of processor parameter metrics, database change parameter metrics, data persistence layer change parameter metrics, message queue parameter metrics, and custom monitoring parameter metrics.

[0092] Among them, the processor parameter metrics may be parameter metrics reflecting the operating status and performance of the processor. The processor parameter metrics focus on metrics related to intensive computing, such as CPU metrics.

[0093] The database change parameter metrics may be parameter metrics reflecting the operating status and performance of the processor. By means of the database change parameter metrics, changes to the Remote Dictionary Server (Redis) are monitored. For example, monitoring the Redis memory water level, Redis response time, CPU situation of Redis, and whether there is eviction in Redis, etc.

[0094] The change parameter metrics of the data persistence layer can be the parameter metrics reflecting the changes in the data persistence layer (also known as the data access layer, Data Access Object, DAO). By focusing on the change parameter metrics of the data persistence layer, the changes in the DAO layer can include, but are not limited to, monitoring the memory water level of the DAO layer, the response time of the DAO layer, the CPU situation of the DAO layer, and whether there is eviction in the DAO layer.

[0095] The message queue parameter metrics can be the parameter metrics for measuring whether the message queue can ensure the reliable delivery of messages. By focusing on the message queue parameter metrics, the relevant performance changes of the message queue can be concerned, such as message backlog, send failure rate, consumption response time, etc.

[0096] The custom monitoring parameter metrics can be the monitoring metrics set by business personnel according to business requirements. For example, the custom monitoring parameter metrics can include the monitoring of the logging system, such as the monitoring of music logs (Music Log, mlog) in a music platform, which is attached to this application; the custom monitoring parameter metrics can also include the trend of page views (Page View, PV) of client-side data collection points, etc.

[0097] Taking a music platform as an example, taking the change of the song library in the gray version as an example, the associated metrics of this change can include, but are not limited to, the Redis capacity of the song list in the playlist and the CPU performance of the Redis of the song list in the playlist. The above monitoring metrics can be jointly configured by the business person in charge and business developers. Through the associated metrics of this change, the monitoring metrics that need to be concerned about in this gray release can be determined, and the coverage rate of metric monitoring can be improved.

[0098] It is easy for those skilled in the art to understand that the present disclosure is applicable to various types of application programs, such as e-commerce platforms, social platforms, and instant messaging platforms, etc. The present disclosure does not make any special limitations on the specific types of application programs.

[0099] In an embodiment of the present disclosure, for step S220, obtaining the gray monitoring metric values and online monitoring metric values corresponding to the gray version and the online version of the application program respectively includes: determining the monitoring scope of gray monitoring, where the monitoring scope includes one or more of the monitored user scope, the monitored business scope, and the specified monitoring time period; determining the gray monitoring metric values corresponding to the gray version based on the monitoring scope and the monitoring metrics; determining the online monitoring metric values corresponding to the online version based on the monitoring scope and the monitoring metrics.

[0100] Among them, the monitoring scope can be the specific scope monitored by the grayscale monitoring system. The monitored user scope can be the scope of users participating in the current version change test. The monitored business scope can be the business scope that the grayscale change system needs to monitor during the current application version change. The specified monitoring time period can be the time interval that the grayscale change system needs to monitor during the current application version change.

[0101] Before performing grayscale monitoring, the application can remind users to participate in the version test of the current application. After the users agree to participate in the current version change test, relevant metrics of the application are analyzed based on the user feedback data to optimize the application according to the user feedback data. In addition, the present disclosure can also perform grayscale monitoring by collecting user feedback data of specified test users on the specified test users. Before performing grayscale monitoring, the grayscale monitoring scope can be determined first. For example, the current grayscale monitoring scope can be determined according to the monitored user scope, the monitored business scope, the specified monitoring time period, etc.

[0102] Regarding the monitored user scope, during the grayscale release process, when a certain business of the application changes, only the usage data of the test users using this part of the business can be analyzed. Therefore, the monitored user scope is determined based on the users using the specific business. Regarding the monitored business scope, during the grayscale release, the specific business that needs to be focused on in the current release can be determined, and the monitored business scope can be determined.

[0103] Regarding the monitored time period, software testing and operation and maintenance personnel need to pre-determine the specific time period that needs to be focused on in the current grayscale release. For example, if it is necessary to analyze the monitoring metric data for the three months before the current time point, at this time, the specified monitoring time period can be determined as the time interval from the most recent three months before the current time point to the end of the grayscale release. In addition, the time period that needs to be focused on every day can also be configured. For example, 18:00-22:00 every day can be used as the time period that needs to be focused on.

[0104] After determining the above monitoring scope, according to the pre-configured monitoring metrics, the grayscale monitoring metric values corresponding to the grayscale version and the online monitoring metric values corresponding to the online version can be obtained respectively, and the monitored monitoring metric values are displayed on the release metric monitoring details page. Refer to Figure 3 , Figure 3 schematically shows a page diagram showing the change trends of the grayscale version and the online version metrics according to some embodiments of the present disclosure. Figure 3 The change trends of the different monitoring metric values in [] are only schematic examples, and the actual change trends of the various monitoring metric values can be displayed according to the specific operation conditions of the application. Through the metric monitoring details page, relevant personnel can timely monitor various operating states of the application and master the real-time operating conditions.

[0105] In one embodiment of the present disclosure, for step S230, the gray - scale monitoring index value and the online monitoring index value are analyzed and compared to determine the root cause of the abnormal event generated during the gray - scale monitoring of the application program, including: analyzing and comparing the gray - scale monitoring index value and the online monitoring index value to determine the initial failure type corresponding to the abnormal event; determining the change event corresponding to the gray - scale release, and obtaining the event - related data of the change event; combining the initial failure type and the event - related data to locate the root cause of the failure.

[0106] Among them, the initial failure type can be the failure type determined by analyzing and comparing the monitoring index values of the gray - scale version and the online version. The change event can be all events that change before and after the version change determined by the monitoring platform during gray - scale monitoring. The event - related data can be the data related to the change event.

[0107] During the gray - scale monitoring process, the monitored gray - scale monitoring index value and the online monitoring index value can be intuitively displayed in a data chart, and the monitoring index values in the data chart are analyzed and compared to determine whether an abnormal event occurs during the program operation through numerical comparison. When an abnormal event occurs during the program operation, the initial failure type corresponding to the abnormal event can be determined by analyzing the monitoring index value.

[0108] In addition, the change event corresponding to the gray - scale release can be determined, and the event - related data of the change event can be obtained. For example, when the music library service is changed during this gray - scale release, the change event can be the music library change. At this time, the monitoring index value related to the music library change can be determined as the event - related data. By analyzing the above - mentioned event - related data and the initial failure type, the root cause of the abnormal event can be finally located. The present disclosure can automatically perform fault diagnosis using existing platform tools and artificial intelligence (AI) large models, so as to realize the automated diagnosis of faults using gray - scale monitoring.

[0109] In one embodiment of the present disclosure, analyzing and comparing the gray - scale monitoring index value and the online monitoring index value to determine the initial failure type corresponding to the abnormal event includes: generating the time - series data to be analyzed according to the gray - scale monitoring index value and the online monitoring index value; performing data pre - processing on the time - series data to be analyzed to obtain the index data to be analyzed; performing a bottom - up pre - pruning process on the index data to be analyzed to obtain the target associated index data; performing a top - down abnormal location process on the target associated index data to obtain the initial abnormal location result; performing root - cause location and voting decision processing on the initial abnormal location result to obtain the initial failure type.

[0110] Among them, the time-series data to be analyzed can be a data set obtained by arranging the monitored metric values collected during the gray-box monitoring process in chronological order. The metric data to be analyzed can be the data to be analyzed obtained after preprocessing the time-series data to be analyzed. The target associated metric data can be the metric data for analyzing the root cause of a fault obtained after pre-cropping processing. The initial anomaly location result can be the initial result of the root cause of the fault determined after analyzing the target associated metric data. The initial fault type can be the fault type determined after analyzing the monitored metric values of the online version and the gray-box version.

[0111] Reference Figure 4 , Figure 4 schematically shows a flowchart for root cause location detection of an abnormal event according to some embodiments of the present disclosure. The present disclosure provides a solution for root cause location of an abnormal event by combining the Squeeze root cause detection algorithm. For the obtained gray-box monitored metric values and online monitored metric values, after arranging the above data in chronological order, time-series data 410 to be analyzed is generated. According to Figure 4 it can be seen that in step S410, the generated time-series data 410 to be analyzed is preprocessed to obtain the metric data to be analyzed. The data preprocessing can include, but is not limited to, data cleaning, data integration, data transformation, and data reduction and other processing methods.

[0112] In step S420, root cause location processing is performed based on the metric data to be analyzed. Specifically, first, the metric data to be analyzed is pre-cropped from bottom to top to obtain the target associated metric data; then, abnormal location processing is performed on the target associated metric data from top to bottom to obtain the initial anomaly location result; and then, root cause location and voting decision processing are performed on the initial anomaly location result to obtain the initial fault type, which is used as the root cause location result 420. The determined initial fault type is used as the root cause location result of the abnormal event determined by the Squeeze root cause detection algorithm. By performing root cause location on the monitored metric values, the initial fault type corresponding to the abnormal event can be determined.

[0113] The present disclosure can also use other root cause location algorithms for root cause location detection, and the present disclosure does not make any special limitations on the specifically used root cause location algorithm.

[0114] In one embodiment of the present disclosure, a change event corresponding to the gray release is determined, and event-related data of the change event is obtained, including: determining an event center associated with the monitoring platform, and obtaining the change content within a specified time interval before and after the change of the gray version based on the event center, where the change content includes server-side change data and client-side change data; determining, through an activity guarantee platform associated with the monitoring platform, the change activity data within a specified time interval before and after the change of the gray version; determining the upstream and downstream applications of the application program, and obtaining the application change data of the upstream and downstream applications within a specified time period of the change of the gray version; determining the abnormal index data corresponding to the online version within a specified time period of the change of the gray version.

[0115] Among them, the event center can be a platform center for obtaining various change events. The change content can be the content that changes when the application program changes in this version. The server-side change data can be the data that changes on the server side. The client-side change data can be the data that changes on the client side. The activity guarantee platform can be a platform for obtaining major activities generated within a specified time interval of the application program version change. The change activity data can be the change index data corresponding to the major activities. The upstream and downstream applications can be the upstream application and the downstream application associated with the application program. The application change data can be the change data generated by the upstream and downstream applications. The abnormal index data can be the abnormal index data corresponding to the online version.

[0116] The gray monitoring platform (also known as the monitoring platform) can obtain all change events before and after this change based on the relevance of the change. Determine the event center associated with the monitoring platform, and obtain the change content within a specified time interval before and after the change of the gray version through the event center. For example, obtain all changes within a one-week time interval before and after the version change through the time center as the content corresponding to the change event. Specifically, it can include server-side changes, such as code and configuration changes on the server side, which are used as server-side change data; it can also include client-side changes, including code and configuration changes on the client side, that is, client-side change data.

[0117] Determine the activity guarantee platform associated with the monitoring platform, and obtain the change activity data within a specified time interval before and after the change of the gray version through the activity guarantee platform associated with the monitoring platform. For example, the activity guarantee platform can obtain major activities generated within a three-day time interval before and after this version change. Taking the music library change of the music platform as an example, the activity guarantee platform can obtain the activity data related to the music library change within three days before and after the gray version release.

[0118] For the upstream and downstream applications related to the application, during the gray-box monitoring period, it is possible to obtain the application change data of the upstream and downstream applications within the specified time period of the current gray-scale version change. For example, obtain the changes of the core dependent applications of the upstream and downstream of the released application within the one-week time interval before and after the version change, and maintain them in the process management tool of the application program (such as the overmind application).

[0119] During the gray-box monitoring process, it is also possible to determine the abnormal metric data corresponding to the online version within the specified time period of the current gray-scale version change. For example, determine whether there are other online problems occurring within the three-day time interval before and after the current version change. By obtaining the event-related data of the above-dimensional change events, the above event-related data can be used as the data basis for finally locating the root cause of the failure.

[0120] In one embodiment of the present disclosure, determine the analysis dimensions corresponding to the event-related data, where the analysis dimensions include one or more of the associated fields, associated applications, and associated resources; based on the associated fields, determine the associated field metric values corresponding to the event-related data; based on the associated applications, determine the associated application metric values corresponding to the event-related data; based on the associated resources, determine the associated resource metric values corresponding to the event-related data; monitor the associated field metric values, associated application metric values, and associated resource metric values through a monitoring list.

[0121] Among them, the associated field metric value can be the parameter value of the core metric corresponding to the field associated with the application program. The associated application metric value can be the parameter value of the core metric corresponding to the upstream and downstream applications or other applications associated with the application program. The associated resource metric value can be the parameter value of the core metric corresponding to various resources associated with the application program.

[0122] Reference Figure 5 , Figure 5 schematically shows a flowchart for analyzing the change events before and after the current version change according to some embodiments of the present disclosure. As can be seen from Figure 5 in the process of analyzing the change, it is possible to first determine the analysis dimensions corresponding to the event-related data. For example, the analysis dimensions can include, but are not limited to, one or more of the associated fields, associated applications, and associated resources.

[0123] Specifically, based on the associated field of the application program, the core associated field metric corresponding to the event-related data can be determined, the associated field metric value corresponding to the core associated field metric can be obtained, and its real-time change trend can be displayed in the monitoring list.

[0124] Based on the associated applications of the application program, the core associated application metric corresponding to the event-related data can be determined, the associated field metric value corresponding to the core associated application metric can be obtained, and the real-time change trend of the above metric values can be displayed in the monitoring list.

[0125] Determine the associated resources of the application. For example, the associated resources may include database resources, etc. According to the associated resources, the associated resource metric values corresponding to the event-related data can be determined, and the associated resource metric values are monitored through a monitoring list to observe the change trend of the associated resource metric values. When performing root cause localization of abnormal events, the final root cause of the abnormality can be located by combining the event-related data associated with the application.

[0126] Reference Figure 6 , Figure 6 Schematically shows a schematic diagram of fault localization according to abnormal metrics according to some embodiments of the present disclosure. Figure 6 It is shown that when an abnormal event of user failure to obtain a song list (playlist) occurs, by analyzing the above abnormal event, it can be determined that the failure to obtain the playlist is caused by the failure of a Remote Procedure Call (RPC). Further analyzing the reason for the RPC call failure, it can be located that the RPC call fails due to the increase in RPC traffic. Therefore, the ultimate root cause of the playlist call failure is the increase in RPC traffic. For example, within the last 5 minutes, relevant experiments were launched, resulting in an increase in traffic, thus causing the failure to retrieve the playlist. Software testers or operation and maintenance personnel can display corresponding abnormal event prompt information through a display device according to business requirements.

[0127] In summary, the gray-scale monitoring method of the present disclosure determines the monitoring metrics required for gray-scale release of the application, and the monitoring metrics are determined based on a pre-configured metric classification and grading strategy; obtains the gray-scale monitoring metric values and online monitoring metric values corresponding to the gray-scale version and the online version of the application respectively; analyzes and compares the gray-scale monitoring metric values and the online monitoring metric values to determine the root cause of the failure of abnormal events generated during the gray-scale monitoring of the application. The present disclosure can classify and grade the monitoring metrics during the gray-scale monitoring process, improving the associated coverage and accuracy of the monitoring metrics; at the same time, it can achieve monitoring anomaly recognition and improve the troubleshooting efficiency. On the one hand, a pre-configured metric classification and grading strategy is adopted to implement a solution for classifying and grading the monitoring metrics, improving the coverage and accuracy of the monitoring metrics. On the other hand, by comparing the monitoring metric values of the online version and the gray-scale version, the root cause of the abnormal events generated during the gray-scale monitoring can be located, effectively improving the troubleshooting efficiency.

[0128] Exemplary Device

[0129] After introducing the method of the exemplary embodiment of the present disclosure, next, reference Figure 7 is made to illustrate the gray-scale monitoring device of the exemplary embodiment of the present disclosure.

[0130] In Figure 7Among them, the grayscale monitoring device 700 may include: a monitoring metric determination module 710, configured to determine the monitoring metrics required for grayscale release of an application, where the monitoring metrics are determined based on a pre-configured metric classification and grading policy; a metric value acquisition module 720, configured to acquire the grayscale monitoring metric values and online monitoring metric values corresponding to the grayscale version and the online version of the application respectively; and a root cause determination module 730, configured to analyze and compare the grayscale monitoring metric values and the online monitoring metric values to determine the root cause of an abnormal event generated during the grayscale monitoring of the application.

[0131] In an embodiment of the present disclosure, the monitoring metric determination module 710 includes a monitoring metric determination unit, configured to: acquire a pre-configured metric classification and grading policy, where the metric classification and grading policy is generated based on the business domain of the application; and determine the monitoring metrics based on the metric classification and grading policy, where the monitoring metrics include one or more of application metrics, associated business metrics, metrics associated with the current change, and public opinion metrics.

[0132] In an embodiment of the present disclosure, the monitoring metric determination unit includes a monitoring metric determination subunit, configured to: determine the basic metrics associated with the application and use the basic metrics as application metrics; determine the changed services corresponding to the grayscale release and the associated businesses of the changed services, and determine the associated business metrics based on the associated businesses and the metric classification and grading policy; determine the change-associated metrics corresponding to the grayscale release based on the metric classification and grading policy as the metrics associated with the current change; and generate public opinion metrics according to the public opinion keywords that need to be concerned about during the grayscale release of the application.

[0133] In an embodiment of the present disclosure, the metric value acquisition module 720 includes a metric value acquisition unit, configured to: determine the monitoring scope of grayscale monitoring, where the monitoring scope includes one or more of a monitoring user scope, a monitoring business scope, and a specified monitoring time period; determine the grayscale monitoring metric values corresponding to the grayscale version based on the monitoring scope and the monitoring metrics; and determine the online monitoring metric values corresponding to the online version based on the monitoring scope and the monitoring metrics.

[0134] In an embodiment of the present disclosure, the root cause determination module 730 includes a root cause determination unit, configured to: analyze and compare the grayscale monitoring metric values and the online monitoring metric values to determine the initial failure type corresponding to the abnormal event; determine the change event corresponding to the grayscale release and acquire the event-related data of the change event; and locate the root cause of the failure by combining the initial failure type and the event-related data.

[0135] In one embodiment of the present disclosure, the fault root cause determination unit includes a fault type determination subunit, which is configured to: generate time series data to be analyzed according to the gray-scale monitoring metric values and the online monitoring metric values; perform data preprocessing on the time series data to be analyzed to obtain the metric data to be analyzed; perform a bottom-up pre-clipping process on the metric data to be analyzed to obtain the target associated metric data; perform a top-down anomaly localization process on the target associated metric data to obtain an initial anomaly localization result; perform a root cause localization and voting decision process on the initial anomaly localization result to obtain an initial fault type.

[0136] In one embodiment of the present disclosure, the fault root cause determination unit includes an event data acquisition subunit, which is configured to: determine the event center associated with the monitoring platform, and obtain the change content within a specified time interval before and after the current gray-scale version change based on the event center, where the change content includes server-side change data and client-side change data; determine the change activity data within a specified time interval before and after the current gray-scale version change through the activity guarantee platform associated with the monitoring platform; determine the upstream and downstream applications of the application program, and obtain the application change data of the upstream and downstream applications within the specified time period of the current gray-scale version change; determine the abnormal metric data corresponding to the online version within the specified time period of the current gray-scale version change.

[0137] In one embodiment of the present disclosure, the gray-scale monitoring device 700 further includes a metric monitoring module, which is configured to: determine the analysis dimension corresponding to the event-related data, where the analysis dimension includes one or more of the associated domain, the associated application, and the associated resource; determine the associated domain metric value corresponding to the event-related data based on the associated domain; determine the associated application metric value corresponding to the event-related data based on the associated application; determine the associated resource metric value corresponding to the event-related data based on the associated resource; monitor the associated domain metric value, the associated application metric value, and the associated resource metric value through the monitoring list.

[0138] Since each functional module of the gray-scale monitoring device in the exemplary embodiments of the present disclosure corresponds to the steps in the exemplary embodiments of the above gray-scale monitoring method, for details not disclosed in the embodiments of the present disclosure's device, please refer to the embodiments of the above gray-scale monitoring method of the present disclosure, which will not be elaborated herein.

[0139] It should be noted that although several modules or units of the gray-scale monitoring device are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0140] Exemplary Medium

[0141] After introducing the device of the exemplary embodiments of the present disclosure, next, reference is made to Figure 8 the storage medium of the exemplary embodiments of the present disclosure is described.

[0142] In some embodiments, various aspects of the present disclosure can also be implemented as a medium having program code stored thereon, which is used to implement the steps in the grayscale monitoring method according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification when the program code is executed by a processor of a device.

[0143] For example, when the processor of the device executes the program code, it can implement steps such as Figure 2 step S210 described in, determining the monitoring metrics required for grayscale release of the application program, where the monitoring metrics are determined based on a pre-configured metrics classification and grading strategy; step S220, obtaining the grayscale monitoring metric values and online monitoring metric values corresponding to the grayscale version and the online version of the application program respectively; step S230, analyzing and comparing the grayscale monitoring metric values and the online monitoring metric values to determine the root cause of the abnormal event generated during the grayscale monitoring of the application program.

[0144] Reference is made to Figure 8 shown, a program product 800 for implementing the above grayscale monitoring method or implementing the above grayscale monitoring method according to an embodiment of the present disclosure is described, which can use a portable compact disc read-only memory (CD-ROM) and include program code, and can run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto.

[0145] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0146] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium.

[0147] Program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN).

[0148] Exemplary Computing Device

[0149] After introducing the grayscale monitoring method, grayscale monitoring device, and storage medium of the exemplary embodiments of the present disclosure, next, reference is made to Figure 9 to describe the electronic device of the exemplary embodiments of the present disclosure.

[0150] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, method, or program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, which can be collectively referred to as "circuitry", "module", or "system" here.

[0151] In some possible embodiments, the electronic device according to the present disclosure may at least include at least one processing unit and at least one storage unit. Among them, the storage unit stores program code, and when the program code is executed by the processing unit, the processing unit is caused to execute the steps in the grayscale monitoring method according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification. For example, the processing unit may execute steps S210 as shown in Figure 2 to determine the monitoring metrics required for grayscale release of the application program, where the monitoring metrics are determined based on a pre-configured metric classification and grading strategy; step S220, to obtain the grayscale monitoring metric values and online monitoring metric values corresponding to the grayscale version and the online version of the application program respectively; step S230, to analyze and compare the grayscale monitoring metric values and the online monitoring metric values to determine the root cause of the abnormal event generated during the grayscale monitoring of the application program.

[0152] Next, reference is made to Figure 9 to describe the electronic device 900 according to the exemplary embodiments of the present disclosure. Figure 9The illustrated electronic device 900 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0153] As Figure 9 shown, the electronic device 900 is presented in the form of a general-purpose computing device. The components of the electronic device 900 may include, but are not limited to: at least one of the above-mentioned processing units 901, at least one of the above-mentioned storage units 902, a bus 903 connecting different system components (including the storage unit 902 and the processing unit 901), and a display unit 907.

[0154] The bus 903 represents one or more of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the multiple bus architectures.

[0155] The storage unit 902 may include a readable medium in the form of volatile memory, such as random access memory (RAM) 921 and / or cache memory 922, and may further include read-only memory (ROM) 923.

[0156] The storage unit 902 may also include a program / utilities 925 having a set (at least one) of program modules 924. Such program modules 924 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0157] The electronic device 900 may also communicate with one or more external devices 904 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 900, and / or communicate with any device that enables the electronic device 900 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through an input / output (I / O) interface 905. And, the electronic device 900 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 906. As shown in the figure, the network adapter 906 communicates with other modules of the electronic device 900 through the bus 903. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 900, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0158] It should be noted that although several units / modules or sub-units / modules of the grayscale monitoring device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / modules. Conversely, the features and functions of one unit / modules described above can be further divided and embodied by multiple units / modules.

[0159] In addition, although the operations of the method of the present disclosure are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.

[0160] Although the spirit and principles of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the specific embodiments disclosed, and the division of each aspect does not mean that the features in these aspects cannot be combined for benefits. This division is only for the convenience of expression. The present disclosure aims to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. A grayscale monitoring method, characterized in that: include: Determine the monitoring indicators required for the grayscale release of the application, where the monitoring indicators are determined based on a pre-configured indicator classification and grading strategy; Obtaining the grayscale monitoring index value and the online monitoring index value corresponding to the grayscale version and the online version of the application program respectively; The grayscale monitoring index value and the online monitoring index value are analyzed and compared to determine the root cause of the abnormal event generated during the grayscale monitoring of the application.

2. The method according to claim 1, characterized in that The monitoring indicators required for the grayscale release of the application include: Obtaining a pre-configured indicator classification and grading strategy, wherein the indicator classification and grading strategy is generated based on the business domain of the application; The monitoring indicators are determined based on the indicator classification and grading strategy, and the monitoring indicators include one or more of application indicators, related business indicators, current change related indicators and public opinion indicators.

3. The method according to claim 2, characterized in that The determining the monitoring indicator based on the indicator classification and grading strategy includes: Determine a basic indicator associated with the application program, and use the basic indicator as the application indicator; Determine the changed service corresponding to the phased release and the associated business of the changed service, and determine the associated business indicator based on the associated business and the indicator classification and grading strategy; Determine the change-related indicator corresponding to the phased release based on the indicator classification and grading strategy as the change-related indicator for this time; The public opinion index is generated according to the public opinion keywords that need to be paid attention to for the grayscale release of the application.

4. The method according to claim 1, characterized in that: The analyzing and comparing the grayscale monitoring index value with the online monitoring index value to determine the root cause of the abnormal event generated during the grayscale monitoring of the application program includes: Analyze and compare the grayscale monitoring index value with the online monitoring index value to determine the initial fault type corresponding to the abnormal event; Determine a change event corresponding to the phased release, and obtain event-related data of the change event; The root cause of the fault is located by combining the initial fault type with the event-related data.

5. The method according to claim 4, characterized in that The analyzing and comparing the grayscale monitoring index value with the online monitoring index value to determine the initial fault type corresponding to the abnormal event includes: Generate time series data to be analyzed according to the grayscale monitoring index value and the online monitoring index value; Performing data preprocessing on the time series data to be analyzed to obtain indicator data to be analyzed; Performing bottom-up pre-clipping processing on the indicator data to be analyzed to obtain target-related indicator data; Performing top-down anomaly location processing on the target-related indicator data to obtain an initial anomaly location result; The root cause location and voting decision processing are performed on the initial abnormality location result to obtain the initial fault type.

6. The method according to claim 4, characterized in that The determining a change event corresponding to the phased release and acquiring event-related data of the change event includes: Determine the event center associated with the monitoring platform, and obtain the change content within a specified time interval before and after the gray version change based on the event center, wherein the change content includes the server change data and the client change data; Determine the change activity data within a specified time interval before and after the gray version change through the activity assurance platform associated with the monitoring platform; Determine the upstream and downstream applications of the application, and obtain application change data of the upstream and downstream applications within a specified time period of this gray version change; Determine the abnormal indicator data corresponding to the online version within a specified time period of this gray version change.

7. The method according to claim 6, characterized in that The method further comprises: Determine an analysis dimension corresponding to the event-related data, where the analysis dimension includes one or more of a related field, a related application, and a related resource; Based on the associated field, determining an associated field indicator value corresponding to the event-related data; Based on the associated application, determining an associated application indicator value corresponding to the event-related data; Based on the associated resources, determining an associated resource indicator value corresponding to the event-related data; The associated field indicator value, the associated application indicator value and the associated resource indicator value are monitored through a monitoring list.

8. A grayscale monitoring device, characterized in that: include: A monitoring indicator determination module is used to determine the monitoring indicators required for the grayscale release of the application program, wherein the monitoring indicators are determined based on a pre-configured indicator classification and grading strategy; An indicator value acquisition module, used to obtain the grayscale monitoring indicator value and the online monitoring indicator value corresponding to the grayscale version and the online version of the application program respectively; The fault root cause determination module is used to analyze and compare the grayscale monitoring index value and the online monitoring index value to determine the fault root cause of the abnormal event generated during the grayscale monitoring of the application.

9. An electronic device, characterized in that: include: processor; as well as A memory having computer-readable instructions stored thereon, wherein the computer-readable instructions, when executed by the processor, implement the grayscale monitoring method as described in any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the grayscale monitoring method according to any one of claims 1 to 7 is implemented.