Strategy adjustment method and device, computer equipment and storage medium
By performing simulation monitoring and policy adjustments within the server historical time period, the problem of low efficiency in monitoring strategy adjustment in the existing technology is solved, and faster and more efficient policy adjustment and fault handling are achieved.
Patent Information
- Application Number
- CN202311614382.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-28
- Publication Date
- 2025-05-30
AI Technical Summary
The existing technology is inefficient when adjusting server monitoring strategies, requiring multiple trial runs and adjustments, resulting in a long monitoring strategy formulation cycle and the inability to locate and handle server failures in a timely manner.
By obtaining the server log and actual failure information in the historical time period of the target server, conduct simulation monitoring and adjust the simulation strategy to improve the adjustment efficiency of the monitoring strategy.
This method can efficiently adjust the monitoring strategy without waiting for multiple cycles, shorten the policy formulation cycle, and improve the efficiency of server fault location and processing.
Smart Images

Figure CN120066907A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a method and apparatus for policy adjustment, a computer device, and a storage medium. Background Art
[0002] During the operation of a server, various fault problems may occur, such as memory faults, hard disk faults, operating system faults, application program faults, and network faults, etc. When a server fails, it will not only cause the website or application program hosted on the server to respond slowly or even be inaccessible, but may also cause various security problems such as data corruption, data loss, and data leakage, thus seriously affecting the normal operation of the network. A reasonable monitoring policy for server faults can help relevant personnel quickly locate and handle server faults. Therefore, how to adjust the monitoring policy for server faults to formulate a reasonable monitoring policy is a problem that needs to be solved.
[0003] Currently, the monitoring policy is usually directly put into trial operation in the production environment and continuously adjusted. After relevant personnel formulate a monitoring policy based on experience, they directly put the monitoring policy into trial operation in the production environment, and evaluate the effectiveness of the monitoring policy and adjust the monitoring policy according to the fault information monitored within a specified period and the actual fault information that occurred, and continuously repeat the above steps until a reasonable monitoring policy is obtained.
[0004] However, when the monitoring policy is always unreasonable, the monitoring policy needs to be adjusted multiple times to finally determine a reasonable monitoring policy. Since the monitoring policy needs to be re-put into trial operation in the production environment after each adjustment, the efficiency of formulating a reasonable monitoring policy is low. At the same time, due to the long cycle of formulating a reasonable monitoring policy, there may be a lack of a reasonable monitoring policy in the production environment, and server faults cannot be located and handled in a timely manner, thereby reducing production efficiency. Summary of the Invention
[0005] Embodiments of this application provide a method and apparatus for policy adjustment, a computer device, and a storage medium. When the monitoring policy is adjusted multiple times, the server logs in the target historical time period can be called, without waiting for multiple cycles, which is more efficient and time-saving. The technical solutions are as follows:
[0006] On the one hand, a method for policy adjustment is provided, and the method includes:
[0007] Obtain multiple server logs and actual fault information of a target server in a target historical time period, where the actual fault information is used to indicate the server faults that actually occurred on the target server in the target historical time period;
[0008] For any type of fault, based on the simulation strategy for the type of fault and the multiple server logs, perform simulation monitoring on the target server to obtain simulation fault information, where the simulation strategy is used to determine whether a server fault of the type of fault has occurred;
[0009] Based on the difference between the simulation fault information and the actual fault information, adjust the simulation strategy for the type of fault.
[0010] On the other hand, a strategy adjustment device is provided, and the device includes:
[0011] A first acquisition module, configured to acquire multiple server logs and actual fault information of a target server within a target historical time period, where the actual fault information is used to indicate server faults that actually occurred on the target server within the target historical time period;
[0012] A simulation module, configured to, for any type of fault, based on the simulation strategy for the type of fault and the multiple server logs, perform simulation monitoring on the target server to obtain simulation fault information, where the simulation strategy is used to determine whether a server fault of the type of fault has occurred;
[0013] An adjustment module, configured to adjust the simulation strategy for the type of fault based on the difference between the simulation fault information and the actual fault information.
[0014] In some embodiments, the simulation module includes:
[0015] A screening unit, configured to, for any type of fault, based on multiple data pre-screening keywords in the simulation strategy for the type of fault, screen out at least one valid log from the multiple server logs, where the valid log is a server log containing any data pre-screening keyword;
[0016] A monitoring unit, configured to perform simulation monitoring on the target server based on the simulation strategy for the type of fault and the valid log to obtain the simulation fault information.
[0017] In some embodiments, the monitoring unit is configured to process the valid log based on an exception keyword extraction rule in the simulation strategy for the type of fault to obtain multiple exception keywords, the occurrence times of each exception keyword, and the occurrence moments of each exception keyword, where the exception keyword extraction rule is used to extract exception keywords in a regular expression manner; determine the simulation fault information based on an alarm determination rule in the simulation strategy for the type of fault, where the alarm determination rule is used to determine whether to trigger a fault alarm based on the occurrence frequency of the exception keyword.
[0018] In some embodiments, the monitoring unit is further configured to, for any moment when an abnormal keyword appears, determine the number of occurrences of the abnormal keyword within the target time range corresponding to the moment, where the end moment of the target time range is the moment; when the number of occurrences is not less than a threshold number, determine to trigger a fault alarm; and add the alarm fault list of the fault alarm to the simulation fault information.
[0019] In some embodiments, the monitoring unit is further configured to, for any moment during the simulation monitoring process, determine the number of occurrences of any abnormal keyword within the target time range corresponding to the moment, where the end moment of the target time range is the moment; when the number of occurrences is not less than a threshold number, determine to trigger a fault alarm; and add the alarm fault list of the fault alarm to the simulation fault information.
[0020] In some embodiments, the adjustment module includes:
[0021] A first determination unit, configured to determine a simulation strategy misjudgment rate based on the simulation fault information and the actual fault information, where the simulation strategy misjudgment rate is used to represent the accuracy of the simulation strategy of the fault type in monitoring the target server;
[0022] A second determination unit, configured to determine a similar strategy misjudgment rate based on the actual fault information and a similar strategy of the simulation strategy, where the similar strategy misjudgment rate is used to represent the accuracy of the similar strategy in monitoring the target server;
[0023] An adjustment unit, configured to adjust the simulation strategy of the fault type when the simulation strategy misjudgment rate is not less than the similar strategy misjudgment rate.
[0024] In some embodiments, the first determination unit is configured to determine a first quantity and a second quantity based on the simulation fault information and the actual fault information, where the first quantity is the number of first fault lists, and the first fault list is an alarm fault list that exists in both the simulation fault information and the actual fault information, and the second quantity is the number of second fault lists, and the second fault list is the first fault list marked as a false alarm; and determine the ratio of the second quantity to the first quantity as the simulation strategy misjudgment rate.
[0025] In some embodiments, the second determination unit is configured to determine a third quantity and a fourth quantity based on the actual fault information and a similar strategy of the simulation strategy. The third quantity is the quantity of third fault tickets, and the third fault tickets are alarm fault tickets determined by the similar strategy in the actual fault information. The fourth quantity is the quantity of fourth fault tickets, and the fourth fault tickets are third fault tickets marked as false alarms. The ratio of the fourth quantity to the third quantity is determined as the misjudgment rate of the similar strategy.
[0026] In some embodiments, the adjustment module is further configured to add the simulation strategy of the fault type to the monitoring strategy library when the misjudgment rate of the simulation strategy is less than the misjudgment rate of the similar strategy. The monitoring strategy library is used to provide monitoring strategies required for real-time fault monitoring.
[0027] In some embodiments, the target server includes multiple servers. One server can correspond to multiple alarm fault tickets, and one server corresponds to one fault address. The apparatus further includes:
[0028] A second acquisition module, configured to acquire a spare part threshold corresponding to the target server. The spare part threshold is the quantity of spare servers.
[0029] A first determination module, configured to determine an absolute quantity of fault tickets based on the simulated fault information. The absolute quantity of fault tickets is the quantity of fault addresses that appear in the simulated fault information.
[0030] A first display module, configured to display a prompt message when the absolute quantity of fault tickets is greater than the spare part threshold. The prompt message is used to prompt the absolute quantity of fault tickets, the spare part threshold, and at least one fault address.
[0031] In some embodiments, the apparatus further includes:
[0032] A second determination module, configured to determine a third quantity and a fifth quantity based on the simulated fault information and the actual fault information. The third quantity is the quantity of third fault tickets, and the third fault tickets are alarm fault tickets determined by the similar strategy of the simulation strategy in the actual fault information. The fifth quantity is the quantity of fifth fault tickets, and the fifth fault tickets are alarm fault tickets determined by the simulation strategy of the fault type in the simulated fault information.
[0033] A third determination module, configured to determine the ratio of the difference between the fifth quantity and the third quantity to the third quantity as the fault ticket change ratio. The fault ticket change ratio is used to represent the change degree of the monitoring effect of the simulation strategy of the fault type compared with the similar strategy.
[0034] A second display module, configured to display the change ratio of the trouble ticket on a display page, where the display page is used to display comparison information between the simulated fault information and the actual fault information.
[0035] In some embodiments, the apparatus further includes:
[0036] A third acquisition module, configured to acquire data of a parsing configuration file and a server cluster, where the server cluster includes the target server;
[0037] A parsing module, configured to parse the data of the server cluster based on the parsing configuration file to obtain server logs of the target server.
[0038] On the other hand, a computer device is provided, where the computer device includes a processor and a memory, and the memory is used to store at least one segment of computer program, and the at least one segment of computer program is loaded and executed by the processor to implement the policy adjustment method in the embodiments of the present application.
[0039] On the other hand, a computer-readable storage medium is provided, where at least one segment of computer program is stored in the computer-readable storage medium, and the at least one segment of computer program is loaded and executed by a processor to implement the policy adjustment method in the embodiments of the present application.
[0040] On the other hand, a computer program product is provided, including a computer program, and the computer program is executed by a processor to implement the policy adjustment method in the embodiments of the present application.
[0041] The present application provides a policy adjustment method, which performs simulation monitoring on a server within a target historical time period according to a simulation policy and multiple server logs, and adjusts the simulation policy according to the difference between the obtained simulated fault information and the actual fault information. When the monitoring policy is adjusted multiple times, compared with the method of directly putting the monitoring policy into trial operation in the live network and adjusting it multiple times in the traditional method, this method can call the server logs within the target historical time period, without waiting for multiple cycles, and is more efficient and time-saving. Description of the Drawings
[0042] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0043] Figure 1 It is a schematic diagram of an implementation environment of a policy adjustment method provided according to an embodiment of the present application;
[0044] Figure 2 It is a flowchart of a policy adjustment method provided according to an embodiment of the present application;
[0045] Figure 3 It is a schematic diagram of a policy simulation system provided according to an embodiment of the present application;
[0046] Figure 4 It is a flowchart of another policy adjustment method provided according to an embodiment of the present application;
[0047] Figure 5 It is a schematic diagram of a log storage module provided according to an embodiment of the present application;
[0048] Figure 6 It is a schematic diagram of a data pre-screening module provided according to an embodiment of the present application;
[0049] Figure 7 It is a schematic diagram of an alarm simulation module provided according to an embodiment of the present application;
[0050] Figure 8 It is a schematic diagram of a data comparison module provided according to an embodiment of the present application;
[0051] Figure 9 It is a schematic diagram of a method comparison provided according to an embodiment of the present application;
[0052] Figure 10 It is a block diagram of a policy adjustment device provided according to an embodiment of the present application;
[0053] Figure 11 It is a block diagram of another policy adjustment device provided according to an embodiment of the present application;
[0054] Figure 12 It is a schematic structural diagram of a terminal provided according to an embodiment of the present application;
[0055] Figure 13 It is a schematic structural diagram of a server provided according to an embodiment of the present application. Detailed implementation manners
[0056] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0057] In the present application, terms such as "first" and "second" are used to distinguish identical items or similar items with basically the same functions and effects. It should be understood that there is no logical or chronological dependency between "first", "second", and "nth", nor are the quantity and execution order limited.
[0058] In this application, the term "at least one" means one or more, and "a plurality" means two or more.
[0059] In this application, the term "cloud technology" refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing. Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. Cloud computing technology will become an important support. The back-end services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the back-end system for logical processing. Data at different levels will be processed separately, and various types of industry data require a powerful system back-end support, which can only be achieved through cloud computing.
[0060] The term "cloud storage" in this application is a new concept extended and developed from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as the storage system) refers to a storage system that combines a large number of different types of storage devices (storage devices are also called storage nodes) in the network through cluster applications, grid technology, and distributed storage file systems, etc., and works together through application software or application interfaces to jointly provide data storage and business access functions to the outside world. Currently, the storage method of the storage system is as follows: Create a logical volume. When creating a logical volume, physical storage space is allocated for each logical volume. This physical storage space may be composed of a certain storage device or the disks of several storage devices. The client stores data on a certain logical volume, that is, stores the data on the file system. The file system divides the data into many parts, and each part is an object. The object not only contains data but also additional information such as data identification (ID, IDentity). The file system writes each object into the physical storage space of the logical volume respectively, and the file system will record the storage location information of each object. Thus, when the client requests to access the data, the file system can enable the client to access the data according to the storage location information of each object. The process of the storage system allocating physical storage space for the logical volume is specifically as follows: According to the capacity estimation of the objects stored in the logical volume (this estimation often has a large margin relative to the actual capacity of the objects to be stored) and the group of redundant arrays of independent disks (RAID, Redundant Array of Independent Disk), the physical storage space is pre-divided into stripes, and a logical volume can be understood as a stripe, thereby allocating physical storage space for the logical volume.
[0061] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, multiple server logs involved in this application are obtained under full authorization.
[0062] Figure 1 It is a schematic diagram of the implementation environment of a strategy adjustment method provided by an embodiment of the present application. Refer to Figure 1 , this implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this.
[0063] In some embodiments, the terminal 101 is various types of terminals such as a mobile phone, a desktop computer, a laptop computer, a tablet computer, a smart watch, etc. An application program can be installed and run on the terminal 101. This application program can perform simulation monitoring on a target server based on a simulation strategy and multiple server logs, and adjust this simulation strategy. This application program is associated with the server 102, and the server 102 provides background services to the terminal 101.
[0064] In some embodiments, the server 102 is an independent physical server, and can also be a server cluster or a distributed system composed of multiple physical servers, and can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0065] In some embodiments, the server 102 undertakes the main computing work, and the terminal 101 undertakes the secondary computing work; or, the server 102 undertakes the secondary computing work, and the terminal 101 undertakes the main computing work; or, the server 102 and the terminal 101 adopt a distributed computing architecture for collaborative computing.
[0066] Figure 2 It is a flowchart of a strategy adjustment method provided by an embodiment of the present application. This method is executed by a computer device. Refer to Figure 2 , this method includes the following steps:
[0067] 201. Obtain multiple server logs and actual fault information of the target server during the target historical period. The actual fault information is used to indicate the server faults that actually occurred on the target server during the target historical period.
[0068] In the embodiments of the present application, the terminal obtains multiple server logs and actual fault information of the target server during the target historical period. Among them, the target server can be a single server or multiple servers. Among them, server faults can be divided into multiple types such as hardware faults, software faults, network faults, and power supply faults. For different types of server faults, the adopted monitoring strategies may be different. Among them, the server logs record information such as the request address, request time, requested web page, user agent, and referral address. Among them, the actual fault information refers to the server faults actually obtained by monitoring the target server during the target historical period.
[0069] 202. For any fault type, based on the simulation strategy of this fault type and multiple server logs, perform simulation monitoring on the target server to obtain simulation fault information. The simulation strategy is used to determine whether a server fault of this fault type occurs.
[0070] In the embodiments of the present application, for any fault type, the terminal obtains the simulation strategy of this fault type and performs simulation monitoring on the target server based on this simulation strategy and multiple server logs, so as to determine the simulation fault information in the target historical period, which is convenient for subsequent determination of whether to adjust this simulation strategy. Among them, the simulation fault information may include information such as the fault address, fault type, and fault occurrence time. The embodiments of the present application do not limit this.
[0071] 203. Adjust the simulation strategy of this fault type based on the difference between the simulation fault information and the actual fault information.
[0072] In the embodiments of the present application, based on the difference between the simulation fault information and the actual fault information, the terminal adjusts the simulation strategy of the fault type. Among them, the difference between the simulation fault information and the actual fault information can not only indicate the simulation monitoring effect of this simulation strategy, but also indicate the change in the monitoring effect of this simulation strategy compared with the monitoring strategy in actual application.
[0073] The embodiments of the present application provide a strategy adjustment method. By performing simulation monitoring on the server during the target historical period according to the simulation strategy and multiple server logs, and adjusting the simulation strategy according to the difference between the obtained simulation fault information and the actual fault information. When adjusting the monitoring strategy multiple times, compared with the method of directly putting the monitoring strategy into trial operation in the live network and adjusting it multiple times in the traditional method, this method can call the server logs in the target historical period, without waiting for multiple cycles, and is more efficient and time-saving.
[0074] It should be noted that, for the convenience of describing the strategy adjustment method in this solution, a strategy simulation system is proposed. Refer to Figure 3 as shown Figure 3 is a schematic diagram of a strategy simulation system provided according to an embodiment of the present application. Among them, the strategy simulation system includes a log storage module, a data pre-screening module, an alarm simulation module, a simulated alarm data storage module, a data comparison module, and a monitoring strategy library.
[0075] Among them, the server collects server logs through two methods: in-band and out-of-band, and transmits a large amount of server logs across the network to the log storage module through a network channel. Among them, the log storage module is used to obtain multiple server logs within a target historical time period. That is to say, the log storage module obtains a large amount of server logs through a message queue, stores the data obtained by parsing the server logs into elasticsearch (ES), and establishes a search index. Among them, the data pre-screening module is used to pre-screen and store valid logs. That is to say, the data pre-screening module retrieves valid logs from a large amount of servers, reducing the number of server logs that need to be simulated and monitored to 1 / 100 of the original number, thereby reducing the pressure on the downstream system and improving the overall simulation efficiency. Among them, the alarm simulation module is used to determine whether to trigger a fault alarm. That is to say, the alarm simulation module loads the simulation strategy configured by the strategy configurator, receives the data sent by the upstream data pre-screening module, and performs strategy matching according to the simulation strategy, and outputs a fault alarm to the simulated alarm data storage module when the simulation strategy is hit. Among them, the simulated alarm data storage module is used to store the relevant data of the fault alarm. Among them, the data comparison module is used to compare the simulated fault information and the actual fault information. That is to say, after a round of simulation monitoring of the server is completed, the data comparison module reads the simulated alarm data storage module and the alarm fault list in the fault repair process respectively for data comparison, and displays the data comparison result. Among them, in the fault repair process, relevant personnel need to diagnose the actual fault information to determine whether the fault alarm needs to be repaired. If the fault alarm does not need to be repaired, a false alarm label is added. Among them, the monitoring strategy library is used to store monitoring strategies and provide monitoring strategies for real-time fault monitoring. After the strategy configurator confirms that the simulation strategy is effective, the simulation strategy is stored in the monitoring strategy library and real-time fault monitoring is performed; if the simulation strategy needs to be optimized, the simulation strategy is adjusted and then a new round of simulation monitoring is started, and the above steps are repeated until the strategy configurator confirms that the simulation strategy is effective.
[0076] Figure 4 is a flowchart of another strategy adjustment method provided according to an embodiment of the present application. This method is executed by a computer device. Refer to Figure 4, the method includes the following steps:
[0077] 401. Obtain multiple server logs and actual fault information of the target server during the target historical period. The actual fault information is used to indicate the server faults that actually occurred on the target server during the target historical period.
[0078] In the embodiments of the present application, the terminal obtains multiple server logs and actual fault information of the target server during the target historical period, which is convenient for subsequently determining whether to adjust the policy based on the above data. Among them, the target server can be a single server or multiple servers, and the embodiments of the present application do not limit this. Among them, the target historical period can be one day or one week, and the embodiments of the present application do not limit this. Among them, the server log refers to one or more log files automatically created and maintained by the server. The server log contains a list of activities performed by the server. The information recorded in the server log includes the client address, request time, requested web page, user agent, and referral address, etc. Among them, server faults can be classified into various types of faults such as hardware faults, software faults, network faults, and power supply faults. In the present application, server hardware faults are taken as an example.
[0079] In some embodiments, the server logs are obtained from the data of the server cluster. Correspondingly, obtain the parsing configuration file and the data of the server cluster. The server cluster includes the target server; based on the parsing configuration file, parse the data of the server cluster to obtain the server logs of the target server. By first obtaining the data of the server cluster and then parsing to obtain the server logs of the target server during the target historical period, it avoids the problem that each time it is necessary to obtain the server logs in different historical periods or the server logs of different servers, it is necessary to re-obtain the data of the server cluster, making the data acquisition process more convenient and fast, and improving the efficiency of obtaining server logs.
[0080] In the policy simulation system, the log storage module obtains and stores the server logs of the target server. Refer to Figure 5 as shown, Figure 5 is a schematic diagram of a log storage module provided according to an embodiment of the present application. Among them, the log storage module loads the parsing configuration file from the database and receives the data of the server cluster; then extracts the relevant data of the server logs such as the server address, log time, and log text from the data of the server cluster according to the parsing configuration file; then stores the above relevant data of the server logs into elasticsearch and establishes a full-text index for subsequent log query and log return.
[0081] Among them, the terminal can obtain the data of the server cluster through two methods: in-band and out-of-band. In-band means that data transmission is carried out on the same channel or transmission medium. For example, when using a network protocol, data packets are transmitted on the same channel as the control information; out-of-band means using a channel or transmission medium different from the main data channel. For example, a dedicated management network is used for management or monitoring. Among them, elasticsearch is a search server that provides a distributed multi-user full-text search engine, which can easily enable a large amount of data to have the capabilities of search, analysis, and exploration. It can also add a search box to an application or website and store and analyze log, metric, and security event data. It should be noted that elasticsearch can be replaced by other storage components that can store relevant data of server logs, and the embodiments of the present application do not limit this.
[0082] Among them, the log storage module can provide a data interface externally. When a query request sent by the downstream data pre-screening module arrives, the log storage module converts the query request into a query command of elasticsearch and returns the queried data to the data pre-screening module. Among them, when the amount of data to be returned is large, paging is required for return.
[0083] Among them, the log storage module creates a full-text index for the relevant data of the stored server logs, which is convenient for quickly searching and returning the queried data when a query request arrives. A full-text index is a special type of token-based functional index that stores the key words and their location information in one or more columns of a database table.
[0084] It should be noted that the log storage module can be a module configured in the terminal or an external module of the terminal, and the embodiments of the present application do not limit this.
[0085] 402. For any failure type, based on multiple data pre-screening keywords in the simulation strategy of this failure type, at least one valid log is screened out from multiple server logs. The simulation strategy is used to determine whether a server failure of this failure type occurs, and a valid log is a server log containing any data pre-screening keyword.
[0086] In the embodiments of the present application, for any failure type, the terminal obtains the simulation strategy of this failure type and screens out at least one valid log from multiple server logs based on multiple data pre-screening keywords in the simulation strategy. Among them, the data pre-screening keyword can be a single keyword or multiple keywords.
[0087] In the policy simulation system, the data pre-screening module screens out valid logs from multiple server logs, as shown in Figure 6 shown.Figure 6 It is a schematic diagram of a data pre-screening module provided according to an embodiment of the present application. Among them, the data pre-screening module can be a module configured in the terminal or an external module of the terminal, and the embodiments of the present application do not limit this. Among them, the data pre-screening keywords in the simulation policy are loaded from mysql; then the query request including the data pre-screening keywords and the target historical time period is sent to the log storage module; then the filtered valid logs are stored in kafka for use by the downstream alarm simulation module. For example, there are about 370 million server logs in elasticsearch every day. After filtering the server logs according to the data pre-screening keywords in the simulation policy, about 4 million valid logs are obtained. That is to say, through the data pre-screening step, the data volume level is reduced to about 1 / 100 of the original data on average, thus avoiding pulling and replaying all server logs and improving the data monitoring efficiency.
[0088] Among them, Kafka is a high-throughput distributed publish-subscribe messaging system that provides distributed, partitioned, and reliable distributed log storage services. It should be noted that mysql and kafka can be replaced by other storage components, and the embodiments of the present application do not limit this.
[0089] 403. Based on the exception keyword extraction rule in the simulation policy for this fault type, process the valid logs to obtain multiple exception keywords, the occurrence times of each exception keyword, and the occurrence moments of each exception keyword. The exception keyword extraction rule is used to extract exception keywords based on regular expressions.
[0090] In the embodiments of the present application, the terminal extracts multiple exception keywords from the valid logs based on the exception keyword extraction rule in the simulation policy for this fault type, and at the same time obtains the occurrence times of each of the above exception keywords and the occurrence moments of each exception keyword. Among them, the exception keyword extraction rule includes data pre-screening keywords and a regular expression for extracting exception keywords. Usually, each exception keyword extraction rule matches a single keyword, and there can be multiple exception keyword extraction rules in the simulation policy. Among them, the regular expression is a language for matching keywords, which is composed of a series of characters and special characters and is used to describe the keywords to be matched. Using regular expressions can find, replace, extract, and verify specific keywords in the text.
[0091] 404. Based on the alarm determination rule of the simulation policy for this fault type, determine the simulation fault information. The alarm determination rule is used to determine whether to trigger a fault alarm based on the occurrence frequency of the exception keyword.
[0092] In an embodiment of the present application, the terminal determines simulation fault information based on the alarm determination rule of the simulation strategy for this fault type. Among them, the alarm determination rule includes the trigger threshold for fault alarms.
[0093] In some embodiments, whether to trigger an alarm is related to the occurrence frequency of abnormal keywords. Correspondingly, for any moment during the simulation monitoring process, determine the number of occurrences of any abnormal keyword within the target time range corresponding to the moment, and the end moment of the target time range is the moment; if the number of occurrences is not less than the number threshold, determine that a fault alarm is triggered; add the alarm fault list of the fault alarm to the simulation fault information. By determining whether to trigger a fault alarm according to the relationship between the number of occurrences of the abnormal keyword within the target time range corresponding to each moment when the abnormal keyword appears and the number threshold, the simulation fault information within the target historical time can be determined, thus facilitating the determination of the effectiveness of this simulation strategy.
[0094] It should be noted that when multiple consecutive moments trigger fault alarms, multiple alarm fault lists of the fault alarms can be generated, or only a single alarm fault list can be generated. The present application does not limit this.
[0095] In some embodiments, it is not necessary to determine whether to trigger a fault alarm for each moment, but only to determine whether to trigger a fault alarm for the moment when an abnormal keyword appears. Correspondingly, for any moment when an abnormal keyword appears, determine the number of occurrences of the abnormal keyword within the target time range corresponding to the moment, and the end moment of the target time range is the moment; if the number of occurrences is not less than the number threshold, determine that a fault alarm is triggered; add the alarm fault list of the fault alarm to the simulation fault information. By determining whether to trigger a fault alarm only when the abnormal keyword appears, the simulation fault information can be determined more quickly. Compared with determining whether to trigger an alarm for each moment, this method is more efficient and faster.
[0096] In the policy simulation system, it is determined whether to trigger an alarm by the alarm simulation module configured in the terminal. See Figure 7 as shown Figure 7It is a schematic diagram of an alarm simulation module provided according to an embodiment of the present application. Among them, first, the simulation policy configured by the policy configurator is loaded, including the abnormal keyword extraction rule and the alarm determination rule. Then, the valid logs sent by the upstream data pre-screening module are received and processed successively by the abnormal keyword extraction component and the alarm judgment component. Among them, the abnormal keyword extraction component matches multiple abnormal keywords in the valid logs by means of regular expression matching, and counts the number and time of occurrence of the above-mentioned multiple abnormal keywords, and then sends them to the alarm judgment component. The alarm judgment component judges whether the number of occurrences of the abnormal keywords within the configured time range reaches the number threshold according to the simulation policy configured by the policy configurator. When the number threshold is reached, a fault alarm is triggered, and the relevant data of the fault alarm is sent to the downstream simulation alarm data storage module. It should be noted that when sending the relevant data of the fault alarm to the downstream simulation alarm data storage module, it can be to send the relevant data of each fault alarm separately, or to package and send the relevant data of all fault alarms. The embodiments of the present application do not limit this.
[0097] Among them, the simulation alarm data storage module stores the relevant data of the fault alarm in the form of a trouble ticket. The relevant data of the stored fault alarm mainly includes the address of the faulty service, the fault occurrence time, and the fault type, etc. Among them, the simulation alarm storage module can be a module configured in the terminal or an external module of the terminal. The embodiments of the present application do not limit this.
[0098] It should be noted that in the policy simulation system, the simulation policy is mainly used in the data pre-screening module and the alarm simulation module. Among them, in order to avoid data transformation and more easily convert the adjusted simulation policy into the monitoring policy library, generally, the data structure of the simulation policy needs to be consistent with the data structure of the monitoring policy in the monitoring policy library. Among them, the abnormal keyword extraction rule of the simulation policy includes the data pre-screening keyword and the regular expression, and the alarm determination rule of the simulation policy includes the trigger threshold.
[0099] For example, when it is necessary to monitor the error alarm of memory CE in the dmesg log, the data pre-screening keyword in the simulation policy is configured as "CE", and the regular expression is set as "EDAC MC\\d{1,2}:(? <eigentype>.*?) CEmemory.*error on(?) <collectslot>.*?)\\("). Among them, the dmesg log mainly records kernel information. Among them, the above regular expression means that it matches the string "EDAC MC" followed by one or two digits, then a space and the string "CE memory error on", followed by any non-greedy characters (i.e., matching as few characters as possible), and finally a space, the string "(" and any non-whitespace characters (i.e., matching as many non-whitespace characters as possible), and then the ")" symbol. At the same time, the regular expression uses two named capture groups "eigenType" and "collectSlot" to match any non-greedy characters after "CE memory erroron" and any non-whitespace characters within the parentheses respectively.
[0100] At the same time, when it is necessary to monitor the error alarm of memory CE in the dmesg log, set the trigger threshold in the simulation strategy to "1000 times in 1 hour". That is to say, when the keyword accumulates 1000 times within 1 hour, a fault alarm is triggered.
[0101] It should be noted that the simulation strategy is not limited to the above configuration methods and configuration contents, as long as it can monitor the server, and the embodiments of the present application do not limit this.
[0102] 405. Based on the simulation fault information and the actual fault information, determine the misjudgment rate of the simulation strategy. The misjudgment rate of the simulation strategy is used to represent the accuracy of the simulation strategy for monitoring the target server for this fault type.
[0103] In the embodiments of the present application, the terminal determines the misjudgment rate of the simulation strategy based on the simulation fault information and the actual fault information, so as to facilitate judging whether the simulation strategy is effective.
[0104] In some embodiments, the accuracy of the simulation strategy is reflected by the misjudgment rate of the simulation strategy. Correspondingly, based on the simulation fault information and the actual fault information, determine the first quantity and the second quantity. The first quantity is the quantity of the first fault tickets, and the first fault tickets are the alarm fault tickets that exist simultaneously in the simulation fault information and the actual fault information. The second quantity is the quantity of the second fault tickets, and the second fault tickets are the first fault tickets marked as false alarms; determine the ratio of the second quantity to the first quantity as the misjudgment rate of the simulation strategy. By determining the simulation misjudgment rate according to the quantities of different types of alarm fault tickets, the accuracy of the simulation strategy can be reflected, so as to facilitate judging whether the simulation strategy is effective.
[0105] 406. Based on the actual fault information and the similar strategy of the simulation strategy, determine the misjudgment rate of the similar strategy. The misjudgment rate of the similar strategy is used to represent the accuracy of the similar strategy for monitoring the target server.
[0106] In an embodiment of the present application, the terminal determines the misjudgment rate of the similarity strategy based on the actual fault information and the similarity strategy of the simulation strategy, so as to be used as a reference for judging whether the simulation strategy is effective.
[0107] In some embodiments, the effect of the simulation strategy is further reflected by determining the misjudgment rate of the similarity strategy. Correspondingly, based on the actual fault information and the similarity strategy of the simulation strategy, a third quantity and a fourth quantity are determined. The third quantity is the quantity of the third fault tickets, and the third fault tickets are the alarm fault tickets determined by the similarity strategy in the actual fault information. The fourth quantity is the quantity of the fourth fault tickets, and the fourth fault tickets are the third fault tickets marked as false alarms; the ratio of the fourth quantity to the third quantity is determined as the misjudgment rate of the similarity strategy. By determining the similarity misjudgment rate according to the quantity of different types of alarm fault tickets, the accuracy of the similarity strategy can be reflected, so as to be used as a reference for judging whether the simulation strategy is effective.
[0108] In the policy simulation system, the data comparison module configured in the terminal determines the effectiveness of the simulation strategy according to the simulation fault information and the actual fault information. Refer to Figure 8 as shown Figure 8 is a schematic diagram of a data comparison module provided according to an embodiment of the present application. Among them, after a round of simulation monitoring of the server is completed, the data comparison service is started. First, the fault address, fault time, and fault type are pulled from the simulation alarm data storage module, and the fault address, fault time, fault type, and whether there is a false alarm annotation are pulled from the fault repair process, and then the data is compared and the comparison result of the data comparison is displayed on the display page of the terminal. Among them, the data in the fault repair process is the alarm fault tickets obtained by real-time fault monitoring of the live network according to the monitoring policy library. Among them, in the fault repair process, relevant personnel need to diagnose the actual fault information to determine whether the fault alarm needs to be repaired. If the fault alarm does not need to be repaired, a false alarm annotation is added. Among them, the live network is the formal production environment.
[0109] It should be noted that according to the different size situations of the misjudgment rate of the simulation strategy and the misjudgment rate of the similarity strategy, the operations performed by the terminal are also different, as shown in 407-408 below.
[0110] 407. In the case where the misjudgment rate of the simulation strategy is not less than the misjudgment rate of the similarity strategy, adjust the simulation strategy of this fault type.
[0111] In the embodiments of the present application, when the misjudgment rate of the simulation policy is not less than the misjudgment rate of the similarity policy, that is, when the accuracy of the simulation policy is lower than that of the similarity policy, the policy configurator adjusts the simulation policy of this fault type through the terminal. Among them, the data pre-screening keywords, regular expressions, and trigger thresholds in the simulation policy can all be adjusted. Among them, the policy configurator can add the missing keywords to the simulation policy through the terminal. The missing keywords refer to the keywords that appear in the actual fault information but do not appear in the simulation fault information. The policy configurator can also modify the regular expression or reconfigure the trigger threshold through the terminal. The embodiments of the present application do not limit this.
[0112] 408. When the misjudgment rate of the simulation policy is less than the misjudgment rate of the similarity policy, add the simulation policy of this fault type to the monitoring policy library, and the monitoring policy library is used to provide the monitoring policies required for real-time fault monitoring.
[0113] In the embodiments of the present application, when the misjudgment rate of the simulation policy is less than the misjudgment rate of the similarity policy, that is, when the accuracy of the simulation policy is higher than that of the similarity policy, the terminal adds the simulation policy of this fault type to the monitoring policy library. It should be noted that the simulation policy can be directly added to the monitoring policy library, or the similarity policy in the monitoring policy library can be replaced with the simulation policy. The embodiments of the present application do not limit this.
[0114] In some embodiments, the results of the simulation monitoring are prompted through prompt messages. Correspondingly, the target server includes multiple servers. One server can correspond to multiple alarm fault tickets, and one server corresponds to one fault address. Obtain the spare part threshold corresponding to the target server. The spare part threshold is the number of standby servers; based on the simulation fault information, determine the absolute number of fault tickets. The absolute number of fault tickets is the number of fault addresses that appear in the simulation fault information; when the absolute number of fault tickets is greater than the spare part threshold, display a prompt message, and the prompt message is used to prompt the absolute number of fault tickets, the spare part threshold, and at least one fault address. By displaying relevant prompt messages when the absolute number of fault tickets is greater than the spare part threshold, it is possible to facilitate relevant personnel to quickly locate and handle server faults, and timely replenish standby servers, thereby avoiding the problem that when there are too many server hardware faults, the server hardware faults cannot be resolved in time due to the lack of standby servers.
[0115] In some embodiments, the effect of the simulation policy is determined by the change ratio of trouble tickets. Accordingly, based on the simulation fault information and the actual fault information, a third quantity and a fifth quantity are determined. The third quantity is the number of third trouble tickets, and the third trouble tickets are the alarm trouble tickets determined by the similar policy of the simulation policy in the actual fault information. The fifth quantity is the number of fifth trouble tickets, and the fifth trouble tickets are the alarm trouble tickets determined by the simulation policy of this fault type in the simulation fault information. The ratio of the difference between the fifth quantity and the third quantity to the third quantity is determined as the change ratio of trouble tickets, and the change ratio of trouble tickets is used to represent the change degree of the monitoring effect of the simulation policy of this fault type compared with the similar policy. The change ratio of trouble tickets is displayed on the display page, and the display page is used to display the comparison information between the simulation fault information and the actual fault information. By determining the change ratio of trouble tickets, the change degree of the monitoring effect of the simulation policy compared with the similar policy is reflected and displayed on the display page of the terminal, so that relevant personnel can easily judge whether it is necessary to continue to adjust the simulation policy and whether it is necessary to replace the similar policy in the monitoring policy library.
[0116] It should be noted that data such as the misjudgment rate of the simulation policy, the misjudgment rate of the similar policy, the absolute number of trouble tickets, and the change ratio of trouble tickets can be considered simultaneously to determine whether the simulation policy needs to be adjusted. The embodiments of the present application do not limit this.
[0117] See Figure 9 as shown Figure 9 is a schematic diagram of a method comparison provided according to an embodiment of the present application. Among them, Figure 9 (a) refers to the steps of policy adjustment in the traditional solution. The policy configurator directly formulates a monitoring policy based on experience, and then monitors the live network according to the monitoring policy, but only collects and counts the number of trouble tickets that occur in the live network every day or week. Then analyze the above data and evaluate the effectiveness of the monitoring policy. If the monitoring policy is reasonable, it is directly put into production. If it is unreasonable, a monitoring policy needs to be formulated again. Figure 9 (b) refers to the steps of policy adjustment in this solution. First, the policy configurator configures an initial simulation policy. The policy simulation system automatically pulls the server logs for a period of time and processes them. Then, the simulation fault information and the actual fault information are compared data, and the comparison results of the data comparison are displayed. The policy configurator analyzes the above comparison results. If the simulation policy is unreasonable, it is adjusted online and the above steps are repeated until the simulation policy is reasonable. Then, the simulation policy is added to the monitoring policy library for use by the real-time fault monitoring system.
[0118] It should be noted that the idea and strategy simulation system of this solution can be applied not only to the simulation monitoring and strategy adjustment of server hardware failures, but also to other similar simulation monitoring and strategy adjustment tasks. For example, server software monitoring simulation, Internet of Things hardware monitoring simulation, and automotive hardware failure monitoring simulation, etc. The embodiments of this application do not limit this. It should be noted that Elasticsearch, mysql, and kafka used in the present invention can all be replaced by other storage components, and the embodiments of this application do not limit this.
[0119] The embodiment of this application provides a strategy adjustment method. By simulating and monitoring the server within a target historical time period according to the simulation strategy and multiple server logs, and adjusting the simulation strategy according to the difference between the obtained simulation failure information and the actual failure information. When adjusting the monitoring strategy multiple times, compared with the method of directly putting the monitoring strategy into trial operation in the live network and adjusting it multiple times in the traditional method, this method can call the server logs within the target historical time period, without waiting for multiple cycles, and is more efficient and time-saving.
[0120] Figure 10 It is a block diagram of a strategy adjustment device provided according to an embodiment of this application. This device is used to execute the steps when the above strategy adjustment method is executed. Refer to Figure 10 As shown, this strategy adjustment device includes: a first acquisition module 1001, a simulation module 1002, and an adjustment module 1003.
[0121] The first acquisition module 1001 is used to acquire multiple server logs and actual failure information of the target server within the target historical time period. The actual failure information is used to indicate the server failures that actually occurred to the target server within the target historical time period.
[0122] The simulation module 1002 is used to, for any failure type, based on the simulation strategy of the failure type and multiple server logs, simulate and monitor the target server to obtain simulation failure information. The simulation strategy is used to determine whether a server failure of the failure type occurs.
[0123] The adjustment module 1003 is used to adjust the simulation strategy of the failure type based on the difference between the simulation failure information and the actual failure information.
[0124] In some embodiments, Figure 11 It is a block diagram of another strategy adjustment device provided according to an embodiment of this application. Refer to Figure 11 As shown, the simulation module 1002 includes:
[0125] The screening unit 10021 is configured to, for any fault type, based on multiple data pre-screening keywords in the simulation strategy of the fault type, screen out at least one valid log from multiple server logs, where the valid log is a server log containing any data pre-screening keyword;
[0126] The monitoring unit 10022 is configured to perform simulation monitoring on the target server based on the simulation strategy of the fault type and the valid log to obtain simulation fault information.
[0127] In some embodiments, the monitoring unit 10022 is configured to process the valid log based on the exception keyword extraction rule in the simulation strategy of the fault type to obtain multiple exception keywords, the occurrence times of each exception keyword, and the occurrence moments of each exception keyword. The exception keyword extraction rule is used to extract exception keywords based on regular expressions; determine the simulation fault information based on the alarm determination rule of the simulation strategy of the fault type, and the alarm determination rule is used to determine whether to trigger a fault alarm based on the occurrence frequency of the exception keyword.
[0128] In some embodiments, the monitoring unit 10022 is further configured to, for any moment when an exception keyword appears, determine the occurrence times of the exception keyword within the target time range corresponding to the moment, where the end moment of the target time range is the moment; in the case that the occurrence times are not less than the times threshold, determine to trigger a fault alarm; add the alarm fault list of the fault alarm to the simulation fault information.
[0129] In some embodiments, the monitoring unit 10022 is further configured to, for any moment during the simulation monitoring process, determine the occurrence times of any exception keyword within the target time range corresponding to the moment, where the end moment of the target time range is the moment; in the case that the occurrence times are not less than the times threshold, determine to trigger a fault alarm; add the alarm fault list of the fault alarm to the simulation fault information.
[0130] In some embodiments, the adjustment module 1003 includes:
[0131] The first determination unit 10031 is configured to determine the simulation strategy misjudgment rate based on the simulation fault information and the actual fault information, and the simulation strategy misjudgment rate is used to represent the accuracy of the simulation strategy of the fault type in monitoring the target server;
[0132] The second determination unit 10032 is configured to determine the similar strategy misjudgment rate based on the actual fault information and the similar strategy of the simulation strategy, and the similar strategy misjudgment rate is used to represent the accuracy of the similar strategy in monitoring the target server;
[0133] The adjustment unit 10033 is configured to adjust the simulation strategy of the fault type in the case that the simulation strategy misjudgment rate is not less than the similar strategy misjudgment rate.
[0134] In some embodiments, the first determination unit 10031 is configured to determine a first quantity and a second quantity based on simulation fault information and actual fault information. The first quantity is the quantity of first fault tickets, and the first fault tickets are alarm fault tickets that exist simultaneously in the simulation fault information and the actual fault information. The second quantity is the quantity of second fault tickets, and the second fault tickets are the first fault tickets marked as false alarms. The ratio of the second quantity to the first quantity is determined as the misjudgment rate of the simulation strategy.
[0135] In some embodiments, the second determination unit 10032 is configured to determine a third quantity and a fourth quantity based on the actual fault information and a similar strategy of the simulation strategy. The third quantity is the quantity of third fault tickets, and the third fault tickets are alarm fault tickets determined by the similar strategy in the actual fault information. The fourth quantity is the quantity of fourth fault tickets, and the fourth fault tickets are the third fault tickets marked as false alarms. The ratio of the fourth quantity to the third quantity is determined as the misjudgment rate of the similar strategy.
[0136] In some embodiments, the adjustment module 1003 is further configured to add the simulation strategy of the fault type to the monitoring strategy library when the misjudgment rate of the simulation strategy is less than the misjudgment rate of the similar strategy. The monitoring strategy library is used to provide the monitoring strategies required for real-time fault monitoring.
[0137] In some embodiments, the target server includes multiple servers. One server can correspond to multiple alarm fault tickets, and one server corresponds to one fault address. The apparatus further includes:
[0138] The second acquisition module 1101 is configured to acquire the spare part threshold corresponding to the target server, and the spare part threshold is the quantity of standby servers.
[0139] The first determination module 1102 is configured to determine the absolute quantity of fault tickets based on the simulation fault information, and the absolute quantity of fault tickets is the quantity of fault addresses that appear in the simulation fault information.
[0140] The first display module 1103 is configured to display a prompt message when the absolute quantity of fault tickets is greater than the spare part threshold. The prompt message is used to prompt the absolute quantity of fault tickets, the spare part threshold, and at least one fault address.
[0141] In some embodiments, the apparatus further includes:
[0142] The second determination module 1104 is configured to determine a third quantity and a fifth quantity based on the simulation fault information and the actual fault information. The third quantity is the quantity of third fault tickets, and the third fault tickets are alarm fault tickets determined by the similar strategy of the simulation strategy in the actual fault information. The fifth quantity is the quantity of fifth fault tickets, and the fifth fault tickets are alarm fault tickets determined by the simulation strategy of the fault type in the simulation fault information.
[0143] A third determination module 1105 is configured to determine the ratio of the difference between the fifth quantity and the third quantity to the third quantity as the fault ticket change ratio, and the fault ticket change ratio is used to represent the change degree of the monitoring effect of the simulation strategy of the fault type compared with the similar strategy.
[0144] A second display module 1106 is configured to display the fault ticket change ratio on a display page, and the display page is used to display the comparison information between the simulated fault information and the actual fault information.
[0145] In some embodiments, the apparatus further includes:
[0146] A third acquisition module 1107 is configured to acquire the data of the parsing configuration file and the server cluster, and the server cluster includes a target server.
[0147] A parsing module 1108 is configured to parse the data of the server cluster based on the parsing configuration file to obtain the server log of the target server.
[0148] A strategy adjustment apparatus provided by the present application performs simulation monitoring on a server within a target historical time period according to a simulation strategy and multiple server logs, and adjusts the simulation strategy according to the difference between the obtained simulated fault information and the actual fault information. When adjusting the monitoring strategy multiple times, compared with the method of directly putting the monitoring strategy into trial operation in the live network and adjusting it multiple times in the traditional method, this method can call the server logs within the target historical time period, without waiting for multiple cycles, and is more efficient and time-saving.
[0149] It should be noted that when the strategy adjustment apparatus provided in the above embodiment runs an application program, only the above-mentioned division of each functional module is used for illustration. In actual application, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the terminal is divided into different functional modules to complete all or part of the functions described above. In addition, the strategy adjustment apparatus provided in the above embodiment and the strategy adjustment method embodiment belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be elaborated here.
[0150] Figure 12 It is a schematic structural diagram of a terminal provided according to an embodiment of the present application. The terminal 1200 may be a portable mobile terminal, such as: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a notebook computer or a desktop computer. The terminal 1200 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.
[0151] Generally, the terminal 1200 includes: a processor 1201 and a memory 1202.
[0152] The processor 1201 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1201 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), PLA (Programmable Logic Array). The processor 1201 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1201 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1201 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0153] The memory 1202 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1202 may also include high-speed random access memory, as well as non-volatile memory, such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1202 is used to store at least one computer program, and the at least one computer program is used to be executed by the processor 1201 to implement the policy adjustment method provided in the method embodiments of the present application.
[0154] In some embodiments, the terminal 1200 may further optionally include: a peripheral device interface 1203 and at least one peripheral device. The processor 1201, the memory 1202, and the peripheral device interface 1203 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1203 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 1204, a display screen 1205, a camera assembly 1206, an audio circuit 1207, and a power supply 1208.
[0155] The peripheral device interface 1203 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 1201 and the memory 1202. In some embodiments, the processor 1201, the memory 1202, and the peripheral device interface 1203 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1201, the memory 1202, and the peripheral device interface 1203 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0156] The radio frequency circuit 1204 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1204 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1204 converts an electrical signal into an electromagnetic signal for transmission, or converts a received electromagnetic signal into an electrical signal. In some embodiments, the radio frequency circuit 1204 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 1204 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1204 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0157] The display screen 1205 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1205 is a touch display screen, the display screen 1205 also has the ability to collect touch signals on or above the surface of the display screen 1205. The touch signals can be input to the processor 1201 as control signals for processing. At this time, the display screen 1205 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there may be one display screen 1205, which is provided on the front panel of the terminal 1200; in other embodiments, there may be at least two display screens 1205, which are respectively provided on different surfaces of the terminal 1200 or are in a foldable design; in other embodiments, the display screen 1205 may be a flexible display screen, which is provided on the curved surface or the folding surface of the terminal 1200. Even further, the display screen 1205 can also be set to an irregular non-rectangular shape, that is, an irregular-shaped screen. The display screen 1205 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0158] The camera module 1206 is used to capture images or videos. In some embodiments, the camera module 1206 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to realize the function of background blurring by fusing the main camera and the depth-of-field camera, the function of panoramic shooting by fusing the main camera and the wide-angle camera, and the VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera module 1206 may further include a flash. The flash can be a single-color-temperature flash or a dual-color-temperature flash. A dual-color-temperature flash refers to a combination of a warm-light flash and a cold-light flash, which can be used for light compensation under different color temperatures.
[0159] The audio circuit 1207 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1201 for processing, or input to the radio frequency circuit 1204 to enable voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 1200. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 1201 or the radio frequency circuit 1204 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1207 may further include a headphone jack.
[0160] The power supply 1208 is used to supply power to each component in the terminal 1200. The power supply 1208 may be alternating current, direct current, a primary battery or a rechargeable battery. When the power supply 1208 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.
[0161] In some embodiments, the terminal 1200 further includes one or more sensors 1209. The one or more sensors 1209 include but are not limited to: an acceleration sensor 1210, a gyroscope sensor 1211, a pressure sensor 1212, an optical sensor 1213, and a proximity sensor 1214.
[0162] The acceleration sensor 1210 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established with the terminal 1200. For example, the acceleration sensor 1210 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 1201 can control the display screen 1205 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 1210. The acceleration sensor 1210 can also be used for collecting game or user's motion data.
[0163] The gyroscope sensor 1211 can detect the body direction and rotation angle of the terminal 1200. The gyroscope sensor 1211 can cooperate with the acceleration sensor 1210 to collect the 3D actions of the user on the terminal 1200. According to the data collected by the gyroscope sensor 1211, the processor 1201 can implement the following functions: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.
[0164] The pressure sensor 1212 can be disposed on the side frame of the terminal 1200 and / or the lower layer of the display screen 1205. When the pressure sensor 1212 is disposed on the side frame of the terminal 1200, it can detect the holding signal of the user on the terminal 1200, and the processor 1201 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 1212. When the pressure sensor 1212 is disposed on the lower layer of the display screen 1205, the processor 1201 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 1205. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0165] The optical sensor 1213 is used to collect the ambient light intensity. In one embodiment, the processor 1201 can control the display brightness of the display screen 1205 according to the ambient light intensity collected by the optical sensor 1213. Optionally, when the ambient light intensity is high, the display brightness of the display screen 1205 is increased; when the ambient light intensity is low, the display brightness of the display screen 1205 is decreased. In another embodiment, the processor 1201 can also dynamically adjust the shooting parameters of the camera module 1209 according to the ambient light intensity collected by the optical sensor 1213.
[0166] The proximity sensor 1214, also known as the distance sensor, is disposed on the front panel of the terminal 1200. The proximity sensor 1214 is used to collect the distance between the user and the front of the terminal 1200. In one embodiment, when the proximity sensor 1214 detects that the distance between the user and the front of the terminal 1200 is gradually decreasing, the processor 1201 controls the display screen 1205 to switch from the lit state to the off state; when the proximity sensor 1214 detects that the distance between the user and the front of the terminal 1200 is gradually increasing, the processor 1201 controls the display screen 1205 to switch from the off state to the lit state.
[0167] Those skilled in the art can understand that Figure 12 the structure shown in
[0168] Figure 13 It is a schematic structural diagram of a server provided by an embodiment of the present application. The server 1300 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 1301 and one or more memories 1302. Among them, at least one computer program is stored in the memory 1302, and the at least one computer program is loaded and executed by the processor 1301 to implement the policy adjustment method provided by each of the above method embodiments. Of course, the server may also have components such as wired or wireless network interfaces, keyboards, and input / output interfaces for input / output. The server may also include other components for implementing the functions of the device, which will not be elaborated here.
[0169] An embodiment of the present application also provides a computer-readable storage medium in which at least one segment of computer program is stored, and the at least one segment of computer program is loaded and executed by a processor to implement the policy adjustment method in the above embodiment. For example, the computer-readable storage medium may be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0170] An embodiment of the present application also provides a computer program product, including a computer program, and the computer program is executed by a processor to implement the policy adjustment method in the embodiment of the present application.
[0171] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiment can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a disk, or an optical disc, etc.
[0172] The above are only optional embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.< / collectslot> < / eigentype>
Claims
1. A strategy adjustment method, characterized in that, the method includes: Obtain multiple server logs and actual fault information of the target server within a target historical time period, where the actual fault information is used to indicate server faults that actually occurred on the target server within the target historical time period; For any fault type, based on the simulation strategy of the fault type and the multiple server logs, perform simulation monitoring on the target server to obtain simulation fault information, where the simulation strategy is used to determine whether a server fault of the fault type occurs; Based on the difference between the simulation fault information and the actual fault information, adjust the simulation strategy of the fault type.
2. The method according to claim 1, characterized in that, the step of, for any fault type, based on the simulation strategy of the fault type and the multiple server logs, performing simulation monitoring on the target server to obtain simulation fault information includes: For any fault type, based on multiple data pre-screening keywords in the simulation strategy of the fault type, screen out at least one valid log from the multiple server logs, where the valid log is a server log containing any data pre-screening keyword; Based on the simulation strategy of the fault type and the valid log, perform simulation monitoring on the target server to obtain the simulation fault information.
3. The method according to claim 2, characterized in that, the step of, based on the simulation strategy of the fault type and the valid log, performing simulation monitoring on the target server to obtain the simulation fault information includes: Based on the abnormal keyword extraction rule in the simulation strategy of the fault type, process the valid log to obtain multiple abnormal keywords, the occurrence times of each abnormal keyword, and the occurrence moments of each abnormal keyword, where the abnormal keyword extraction rule is used to extract abnormal keywords based on regular expressions; Based on the alarm determination rule of the simulation strategy of the fault type, determine the simulation fault information, where the alarm determination rule is used to determine whether to trigger a fault alarm based on the occurrence frequency of abnormal keywords.
4. The method according to claim 3, characterized in that, the step of, based on the alarm determination rule of the simulation strategy of the fault type, determining the simulation fault information includes: For any moment when an abnormal keyword appears, determine the occurrence times of the abnormal keyword within the target time range corresponding to the moment, where the end moment of the target time range is the moment; If the occurrence times are not less than the times threshold, determine that a fault alarm is triggered; Add the alarm fault list of the fault alarm to the simulation fault information.
5. The method according to claim 3, characterized in that, the step of, based on the alarm determination rule of the simulation strategy of the fault type, determining the simulation fault information includes: For any moment during the simulation monitoring process, determine the occurrence times of any abnormal keyword within the target time range corresponding to the moment, where the end moment of the target time range is the moment; When the occurrence times is not less than the times threshold, determine to trigger a fault alarm; Add the alarm fault list of the fault alarm to the simulation fault information.
6. The method according to claim 1, wherein, The adjusting the simulation strategy of the fault type based on the difference between the simulation fault information and the actual fault information includes: Based on the simulation fault information and the actual fault information, determine the misjudgment rate of the simulation strategy, and the misjudgment rate of the simulation strategy is used to represent the accuracy of the simulation strategy of the fault type in monitoring the target server; Based on the actual fault information and the similar strategy of the simulation strategy, determine the misjudgment rate of the similar strategy, and the misjudgment rate of the similar strategy is used to represent the accuracy of the similar strategy in monitoring the target server; When the misjudgment rate of the simulation strategy is not less than the misjudgment rate of the similar strategy, adjust the simulation strategy of the fault type.
7. The method according to claim 6, wherein, The determining the misjudgment rate of the simulation strategy based on the simulation fault information and the actual fault information includes: Based on the simulation fault information and the actual fault information, determine a first quantity and a second quantity. The first quantity is the quantity of the first fault list, and the first fault list is the alarm fault list that exists in both the simulation fault information and the actual fault information. The second quantity is the quantity of the second fault list, and the second fault list is the first fault list marked as a false alarm; Determine the ratio of the second quantity to the first quantity as the misjudgment rate of the simulation strategy.
8. The method according to claim 6, wherein, The determining the misjudgment rate of the similar strategy based on the actual fault information and the similar strategy of the simulation strategy includes: Based on the actual fault information and the similar strategy of the simulation strategy, determine a third quantity and a fourth quantity. The third quantity is the quantity of the third fault list, and the third fault list is the alarm fault list determined by the similar strategy in the actual fault information. The fourth quantity is the quantity of the fourth fault list, and the fourth fault list is the third fault list marked as a false alarm; Determine the ratio of the fourth quantity to the third quantity as the misjudgment rate of the similar strategy.
9. The method according to claim 6, wherein, The method further includes: When the misjudgment rate of the simulation strategy is less than the misjudgment rate of the similar strategy, add the simulation strategy of the fault type to the monitoring strategy library, and the monitoring strategy library is used to provide the monitoring strategies required for real-time fault monitoring.
10. The method according to claim 1, wherein, The target server includes multiple servers. One server can correspond to multiple alarm fault lists, and one server corresponds to one fault address. The method further includes: Obtain the spare part threshold corresponding to the target server, and the spare part threshold is the quantity of the standby servers; Based on the simulation fault information, determine the absolute quantity of the fault lists, and the absolute quantity of the fault lists is the quantity of the fault addresses that appear in the simulation fault information; When the absolute number of the trouble tickets is greater than the spare part threshold, a prompt message is displayed, and the prompt message is used to prompt the absolute number of the trouble tickets, the spare part threshold, and at least one trouble address.
11. The method according to claim 1, wherein, the method further includes: Based on the simulated fault information and the actual fault information, a third quantity and a fifth quantity are determined. The third quantity is the quantity of third trouble tickets, and the third trouble tickets are the alarm trouble tickets determined by the similar strategy of the simulation strategy in the actual fault information. The fifth quantity is the quantity of fifth trouble tickets, and the fifth trouble tickets are the alarm trouble tickets determined by the simulation strategy of the fault type in the simulated fault information; The ratio of the difference between the fifth quantity and the third quantity to the third quantity is determined as the trouble ticket change ratio, and the trouble ticket change ratio is used to represent the change degree of the monitoring effect of the simulation strategy of the fault type compared with the similar strategy; The trouble ticket change ratio is displayed on a display page, and the display page is used to display the comparison information between the simulated fault information and the actual fault information.
12. The method according to claim 1, wherein, the method further includes: Obtain the data of the parsing configuration file and the server cluster, and the server cluster includes the target server; Based on the parsing configuration file, parse the data of the server cluster to obtain the server log of the target server.
13. A strategy adjustment device, wherein, the device includes: A first acquisition module, configured to acquire multiple server logs and actual fault information of a target server within a target historical time period, and the actual fault information is used to indicate the server faults that actually occurred on the target server within the target historical time period; A simulation module, configured to, for any fault type, perform simulation monitoring on the target server based on the simulation strategy of the fault type and the multiple server logs to obtain simulated fault information, and the simulation strategy is used to determine whether a server fault of the fault type occurs; An adjustment module, configured to adjust the simulation strategy of the fault type based on the difference between the simulated fault information and the actual fault information.
14. A computer device, wherein, the computer device includes a processor and a memory, and the memory is used to store at least one segment of computer program, and the at least one segment of computer program is loaded and executed by the processor to perform the strategy adjustment method according to any one of claims 1 to 12.
15. A computer-readable storage medium, wherein, the computer-readable storage medium is used to store at least one segment of computer program, and the at least one segment of computer program is used to perform the strategy adjustment method according to any one of claims 1 to 12.