Abnormality detection method and device, equipment and storage medium
By responding to exception alarm events in the back-end service, obtaining the client's exception reporting information and analyzing the call records of the service instance, the problem of difficulty in abnormal positioning after the service upgrade is solved, and more efficient and accurate abnormal change positioning is achieved.
Patent Information
- Application Number
- CN202311707637.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-12
- Publication Date
- 2025-06-13
AI Technical Summary
After the back-end service upgrade and change, it is difficult to locate abnormal service changes efficiently and accurately, especially because there may be time difference in changes in multiple server instances, which leads to difficulty in positioning abnormally.
By responding to the alarm event of the preset exception, obtain the exception reporting information of the client, determine the call record list of the target preset service and its called service instance, determine the set of changed service instances associated with the alarm event, and determine whether the target change event is related to the preset exception based on this information.
It improves the positioning efficiency and accuracy of abnormal changes in service, and can quickly and accurately identify the correlation between change events and abnormalities.
Smart Images

Figure CN120144392A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technologies, and in particular, to an anomaly detection method, apparatus, device, and storage medium. Background Art
[0002] With the rapid development of computer technologies and Internet technologies, the application functions in a client are becoming increasingly rich. Different application functions usually require support from different backend services. To ensure the service quality of the backend services, multiple server instances may be required to implement the same backend service.
[0003] As the application functions are continuously improved, the backend services often need to be upgraded and changed, such as changing the implementation code corresponding to the service. There may be a time difference in the changes of multiple server instances corresponding to the same backend service. After the service is changed, there may be some reasons that cause some clients to crash or have other anomalies. Since the server instances providing services to each client are not fixed, it is difficult to efficiently and accurately locate the abnormal service changes. Summary of the Invention
[0004] Embodiments of the present disclosure provide an anomaly detection method, apparatus, device, and storage medium, which can optimize the existing anomaly detection solution for service changes.
[0005] In a first aspect, embodiments of the present disclosure provide an anomaly detection method, including:
[0006] In response to a preset anomaly alarm event being triggered, obtaining anomaly reporting information of the client where the preset anomaly occurs;
[0007] Determining a target preset service called by the client before the alarm event is triggered according to the anomaly reporting information, and a call record list of the service instance that calls the target preset service;
[0008] Determining a set of changed service instances corresponding to the target preset service, where the set of changed service instances is a set of service instances that have changed in a target change event of the target preset service, and the target change event is associated with the trigger time of the alarm event;
[0009] Determining whether the target change event is related to the preset anomaly according to the relationship between the set of changed service instances and the call record list.
[0010] In a second aspect, embodiments of the present disclosure further provide an anomaly detection apparatus, including:
[0011] An anomaly reporting information acquisition module, configured to obtain anomaly reporting information of the client where the preset anomaly occurs in response to a preset anomaly alarm event being triggered;
[0012] A call record determination module, configured to determine a target preset service called by the client before the alarm event is triggered according to the exception reporting information, and a call record list of the called service instances of the target preset service;
[0013] An instance set determination module, configured to determine a set of changed service instances corresponding to the target preset service, where the set of changed service instances is a set of changed service instances of a target change event of the target preset service, and the target change event is associated with the triggering time of the alarm event;
[0014] An exception change determination module, configured to determine whether the target change event is related to the preset exception according to the relationship between the set of changed service instances and the call record list.
[0015] In a third aspect, an embodiment of the present disclosure further provides an electronic device, where the electronic device includes:
[0016] One or more processors;
[0017] A storage device, configured to store one or more programs,
[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement the exception detection method provided by the embodiment of the present disclosure.
[0019] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium containing computer-executable instructions, where the computer-executable instructions are used to execute the exception detection method provided by the embodiment of the present disclosure when executed by a computer processor.
[0020] The exception detection solution provided by the embodiment of the present disclosure, in response to the triggering of an alarm event of a preset exception, obtains exception reporting information of a client where the preset exception occurs, determines a target preset service called by the client before the alarm event is triggered according to the exception reporting information, and a call record list of the called service instances of the target preset service, determines a set of changed service instances of a target change event of the target preset service associated with the triggering time of the alarm event, and determines whether the target change event is related to the preset exception according to the set of changed service instances and the call record list. By adopting the above technical solution, the positioning efficiency and positioning accuracy of service exception changes can be improved. Description of the Drawings
[0021] In conjunction with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and that the elements and components are not necessarily drawn to scale.
[0022] Figure 1 It is a schematic flowchart of an anomaly detection method provided by an embodiment of the present disclosure;
[0023] Figure 2 It is a schematic flowchart of another anomaly detection method provided by an embodiment of the present disclosure;
[0024] Figure 3 It is a schematic structural diagram of an anomaly detection device provided by an embodiment of the present disclosure;
[0025] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Specific Embodiments
[0026] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0027] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0028] As used herein, the term "comprising" and its variations are open-ended, i.e., "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0029] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules, or units.
[0030] It should be noted that the modifications of "one" and "multiple" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the context, it should be understood as "one or more".
[0031] Figure 1 The following is a schematic flowchart of an anomaly detection method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to scenarios of anomaly detection. This method can be executed by an anomaly detection device, which can be implemented in the form of software and / or hardware. Optionally, it is implemented by an electronic device, which can be a device such as a server. The server can be configured as an anomaly detection platform.
[0032] As Figure 1 shown, the method includes:
[0033] Step 101, in response to the triggering of an alarm event for a preset anomaly, obtain the anomaly reporting information of the client where the preset anomaly occurs.
[0034] Exemplarily, the preset anomaly can be a preset type of anomaly that occurs in the client, such as a crash. Specifically, it can be an error when the client accesses a preset page. For example, when the number of clients with a preset anomaly reaches a preset threshold, and / or when the frequency of the preset anomaly occurring in the client is greater than a preset frequency threshold, an alarm event for the preset anomaly can be triggered. There can also be other triggering conditions, which are not specifically limited.
[0035] In the embodiment of the present disclosure, when the client determines that a preset anomaly has occurred in itself, it can cache the anomaly information when the preset anomaly occurs. When a preset reporting event is triggered, it reports one or more pieces of cached anomaly information to the anomaly detection platform, so it can be recorded as anomaly reporting information. Optionally, the anomaly reporting information can be added to the crash log. When the reporting event (preset reporting event) of the crash log is triggered, the anomaly reporting information can be reported together with the crash log for change attribution analysis. Optionally, the anomaly reporting information can include information about the most recent network request (which can be recorded as the target network request) when the preset anomaly occurs, such as the request link and request identifier, etc. The request link can specifically be a Uniform Resource Locator (URL), and the request identifier can be a unique identifier used to represent a network request initiated by the client, that is, different network requests correspond to different request identifiers.
[0036] Exemplarily, obtain the exception reporting information of the client where the preset exception occurs from the local cache of the exception detection platform. Optionally, there may be multiple types of preset exceptions, and the specific classification method is not limited. For example, it can be divided based on the preset pages accessed. For example, the preset exceptions corresponding to different preset pages can be considered different preset exceptions. For example, when the client reports an error while accessing preset page a, it can correspond to preset exception A; when the client reports an error while accessing preset page b, it can correspond to preset exception B. The exception detection platform can cache the exception reporting information reported by the client. For example, classify and cache the exception reporting information corresponding to different preset exceptions, such as caching the exception reporting information corresponding to preset exception A and caching the exception reporting information corresponding to preset exception B separately. After the alarm event for a certain preset exception is triggered, obtaining the exception reporting information of the client where the preset exception occurs from the cache can improve the acquisition efficiency of the exception reporting information. Optionally, obtain the exception reporting information of the client where the preset exception occurs within a preset time period. The preset time period can be within the target duration range between the triggering of the alarm event, and the target duration can be set according to actual requirements, such as half an hour or 1 hour, etc.
[0037] Step 102: Determine the target preset service called by the client before the alarm event is triggered according to the exception reporting information, and the call record list of the service instance of the target preset service that is called.
[0038] Exemplarily, the preset service can be a service provided by the server corresponding to the client, such as a video playback service, a comment service, a session service, etc. Different preset services can be provided by different service clusters. A service cluster can include multiple distributed service instances, and each service instance can correspond to a server, that is, the service instance can be the smallest unit of the backend server. Determine which preset service or services each client called before the alarm event of the preset exception was triggered according to the exception reporting information, and record the determined preset service as the target preset service. The number of target preset services can be one or more. When calling a preset service, generally, according to the actual situation, such as the load conditions of the service instances in the service cluster, etc., dynamically determine the service instance to be called. When calling the preset service twice, the actually called service instance may be different. The call record list can include a set of service instances called when calling the target preset service (the set does not include duplicate service instances); the call record list can also include the specific service instance called each time when calling the target preset service (denoted as the called service instance). Since different clients may call the same service instance when calling the target preset service, and the same client may also call the same service instance when calling the target preset service multiple times, the call record list can include multiple call records corresponding to the same service instance. Exemplarily, each call record can include the instance identifier of the called service instance, and this instance identifier is used to represent the unique identity of the service instance, such as the instance ID or pod information, etc.
[0039] Exemplarily, when there are multiple target preset services, taking the target preset service as the division dimension, sequentially determine the call record list corresponding to each target preset service respectively. That is, different target preset services can correspond to different call record lists, so as to perform targeted analysis on each target preset service subsequently.
[0040] Optionally, after the client sends the target network request, the server receives the target network request, and can generate a corresponding request identifier for this request, and send the request identifier to the client. When the client has a preset exception, it can cache the received request identifier into the exception information. Optionally, after the server receives the target network request, it can also determine the target preset service to be called and the service instance to be called in response to the target network request, and return the service identifier of the target preset service (representing the unique identifier of the preset service) and the instance identifier of the called service instance to the client. When the client has a preset exception, it can cache the received service identifier and instance identifier into the exception reporting information.
[0041] Step 103: Determine the set of changed service instances corresponding to the target preset service, where the set of changed service instances is the set of service instances that have changed for the target change event of the target preset service, and the target change event is associated with the trigger time of the alarm event.
[0042] Exemplarily, after determining the target preset service, the service change event associated with the trigger time of the alarm event can be retrieved according to the service identifier of the target preset service, which is denoted as the target change event. Optionally, the trigger time of the target change event is within a preset duration range between the triggering of the alarm event, and the preset duration can be set according to actual requirements, such as half an hour or 1 hour, etc.
[0043] Exemplarily, after the service change event is triggered, the service instances in the service instance set of the corresponding preset service generally do not change synchronously at one time, but change in batches. For example, if there are 10 service instances in the service instance set, 2 of them may change first, then 3, and finally 5, etc. Therefore, when the alarm event is triggered, there may be changed service instances or unchanged service instances in the service instance set of the target preset service. If the preset anomaly is related to a certain service change event, it is generally related to the changed service instances. Therefore, in this step, determining the set of service instances that have changed for the target change event of the target preset service, that is, the set of changed service instances, facilitates subsequent change attribution analysis.
[0044] Step 104: Determine whether the target change event is related to the preset anomaly according to the relationship between the set of changed service instances and the call record list.
[0045] Exemplarily, it can be determined whether the target change event is related to the preset anomaly according to whether the changed service instances in the set of changed service instances appear in the call record list. For example, if they appear, it is related to the preset anomaly, and if they do not appear, it is not related to the preset anomaly. Optionally, it can also be determined whether the target change event is related to the preset anomaly according to the number of changed service instances that appear in the call record list in the set of changed service instances. For example, if the number of appearances is greater than the preset appearance number threshold, it is related to the preset anomaly, otherwise, it is not related to the preset anomaly. Optionally, it can also be determined according to whether the called service instances in each call record in the call record list appear in the set of changed service instances. If they appear, it is related to the preset anomaly, and if they do not appear, it is not related to the preset anomaly.
[0046] In the exception detection method according to the embodiments of the present disclosure, after a preset exception alarm event is triggered, the target preset service and service instance to be called can be accurately determined according to the exception reporting information of the client where the preset exception occurs. According to the relationship between the changed service instances of the change event associated with the triggering time of the alarm event of the target preset service and the call record list of the called service instances, it can be quickly and accurately determined whether the change event is related to the preset exception, thereby improving the positioning efficiency and positioning accuracy of service exception changes.
[0047] In some embodiments, each call record in the call record list corresponds to a call of the called service instance; determining whether the target change event is related to the preset exception according to the relationship between the set of changed service instances and the call record list includes: determining whether the target change event is related to the preset exception according to the ratio of a first value to a second value, where the first value is the number of times the called service instance in each call record in the call record list appears in the set of changed service instances, and the second value is the number of call records in the call record list. Thus, it can be more accurately determined whether the target change event is related to the preset exception.
[0048] For ease of explanation, by way of example, assume that the target preset service is Service A, and the set of service instances corresponding to Service A includes 10 service instances, assumed to be Service Instance 1 to Service Instance 10. Within half an hour before the triggering time of the preset exception alarm event, Service A triggered a target change event, and the set of changed service instances corresponding to the target change event includes Service Instance 1, Service Instance 2, and Service Instance 3. Suppose there are 8 call records in the call record list, that is, the second value is 8, corresponding to Service Instance 1, Service Instance 4, Service Instance 2, Service Instance 3, Service Instance 5, Service Instance 3, Service Instance 9, and Service Instance 7 respectively. Then, the number of times each called service instance in the call record list appears in the set of changed service instances is 4 times, that is, the first value is 4. Then, the ratio of the first value to the second value is 4 / 8 = 1 / 2. According to this ratio, it is determined whether the target change event is related to the preset exception.
[0049] Exemplarily, if the ratio is relatively high, it can be explained that the correlation between the preset exception and the target change event is relatively high. Optionally, it is determined whether the ratio of the first value to the second value is greater than or equal to a preset ratio threshold. If so, it is determined whether the target change event is related to the preset exception. Optionally, if not, it can be determined whether the target change event is not related to the preset exception. As in the above example, assume that the preset ratio threshold is 40%. Then, the target change event is related to the preset exception.
[0050] Optionally, when the number of target preset services is multiple, the ratios corresponding to each target preset service can also be sorted in descending order, and the target preset services with ranking serial numbers greater than or equal to the preset serial number value are determined as the hit services. Then, the target change events corresponding to the hit services are related to the preset exception. The preset serial number value can be, for example, 1 or 2, etc., which is set according to the actual situation.
[0051] Optionally, the target change event related to the preset exception can be determined by combining the preset ratio threshold and the preset serial number value. For example, the target preset service with a ratio greater than or equal to the preset ratio threshold and a serial number greater than or equal to the preset serial number value is determined as the hit service. Then, the target change event corresponding to the hit service is related to the preset exception.
[0052] Figure 2 The figure is a schematic flowchart of another anomaly detection method provided by an embodiment of the present disclosure. The embodiments of the present disclosure are optimized based on each of the above optional solutions. The anomaly reporting information includes the request identifier of the target network request reported by the client. The target network request is the network request initiated by the client for the last time each time a preset anomaly occurs before the alarm event is triggered. According to the request identifier, the target preset service and the call record list can be quickly and accurately determined. Specifically, as Figure 2 shown, the method includes the following steps:
[0053] Step 201: In response to the triggering of the alarm event of the preset anomaly, obtain the anomaly reporting information of the client where the preset anomaly occurs.
[0054] Step 202: For each request identifier in the anomaly reporting information, obtain the call chain information corresponding to the current request identifier. The call chain information includes the preset service called by the target network request to which the current request identifier belongs and the called service instance of the called preset service.
[0055] Optionally, the request identifier is generated by the server that receives the target network request and sent to the client. The server records the corresponding relationship between the request identifier and the call chain information. For example, after the server receives the target network request of client C, a unique request identifier is generated for this request and sent to client C. After the server generates the request identifier, it can also record the corresponding relationship between the request identifier and the call chain information. This corresponding relationship can be stored in the server or sent by the server to the anomaly detection platform. In this step, the anomaly detection platform can obtain the corresponding call chain information from the server or locally according to the request identifier in the anomaly reporting information.
[0056] Exemplarily, to process a network request, multiple preset services may be called sequentially. For example, when processing network request 1, preset services E, F, and G are called respectively; when processing network request 2, preset services E and G are called respectively; when processing network request 3, preset services G and H are called respectively. Taking network request 1 as an example, the call chain information may include the service identifiers corresponding to preset services E, F, and G respectively. When calling preset services E, F, and G respectively, specific service instances need to be called, such as service instances E1, F2, and G1, etc. The call chain information may also include the instance identifiers of the called service instances.
[0057] Optionally, the exception reporting information further includes the request link of the target network request. Before obtaining the call chain information corresponding to the current request identifier for each request identifier in the exception reporting information, it further includes: aggregating the request identifiers in the exception reporting information with the request link as the aggregation dimension to obtain a set of request identifiers corresponding to each request link respectively; determining the number of request identifiers in each set of request identifiers; and determining the set of request identifiers with the number of request identifiers greater than the preset number threshold as the target request representation set. Among them, obtaining the call chain information corresponding to the current request identifier for each request identifier in the exception reporting information includes: obtaining the call chain information corresponding to the current request identifier for each request identifier in the target request representation set. Thus, preset services or service instances corresponding to request identifiers with relatively low probabilities can be filtered out in advance, improving the exception detection efficiency.
[0058] Exemplarily, the request link is a URL, and the URL contains path information. The path can be, for example, the part after the domain name (such as.com) in the URL. When different clients request the same page or the same data at different times, the path is the same, that is, the URL is the same. Therefore, aggregation can be performed based on the URL. Since there are cases where the same client requests the same URL multiple times and different clients request the same URL respectively, for the same URL, multiple request identifiers can correspond. Assuming the request identifier is denoted as logid, the exception reporting information can be aggregated according to the URL into a format such as {URL1+logid1 / logid2 / logid3}, that is, each request link can correspond to a set of request identifiers. Count the number of request identifiers in each set of request identifiers. As in the above example, the number of request identifiers in the set of request identifiers corresponding to URL1 is 3. If the number is small, it indicates that the situation of preset exceptions occurring when requesting the corresponding URL is less. Therefore, the relevance between the preset services included in the call chain information corresponding to the request identifier and the current alarm event is small and can be filtered out. That is, the set of request identifiers with the number of request identifiers greater than the preset number threshold is determined as the target request representation set and used for subsequent analysis. The preset number threshold can be set according to actual needs, such as 1 or 2, etc.
[0059] Step 203: Dedup the preset services in the call chain information to obtain the target preset services called by the client before the alarm event is triggered.
[0060] Exemplarily, as in the above example, after deduping the preset services in the call chain information, the target preset services can be obtained as preset service E, preset service F, preset service G, and preset service H.
[0061] Step 204: For each target preset service, record one by one the called service instances of the current target preset service in the call chain information to obtain a call record list of the called service instances of the current target preset service.
[0062] Taking preset service E as an example, every time a called service instance appears in the call chain information, a call record is added to the call record list. For example, it may be service instance E1, service instance E2, service instance E1, service instance E3, service instance E3, service instance E2, and service instance E1, etc. Denote the call record list as logid_pod_list.
[0063] Step 205: Determine the set of changed service instances corresponding to the target preset service, where the set of changed service instances is the set of changed service instances of the target change event of the target preset service, and the triggering time of the target change event is within the preset time range between the triggering of the alarm event.
[0064] Exemplarily, the set of changed service instances can be denoted as deployed_pod_list.
[0065] Step 206: Determine the ratio of the first value to the second value, where the first value is the number of times the called service instance in each call record in the call record list appears in the set of changed service instances, and the second value is the number of call records in the call record list.
[0066] Exemplarily, loop through the called service instances in each call record in logid_pod_list, determine whether the current called service instance appears in deployed_pod_list. If it appears, increment the counter by 1. Finally, use the value of the counter as the first value, use len(logid_pod_list) (the length of the call record list, that is, the total number of call records) as the second value, and denote the ratio of the first value to the second value as the grouping rate. The higher the grouping rate, the more suspicious the corresponding service change is considered.
[0067] Step 207: Determine whether the ratio is greater than or equal to a preset ratio threshold. If so, execute Step 208; otherwise, execute Step 209.
[0068] Exemplarily, if there are multiple target preset services, determine in sequence whether each ratio is greater than or equal to the preset ratio threshold. If all are less than, execute Step 209. If there is at least one ratio greater than or equal to the preset ratio threshold, execute Step 208.
[0069] Step 208: Determine that the target change event is related to a preset anomaly.
[0070] Exemplarily, determine the target change event of the target preset service whose corresponding ratio is greater than or equal to the preset ratio threshold as a service change event related to the preset anomaly. Optionally, after determining the service change event related to the preset anomaly, a prompt operation can be triggered to prompt relevant personnel (such as personnel who release the change or testers, etc.) or the system to take corresponding measures for further problem troubleshooting or targeted repair and other processing.
[0071] Step 209: Determine that the target change event has nothing to do with the preset anomaly.
[0072] The anomaly detection method provided by the embodiments of the present disclosure, after a preset anomaly alarm event is triggered, obtains the anomaly reporting information of the client where the preset anomaly occurs, obtains the corresponding call chain information according to the request identifier in the anomaly reporting information, and can accurately determine the called target preset service and the call record list of the called service instance according to the call chain information. According to the trigger time of the alarm event, the target change event associated with the target preset service is found, and then the set of changed service instances corresponding to the target change event is determined. Calculate the ratio of the number of times the changed service instance appears in the call record list to the number of call records. If the ratio is greater than or equal to the preset ratio threshold, it can be quickly and accurately determined that the target change event is related to the preset anomaly, further improving the positioning efficiency and positioning accuracy of service anomaly changes.
[0073] Figure 3 The following is a schematic structural diagram of an anomaly detection device provided by the embodiments of the present disclosure, as Figure 3 shown, the device includes:
[0074] An anomaly reporting information acquisition module 301, configured to obtain the anomaly reporting information of the client where the preset anomaly occurs in response to the triggering of a preset anomaly alarm event;
[0075] A call record determination module 302, configured to determine the target preset service called by the client before the alarm event is triggered according to the anomaly reporting information, and the call record list of the called service instance of the target preset service;
[0076] An instance set determination module 303, configured to determine the set of changed service instances corresponding to the target preset service, where the set of changed service instances is the set of changed service instances of the target change event of the target preset service, and the target change event is associated with the trigger time of the alarm event;
[0077] An anomaly change determination module 304, configured to determine whether the target change event is related to the preset anomaly according to the relationship between the set of changed service instances and the call record list.
[0078] The anomaly detection device provided by the embodiments of the present disclosure, after a preset anomaly alarm event is triggered, can accurately determine the called target preset service and service instance according to the anomaly reporting information of the client where the preset anomaly occurs. According to the relationship between the changed service instances of the change event associated with the trigger time of the alarm event of the target preset service and the call record list of the called service instances, it can be quickly and accurately determined whether the change event is related to the preset anomaly, thereby improving the positioning efficiency and positioning accuracy of service anomaly changes.
[0079] Optionally, each call record in the call record list corresponds to one call of the called service instance; wherein, the exception change determination module is specifically configured to: determine whether the target change event is related to the preset exception according to the ratio of the first value to the second value, where the first value is the number of times the called service instance in each call record in the call record list appears in the set of changed service instances, and the second value is the number of call records in the call record list.
[0080] Optionally, the exception reporting information includes the request identifier of the target network request reported by the client, and the target network request is the network request initiated by the client each time the preset exception occurs and is the most recent network request before the alarm event is triggered;
[0081] Among them, the call record determination module includes:
[0082] The call chain information acquisition unit is configured to, for each request identifier in the exception reporting information, acquire the call chain information corresponding to the current request identifier, where the call chain information includes the preset service called by the target network request to which the current request identifier belongs and the called service instance of the called preset service;
[0083] The target preset service determination unit is configured to perform deduplication processing on the preset services in the call chain information to obtain the target preset services called by the client before the alarm event is triggered;
[0084] The call record list determination unit is configured to, for each target preset service, record one by one the called service instances of the current target preset service in the call chain information to obtain the call record list of the called service instances of the current target preset service.
[0085] Optionally, the request identifier is generated by the server that receives the target network request and sent to the client, and the server records the corresponding relationship between the request identifier and the call chain information.
[0086] Optionally, the exception reporting information further includes the request link of the target network request; the apparatus further includes:
[0087] The request identifier set determination module is configured to, before acquiring the call chain information corresponding to each request identifier in the exception reporting information, aggregate the request identifiers in the exception reporting information with the request link as the aggregation dimension to obtain a request identifier set corresponding to each request link respectively;
[0088] An identification quantity determination module, configured to determine the quantity of request identifications in each of the said request identification sets;
[0089] A target request representation set determination module, configured to determine, as the target request representation set, the request identification set in which the quantity of request identifications is greater than a preset quantity threshold;
[0090] Wherein, the call record list determination unit is specifically configured to: for each of the request identifications in the target request representation set, obtain the call chain information corresponding to the current request identification.
[0091] Optionally, the triggering time of the target change event is within a preset time range between the triggering of the alarm event.
[0092] Optionally, determining whether the target change event is related to the preset anomaly according to the ratio of the first value to the second value includes: determining whether the ratio of the first value to the second value is greater than or equal to a preset ratio threshold, and if so, determining whether the target change event is related to the preset anomaly.
[0093] The anomaly detection device provided by the embodiments of the present disclosure can execute the anomaly detection method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.
[0094] It should be noted that the various units and modules included in the above device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present disclosure.
[0095] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. The following refers to Figure 4 , which shows a schematic structural diagram of an electronic device (such as Figure 4 the terminal device or server in) 400 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The electronic device shown is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.
[0096] Such as Figure 4As shown, the electronic device 400 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 401, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 are also stored. The processing device 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An editing / output (I / O) interface 405 is also connected to the bus 404.
[0097] Generally, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 4 an electronic device 400 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. More or fewer devices may be implemented or included alternatively.
[0098] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0099] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0100] The electronic device provided in the embodiment of the present disclosure and the anomaly detection method provided in the above embodiment belong to the same inventive concept. Technical details not described in detail in this embodiment may be referred to in the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0101] The embodiment of the present disclosure provides a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, the anomaly detection method provided in the above embodiment is implemented.
[0102] An embodiment of the present disclosure provides a computer program product, including a computer program, which implements the anomaly detection method provided in the above embodiment when executed by a processor.
[0103] It should be noted that the computer-readable medium in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium may be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0104] In some embodiments, the client and the server may communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed network.
[0105] The above computer-readable medium may be included in the above electronic device; or may exist separately without being assembled into the electronic device.
[0106] The above computer-readable medium carries one or more programs which, when executed by the electronic device, cause the electronic device to: in response to a preset abnormal alarm event being triggered, obtain the abnormal reporting information of the client where the preset abnormality occurs; determine, according to the abnormal reporting information, the target preset service called by the client before the alarm event is triggered, and the call record list of the called service instance of the target preset service; determine the set of changed service instances corresponding to the target preset service, where the set of changed service instances is the set of changed service instances of the target change event of the target preset service, and the target change event is associated with the triggering time of the alarm event; determine whether the target change event is related to the preset abnormality according to the relationship between the set of changed service instances and the call record list.
[0107] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0108] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that, in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0109] The units involved in the embodiments of the present disclosure can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the module itself in some cases. For example, the abnormal reporting information acquisition module can also be described as "a module that acquires the abnormal reporting information of the client where the preset abnormality occurs in response to the triggering of an alarm event for the preset abnormality".
[0110] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0111] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable Read Only Memory (EPROM or Flash Memory), optical fibers, portable compact disc read only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0112] According to one or more embodiments of the present disclosure, an abnormal detection method is provided, including:
[0113] In response to the triggering of an alarm event for a preset abnormality, acquiring the abnormal reporting information of the client where the preset abnormality occurs;
[0114] Determining the target preset service called by the client before the triggering of the alarm event according to the abnormal reporting information, and the call record list of the called service instance of the target preset service;
[0115] Determining the set of changed service instances corresponding to the target preset service, where the set of changed service instances is the set of changed service instances of the target change event of the target preset service, and the target change event is associated with the triggering time of the alarm event;
[0116] Determine whether the target change event is related to the preset exception according to the relationship between the set of changed service instances and the call record list.
[0117] According to one or more embodiments of the present disclosure, each call record in the call record list corresponds to a call to the called service instance.
[0118] Among them, determining whether the target change event is related to the preset exception according to the relationship between the set of changed service instances and the call record list includes:
[0119] Determine whether the target change event is related to the preset exception according to the ratio of the first value to the second value, where the first value is the number of times the called service instance in each call record in the call record list appears in the set of changed service instances, and the second value is the number of call records in the call record list.
[0120] According to one or more embodiments of the present disclosure, the exception reporting information includes the request identifier of the target network request reported by the client, and the target network request is the network request initiated by the client each time the preset exception occurs and is the most recent one before the alarm event is triggered.
[0121] Among them, determining the target preset service called by the client before the alarm event is triggered and the call record list of the called service instances of the target preset service according to the exception reporting information includes:
[0122] For each request identifier in the exception reporting information, obtain the call chain information corresponding to the current request identifier, where the call chain information includes the preset service called by the target network request to which the current request identifier belongs and the called service instance of the called preset service.
[0123] Deduplicate the preset services in the call chain information to obtain the target preset service called by the client before the alarm event is triggered.
[0124] For each target preset service, record one by one the called service instances of the current target preset service in the call chain information to obtain the call record list of the called service instances of the current target preset service.
[0125] According to one or more embodiments of the present disclosure, the request identifier is generated by the server that receives the target network request and sent to the client, and the server records the corresponding relationship between the request identifier and the call chain information.
[0126] According to one or more embodiments of the present disclosure, the abnormal reporting information further includes the request link of the target network request; before obtaining the call chain information corresponding to the current request identifier for each request identifier in the abnormal reporting information, it further includes:
[0127] Aggregate the request identifiers in the abnormal reporting information with the request link as the aggregation dimension to obtain a set of request identifiers corresponding to each request link;
[0128] Determine the number of request identifiers in each set of request identifiers;
[0129] Determine the set of target request representations with the number of request identifiers greater than the preset number threshold as the target request representation set;
[0130] Wherein, obtaining the call chain information corresponding to the current request identifier for each request identifier in the abnormal reporting information includes:
[0131] For each request identifier in the set of target request representations, obtain the call chain information corresponding to the current request identifier.
[0132] According to one or more embodiments of the present disclosure, the triggering time of the target change event is within a preset duration range between the triggering of the alarm event.
[0133] According to one or more embodiments of the present disclosure, determining whether the target change event is related to the preset abnormality according to the ratio of the first value to the second value includes:
[0134] Judge whether the ratio of the first value to the second value is greater than or equal to the preset ratio threshold. If so, determine whether the target change event is related to the preset abnormality.
[0135] According to one or more embodiments of the present disclosure, an abnormal detection device is provided, including:
[0136] An abnormal reporting information acquisition module, configured to acquire abnormal reporting information of a client where the preset abnormality occurs in response to the triggering of an alarm event of the preset abnormality;
[0137] A call record determination module, configured to determine a target preset service called by the client before the triggering of the alarm event according to the abnormal reporting information, and a call record list of the called service instance of the target preset service;
[0138] An instance set determination module, configured to determine a set of changed service instances corresponding to the target preset service, where the set of changed service instances is a set of changed service instances of a target change event of the target preset service, and the target change event is associated with the triggering time of the alarm event;
[0139] An abnormal change determination module, configured to determine whether the target change event is related to the preset abnormality according to the relationship between the set of changed service instances and the call record list.
[0140] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, a technical solution formed by mutually replacing the above features with technical features having similar functions disclosed in the present disclosure (but not limited to).
[0141] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0142] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.
Claims
1. An anomaly detection method, characterized in that, it includes: In response to the triggering of an alarm event for a preset anomaly, obtaining anomaly reporting information of the client where the preset anomaly occurs; Determining a target preset service called by the client before the alarm event is triggered according to the anomaly reporting information, and a call record list of the called service instances of the target preset service; Determining a set of changed service instances corresponding to the target preset service, where the set of changed service instances is a set of changed service instances of a target change event of the target preset service, and the target change event is associated with the triggering time of the alarm event; Determining whether the target change event is related to the preset anomaly according to the relationship between the set of changed service instances and the call record list.
2. The method according to claim 1, characterized in that, Each call record in the call record list corresponds to one call of the called service instance; wherein, the determining whether the target change event is related to the preset anomaly according to the relationship between the set of changed service instances and the call record list includes: Determining whether the target change event is related to the preset anomaly according to the ratio of a first value to a second value, where the first value is the number of times the called service instance in each call record in the call record list appears in the set of changed service instances, and the second value is the number of call records in the call record list.
3. The method according to claim 2, characterized in that, The anomaly reporting information includes a request identifier of a target network request reported by the client, and the target network request is the network request initiated most recently by the client each time the preset anomaly occurs before the alarm event is triggered; wherein, the determining the target preset service called by the client before the alarm event is triggered according to the anomaly reporting information, and the call record list of the called service instances of the target preset service includes: For each request identifier in the anomaly reporting information, obtaining call chain information corresponding to the current request identifier, where the call chain information includes a preset service called by the target network request to which the current request identifier belongs and the called service instances of the preset service called; Performing deduplication processing on the preset services in the call chain information to obtain the target preset service called by the client before the alarm event is triggered; For each target preset service, recording one by one the called service instances of the current target preset service in the call chain information to obtain a call record list of the called service instances of the current target preset service.
4. The method according to claim 3, characterized in that, The request identifier is generated by the server that receives the target network request and sent to the client, and the server records the corresponding relationship between the request identifier and the call chain information.
5. The method according to claim 3, characterized in that, The abnormal reporting information further includes the request link of the target network request; Before obtaining the call chain information corresponding to the current request identifier for each of the request identifiers in the abnormal reporting information, it further includes: Aggregating the request identifiers in the abnormal reporting information with the request link as the aggregation dimension to obtain a set of request identifiers corresponding to each request link; Determining the number of request identifiers in each set of request identifiers; Determining the set of request identifiers with the number of request identifiers greater than a preset number threshold as the target request representation set; Among them, obtaining the call chain information corresponding to the current request identifier for each of the request identifiers in the abnormal reporting information includes: For each of the request identifiers in the target request representation set, obtaining the call chain information corresponding to the current request identifier.
6. The method according to claim 1, characterized in that The triggering time of the target change event is within a preset duration range between the triggering of the alarm event.
7. The method according to any one of claims 2-6, characterized in that Determining whether the target change event is related to the preset abnormality according to the ratio of the first value to the second value includes: Judging whether the ratio of the first value to the second value is greater than or equal to a preset ratio threshold. If so, determining whether the target change event is related to the preset abnormality.
8. An abnormal detection device, characterized in that including: An abnormal reporting information acquisition module, configured to acquire abnormal reporting information of a client where the preset abnormality occurs in response to the triggering of an alarm event of the preset abnormality; A call record determination module, configured to determine a target preset service called by the client before the triggering of the alarm event according to the abnormal reporting information, and a call record list of the called service instance of the target preset service; An instance set determination module, configured to determine a set of changed service instances corresponding to the target preset service, where the set of changed service instances is a set of changed service instances of a target change event of the target preset service, and the target change event is associated with the triggering time of the alarm event; An abnormal change determination module, configured to determine whether the target change event is related to the preset abnormality according to the relationship between the set of changed service instances and the call record list.
9. An electronic device, characterized in that The electronic device includes: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the abnormal detection method according to any one of claims 1-7.
10. A storage medium containing computer-executable instructions, characterized in that The computer-executable instructions are used to execute the abnormal detection method according to any one of claims 1-7 when executed by a computer processor.