Method, device and computer program product for abnormality detection

By obtaining and analyzing information about the client's target request and its contextual request, and using machine learning models for anomaly detection, the problem of abnormal request detection in public computing environments is solved and security is improved.

CN114091016BActive Publication Date: 2025-09-19EMC IP HLDG CO LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010790880.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-07
Publication Date
2025-09-19
Estimated Expiration
2040-08-07

AI Technical Summary

Technical Problem

In public computing environments, existing technologies are difficult to effectively detect and prevent abnormal requests, leading to security issues such as data theft and data leakage.

Method used

By obtaining information about the client's target request and its context request, anomaly detection is performed using a machine learning model, converted into a vectorized feature representation, and the anomaly detection model is used to determine the abnormality of the request.

Benefits of technology

It improves the accuracy of detecting abnormal requests, effectively prevents the execution of abnormal requests, and improves the security of applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114091016B_ABST
    Figure CN114091016B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, device, and computer program product for abnormality detection. The method provided by an embodiment of the present disclosure includes obtaining information related to a target request and at least one context request initiated by a client to an application, the information indicating at least the type and initiation time of the target request and the type and initiation time of at least one context request; converting the obtained information into a vectorized feature representation for the target request; and using an abnormality detection model to determine an abnormality detection result for the target request based on the vectorized feature representation, the abnormality detection result indicating whether the target request is an abnormal request, and the abnormality detection model characterizing the correlation between the vectorized feature representation for the request and the abnormality detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to the field of artificial intelligence (AI), and more particularly to methods, apparatuses, devices, and computer program products for abnormality detection. Background Art

[0002] Currently, many applications are deployed in public computing environments, such as public clouds. Users can initiate requests to applications through clients to obtain corresponding services. This deployment allows services to be provided to a wider range of users. However, in public computing environments, applications may face a variety of attacks, which can lead to data theft, data leakage, malicious data deletion or modification, and other problems. Injection hijacking and token hijacking are common attack methods. Therefore, application security is a critical issue. Summary of the Invention

[0003] According to some embodiments of the present disclosure, a solution for abnormality detection is provided.

[0004] In a first aspect of the present disclosure, a method for abnormality detection is provided. The method includes obtaining information related to a target request and at least one context request initiated by a client to an application, the information indicating at least the type and initiation time of the target request and the type and initiation time of at least one context request; converting the obtained information into a vectorized feature representation for the target request; and determining an abnormality detection result for the target request based on the vectorized feature representation using an abnormality detection model, the abnormality detection result indicating whether the target request is an abnormal request, the abnormality detection model characterizing a correlation between the vectorized feature representation of the request and the abnormality detection result.

[0005] In a second aspect of the present disclosure, an electronic device is provided. The electronic device includes: at least one processor; and at least one memory storing computer program instructions, wherein the at least one memory and the computer program instructions are configured to, together with the at least one processor, enable the electronic device to perform an action. The action includes: obtaining information related to a target request and at least one context request initiated by a client to an application, the information at least indicating the type and initiation time of the target request, and the type and initiation time of at least one context request; converting the obtained information into a vectorized feature representation for the target request; and using an anomaly detection model to determine an anomaly detection result of the target request based on the vectorized feature representation, the anomaly detection result indicating whether the target request is an abnormal request, and the anomaly detection model characterizing the correlation between the vectorized feature representation for the request and the anomaly detection result.

[0006] In a third aspect of the present disclosure, a computer program product is provided, the computer program product being tangibly stored on a non-volatile computer-readable medium and comprising computer-executable instructions, which, when executed, cause a device to perform an action. The action comprises: obtaining information related to a target request and at least one context request initiated by a client to an application, the information indicating at least the type and initiation time of the target request and the type and initiation time of at least one context request; converting the obtained information into a vectorized feature representation for the target request; and determining an abnormality detection result for the target request based on the vectorized feature representation using an abnormality detection model, the abnormality detection result indicating whether the target request is an abnormal request, the abnormality detection model characterizing a correlation between the vectorized feature representation for the request and the abnormality detection result.

[0007] This summary is provided to introduce a selection of concepts in a simplified form that are further described in the detailed description below. This summary is not intended to identify key features or essential features of the disclosure, nor is it intended to limit the scope of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The above and other objects, features and advantages of the embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an illustrative and non-limiting manner, in which:

[0009] Figure 1 shows an example environment in which embodiments of the present disclosure can be implemented;

[0010] Figure 2 A block diagram of a system for abnormality detection according to some embodiments of the present disclosure is shown;

[0011] Figure 3 shows an example representation of a mapping between types of requests and predetermined values ​​according to some embodiments of the present disclosure;

[0012] Figure 4 shows an example representation of a one-hot encoding of a parameter description according to some embodiments of the present disclosure;

[0013] Figure 5 shows an example of sample boundaries learned by an anomaly detection model according to some embodiments of the present disclosure;

[0014] Figure 6 A flowchart illustrating a process of abnormality detection according to some embodiments of the present disclosure; and

[0015] Figure 7 A block diagram is shown of a computing device in which one or more embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION

[0016] The following describes preferred implementations of the present disclosure in more detail with reference to the accompanying drawings. Although preferred implementations of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the implementations described herein. Rather, these implementations are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.

[0017] As used herein, the term "including" and its variations represent open inclusion, i.e., "including but not limited to." Unless otherwise stated, the term "or" means "and / or." The term "based on" means "based at least in part on." The terms "an example implementation" and "an implementation" mean "at least one example implementation." The term "another implementation" means "at least one additional implementation." The terms "first," "second," etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0018] As used herein, the term "model" refers to a model that learns the relationships between inputs and outputs from training data, thereby generating corresponding outputs for given inputs after training. Therefore, the trained model can be considered to be able to characterize this relationship between inputs and outputs.

[0019] Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that uses multiple layers of processing units to process inputs and provide corresponding outputs. A neural network model is an example of a deep learning-based model. In this document, a "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.

[0020] Generally, machine learning can include three stages, namely the training stage, the testing stage, and the use stage (also called the inference stage). In the training stage, a given model can be trained using a large amount of training data, and iterated continuously until the model can obtain consistent inferences from the training data that are similar to those that human intelligence can make. Through training, the model can be considered to be able to learn the association between input and output (also called input-to-output mapping) from the training data. The model can be represented as a function that maps input to output. The parameter values ​​of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. In the use stage, the model can be used to process the actual input based on the parameter values ​​obtained through training to determine the corresponding output.

[0021] Figure 11 shows an example environment in which embodiments of the present disclosure can be implemented. Figure 1 As shown, example environment 100 includes a server 110 on which an application 112 is deployed. Application 112 may be an application that provides any type of service. One example of application 112 includes a data protection application, which is used to provide data protection-related services to customers, including data storage and access, maintenance, backup, and recovery. Other examples of application 112 include data processing applications, data analysis applications, and the like.

[0022] In example environment 100, one or more clients 120-1, 120-2, ..., 120-N can access application 112. For ease of discussion, clients 120-1, 120-2, ..., 120-N are collectively or individually referred to herein as client 120, where N is an integer greater than or equal to 1. Client 120 can send a target request 102 for application 112 to server 110 to request a corresponding service. For example, for a data protection application, target request 102 can be a request to back up data from client 120. After receiving target request 102, server 110 can process target request 102 to run application 112 to provide the corresponding service.

[0023] In some embodiments, server 110 and / or client 120 may be deployed in a public computing environment, such as a cloud computing environment. They may be any device with computing capabilities. Server 110 may be any physical or virtual device with computing capabilities, such as a centralized server, a distributed server, a mainframe, an edge computing device, etc. Although shown as a single device, server 110 or client 120 may be implemented by one or more physical or virtual devices.

[0024] It should be understood that although Figure 1 A single application deployed in a server is shown, but multiple identical or different applications may exist. Environment 110 may also include more servers and clients. The embodiments of the present disclosure are not limited in this respect.

[0025] When processing requests to an application, some may be malicious requests from an attacker. The execution of malicious requests can lead to various security issues, such as data theft, data leakage, and malicious data deletion or modification. To improve security, it is desirable to be able to identify whether a client request is normal or abnormal. However, currently, there is no suitable solution for detecting such abnormal requests, especially for applications deployed in public computing environments that may receive requests from multiple clients.

[0026] An embodiment of the present disclosure proposes a scheme for abnormality detection. In this scheme, for a target request initiated by a client to an application, the context of the target request is obtained, including one or more context requests initiated previously or later by the same client that initiated the target request. For the purpose of abnormality detection, information related to the target request and one or more context requests is obtained, including at least the type and initiation time of these requests. This scheme also uses a machine learning model to implement abnormality detection. Specifically, the obtained information is converted into a vectorized feature representation. The vectorized feature representation is processed using an abnormality detection model to determine the abnormality detection result of the target request, which indicates whether the target request is an abnormal request.

[0027] According to the embodiments of the present disclosure, by considering the context of requests initiated by the same client, it is possible to more accurately determine whether the current request is normal or abnormal. Detecting abnormal requests can effectively avoid potential dangers brought about by the execution of abnormal requests and improve the security level of the application.

[0028] Hereinafter, some example embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.

[0029] Figure 2 A block diagram of an abnormality detection system 200 according to some embodiments of the present disclosure is shown. Abnormality detection system 200 can be deployed to detect whether a target request 102 from one or more clients 120 to an application 112 is an abnormal request. In some embodiments, the abnormality detection result regarding whether the target request 102 is a normal request or an abnormal request can be used by server 110 to determine how to handle the target request 102.

[0030] As shown in the figure, the abnormality detection system 200 includes an information collector 210, an information processor 220, and an abnormality detector 230. The abnormality detection system 200 can be implemented on a single or multiple computing devices with computing capabilities. In some embodiments, the abnormality detection system 200 can be implemented together with the server 110 in a public computing environment, such as a cloud computing environment. The various components of the abnormality detection system 220 can also be implemented by a single or multiple computing devices. Although in Figure 2 Although shown as a separate system in FIG, in some embodiments, the abnormality detection system 200 may also be implemented in the server 110.

[0031] According to an embodiment of the present disclosure, when performing abnormality detection on a target request, the context of the target request is simultaneously obtained, thereby extracting more features for abnormality detection. Specifically, during operation, the information collector 210 is configured to obtain information 203 related to the target request 102 and at least one context request initiated by the client 120 to the application 112 from the server 110.

[0032] In this document, the context request of a target request refers to a request whose initiation time is before or after the initiation time of the target request. The target request and its context request may be requests initiated by the same client 120 and directed to the same application 112. In some embodiments, the context request of the target request 102 may include one or more requests whose initiation time is before the initiation time of the target request 102. Such requests are also referred to as historical requests. This is advantageous in many scenarios, especially when fast real-time monitoring of the target request 102 is required. Of course, in some embodiments, if the processing delay requirement of the target request 102 is not high, or the speed at which the corresponding client 120 initiates the request is fast, one or more subsequent requests whose initiation time is after the initiation time of the target request 102 may also be collected.

[0033] The server 110 may record each request initiated by the client 120 and record it as a context request for the client 120. Figure 2 As shown, server 110 may record context requests 202-1 for client 120-1, context requests 202-2 for client 120-2, ..., and context requests 202-N for client 120-N. For ease of discussion, context requests 202-1, 202-2, ..., 202-N may be collectively or individually referred to as context requests 202. Server 110 may store each request from client 120 in a context request record corresponding to that client.

[0034] The information 203 includes the type and initiation time of the target request 102, and also includes the type and initiation time of each of the one or more context requests 202. The request initiated by the client 120 to the application 112 may include the type of the request, the initiation time and a description of the parameters associated with the request. Depending on the services that the application 112 can provide, the type of request initiated by the client 120 may be different. Different types of requests may require the application 112 to perform different tasks. For the same application 112, the requests initiated by the client 120 may also be divided into multiple types. For example, for a data protection application, the types of requests initiated by the client 120 may include client settings, client registration, data backup, data recovery, data maintenance, and the like.

[0035] In some embodiments, the parameter description of a request initiated by client 120 may include a definition of the parameters required to execute the request. The parameter description is sometimes also referred to as the "command line" specified in the request. In some examples, the parameter description associated with a request may include the identifier (ID) of client 120, a user account, parameter specifications for the task to be performed, etc. Parameters may vary in different types of requests. Some parameters in the same type of request may be set differently.

[0036] As an example only, examples of data backup and data recovery requests are given below. A data backup request can be defined as: 2018-03-27 02:18:45 avtar --id=xxxx --account= / xxxxx / xxxxxx-c / xxxx / --hfsaddr=xxx.xxx.xxx. In this request, "2018-03-27 02:18:45" indicates the time when the request was initiated, and parameter fields such as "avtar --id", "--account", "-c", and "--hfsaddr" are all defined (for illustrative purposes only, the "x" sequence is used to indicate the parameter description). Similarly, a data recovery request can be defined as: 2018-03-28 09:06:21 avtar --id=xxxx --account= / clients / xxxxx-x --hfsaddr=xxx.xxx.xxx --debug --logfile= / xxxx / xxxxx. In this request, some parameter fields "avtar --id" and "--account" are the same as those of the data backup request, while parameter fields "-x --hfsaddr" and "--debug --logfile" may be specific to the data recovery type request.

[0037] It should be understood that the above only provides some examples of request and parameter description thereof. In other examples, the request initiated to the application 112 can be defined in other different forms.

[0038] In some embodiments, server 110 may record requests initiated from various clients 120 as client-specific context requests 202 for subsequent anomaly detection by anomaly detection 200. In some embodiments, if the information 203 to be collected by information collector 210 for anomaly detection of the current target request 102 includes the type and initiation time of the context request 202 (as discussed below), server 110 may record the type and initiation time of the context request (particularly historical requests) 202 without recording the associated parameter description. In some embodiments, if anomaly detection 200 requires parameter descriptions associated with the context request 202, server 110 also records such information. For target request 102, since this request may need to be executed later, upon receiving the target request 102, server 110 also records the parameter descriptions associated with the target request 102 in addition to the type and initiation time.

[0039] In some embodiments, information collection by the information collector 210 may be triggered by the server 110. For example, after receiving the target request 102 from the client 120, the server 110 may initiate a request to the abnormality detection system 200 to detect abnormalities in the target request 102. In response to the request, the information collector 210 may perform information collection for subsequent detection purposes.

[0040] In some embodiments, information collector 210 may also directly collect information related to the target request and / or information related to at least one context request 202 from client 120. For example, client 120 may be required to send target request 102 to abnormality detection system 200. After abnormality detection system 200 determines that target request 102 is a normal request, target request 102 is sent to server 110 for processing. In this case, information collector 210 may collect and record information related to context request 202 of each client 120 for subsequent use.

[0041] After collecting information 203 for a specific target request 102, information collector 210 provides information 203 to information processor 220. Information processor 220 is configured to further process information 203 to extract detection information suitable for use by anomaly detector 230. Specifically, information processor 220 is configured to convert the acquired information 203 into a vectorized feature representation 222 for target request 102. Vectorized feature representation 222 can be considered as a vector including multiple dimensional values ​​to characterize characteristics of information 203.

[0042] In some embodiments, the information processor 220 is configured to map the information 203 into corresponding vectorized feature representations 222 based on a predetermined mapping rule. The predetermined mapping rule may indicate how to perform feature engineering extraction on the information 203.

[0043] As described above, information 203 includes various forms of information, such as initiation time, type, and possible parameter descriptions. These forms of information are each converted into various parts in vectorized feature representation 222. In some examples, vectorized feature representation 222 may include: a context part corresponding to at least one context request 202, including various parts converted from the type and initiation time of context request 202; and a target part corresponding to target request 102, including various parts converted from the type, initiation time, and parameter description of target request 102. In some embodiments, the values ​​corresponding to these parts may be arranged in a predetermined order to form vectorized feature representation 222.

[0044] To better characterize information 203, the predetermined mapping rules may include mapping rules corresponding to different information formats within information 203. In some embodiments, for the types of target request 102 and context request 202, the predetermined mapping rules may include a first mapping of multiple request types to a first plurality of predetermined values. In other words, the first mapping maps multiple potential request types for application 112 to different predetermined values. The predetermined values ​​may be in numerical form to facilitate subsequent processing.

[0045] Figure 3 An example of a first mapping 300 of a plurality of request types to a first plurality of predetermined values ​​is shown. For ease of understanding, the first mapping 300 can be represented as a tree structure, with a root node 302 indicating that the first mapping 300 is for the application 112. Requests that can be initiated to the application 112 are divided into a plurality of request types, each of which is mapped to a predetermined value. Figure 3 In the example, node 310 indicates that the data backup request of application 112 is mapped to the value "1", node 320 indicates that the data recovery request is mapped to the value "2", node 330 indicates that the data maintenance request is mapped to the value "3", and so on.

[0046] In some embodiments, a certain request type can be further divided into multiple refined request types. The first mapping 300 can also indicate the mapping between these refined request types and corresponding predetermined values. Figure 3As shown, data backup requests can be further categorized into full backup requests, incremental backup requests, and so on. Thus, child node 311 following node 310 indicates that a full backup request is mapped to the value "1.1," and child node 312 indicates that a full backup request is mapped to the value "1.2." Similarly, data restore requests can sometimes be categorized into requests to restore to a new version, requests to restore to the original version, and so on. Thus, child node 321 following node 320 indicates that a restore to a new version request is mapped to the value "2.1," and node 322 indicates that a restore to the original version request is mapped to the value "2.2."

[0047] In some embodiments, when defining the first mapping, the predetermined values ​​corresponding to the refined request types under the same major type can be configured to be closer to each other (compared to other predetermined values ​​of the refined request types under another major type). For example, the predetermined values ​​(1.1 and 1.2) corresponding to the full backup request and the incremental backup request under the data backup request are closer, while the predetermined values ​​of the two requests under the data recovery request are closer to each other (compared to the predetermined values ​​1.1 and 1.2).

[0048] It should be understood that Figure 3 Only one example mapping between request types and predetermined values ​​is shown. Figure 3 The predetermined values ​​shown in the example are only examples and do not imply any limitation to the embodiments of the present disclosure. In other examples, any other predetermined values ​​may be set for mapping to different request types of the application 112. In some embodiments, in addition to Figure 3 In addition to using the tree structure shown to represent the first mapping, other forms, such as a list form, may be used to represent the mapping between the request type of the application 112 and the predetermined value.

[0049] Based on the first mapping between the request type and the predetermined value, the information processor 220 can map the types of the target request 102 and the one or more context requests 202 indicated in the information 203 to corresponding predetermined values. As an example, assume that the information 203 indicates that the target request 102 is a full backup request, and the two historical requests of the target 102 are a restore to original version request and a restore to new version request. Figure 3 In the first mapping shown in the example of , the information processor 220 may determine that the full backup request corresponds to the predetermined value "1.1," while the restore to original version request and the restore to new version request correspond to the predetermined values ​​"2.2" and "2.1," respectively.

[0050] The partially processed information 203 can be represented as follows:

[0051] (2.2,[2018-03-27 02:18:45]+2.1,[2018-03-28 09:06:21]+1.1,[2018-03-2912:06:21]+[id:xxxx,account: / xxxx / xxxx,c: / xxx / ,hfsaddr:xxxx.xxx.xxx])

[0052] Among them, "2.2", "2.1" and "1.1" represent the predetermined values ​​corresponding to the first historical request 202, the second historical request 202 and the target request 102, respectively; [2018-03-27 02:18:45], [2018-03-28 09:06:21] and [2018-03-2912:06:21] represent the initiation times corresponding to the first historical request 202, the second historical request 202 and the target request 102, respectively; and [id:xxxx,account: / xxxx / xxxx,c: / xxx / ,hfsaddr:xxxx.xxx.xxx] represents the parameter description of the target request.

[0053] In some embodiments, for the initiation time of the target request 102 and the context request 202, the predetermined mapping rule may include a second mapping of multiple time intervals to a second plurality of predetermined values. The multiple time intervals may be divided from the request time period of the client 120 to the application 112. In some cases, different clients 120 have certain request patterns, and such request patterns can be reflected in the timing of the requests. Therefore, the initiation time of the request helps to detect whether the target request 102 is abnormal. In order to better characterize the time characteristics, a request time period can be divided into multiple time intervals for converting the specific initiation time of the target request 102 and the context request 202. The request time period can be set according to the client and application, for example, it can be set to 1 day, 1 week, 1 month, etc.

[0054] As an example, a 24-hour request time period can be divided into a predetermined number of time intervals, such as 48 time intervals, each of which is half an hour long. For example, 00:00-00:30 is one time interval, 00:30-01:00 is another time interval, and so on, up to the time interval of 23:30-00:00. Each time interval can be mapped to a different preset value, such as a value from 1 to 48. In other examples, the length of the time interval is set to other values, or the request time period can also be other values.

[0055] When performing information processing, the information processor 220 can determine the time interval in which the execution time of the target request 102 and the context request 202 respectively falls, and determine the predetermined value corresponding to the time interval in which the request falls based on a second mapping between the time interval and the second predetermined value.

[0056] If only a one-day request time period is considered, the initiation time of the target request 102 and the context request 202 in the information 203 may be considered at a specific time within the day, not the date. In the above example, for the initiation time [2018-03-27 02:18:45] of the first historical request 202, it can be determined that the time "02:18:45" falls within the time interval 02:00-02:30; for the initiation time [2018-03-28 09:06:21] of the second historical request 202, it can be determined that the time "09:06:21" falls within the time interval 09:00-09:30; and for the initiation time [2018-03-29 12:06:21] of the target request 102, it can be determined that the time "12:06:21" falls within the time interval 12:00-12:30. Accordingly, the information processor 220 may determine predetermined values ​​corresponding to these time intervals, such as predetermined values ​​5, 19, and 25.

[0057] If the initiation time is processed, the partially processed information 203 can be represented as follows:

[0058] (2.2,5+2.1,19+1.1,25+[id:xxxx,account: / xxxx / xxxx,c: / xxx / ,hfsaddr:xxxx.xxx.xxx])

[0059] It should be understood that the above only gives an example of converting the initiation time of the request into a part of the vectorized feature representation. The initiation time can also be converted in other ways, as long as different numerical values ​​can be used to indicate different initiation times. For example, no pre-mapping can be set, but the initiation time can be directly represented as a numerical value in a predetermined format. For example, the initiation time [2018-03-27 02:18:45] of the first historical request 202 can be represented as the numerical value 20180327021845, or if the date is not considered and only the time within a day is considered, it can be represented as 021845, and so on. Of course, by dividing the time interval, the numerical representation dimension of the initiation time can be better reduced.

[0060] In some embodiments, the number of context requests 202 used for abnormality monitoring of the target request 102 can be a predetermined number. If the predetermined number of context requests for the target request 102 cannot be collected in some cases, the information processor 220 can also set the other context parts in the vectorized feature representation, except for the context parts corresponding to the available context requests, to preset values. For example, if the required number of context requests 202 is 2 historical requests, but the client 120 only issues one historical request before initiating the target request 102, then the type and initiation time of the other historical request are both determined to be a predetermined value of 0 in the vectorized feature representation 222. In some embodiments, if the client 120 does not initiate any historical request before initiating the target request 102, then the type and initiation time of the two historical requests are both determined to be a predetermined value of "0" in the vectorized feature representation 222.

[0061] By determining the unavailable context request as a predetermined value, the context persistence of the target request 102 initiated on the client can be well characterized. This can also be used to distinguish normal requests from abnormal requests. For example, if the attacker steals the token of the legitimate client 120, he may use the token to directly initiate a backup request from the illegal client 120. However, from the context of the request of the illegal client 120, there is no historical request for the currently initiated backup request, that is, the context portion corresponding to the historical request in the vectorized feature representation 222 is represented as 0. This may imply the abnormality of the current target request, because the legitimate client usually starts the formal data protection task after requests such as setup and registration.

[0062] In order to obtain the vectorized feature representation 222, the information processor 220 may also continue to process the parameter description associated with the target request 102 included in the information 203. In some embodiments, the information processor 220 converts the parameter description associated with the target request 102 into a one-hot encoding representation. The parameter description is helpful in determining whether the parameters in the request are all valid parameters during abnormality detection. Because for malicious requests from the attacker (which should be considered abnormal requests), the parameter descriptions therein may have some abnormal characteristics. Therefore, by performing one-hot encoding on the parameter descriptions, the difference between normal parameters and abnormal parameters can be reflected more quickly.

[0063] The one-hot encoding representation may include numerical values ​​for predetermined dimensions, with each dimension's numerical value corresponding to a valid symbol. If the parameter description includes the valid symbol, the numerical value for the dimension may be represented as 1; if the valid symbol is not present, the numerical value for the dimension may be represented as 0. In some embodiments, the parameter description may include settings for multiple parameters (such as user accounts, request parameters, etc.), each of which may be mapped to a one-hot encoding representation, and the one-hot encoding representations of the multiple parameters may be combined into an overall one-hot encoding representation of the parameter description associated with the target request 102.

[0064] Figure 4 An example of a one-hot encoding representation 400 is shown. In this example, it is assumed that the symbols of the parameter description may include a set of 26 letters 410, a set of 10 numbers 420, and a special symbol 430 (such as *, !, <, etc.). Parameters 1 to n in the parameter description can each be represented as a string of 0s and 1s, with the value 1 or 0 of the corresponding bit indicating whether the character at that position exists or not. Figure 4 Although the example shows the bit corresponding to one special symbol, in some embodiments, the presence or absence of different special symbols can be indicated by their respective corresponding bits.

[0065] The parameter description associated with abnormal request 440 may contain some anomalies, causing the values ​​of certain bits in its one-hot encoding to differ from those of normal parameters. For example, an attacker may forge the user account "admin=>", and the one-hot encoding corresponding to this parameter may indicate the presence of the unusual symbols "=" and ">". This helps to subsequently identify such abnormal requests.

[0066] In some embodiments, in addition to one-hot encoding, other encoding techniques, particularly techniques suitable for text or character encoding, may be used to convert the parameter description into a multi-dimensional vector representation.

[0067] In some embodiments, after the parameter description is also processed, the processed information 203 may be represented as follows:

[0068] (2.2,5+2.1,19+1.1,25+10..01000010..010101...010010..010101...01) where "10..01000010..010101...010010..010101...01" is the one-hot encoding of the parameter description associated with the target request.

[0069] In some embodiments, the processed information belongs to a vectorized representation and can be directly determined as a vectorized feature representation 222. In some embodiments, the information processor 220 can also perform normalization processing on the converted information to obtain a vectorized feature representation 222. After normalization, numerical values ​​with different value ranges can be normalized to the same value interval (for example, a value interval of 0 to 1). For example, in the above processing, the range of values ​​to which the types of requests are mapped is different from the range of values ​​corresponding to the initiation time. In addition, the one-hot encoding representation is represented as a binary sequence string. These different value ranges can be unified through normalization.

[0070] In some embodiments, the information processor 220 may utilize various data normalization methods to perform normalization processing. Data normalization methods may include, for example, the minimum-maximum method, standard scoring, standardization, etc. As an example only, data normalization based on the minimum-maximum method may be expressed as:

[0071]

[0072] Where X represents a value of the same type, such as a value corresponding to the type of the target request 102 and the context request 202 (e.g., 2.2, 2.1, or 1.1), a value corresponding to the initiation time (e.g., 5, 19, or 25), or a value corresponding to the parameter description (i.e., a one-hot encoding representation). X' represents a normalized value. Xmax and Xmin represent the maximum and minimum values ​​under this type of value, such as the maximum and minimum values ​​mapped to the request type in the first mapping, the maximum and minimum values ​​mapped to the time interval in the second mapping, and the maximum value (e.g., a sequence of all ones) and minimum value (e.g., a sequence of all zeros) represented by the one-hot encoding.

[0073] Through normalization, the values ​​in the vectorized representation obtained after numerical mapping and one-hot encoding can be processed into values ​​within a uniform range, such as the range of 0 to 1. This helps in subsequent anomaly detection.

[0074] The vectorized feature representation 222 generated by the information processor 220 is provided to the abnormality detector 230. In an embodiment of the present disclosure, the abnormality detector 230 is configured to use the abnormality detection model 232 to determine an abnormality detection result of the target request 102 based on the vectorized feature representation 222 to indicate whether the target request 102 is a normal request or an abnormal request.

[0075] Anomaly detection model 232 is a trained model that can characterize the correlation between the vectorized feature representation of a request (i.e., the model's input) and the anomaly detection result (i.e., the model's output). The output of anomaly detection model 232 is a classification task, that is, it classifies the input into two categories ("normal request" or "abnormal request"). Anomaly detection model 232 can be designed as any machine learning model or neural network model.

[0076] In some embodiments, the abnormality detection model 232 can be designed as a support vector machine (SVM) model, such as a single-class SVM. In the scenario of abnormality detection, since the negative samples (i.e., abnormal requests) used for model training are usually small, most of the samples that can be collected are positive samples (i.e., normal requests, also called white samples). The SVM model can support the training of small data samples and can achieve good optimization. Among various SVM models, the single-class SVM can better complete model training based on a large number of positive samples and a small number of negative samples. Therefore, constructing the abnormality detection model 232 based on the SVM model, especially the single-class SMV model, is very conducive to model optimization.

[0077] The single-class SMV model can be trained to identify the classification principle of data in one category from the training data, and data that does not meet the classification principle is considered to be of another category. Such model learning is also suitable for anomaly detection. Figure 5 FIG. 5 shows an example of learning and applying the single-class SMV-based abnormality detection model 232. Figure 5 As shown, a sample division boundary 510 can be learned from the training data, because most training samples 520 can be clustered together at this boundary. After training is completed, the abnormality detection model 232 can be used to determine requests falling within the boundary 510 as normal requests 530, and to determine requests falling outside the boundary 520 as abnormal requests 540.

[0078] In some embodiments, the abnormality detection model 232 can also be referred to as any other type of model, and the embodiments of the present disclosure are not limited in this respect. In some embodiments, when training the abnormality detection model 232, in order to improve the accuracy of abnormality detection, corresponding models can also be trained for the application policies of different clients of the application 112. For example, if the application 112 is a data protection application, for the clients that can request the application 112, one or more clients 120 may be assigned or subscribe to different data protection policies of the application 112. Due to different data protection policies, the characteristics of the requests issued by the client 120 may be different. Therefore, multiple abnormality detection models 232 are trained through the different data protection policies that the application 112 can provide. When performing abnormality detection, the abnormality detector 230 can select the corresponding abnormality detection model 232 to implement detection for the data protection policy applied by the client that currently issues the target request 102.

[0079] Continue to refer back Figure 2 In some embodiments, after the abnormality detector 230 determines whether the target request 102 is a normal request or an abnormal request, it can send an indication to the server 110 to indicate the abnormality detection result of the target request 102.

[0080] For example, if abnormality detector 230 determines that target request 102 is a normal request, it sends an indication 232 to server 110, indicating that target request 102 is a "normal request." If abnormality detector 230 determines that target request 102 is an abnormal request, it sends an indication 234 to server 110, indicating that target request 102 is an "abnormal request." Alternatively or additionally, indication 234 may also be sent to other devices, such as client 120, or used to notify a system administrator.

[0081] In some embodiments, the abnormality detection system 200 may further include a secondary verification module 240, such as Figure 2 If abnormality detector 230 determines that target request 102 is an abnormal request, it sends an indication 236 to secondary verification module 240, indicating that target request 102 is an "abnormal request." Secondary verification module 240 may also employ other verification methods to further verify whether target request 102 is an abnormal request. Any other verification method may be employed, and even manual confirmation may be introduced to perform secondary verification.

[0082] If secondary verification module 240 verifies that target request 102 is a normal request, it sends an indication 242 to server 110, indicating that target request 102 is a "verified normal request." Otherwise, if secondary verification module 240 verifies that target request 102 is an abnormal request, it sends an indication 244 to server 110, indicating that target request 102 is a "verified abnormal request." Indication 244 may be sent to server 110 and / or to other devices, such as client 120, or used to notify a system administrator.

[0083] If the server 110 confirms that the target request 102 is a normal request through the indication 232 or the indication 242, the request may be processed normally. Otherwise, if the target request 102 is confirmed to be an abnormal request, the server 110 may refuse to process the request.

[0084] Figure 6 FIG. 6 is a flow chart showing a process 600 for abnormality detection according to some embodiments of the present disclosure. The process 600 may be performed by Figure 1 For ease of discussion, the following will refer to Figure 2 6. The process 600 is described below.

[0085] At block 610, the abnormality detection system 200 obtains information related to a target request and at least one context request initiated by a client to an application. The information indicates at least the type and initiation time of the target request, and the type and initiation time of at least one context request. At block 620, the abnormality detection system 200 converts the obtained information into a vectorized feature representation of the target request. At block 630, the abnormality detection system 200 utilizes an abnormality detection model to determine an abnormality detection result for the target request based on the vectorized feature representation. The abnormality detection result indicates whether the target request is an abnormal request. The abnormality detection model characterizes the correlation between the vectorized feature representation of the request and the abnormality detection result.

[0086] In some embodiments, converting the acquired information into a vectorized feature representation includes: determining, from the plurality of predetermined values, predetermined values ​​corresponding to the type of the target request and the type of at least one context request, respectively, based on a first mapping of the plurality of request types to the first plurality of predetermined values; and determining a portion of the vectorized feature representation based on the determined predetermined values.

[0087] In some embodiments, converting the acquired information into a vectorized feature representation includes: determining, from a plurality of time intervals, a first time interval into which the initiation time of the target request falls and at least one second time interval into which the initiation time of at least one context request falls, the plurality of time intervals being divided from a time period of a request of the client to the application; based on a second mapping of the plurality of time intervals to a second plurality of predetermined values, determining, from the second plurality of predetermined values, predetermined values ​​corresponding to each of the first time interval and the at least one second time interval; and determining another part of the vectorized feature representation based on the determined predetermined values.

[0088] In some embodiments, the information further indicates a parameter description associated with the target request. In some embodiments, converting the obtained information into a vectorized feature representation includes: converting the parameter description into a one-hot encoded representation; and determining another portion of the vectorized feature representation based on the one-hot encoded representation.

[0089] In some embodiments, the vectorized feature representation is configured to include context portions corresponding to a predetermined number of context requests. In some embodiments, converting the acquired information into the vectorized feature representation includes: if it is determined that the number of at least one context request indicated by the information is less than the predetermined number, setting the context portions in the vectorized feature representation other than the context portion corresponding to the at least one context request to a preset value.

[0090] In some embodiments, the application includes a data protection application, and process 600 further includes selecting an anomaly detection model based on a data protection policy applied by the client in the data protection application, the anomaly detection model being trained based on training data related to the data protection policy.

[0091] In some embodiments, the at least one context request includes at least one historical request whose initiation time is before the initiation time of the target request.

[0092] Figure 7 Schematically shows a block diagram of a device 700 that can be used to implement an embodiment of the present disclosure. It should be understood that Figure 7 The device 700 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 7 The device 700 shown can be used to implement Figure 6 Process 600. Figure 7 The device 700 shown may be implemented as or included in Figure 2 Abnormality detection system 200, or a part of abnormality detection system 200.

[0093] like Figure 7As shown in FIG, device 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory device (ROM) 702 or loaded from a storage unit 708 into a random access memory device (RAM) 703. Various programs and data required for the operation of device 700 can also be stored in RAM 703. CPU 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to bus 704.

[0094] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0095] The various processes and processing described above, such as process 600, may be performed by processing unit 701. For example, in some embodiments, process 600 may be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed onto device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by CPU 701, one or more steps of process 600 described above may be performed.

[0096] Embodiments of the present disclosure may also provide a computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, wherein the computer-executable instructions are executed by a processor to implement the method described above.

[0097] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, computer-readable media, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0098] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0099] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0100] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0101] As used herein, the term "determine" encompasses a wide variety of actions. For example, "determine" may include computing, calculating, processing, deriving, investigating, searching (e.g., searching in a table, database, or another data structure), ascertaining, etc. Furthermore, "determine" may include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), etc. Furthermore, "determine" may include resolving, selecting, choosing, establishing, etc.

[0102] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for abnormality detection, comprising: Obtaining information related to a target request and at least one context request initiated by a client to an application, the information indicating at least a type and initiation time of the target request and a type and initiation time of the at least one context request; Converting the acquired information into a vectorized feature representation for the target request; as well as An abnormality detection model is used to determine an abnormality detection result of the target request based on the vectorized feature representation, wherein the abnormality detection result indicates whether the target request is an abnormal request, and the abnormality detection model characterizes the correlation between the vectorized feature representation of the request and the abnormality detection result.

2. The method according to claim 1, wherein converting the acquired information into the vectorized feature representation comprises: determining, based on a first mapping of a plurality of request types to a first plurality of predetermined values, predetermined values ​​corresponding to the type of the target request and the type of the at least one context request, respectively, from the plurality of predetermined values; as well as A portion of the vectorized feature representation is determined based on the determined predetermined value.

3. The method according to claim 1, wherein converting the acquired information into the vectorized feature representation comprises: Determining, from a plurality of time intervals, a first time interval into which the initiation time of the target request falls and at least one second time interval into which the initiation time of the at least one context request falls, the plurality of time intervals being divided from a time period of a request from the client to the application; determining, based on a second mapping of the plurality of time intervals to a second plurality of predetermined values, predetermined values ​​corresponding to each of the first time interval and the at least one second time interval from the second plurality of predetermined values; and Another portion of the vectorized feature representation is determined based on the determined predetermined value.

4. The method of claim 1 , wherein the information further indicates a parameter description associated with the target request, and converting the obtained information into the vectorized feature representation comprises: Convert the parameter description into a one-hot encoding representation; as well as Another portion of the vectorized feature representation is determined based on the one-hot encoded representation.

5. The method according to claim 1 , wherein the vectorized feature representation is configured to include context portions corresponding to a predetermined number of context requests, and converting the acquired information into the vectorized feature representation comprises: If it is determined that the number of the at least one context request indicated by the information is less than the predetermined number, other context parts in the vectorized feature representation except the context part corresponding to the at least one context request are set to preset values.

6. The method of claim 1 , wherein the application comprises a data protection application, the method further comprising: The abnormality detection model is selected based on a data protection policy applied by the client in the data protection application, and the abnormality detection model is trained based on training data related to the data protection policy. The method according to claim 1 , wherein the at least one context request comprises at least one historical request whose initiation time is before the initiation time of the target request.

8. An electronic device comprising: at least one processor; as well as at least one memory storing computer program instructions, the computer program instructions, when executed by the at least one processor, causing the electronic device to perform actions, the actions comprising: Obtaining information related to a target request and at least one context request initiated by a client to an application, the information indicating at least a type and initiation time of the target request and a type and initiation time of the at least one context request; Converting the acquired information into a vectorized feature representation for the target request; and An abnormality detection model is used to determine an abnormality detection result of the target request based on the vectorized feature representation, wherein the abnormality detection result indicates whether the target request is an abnormal request, and the abnormality detection model characterizes the correlation between the vectorized feature representation of the request and the abnormality detection result.

9. The apparatus according to claim 8, wherein converting the acquired information into the vectorized feature representation comprises: determining, based on a first mapping of a plurality of request types to a first plurality of predetermined values, predetermined values ​​corresponding to the type of the target request and the type of the at least one context request, respectively, from the plurality of predetermined values; as well as A portion of the vectorized feature representation is determined based on the determined predetermined value.

10. The apparatus according to claim 8, wherein converting the acquired information into the vectorized feature representation comprises: Determining, from a plurality of time intervals, a first time interval into which the initiation time of the target request falls and at least one second time interval into which the initiation time of the at least one context request falls, the plurality of time intervals being divided from a time period of a request from the client to the application; determining, based on a second mapping of the plurality of time intervals to a second plurality of predetermined values, predetermined values ​​corresponding to each of the first time interval and the at least one second time interval from the second plurality of predetermined values; and Another portion of the vectorized feature representation is determined based on the determined predetermined value.

11. The apparatus of claim 8, wherein the information further indicates a parameter description associated with the target request, and converting the acquired information into the vectorized feature representation comprises: Convert the parameter description into a one-hot encoding representation; as well as Another portion of the vectorized feature representation is determined based on the one-hot encoded representation.

12. The apparatus according to claim 8, wherein the vectorized feature representation is configured to include context portions corresponding to a predetermined number of context requests, and converting the acquired information into the vectorized feature representation comprises: If it is determined that the number of the at least one context request indicated by the information is less than the predetermined number, other context parts in the vectorized feature representation except the context part corresponding to the at least one context request are set to preset values.

13. The device of claim 8, wherein the application comprises a data protection application, the actions further comprising: The abnormality detection model is selected based on a data protection policy applied by the client in the data protection application, and the abnormality detection model is trained based on training data related to the data protection policy. 14 . The apparatus according to claim 8 , wherein the at least one context request comprises at least one historical request whose initiation time is before the initiation time of the target request.

15. A computer program product tangibly stored on a non-transitory computer-readable medium and comprising computer-executable instructions that, when executed by a processor, cause the processor to perform actions comprising: Obtaining information related to a target request and at least one context request initiated by a client to an application, the information indicating at least a type and initiation time of the target request and a type and initiation time of the at least one context request; Converting the acquired information into a vectorized feature representation for the target request; as well as An abnormality detection model is used to determine an abnormality detection result of the target request based on the vectorized feature representation, wherein the abnormality detection result indicates whether the target request is an abnormal request, and the abnormality detection model characterizes the correlation between the vectorized feature representation of the request and the abnormality detection result.

16. The computer program product of claim 15, wherein converting the acquired information into the vectorized feature representation comprises: determining, based on a first mapping of a plurality of request types to a first plurality of predetermined values, predetermined values ​​corresponding to the type of the target request and the type of the at least one context request, respectively, from the plurality of predetermined values; as well as A portion of the vectorized feature representation is determined based on the determined predetermined value.

17. The computer program product of claim 15, wherein converting the acquired information into the vectorized feature representation comprises: Determining, from a plurality of time intervals, a first time interval into which the initiation time of the target request falls and at least one second time interval into which the initiation time of the at least one context request falls, the plurality of time intervals being divided from a time period of a request from the client to the application; determining, based on a second mapping of the plurality of time intervals to a second plurality of predetermined values, predetermined values ​​corresponding to each of the first time interval and the at least one second time interval from the second plurality of predetermined values; and Another portion of the vectorized feature representation is determined based on the determined predetermined value.

18. The computer program product of claim 15, wherein the information further indicates a parameter description associated with the target request, and converting the obtained information into the vectorized feature representation comprises: Convert the parameter description into a one-hot encoding representation; as well as Another portion of the vectorized feature representation is determined based on the one-hot encoded representation.

19. The computer program product of claim 15, wherein the vectorized feature representation is configured to include context portions corresponding to a predetermined number of context requests, and converting the acquired information into the vectorized feature representation comprises: If it is determined that the number of the at least one context request indicated by the information is less than the predetermined number, other context parts in the vectorized feature representation except the context part corresponding to the at least one context request are set to preset values.

20. The computer program product of claim 15, wherein the application comprises a data protection application, the actions further comprising: The abnormality detection model is selected based on a data protection policy applied by the client in the data protection application, and the abnormality detection model is trained based on training data related to the data protection policy.

21. The computer program product of claim 15, wherein the at least one context request comprises at least one historical request whose initiation time precedes the initiation time of the target request.

Citation Information

Patent Citations

  • Ipfix-based detection of amplification attacks on databases

    CN110313161A

  • Unsupervised anomaly detection system and method

    CN111277603A