Method, apparatus and device for identifying network attack
By decrypting API requests, removing redundant data, semantic processing and digital sequence conversion, combined with the attack recognition model, the accuracy and coverage problems of API injection attack detection in the prior art are solved, and higher recognition accuracy and lower false alarm rate are achieved.
Patent Information
- Application Number
- PCT/CN2024/133639
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2024-11-21
- Publication Date
- 2025-06-26
AI Technical Summary
Existing API injection attack detection solutions cannot effectively detect and defend against all types of API injection attacks, especially unknown attack types, and are less accurate when distinguishing between legitimate API requests and illegal API requests, and have problems with false positives and missed reports.
By decrypting and redundant data removal of the received API request data, converting it into semantic text data, and converting it into a sequence of numbers, the attack identification model is entered to determine the legitimacy of the API request.
It improves the accuracy of legality identification of different types of API requests, reduces the probability of false positives, and can more effectively identify potential illegal API requests.
Smart Images

Figure CN2024133639_26062025_PF_FP_ABST
Abstract
Description
Method, device and equipment for identifying network attacks
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of the People's Republic of China on December 20, 2023, with application number 202311767799.2 and invention name "A method, device and equipment for identifying network attacks", the entire contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of network security technology, and in particular to a method, apparatus, and device for identifying network attacks. Background Art
[0004] In the field of network security, there are various types of Application Programming Interface (API) injection attacks, including Structured Query Language (SQL) injection, Cross-Site Scripting (XSS), and Operating System (OS) command injection. These attacks all share a common characteristic: attackers disguise illegal API requests as legitimate ones through malicious code, tricking applications into performing improper operations and bypassing normal security mechanisms. A successful attack can lead to potentially catastrophic consequences, such as database theft, confidential data leakage, and application crashes.
[0005] Existing API injection attack detection solutions usually detect and defend against a single attack method and cannot effectively detect and defend against all types of API injection attacks, especially unknown attack types. In addition, existing API injection attack detection methods have low accuracy in distinguishing between legitimate and illegal API requests, and have high rates of false positives and false negatives. They also perform poorly when processing special symbols in API requests. Attackers can manipulate request parameters by encoding special symbols to bypass detection, resulting in the inability to effectively identify potential illegal API requests. Summary of the Invention
[0006] To address the above issues, the present application provides a method, apparatus, and device for identifying network attacks, which distinguish between legitimate API requests and illegal API requests based on semantically processed API requests, thereby improving the accuracy of detecting different types of illegal API requests.
[0007] In a first aspect, the present application provides a method for identifying network attacks, the method comprising:
[0008] Decrypting the received application program interface (API) request data to obtain original data, and removing redundant data irrelevant to attack detection from the original data to obtain original feature data related to attack detection in the API request;
[0009] Based on the preset correspondence between the symbol and the identification text, the symbol in the original feature data is replaced with the corresponding identification text to obtain text data corresponding to the API request;
[0010] Based on a preset correspondence between text data and digital sequences, the text data is converted into a digital sequence to obtain a digital sequence corresponding to the API request;
[0011] The digital sequence is input into an attack recognition model, and whether the API request is legal is determined based on a recognition result output by the attack recognition model.
[0012] In one or more embodiments, removing redundant data irrelevant to attack detection from the original data includes:
[0013] Determining parameters irrelevant to attack detection based on a set data screening rule, and removing the parameters irrelevant to attack detection and parameter values corresponding to the parameters from the original data; and
[0014] The numerical parameter values of the parameters related to attack detection in the original data are removed.
[0015] In one or more embodiments, after obtaining the raw feature data related to attack detection in the API request and before converting the symbols of the raw feature data into a textual expression, the method further includes:
[0016] If the API request belongs to a file request type, the file type suffix and file size of the API request are determined, and the determined file type suffix and file size are added to the original feature data.
[0017] In one or more embodiments, the number sequence includes a first number sequence corresponding to a word and a second number sequence corresponding to a letter;
[0018] The converting of the text data into a digital sequence based on a preset correspondence between the text data and the digital sequence to obtain a digital sequence corresponding to the API request includes:
[0019] Determining, based on a correspondence between words in the text data and the first numerical sequence, the first numerical sequence corresponding to the words in the text data, and replacing the words in the text data with the corresponding first numerical sequence to obtain a word encoding sequence corresponding to the API request;
[0020] Determining, based on a correspondence between letters in the text data and a second numeric sequence, the second numeric sequence corresponding to the letters in the text data, and converting the letters in the text data into the corresponding second numeric sequence to obtain a letter encoding sequence corresponding to the API request;
[0021] Based on the word encoding sequence and the letter encoding sequence, a numeric sequence corresponding to the API request is obtained.
[0022] In one or more embodiments, the attack identification model is obtained in the following manner:
[0023] Convert the obtained API request samples into digital sequences corresponding to the API request samples as a training sample set, wherein the API request samples include legal API requests and illegal API requests;
[0024] Constructing a basic attack identification model according to the received model construction instruction, wherein a loss function of the basic attack identification model is a weighted loss function set based on a relative ratio of legitimate API requests to illegitimate API requests;
[0025] The constructed basic attack recognition model is iteratively trained according to the training sample set until the model accuracy reaches a set value to obtain the attack recognition model.
[0026] In one or more embodiments, converting the acquired API request sample into a digital sequence corresponding to the API request sample includes:
[0027] Obtain API request samples and remove redundant data irrelevant to attack detection in the API request sample data to obtain the original feature data of the sample related to attack detection in the API request sample;
[0028] Based on the preset correspondence between symbols and identification texts, the symbols in the sample original feature data are replaced with corresponding identification texts to obtain sample text data corresponding to the API request sample;
[0029] Based on a preset correspondence between text data and digital sequences, the sample text data is converted into a digital sequence to obtain a sample digital sequence corresponding to the API request.
[0030] In one or more embodiments, determining whether the API request is legitimate based on the identification result output by the attack identification model includes:
[0031] Input the digital sequence into an attack identification model and output a probability that the API request is legitimate;
[0032] The API request with a legal probability exceeding a first preset value is regarded as the first legal API request;
[0033] Sort the API requests whose legal probability is between a first preset value and a second preset value by their legal probability from largest to smallest, and select the API requests with the highest set proportion as the second legal API requests, where the first preset value is higher than the second preset value;
[0034] The first legal API request and the second legal API request are identified as legal API requests, and the remaining API requests are identified as illegal API requests.
[0035] In a second aspect, the present application provides a device for identifying network attacks, the device comprising:
[0036] A data screening module is used to decrypt the received application program interface (API) request data to obtain the original data, and remove redundant data irrelevant to attack detection from the original data to obtain the original feature data related to attack detection in the API request;
[0037] A first conversion module is configured to replace the symbols in the original feature data with corresponding identification texts based on a preset correspondence between the symbols and the identification texts, thereby obtaining text data corresponding to the API request;
[0038] A second conversion module is used to convert the text data into a digital sequence based on a preset correspondence between the text data and the digital sequence, so as to obtain a digital sequence corresponding to the API request;
[0039] The attack recognition module is used to input the digital sequence into an attack recognition model and determine whether the API request is legal based on the recognition result output by the attack recognition model.
[0040] In one or more embodiments, the data screening module removes redundant data irrelevant to attack detection from the original data, specifically including:
[0041] Determining parameters irrelevant to attack detection based on a set data screening rule, and removing the parameters irrelevant to attack detection and parameter values corresponding to the parameters from the original data; and
[0042] The numerical parameter values of the parameters related to attack detection in the original data are removed.
[0043] In one or more embodiments, the apparatus further includes a suffix extraction module 305. After obtaining the raw feature data related to attack detection in the API request and before converting the symbols of the raw feature data into a textual expression, the suffix extraction module is specifically configured to:
[0044] If the API request belongs to a file request type, the file type suffix and file size of the API request are determined, and the determined file type suffix and file size are added to the original feature data.
[0045] In one or more embodiments, the number sequence includes a first number sequence corresponding to a word and a second number sequence corresponding to a letter;
[0046] The second conversion module converts the text data into a digital sequence based on a preset correspondence between the text data and the digital sequence to obtain a digital sequence corresponding to the API request, specifically including:
[0047] Determining, based on a correspondence between words in the text data and the first numerical sequence, the first numerical sequence corresponding to the words in the text data, and replacing the words in the text data with the corresponding first numerical sequence to obtain a word encoding sequence corresponding to the API request;
[0048] Determining, based on a correspondence between letters in the text data and a second numeric sequence, the second numeric sequence corresponding to the letters in the text data, and converting the letters in the text data into the corresponding second numeric sequence to obtain a letter encoding sequence corresponding to the API request;
[0049] Based on the word encoding sequence and the letter encoding sequence, a numeric sequence corresponding to the API request is obtained.
[0050] In one or more embodiments, the apparatus further includes a model training module, specifically configured to:
[0051] Convert the obtained API request samples into digital sequences corresponding to the API request samples as a training sample set, wherein the API request samples include legal API requests and illegal API requests;
[0052] Constructing a basic attack identification model according to the received model construction instruction, wherein a loss function of the basic attack identification model is a weighted loss function set based on a relative ratio of legitimate API requests to illegitimate API requests;
[0053] The constructed basic attack recognition model is iteratively trained according to the training sample set until the model accuracy reaches a set value to obtain the attack recognition model.
[0054] In one or more embodiments, the model training module converts the acquired API request sample into a digital sequence corresponding to the API request sample, specifically including:
[0055] Obtain API request samples and remove redundant data irrelevant to attack detection in the API request sample data to obtain the original feature data of the sample related to attack detection in the API request sample;
[0056] Based on the preset correspondence between symbols and identification texts, the symbols in the sample original feature data are replaced with corresponding identification texts to obtain sample text data corresponding to the API request sample;
[0057] Based on a preset correspondence between text data and digital sequences, the sample text data is converted into a digital sequence to obtain a sample digital sequence corresponding to the API request.
[0058] In one or more embodiments, the attack identification module determines whether the API request is legal based on the identification result output by the attack identification model, specifically including:
[0059] Input the digital sequence into an attack identification model and output a probability that the API request is legitimate;
[0060] The API request with a legal probability exceeding a first preset value is regarded as the first legal API request;
[0061] Sort the API requests whose legal probability is between a first preset value and a second preset value by their legal probability from largest to smallest, and select the API requests with the highest set proportion as the second legal API requests, where the first preset value is higher than the second preset value;
[0062] The first legal API request and the second legal API request are identified as legal API requests, and the remaining API requests are identified as illegal API requests.
[0063] In a third aspect, an embodiment of the present application provides a device comprising at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a method for identifying network attacks as described in any one of the items provided in the first aspect of the present application.
[0064] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which, when the instructions in the computer-readable storage medium are executed by the processor of the terminal device, enables the terminal device to execute a method for identifying network attacks as described in any one of the items provided in the first aspect of the present application.
[0065] The technical solutions provided by the embodiments of this application bring at least the following beneficial effects:
[0066] The present application provides a method, apparatus, and device for identifying network attacks. By semantically processing API requests and converting them into text data in a natural language format, attackers can be prevented from controlling special symbols in the requests to bypass detection, while capturing more accurate attack semantics. The text data corresponding to the API request is then converted to numbers at the word level and to numbers at the letter level to obtain a number sequence corresponding to the API request. The number sequence corresponding to the API request is then input into an attack recognition model, enabling the attack recognition model to capture features in the text more meticulously. At the same time, a tolerance interval is set for the recognition results output by the model, thereby accepting legal and valid requests within a certain error range, effectively improving the accuracy of identifying whether different types of API requests are legal and reducing the probability of false alarms. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. Obviously, the drawings introduced below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0068] FIG1 is a flow chart of a method for identifying network attacks provided in an embodiment of the present application;
[0069] FIG2 is a schematic diagram of an attack identification model architecture provided in an embodiment of the present application;
[0070] FIG3 is a schematic diagram of a network attack identification device provided in an embodiment of the present application;
[0071] FIG4 is a schematic diagram of a device for identifying network attacks provided in an embodiment of the present application. DETAILED DESCRIPTION
[0072] To make the purpose, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Among them, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0073] In the field of network security, there are various types of Application Programming Interface (API) injection attacks, including Structured Query Language (SQL) injection, Cross-Site Scripting (XSS), and Operating System (OS) command injection. These attacks all share a common characteristic: attackers disguise illegal API requests as legitimate ones through malicious code, tricking applications into performing improper operations and bypassing normal security mechanisms. A successful attack can lead to potentially catastrophic consequences, such as database theft, confidential data leakage, and application crashes.
[0074] Existing API injection attack detection solutions usually detect and defend against a single attack method and cannot effectively detect and defend against all types of API injection attacks, especially unknown attack types. In addition, existing API injection attack detection methods have low accuracy in distinguishing between legitimate and illegal API requests, and have high rates of false positives and false negatives. They also perform poorly when processing special symbols in API requests. Attackers can manipulate request parameters by encoding special symbols to bypass detection, resulting in the inability to effectively identify potential illegal API requests.
[0075] In view of the above problems, the present application provides a method, apparatus and equipment for identifying network attacks. By semantically processing API requests and converting them into text data in natural language format, it can prevent attackers from controlling special symbols in the requests to bypass detection, while capturing more accurate attack semantics. The text data corresponding to the API request is then converted into words at the digital level and letters at the digital level to obtain the digital sequence corresponding to the API request. The digital sequence corresponding to the API request is then input into an attack recognition model, which enables the attack recognition model to capture the features in the text more carefully. At the same time, a tolerance interval is set for the recognition results output by the model, so that legal and valid requests are accepted within a certain error range, effectively improving the accuracy of identifying whether different types of API requests are legal and reducing the probability of false alarms.
[0076] The embodiments of the present application are described in further detail below with reference to the accompanying drawings.
[0077] As shown in FIG1 , a flowchart of a method for identifying network attacks provided in an embodiment of the present application is provided. The method includes the following steps S101 to S104:
[0078] Step S101: decrypting the received API request data to obtain original data, and removing redundant data irrelevant to attack detection from the original data to obtain original feature data related to attack detection in the API request;
[0079] As a feasible implementation method, decrypting the received application program interface (API) request data to obtain the original data includes:
[0080] Based on the set format unification rules, the format of the received application program interface API request data is unified;
[0081] The API request data in the unified format is decoded at least once until the original data of the API request is obtained.
[0082] The received API request data may be legitimate API requests generated by normal application operations captured by the application monitor. These requests are typically requests sent to the target API by users, other applications, or services. Alternatively, they may be illegitimate API requests containing malicious code generated by attackers using malicious scripts. Specifically, the set formatting rule may require that all characters in each API request be lowercase.
[0083] API requests are typically encoded using Uniform Resource Locator (URL) parameters or code sequences. Certain characters in API request data have special meanings, such as "&" (&) separating query parameters. Furthermore, because URL data can only contain letters from the ASCII set, parameters outside of this set must be encoded using URL encoding techniques.
[0084] In related technologies, browsers and web servers typically only decode API request data once. However, attackers may bypass decoding by performing a double encoding. This application uses a decoding technique that includes at least one decoding operation to decode the API request data until the original data of the API request is obtained. For example, when the decoding method is set to double decoding, the first decoding restores the API request data to the first layer of encoding, and the second decoding restores the original data of the API request data. This prevents attackers from using multiple encodings to deceive the backend processing logic.
[0085] As a feasible implementation manner, removing redundant data irrelevant to attack detection in the original data includes:
[0086] Determine parameters irrelevant to attack detection based on set data screening rules, and remove parameters irrelevant to attack detection and parameter values corresponding to the parameters in the original data; and remove digital parameter values of parameters related to attack detection in the original data.
[0087] Consider that an API request contains many parameters, not all of which are relevant to attack detection. If redundant data is not removed and all parameters are retained, subsequent analysis may become complicated and error-prone, reducing the efficiency and accuracy of the analysis. Therefore, after obtaining the raw data of the API request in this application, it is necessary to remove redundant data that is not relevant to attack detection in the raw data, thereby obtaining the raw feature data in the API request that is relevant to attack detection, which is convenient for subsequent analysis.
[0088] The aforementioned redundant data irrelevant to attack detection includes parameters and parameter values unrelated to attack detection, as well as the numeric parameter values of parameters relevant to attack detection. Parameters whose values need to be removed can be parameters determined based on predefined data filtering rules. Specifically, these can be common URL parameters such as pagenumber and pagesize. Their numeric parameter values need to be removed, while retaining the parameter names pagenumber and pagesize. This is because subsequent attack identification focuses on the parameters pagenumber and pagesize, rather than the numerical values corresponding to the specific parameters.
[0089] Step S102: Based on a preset correspondence between symbols and identification texts, the symbols in the original feature data are replaced with corresponding identification texts to obtain text data corresponding to the API request;
[0090] As described in the above embodiment, after removing redundant data irrelevant to attack detection in the original data corresponding to each API request, the original feature data related to attack detection in each API request will be obtained. The original feature data includes words and symbols related to attack detection, but does not include numerical values corresponding to specific parameter values.
[0091] In some embodiments, after obtaining the original feature data related to attack detection in the API request, before converting the symbols of the original feature data into a textual expression, the method further includes:
[0092] If the API request belongs to a file request type, the file type suffix and file size of the API request are determined, and the determined file type suffix and file size are added to the original feature data.
[0093] For API requests that are file requests (such as requests for image, audio, and video resources), this application needs to extract the file suffix and the requested resource size. For example, if the API request requests the file "2.PNG", the extracted file suffix is PNG, and the requested resource size is 1024KB, then "PNG+size:1024" is added to the original feature data. It should be noted that the numerical parameter value "1024" here is a parameter value related to attack behavior analysis, so this value is retained.
[0094] In an embodiment of the present application, after obtaining the original feature data related to attack detection in each API request, the symbols in the original feature data of each API request are semantically processed based on the correspondence between preset symbols and identification texts, thereby converting the symbols of the original feature data into a textual expression to obtain text data corresponding to each API request, which text data only includes letters (in particular, when the API request is a file request type, the text data includes letters and numbers) and does not include symbols.
[0095] Specifically, the identification text may be a custom word or letter string, and the API request may be converted into text data similar to a natural language format based on the correspondence between preset symbols and the identification text.
[0096] Exemplarily, the timestamp in the original feature data corresponding to each API request is converted into the identification text "timestamp", all click identifiers are converted into the identification text "clicktag", and the resource id is converted into the identification text "resource".
[0097] In an embodiment of the present application, the storage form of the correspondence between the above-mentioned symbols and the identification text can be a data table stored in the form of a dictionary mapping of symbols and identification texts. Through this data table, the symbols in the original feature data corresponding to each API request can be semantically processed and converted into a text expression. Therefore, in addition to realizing the detection of common injection attacks, the present application can also summarize the correspondence between attack-specific symbols and set identification texts, identify attack behaviors other than common injection attacks, and provide comprehensive security detection for API requests.
[0098] The following is an example to illustrate the semantic processing process:
[0099] For example, SQL injection attacks modify the queries that a database engine will execute to retrieve the requested information. For example, the API request "url: / users?" can be used to view user information. If the application does not have any mechanisms to prevent this type of attack, an attacker can modify the API request to obtain more information. For example, an attacker could modify the above API request to an illegal one: " / users?category=navigation'+OR+1=1". After removing the redundant data, the original feature data would be " / users?category=navigation'+OR+=".
[0100] The special symbols " / ", "?", "=", "'", and "+" in the raw feature data of the illegal API request are eliminated in traditional natural language processing applications. However, their presence in attack detection scenarios can help with attack detection. Therefore, we convert special symbols into specific identifier text to represent deeper semantics. For example, the illegal API request above becomes:
[0101] “slash users question category equality navigation tick plus OR plus equality”.
[0102] For another example, an XSS injection attack is when an attacker manipulates a vulnerable website and returns malicious script content to ordinary users to perform malicious operations, such as stealing users' personal information. For each Hypertext Markup Language (HTML) tag in the malicious script returned in an XSS injection attack, it is replaced with the markup text, for example becomes "chevrons p chevrons".
[0103] Step S103: based on a preset correspondence between text data and digital sequences, convert the text data into a digital sequence to obtain a digital sequence corresponding to the API request;
[0104] In some embodiments, before converting the text data into a digital sequence, the method further includes:
[0105] When it is determined that duplicate text data exists in the text data corresponding to the API request, a deduplication operation is performed, and then the text data after the deduplication operation is performed is converted into a digital form.
[0106] It should be noted that after obtaining the text data corresponding to the API request, this application performs a deduplication operation on the text data corresponding to the API request. This is because there may be cases where the original data of some API requests are different, and the text data obtained after processing in the aforementioned steps S101 and S102 are the same.
[0107] As a feasible implementation manner, the digital sequence includes a first digital sequence corresponding to words and a second digital sequence corresponding to letters;
[0108] The converting of the text data into a digital sequence based on a preset correspondence between the text data and the digital sequence to obtain a digital sequence corresponding to the API request includes:
[0109] Determining, based on a correspondence between words in the text data and the first numerical sequence, the first numerical sequence corresponding to the words in the text data, and replacing the words in the text data with the corresponding first numerical sequence to obtain a word encoding sequence corresponding to the API request;
[0110] Determining, based on a correspondence between letters in the text data and a second numeric sequence, the second numeric sequence corresponding to the letters in the text data, and converting the letters in the text data into the corresponding second numeric sequence to obtain a letter encoding sequence corresponding to the API request;
[0111] Based on the word encoding sequence and the letter encoding sequence, a numeric sequence corresponding to the API request is obtained.
[0112] In the embodiment of the present application, in order to convert the text data corresponding to the API request into a digital form that can be input into the attack recognition model, two numerical conversions are performed on the text data corresponding to the API request.
[0113] First, the text data is converted into numerical values at the word level. During this process, based on the correspondence between words and the first digital sequence, the first digital sequence corresponding to each word in the text is determined, and the words in the text data are replaced with the corresponding first digital sequence to obtain the word encoding sequence corresponding to the API request.
[0114] Secondly, the text data is converted into numerical values at the letter level. Based on the correspondence between letters and the second digital sequence, the second digital sequence corresponding to each letter in the text is determined, and the letters in the text data are converted into the corresponding second digital sequence to obtain the letter coding sequence corresponding to the API request.
[0115] Optionally, the storage form of the correspondence between the above-mentioned words and the first digital sequence, and the storage form of the correspondence between letters and the second digital sequence, can be a data table stored in the form of a dictionary mapping, and the data table can be a total data table including the correspondence between all words and the first digital sequence and the correspondence between letters and the second digital sequence, or it can be a common data table including the correspondence between common words (with an occurrence frequency higher than a set value) and the first digital sequence and the correspondence between letters and the second digital sequence.
[0116] Specifically, in the process of converting the text data into a digital sequence based on the preset correspondence between the text data and the digital sequence, you can first check the common data table with a smaller amount of data. When no correspondence is found in the common data table, you can then check the total data table with a larger amount of data to improve the search efficiency.
[0117] After these two encodings, the API request can be converted into a numerical form for input into the model. Specifically, the parameters of each API request input attack identification model include a numerical sequence consisting of a word encoding sequence and a letter encoding sequence.
[0118] Optionally, the digital sequence corresponding to the API request is truncated and converted into a digital sequence of appropriate length before being input into the attack recognition model to facilitate recognition by the attack recognition model.
[0119] As a feasible implementation method, the attack identification model is obtained by using the following steps S201 to S203:
[0120] Step S201: converting the acquired API request samples into digital sequences corresponding to the API request samples as a training sample set, wherein the API request samples include legal API requests and illegal API requests;
[0121] Specifically, the legitimate API requests mentioned above may be API requests generated by normal application operations captured through sources such as historical logs or application monitors. These requests are usually requests sent to the target API by users, other applications, or services.
[0122] The illegal API request may be an API request containing malicious code generated by using security testing tools, vulnerability scanning tools, or scripts to simulate the bad intentions that an attacker may use.
[0123] Optionally, converting the acquired API request sample into a digital sequence corresponding to the API request sample includes:
[0124] Obtain API request samples and remove redundant data irrelevant to attack detection in the API request sample data to obtain the original feature data of the sample related to attack detection in the API request sample;
[0125] Based on the preset correspondence between symbols and identification texts, the symbols in the sample original feature data are replaced with the corresponding identification texts to obtain the sample text data corresponding to the API request sample;
[0126] Based on the preset correspondence between text data and digital sequences, the sample text data is converted into a digital sequence to obtain a sample digital sequence corresponding to the API request.
[0127] It should be noted that the specific implementation process of converting the obtained API request sample into a digital sequence corresponding to the API request sample can be referred to the aforementioned embodiment and will not be repeated here.
[0128] Step S202: constructing a basic attack recognition model according to the received model construction instruction, wherein the loss function of the basic attack recognition model is a weighted loss function set based on the relative proportion of legal API requests and illegal API requests;
[0129] Specifically, referring to Figure 2, which is a schematic diagram of the attack identification model architecture provided in an embodiment of the present application, the attack identification model constructed in the embodiment of the present application includes an embedding layer 10, a spatial loss layer 20, a long short-term memory network (Long Short-Term Memory, LSTM) layer 30 and a fully connected layer 40 in sequence. The role of each layer is introduced below.
[0130] Embedding layer 10: used to convert the digital sequence consisting of word encoding sequence and letter encoding sequence into an embedding vector.
[0131] In the embedding matrix composed of these embedding vectors, each first / second digit sequence in the digit sequence corresponds to a specific embedding vector. By capturing the semantic similarity between words or characters, the representations (i.e., embedding vectors) of semantically similar words or characters (i.e., first / second digit sequences with similar meanings) in the embedding space are as similar as possible, thereby better understanding API requests. In this way, when given a first / second digit sequence, we can understand its semantic similarity with other first / second digit sequences by looking up its corresponding embedding vector.
[0132] Spatial Dropout Layer 20: As the next layer after Embedding Layer 10, it regularizes and reduces the dependencies between elements in the embedding vector, thereby improving the generalization performance of the model.
[0133] LSTM layer 30: Accepts the output of the spatial loss layer 20, performs the necessary calculations to learn features from the sequence of numbers corresponding to the API request, and passes the learned feature vector to the next layer.
[0134] Fully connected layer 40: This layer is used to obtain recognition results. It uses the feature vector output from the LSTM layer 30 and uses the Sigmoid activation function to return the probability of the API request being valid and invalid.
[0135] In addition, in actual applications, the number of samples in the legal API request sample set in the API request sample is usually more than that of illegal API request samples. In the case of class imbalance, ordinary cross-entropy loss may cause the model to be overly biased towards the class with more samples. Therefore, in order to solve this data imbalance problem, the embodiment of the present application assigns weights to legal API request samples and illegal API request samples based on the relative proportion of legal API request samples and illegal API request samples in the weighted loss function.
[0136] Step S203: Iteratively train the constructed basic attack recognition model according to the training sample set until the model accuracy reaches a set value to obtain the attack recognition model.
[0137] In the embodiments of the present application, the attack recognition model aims to predict the probability of an API request being legitimate. By constructing a neural network structure comprising the aforementioned embedding layer 10, spatial loss layer 20, LSTM layer 30, and fully connected layer 40, and using a weighted loss function to address the class imbalance problem, the trained attack recognition model can predict the probability of an API request being legitimate, thereby effectively detecting malicious API requests.
[0138] Step S104: input the digital sequence into an attack recognition model, and determine whether the API request is legal based on the recognition result output by the attack recognition model.
[0139] As a feasible implementation, the determination of whether the API request is legitimate based on the identification result output by the attack identification model includes the following steps S301 to S304:
[0140] S301, input the digital sequence into the attack recognition model and output the legitimacy probability of the API request.
[0141] S302: The API request with a legal probability exceeding a first preset value is regarded as a first legal API request.
[0142] S303, arranging the API requests whose legal probability is between the first preset value and the second preset value from large to small according to the legal probability, and selecting the API request with the highest ranking according to the set proportion as the second legal API request, wherein the first preset value is higher than the second preset value.
[0143] S304: Identify the first legal API request and the second legal API request as legal API requests, and identify the remaining API requests as illegal API requests.
[0144] In an embodiment of the present application, if the attack identification model outputs a 40% probability that an API request is legitimate, then the probability that the request is malicious is 60%. If 40% is less than a first preset value, the API request will be identified as an illegal API request. In fact, this may be a valid request that was misclassified by the attack identification model. In order to reduce the probability of false positives and increase the fault tolerance interval of the attack identification model, the present application adds a set ratio p to classify a set ratio of illegal API requests as legitimate API requests.
[0145] For example, when the ratio p is set to 10%, if the user sends 20 API requests to be identified, among which 10 have a legal probability exceeding a first preset value and are classified as legal API requests, then for the other 10 API requests whose legal probability is between the first preset value and the second preset value, they are arranged from large to small according to the legal probability, and the API requests ranked in the top 10% of the set ratio are also classified as legal API requests, and the remaining API requests are identified as illegal API requests.
[0146] In some embodiments, the network attack identification method provided in the embodiments of the present application can be applied to network attack identification in a variety of scenarios, such as Web application security scenarios and cloud service security scenarios. The method provided in the present application can be used to detect potential API injection attacks in the API requests received by Web applications or the data of cloud API requests. The API request data is processed through steps such as semantic processing and digital sequence conversion. At the same time, based on the attack identification model trained using a weighted loss function and the set fault tolerance interval, the probability of false alarms can be reduced while ensuring the detection of potential attacks, effectively protecting Web applications or cloud services from malicious attacks.
[0147] In some embodiments, this application can also be applied to log analysis tools, using the methods provided in this application to analyze log files generated by applications to detect API injection attacks. Processing log data through steps such as semantic processing and digital sequence conversion helps attack recognition models identify potential attacks, enabling log analysis tools to better discover and report potential security threats.
[0148] Based on the method for network attack identification provided by the embodiment of the present application, by semantically processing API requests and converting them into text data in natural language format, it is possible to prevent attackers from controlling special symbols in the requests to bypass detection, while capturing more accurate attack semantics. Subsequently, the text data corresponding to the API request is converted into numbers at the word level and the letter level respectively to obtain the number sequence corresponding to the API request. The number sequence corresponding to the API request is then input into the attack identification model, which enables the attack identification model to capture the features in the text more carefully. At the same time, a tolerance interval is set for the recognition results output by the model, so that legal and valid requests are accepted within a certain error range, effectively improving the accuracy of identifying whether different types of API requests are legal and reducing the probability of false alarms.
[0149] Based on the same inventive concept, an embodiment of the present application further provides a device for identifying network attacks, as shown in FIG3 , the device comprising:
[0150] The data screening module 301 is configured to decrypt the received API request data to obtain the original data, and remove redundant data irrelevant to attack detection from the original data to obtain the original feature data related to attack detection in the API request;
[0151] A first conversion module 302 is configured to replace the symbols in the original feature data with corresponding identification texts based on a preset correspondence between the symbols and the identification texts, thereby obtaining text data corresponding to the API request;
[0152] A second conversion module 303 is configured to convert the text data into a digital sequence based on a preset correspondence between text data and digital sequences, thereby obtaining a digital sequence corresponding to the API request;
[0153] The attack identification module 304 is configured to input the digital sequence into an attack identification model and determine whether the API request is legal based on an identification result output by the attack identification model.
[0154] In one or more embodiments, the data screening module 301 removes redundant data irrelevant to attack detection from the original data, specifically including:
[0155] Determining parameters irrelevant to attack detection based on a set data screening rule, and removing the parameters irrelevant to attack detection and parameter values corresponding to the parameters from the original data; and
[0156] The numerical parameter values of the parameters related to attack detection in the original data are removed.
[0157] In one or more embodiments, the apparatus further includes a suffix extraction module 305. After obtaining the raw feature data related to attack detection in the API request and before converting the symbols of the raw feature data into a textual expression, the suffix extraction module 305 is specifically configured to:
[0158] If the API request belongs to a file request type, the file type suffix and file size of the API request are determined, and the determined file type suffix and file size are added to the original feature data.
[0159] In one or more embodiments, the number sequence includes a first number sequence corresponding to a word and a second number sequence corresponding to a letter;
[0160] The second conversion module 303 converts the text data into a digital sequence based on a preset correspondence between text data and digital sequences to obtain a digital sequence corresponding to the API request, specifically including:
[0161] Determining, based on a correspondence between words in the text data and the first numerical sequence, the first numerical sequence corresponding to the words in the text data, and replacing the words in the text data with the corresponding first numerical sequence to obtain a word encoding sequence corresponding to the API request;
[0162] Determining, based on a correspondence between letters in the text data and a second numeric sequence, the second numeric sequence corresponding to the letters in the text data, and converting the letters in the text data into the corresponding second numeric sequence to obtain a letter encoding sequence corresponding to the API request;
[0163] Based on the word encoding sequence and the letter encoding sequence, a numeric sequence corresponding to the API request is obtained.
[0164] In one or more embodiments, the apparatus further includes a model training module 306, specifically configured to:
[0165] Convert the obtained API request samples into digital sequences corresponding to the API request samples as a training sample set, wherein the API request samples include legal API requests and illegal API requests;
[0166] Constructing a basic attack identification model according to the received model construction instruction, wherein a loss function of the basic attack identification model is a weighted loss function set based on a relative ratio of legitimate API requests to illegitimate API requests;
[0167] The constructed basic attack recognition model is iteratively trained according to the training sample set until the model accuracy reaches a set value to obtain the attack recognition model.
[0168] In one or more embodiments, the model training module 306 converts the acquired API request sample into a digital sequence corresponding to the API request sample, specifically including:
[0169] Obtain API request samples and remove redundant data irrelevant to attack detection in the API request sample data to obtain the original feature data of the sample related to attack detection in the API request sample;
[0170] Based on the preset correspondence between symbols and identification texts, the symbols in the sample original feature data are replaced with corresponding identification texts to obtain sample text data corresponding to the API request sample;
[0171] Based on a preset correspondence between text data and digital sequences, the sample text data is converted into a digital sequence to obtain a sample digital sequence corresponding to the API request.
[0172] In one or more embodiments, the attack identification module 304 determines whether the API request is legitimate based on the identification result output by the attack identification model, specifically including:
[0173] Input the digital sequence into an attack identification model and output a probability that the API request is legitimate;
[0174] The API request with a legal probability exceeding a first preset value is regarded as the first legal API request;
[0175] Sort the API requests whose legal probability is between a first preset value and a second preset value by their legal probability from largest to smallest, and select the API requests with the highest set proportion as the second legal API requests, where the first preset value is higher than the second preset value;
[0176] The first legal API request and the second legal API request are identified as legal API requests, and the remaining API requests are identified as illegal API requests.
[0177] Based on the same inventive concept, the present application also provides a device 400 for identifying network attacks, as shown in Figure 4, comprising at least one processor 402; and a memory 401 communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for identifying network attacks provided in the above embodiment.
[0178] Memory 401 is used to store programs. Specifically, the program may include program code, which includes computer operating instructions. Memory 401 may be volatile memory, such as random-access memory (RAM); or non-volatile memory, such as flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or a combination of any one or more of the above volatile and non-volatile memories.
[0179] Processor 402 may be a central processing unit (CPU), a network processor (NP), or a combination of a CPU and an NP. It may also be a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0180] An embodiment of the present invention further provides a computer-readable storage medium comprising instructions, which, when executed on a computer, enables the computer to execute the network attack identification method provided in the above embodiment.
[0181] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0182] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.
[0183] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected to achieve the purpose of the present embodiment according to actual needs.
[0184] In addition, the functional modules in the various embodiments of the present application may be integrated into a single processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may be stored in a computer-readable storage medium.
[0185] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0186] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a server, or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a server or a data center that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0187] The above is a detailed introduction to the technical solution provided by the present application. Specific examples are used in the present application to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
[0188] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0189] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each flow and / or box in the flow chart and / or block diagram, as well as the combination of the flow chart and / or box in the flow chart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the function specified in one or more flow charts and / or one or more boxes in the block diagram.
[0190] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0191] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0192] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. A method for identifying network attacks, characterized in that: include: Decrypting the received application program interface (API) request data to obtain original data, and removing redundant data irrelevant to attack detection from the original data to obtain original feature data related to attack detection in the API request; Based on the correspondence between the preset symbol and the identification text, the symbol in the original feature data is replaced with the corresponding identification text to obtain text data corresponding to the API request; Based on a preset correspondence between text data and digital sequences, the text data is converted into a digital sequence to obtain a digital sequence corresponding to the API request; The digital sequence is input into an attack recognition model, and whether the API request is legal is determined based on a recognition result output by the attack recognition model.
2. The method according to claim 1, characterized in that The decrypting of the received application program interface (API) request data to obtain the original data includes: Unify the format of received API request data based on the set format unification rules; The API request data in the unified format is decoded at least once until the original data of the API request is obtained.
3. The method according to claim 1, characterized in that The removing of redundant data irrelevant to attack detection in the original data includes: Determine parameters irrelevant to attack detection based on set data screening rules, and remove parameters irrelevant to attack detection and parameter values corresponding to the parameters in the original data; and The digital parameter values of the parameters related to attack detection in the original data are removed.
4. The method according to claim 1, characterized in that: After obtaining the original feature data related to attack detection in the API request, and before converting the symbols of the original feature data into a textual expression, the method further includes: If the API request belongs to a file request type, the file type suffix and file size of the API request are determined, and the determined file type suffix and file size are added to the original feature data.
5. The method according to claim 1, characterized in that The digital sequence includes a first digital sequence corresponding to a word and a second digital sequence corresponding to a letter; The converting of the text data into a digital sequence based on a preset correspondence between the text data and the digital sequence to obtain a digital sequence corresponding to the API request includes: Based on the correspondence between the words in the text data and the first digital sequence, determine the first digital sequence corresponding to the words in the text data, and replace the words in the text data with the corresponding first digital sequence to obtain a word encoding sequence corresponding to the API request; Based on the correspondence between the letters in the text data and the second digital sequence, determine the second digital sequence corresponding to the letters in the text data, and convert the letters in the text data into the corresponding second digital sequence to obtain the letter code sequence corresponding to the API request; Based on the word encoding sequence and the letter encoding sequence, a digital sequence corresponding to the API request is obtained.
6. The method according to claim 1, characterized in that The attack identification model is obtained in the following way: Convert the obtained API request samples into digital sequences corresponding to the API request samples as training sample sets, wherein the API request samples include legal API requests and illegal API requests; Constructing a basic attack identification model according to the received model construction instruction, wherein a loss function of the basic attack identification model is a weighted loss function set based on a relative ratio of legal API requests to illegal API requests; According to the training sample set, the constructed basic attack recognition model is iteratively trained until the model accuracy reaches a set value to obtain the attack recognition model.
7. The method according to claim 6, characterized in that The step of converting the acquired API request sample into a digital sequence corresponding to the API request sample includes: Obtain API request samples, remove redundant data irrelevant to attack detection in the API request sample data, and obtain original feature data of the samples related to attack detection in the API request sample; Based on the correspondence between the preset symbol and the identification text, the symbol in the original feature data of the sample is replaced with the corresponding identification text to obtain sample text data corresponding to the API request sample; Based on the preset correspondence between text data and digital sequences, the sample text data is converted into a digital sequence to obtain a sample digital sequence corresponding to the API request.
8. The method according to any one of claims 1 to 7, characterized in that: The determination of whether the API request is legal based on the identification result output by the attack identification model includes: Input the digital sequence into an attack identification model and output a probability of the API request being legitimate; The API request whose legal probability exceeds a first preset value is regarded as a first legal API request; The API requests with a legal probability between a first preset value and a second preset value are arranged from large to small according to the legal probability, and the API requests with a set proportion ranking high are selected as the second legal API requests, wherein the first preset value is higher than the second preset value; The first legal API request and the second legal API request are identified as legal API requests, and the remaining API requests are identified as illegal API requests.
9. A network attack identification device, characterized in that: include: A data screening module, used to decrypt received application program interface (API) request data to obtain original data, and remove redundant data irrelevant to attack detection from the original data to obtain original feature data related to attack detection in the API request; A first conversion module, used to replace the symbols in the original feature data with corresponding identification texts based on a preset correspondence between the symbols and the identification texts, so as to obtain text data corresponding to the API request; A second conversion module, configured to convert the text data into a digital sequence based on a preset correspondence between the text data and the digital sequence, to obtain a digital sequence corresponding to the API request; The attack identification module is used to input the digital sequence into the attack identification model, and determine whether the API request is legal based on the identification result output by the attack identification model.
10. A network attack identification device, characterized in that: It comprises at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described in any one of claims 1-8.
Citation Information
Patent Citations
LSTM loop neural network model and network attack identification method based on the model
CN109308494A
API (Application Program Interface) abnormal access detection method and device, equipment and medium
CN117112395A
Network attack identification method, device and equipment
CN117792720A
Method, electronic device and computer program product for detecting abnormal network request
US20210021624A1