Network traffic logic vulnerability analysis method and system
By collecting and analyzing network traffic data in real time, using machine learning models to automatically detect logical vulnerabilities in network traffic, solving the problem of inefficiency of traditional methods, and achieving efficient and accurate analysis of unauthorized and overprivileged vulnerabilities.
Patent Information
- Application Number
- CN202510146544.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-07-18
AI Technical Summary
It is difficult to efficiently explore and analyze logical vulnerabilities in network traffic in traditional methods, especially unauthorized logic vulnerabilities and unauthorized logic vulnerabilities in business systems. It is difficult for existing technology to obtain API request paths and parameter characteristics through front-end pages and JS. pure manual detection is inefficient and the results are inaccurate.
By collecting and preprocessing mirror traffic data in real time, extracting basic API information and building a dictionary, using trained API business, behavioral analysis models and session analysis models, generate logical vulnerability detection tasks, perform logical vulnerability detection and analyze response data, and realize automated detection of unauthorized and overprivileged vulnerabilities.
It improves the efficiency of logical vulnerability mining and analysis, ensures the accuracy of detection results, simplifies the processing of complex API interfaces and diversified interface parameters, and supports data security analysis in different industries and businesses.
Smart Images

Figure CN120342650A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of analysis of logical vulnerabilities in network traffic, and particularly to a method and system for analyzing logical vulnerabilities in network traffic. Background Art
[0002] Currently, new-generation power system and energy Internet construction technologies represented by artificial intelligence, big data, cloud computing, 5G, etc. are being applied with each passing day. The digital economy has become the most dynamic and promising new economic form, and the importance of building security protection for enterprise digital assets has become increasingly prominent.
[0003] Currently, the global cybersecurity situation remains severe, and cybersecurity threat incidents targeting key industries and new technologies and new scenarios occur frequently. More and more security incidents have proven that traditional protection means based on situation awareness, firewalls, and IPSs are ineffective in actual scenarios, and half of the top ten vulnerabilities in the OWASP ranking are logical vulnerabilities. At the same time, with the explosion of ChatGPT, artificial intelligence has received more and more attention from various industries. Therefore, using large models to empower the field of cybersecurity and solve the difficult problem of mining logical vulnerabilities in network traffic has become a topic of concern in the field of cybersecurity.
[0004] A network traffic logical vulnerability, usually simply referred to as a logical vulnerability, refers to an attacker exploiting a defect in business or function design to obtain sensitive information or damage the integrity of the business through legitimate network traffic. Such vulnerabilities utilize the lack of strict logical processing, code problems, or design deficiencies to successfully execute attack operations logically. Currently, the more common logical vulnerabilities in the network are unauthorized logical vulnerabilities in business systems and logical vulnerabilities related to overprivilege.
[0005] Among them, the unauthorized logical vulnerability in the business system refers to the fact that due to logical design defects or improper implementation in the business system, an attacker can access or operate sensitive information or functions in the system without obtaining appropriate authorization. The traditional method usually obtains the API request path and parameter characteristics through the front-end page and JS to achieve the mining of unauthorized logical vulnerabilities in the business system. However, in the digital economy era, data transactions between different industries and different businesses in the same industry emerge frequently. To ensure data security, different protocols and different authentication methods are often used to provide services externally. The API interface has the complexity and diversity of interface parameters and interface dependencies, and the interface changes relatively frequently. It has become particularly difficult to obtain the API request path and parameter characteristics through the front-end page and JS in the traditional way, which is not conducive to the mining and analysis of unauthorized logical vulnerabilities in the business system.
[0006] The logic vulnerability related to unauthorized access refers to the vulnerability that occurs when a user in a system or application can perform operations beyond the scope of their authorization. Currently, the following methods are commonly used in the industry to detect logic vulnerabilities related to unauthorized access: Manually search for interfaces with parameters such as id in the front-end access to the back-end interfaces, and then perform access detection by manually replacing the parameters. The pure manual detection method is inefficient, and it is also impossible to confirm whether the detection results are accurate without knowing the background logic. Therefore, it is not conducive to the detection and analysis of logic vulnerabilities related to unauthorized access. Summary of the Invention
[0007] The object of the present invention is to provide a method and system for analyzing logic vulnerabilities in network traffic to solve at least one of the above problems.
[0008] In a first aspect, the present invention provides a method for analyzing logic vulnerabilities in network traffic, the method comprising: Real-time collect the mirror traffic data of the target network, and preprocess the collected mirror traffic data to obtain the preprocessed mirror traffic data as the original data; the data format of the original data is in json format, and the parameters are url, requestHeader request header, requestBody request body, responseHeader, and responseBody response body respectively; Extract and store the basic information of the APIs in the original data, and extract and store the business system to which the collected mirror traffic data of the target network belongs according to the extracted basic information of the APIs; Extract requestHeader, requestBody, responseHeader, and responseBody from the original data, and extract words from the extracted requestHeader, requestBody, responseHeader, and responseBody of the original data to construct a first dictionary; extract the url from the original data, and extract words from the extracted url of the original data to construct a second dictionary; Extract url, requestHeader, requestBody, and responseBody from the original data, then segment the extracted url, requestHeader, requestBody, and responseBody according to case, and set different weights for each obtained segment to form a feature vector based on the first dictionary and the second dictionary. Then input the feature vector into the trained API business analysis model for calculation to output the business labels of the APIs in the original data, and input the feature vector into the trained API behavior analysis model for calculation to output the behavior labels of the APIs in the original data; Extract the requestHeader data from the original data, input the extracted requestHeader data into the trained session analysis model for calculation, and output the user credentials in the original data; Extract the url, requestBody, and responseBody from the original data. Then, tokenize the extracted url, requestBody, and responseBody by case, and set different weights for each obtained token to form a feature vector based on the first dictionary and the second dictionary. Then, input the formed feature vector into the trained login interface analysis model for calculation, and output whether the API corresponding to the original data is a login interface; The method further includes: When the login interface analysis model outputs that the API corresponding to the original data is a login interface: obtain and record the user credentials, username, and user login status of the login interface from the original data; then, match the user credentials of the login interface obtained with the user credentials output by the session analysis model. If the match is successful, cache the API corresponding to the user credentials; if the match is unsuccessful, do not cache the API corresponding to the user credentials; The method further includes: Real-time statistics of the users and APIs of the stored business systems; Real-time statistics of the API permissions and API parameters of the users in the stored business systems; The method further includes: Generate a logical vulnerability detection task based on the data output by the session analysis model, the API business analysis model, and the API behavior analysis model, as well as the basic information of the stored APIs and the various statistics; Execute the logical vulnerability detection task to detect logical vulnerabilities and obtain detection response data; Perform logical vulnerability analysis of network traffic based on the response data.
[0009] Further, generating a logical vulnerability detection task based on the data output by the session analysis model, the API business analysis model, and the API behavior analysis model, as well as the basic information of the stored APIs and the various statistics, includes: Real-time judge whether the user credentials output by the session analysis model are empty. If not, collect the url, requestHeader request headers, and requestBody request bodies in the original data to generate an unauthorized detection task; if so, do not generate an unauthorized detection task; Based on the data output by the session analysis model, the API business analysis model, and the API behavior analysis model, as well as the basic information of the stored APIs and the various statistics, generate logical vulnerability detection tasks, further including: Regularly determine whether the basic information of the currently stored APIs meets the conditions for generating horizontal privilege escalation tasks. If not, do not generate horizontal privilege escalation tasks. If so, on the one hand, based on the API permissions and API parameters of the users in the currently statistical business system, obtain the data permissions of the users corresponding to the basic information of each currently stored API; on the other hand, according to the API business labels and API behavior labels output by the trained API business analysis model and the trained API behavior analysis model, obtain the data permissions of the users corresponding to the basic information of each currently stored API; then gather the data permissions obtained from the above two aspects to obtain all the data permissions of the users corresponding to the basic information of each currently stored API; then, according to the differences in data permissions among different users corresponding to the basic information of the currently stored APIs, generate horizontal privilege escalation tasks for each API corresponding to the basic information of the currently stored APIs. Regularly cluster the users in the currently statistical business system according to the API permissions of the users in the currently statistical business system to obtain several user clusters; obtain the API permissions of the users within each user cluster to obtain the API permissions corresponding to each user cluster; regard each user cluster as a role to obtain the role corresponding to each user cluster, and then obtain the API permissions corresponding to each role; according to the differences in API permissions among each role, generate vertical privilege escalation tasks for each role.
[0010] Furthermore, according to the differences in data permissions among different users corresponding to the basic information of the currently stored APIs, generate horizontal privilege escalation tasks for each API corresponding to the basic information of the currently stored APIs. The specific implementation steps are as follows: For each API corresponding to the basic information of each currently stored API, perform the following steps: Randomly select two users from the API permissions of the users in the currently statistical business system who have the same API permissions as the users corresponding to the basic information of the current API, and record them as the first user and the second user; Obtain the data permissions that the first user has but the second user does not have as the first difference part; Obtain the data permissions that the second user has but the first user does not have as the second difference part; Based on the first difference part, generate a horizontal privilege escalation task for the second user; Based on the second difference part, generate a horizontal privilege escalation task for the first user.
[0011] Furthermore, according to the different API permissions between each role, vertical privilege escalation tasks corresponding to each role are generated. The method steps are as follows: Traverse each role in the clustering that has not generated a vertical privilege escalation task; Take the currently traversed role as the first role; Randomly select another different role from the roles in the clustering that have not generated a vertical privilege escalation task, and denote it as the second role; Obtain the API permissions that the first role has but the second role does not have as the first differential permission; Obtain the API permissions that the second role has but the first role does not have as the second differential permission; Generate a vertical privilege escalation task for the second role based on the first differential permission; Generate a vertical privilege escalation task for the first role based on the second differential permission.
[0012] Furthermore, execute the logical vulnerability detection task to detect logical vulnerabilities and obtain detection response data, including: Judge whether there is a logical vulnerability detection task generated currently: If not, no logical vulnerability detection is performed; If so, then: when there is an unauthorized detection task generated, execute the generated unauthorized detection task to detect logical vulnerabilities and obtain detection response data; when there is a horizontal authorization detection task generated, execute the generated horizontal authorization detection task to detect logical vulnerabilities and obtain detection response data; when there is a vertical privilege escalation detection task generated, execute the generated vertical privilege escalation detection task to detect logical vulnerabilities and obtain detection response data.
[0013] Furthermore, execute the unauthorized detection task to detect logical vulnerabilities and obtain detection response data, including: Generate an HTTP request based on the url, requestHeader, requestBody, and request method in the unauthorized detection task; Send the HTTP request to obtain detection response data.
[0014] Furthermore, execute the generated horizontal privilege escalation detection task to detect logical vulnerabilities and obtain detection response data, including: For each generated horizontal authorization detection task, monitor whether the corresponding user is in the user login state; For each horizontal authorization detection task whose corresponding user is not in the user login state, no detection is performed; For each horizontal authorization detection task where the corresponding user is in the user logged-in state, respectively put the user credentials of the corresponding user into its requestHeader request header, and then generate an HTTP request together with the url, requestBody request body, and request method of the horizontal authorization detection task for the requestHeader request header with the user credentials. Then send this HTTP request using the corresponding request method to obtain the detection response data.
[0015] Further, execute the generated vertical privilege escalation detection task to detect logical vulnerabilities and obtain the detection response data, including: For each generated vertical privilege escalation detection task, respectively execute the following steps: Monitor whether there is a user in the user logged-in state under the role corresponding to the vertical privilege escalation detection task: If not, do not perform detection; If there is, randomly select one of the users in the user logged-in state, obtain its user credentials, and put the user credentials into the requestHeader request header of the vertical privilege escalation detection task. Then generate an HTTP request together with the url, requestBody request body, and request method of the vertical privilege escalation detection task for the requestHeader request header, and then send this HTTP request to obtain the detection response data.
[0016] Further, perform network traffic logical vulnerability analysis based on the response data, including: For the detection response data of each unauthorized detection task obtained, input the detection response data into the vulnerability analysis model for analysis, and output whether there is a privilege escalation vulnerability in the API corresponding to the unauthorized detection task; For the detection response data of each horizontal privilege escalation task obtained, input the detection response data into the vulnerability analysis model for analysis, and output whether there is a privilege escalation vulnerability in the API corresponding to the horizontal privilege escalation task; For the detection response data of each vertical privilege escalation task obtained, input the detection response data into the vulnerability analysis model for analysis, and output whether there is a privilege escalation vulnerability in the API corresponding to the vertical privilege escalation task.
[0017] In a second aspect, the present invention provides a network traffic logical vulnerability analysis system, and the system includes: The first module is used to collect the mirror traffic data of the target network in real time, preprocess the collected mirror traffic data to obtain the preprocessed mirror traffic data as the original data. The data format of the original data is in JSON format, and the parameters are url, requestHeader request header, requestBody request body, responseHeader, and responseBody response body respectively. It is also used to extract and store the basic information of the APIs in the original data, and correspondingly extract and store the business systems to which the collected mirror traffic data of the target network belongs based on the extracted basic information of the APIs. The second module is used to extract requestHeader, requestBody, responseHeader, and responseBody from the original data, and perform word extraction on the extracted requestHeader, requestBody, responseHeader, and responseBody of the original data to construct the first dictionary. It extracts the url from the original data and performs word extraction on the extracted url of the original data to construct the second dictionary. The third module is used to extract url, requestHeader, requestBody, and responseBody from the original data, then perform word segmentation on the extracted url, requestHeader, requestBody, and responseBody according to upper and lower cases, and set different weights for each obtained word segment to form a feature vector based on the first dictionary and the second dictionary. Then, input the feature vector into the trained API business analysis model for calculation to output the business labels of the APIs in the original data, and input the feature vector into the trained API behavior analysis model for calculation to output the behavior labels of the APIs in the original data. The fourth module is used to extract the requestHeader data from the original data, input the extracted requestHeader data into the trained session analysis model for calculation, and output the user credentials in the original data. The fifth module is used to extract url, requestBody, and responseBody from the original data, then perform word segmentation on the extracted url, requestBody, and responseBody according to upper and lower cases, and set different weights for each obtained word segment to form a feature vector based on the first dictionary and the second dictionary. Then, input the formed feature vector into the trained login interface analysis model for calculation to output whether the API corresponding to the original data is a login interface. The sixth module is used to, when the API corresponding to the original data output by the login interface analysis model is the login interface: obtain and record the user credentials, username, and user login status of the login interface from the original data; then match the user credentials of the login interface obtained with the user credentials output by the session analysis model. If the match is successful, cache the API corresponding to the user credentials; if the match is unsuccessful, do not cache the API corresponding to the user credentials. The seventh module is used to statistically calculate the users and APIs of the stored business systems in real time; it is also used to statistically calculate the API permissions and API parameters of the users within the stored business systems in real time. The eighth module is used to generate a logical vulnerability detection task based on the data output by the session analysis model, the API business analysis model, and the API behavior analysis model, as well as the basic information of the stored APIs and the various statistically calculated data. The ninth module is used to execute the logical vulnerability detection task to perform logical vulnerability detection and obtain detection response data. The tenth module is used to perform network traffic logical vulnerability analysis based on the response data.
[0018] The present invention has the following beneficial effects compared with the prior art: The present invention helps to solve the problem that it is particularly difficult to obtain the API request path and parameter characteristics through the front-end page and JS in the traditional method, and is convenient for realizing the mining and analysis of unauthorized logical vulnerabilities in the business system.
[0019] The present invention helps to avoid detection in a pure manual manner and helps to avoid detection based on the background logic, thereby helping to solve the problems of low detection efficiency in the pure manual manner and the inability to confirm whether the detection result is accurate without knowing the background logic, and is convenient for realizing the mining and analysis of relevant logical vulnerabilities of overprivilege.
[0020] The design principle of the present invention is reliable, the structure is simple, and it has a very wide application prospect. Description of the Drawings
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 is a schematic flowchart of the method of an embodiment of the present invention; Figure 2 is a schematic block diagram of the system of an embodiment of the present invention. Detailed Embodiments
[0023] To make the technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0024] The following description of at least one exemplary embodiment is actually only illustrative and in no way limits the present invention and its application or use. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0025] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0026] It should be noted that the present invention uses terms such as "first" and "second" to limit the corresponding parts only for the convenience of distinguishing the corresponding parts. Without additional statements, the above terms have no special meanings and should not be construed as limiting the scope of protection of the present invention.
[0027] Next, the technical solutions of the present invention will be introduced and described through several embodiments.
[0028] In one embodiment, the present invention provides a method for analyzing logical vulnerabilities in network traffic.
[0029] Please refer to Figure 1 , and the method 100 includes the following steps 110 to 200.
[0030] Step 110: Real-time collect the mirror traffic data of the target network, preprocess the collected mirror traffic data to obtain the preprocessed mirror traffic data as the original data; extract and store the basic information of the APIs in the original data, and correspondingly extract and store the business system to which the collected mirror traffic data of the target network belongs based on the extracted basic information of the APIs.
[0031] The data format of the original data is in json format, and the parameters are respectively url, requestHeader request header, requestBody request body, responseHeader, and responseBody response body.
[0032] The implementation method of the preprocessing includes: Filter out the traffic volume unrelated to the API in the mirror traffic data, and then remove duplicate traffic data. Only one piece of data for each different URL is retained as the API of the system. Then, handle the missing data by deleting the traffic data with missing request headers and response bodies. After that, perform data format conversion to convert it into JSON format.
[0033] According to the basic information of the extracted API, extract the business systems to which the mirror traffic data of the target network collected belongs, including: Extract the business systems to which the corresponding mirror traffic data of the target network collected belongs according to the IP and port in the basic information of the extracted API.
[0034] In the basic information of the API, one IP plus one port is regarded as one business system.
[0035] Step 120: Extract requestHeader, requestBody, responseHeader, and responseBody from the original data, and extract words from the requestHeader, requestBody, responseHeader, and responseBody of the extracted original data to construct the first dictionary; extract the url from the original data, and extract words from the extracted url of the original data to construct the second dictionary.
[0036] In this embodiment, the construction method of the first dictionary is as follows: Extract requestHeader, requestBody, responseHeader, and responseBody from the original data to obtain the first extraction data. Then, extract all the letters in the first extraction data to form a string. Then, perform word segmentation on the extracted string according to case, obtain several word segments, and remove duplicates, convert to lowercase, sort by letters, and add an index to form the first dictionary.
[0037] The construction method of the second dictionary is as follows: Extract the url from the original data to obtain the second extraction data. Then, extract all the letters in the second extraction data to form a string. Then, perform word segmentation on the extracted string according to case, obtain several word segments, and then remove duplicates, convert to lowercase, sort by letters, and add an index to form the second dictionary.
[0038] Step 130: Extract the url, requestHeader, requestBody, and responseBody from the original data to form a feature vector based on the first dictionary and the second dictionary. Then, input this feature vector into the trained API business analysis model for calculation to output the business label of the API in the original data, and input this feature vector into the trained API behavior analysis model for calculation to output the behavior label of the API in the original data.
[0039] In this embodiment, extracting the url, requestHeader, requestBody, and responseBody from the original data to form a feature vector based on the first dictionary and the second dictionary includes: Extract the url, requestHeader, requestBody, and responseBody from the original data. Then, perform word segmentation on the extracted url, requestHeader, requestBody, and responseBody according to case, and set different weights for each resulting word segment to form a feature vector based on the first dictionary and the second dictionary.
[0040] Step 140: Extract the requestHeader data from the original data, and input the extracted requestHeader data into the trained session analysis model for calculation to output the user credentials in the original data.
[0041] It can be understood that in specific implementation, if there are no user credentials in the original data, then input the extracted requestHeader data into the trained session analysis model for calculation, and the session analysis model outputs that the user credentials in the original data are empty.
[0042] Step 150: Extract the url, requestBody, and responseBody from the original data. Then, perform word segmentation on the extracted url, requestBody, and responseBody according to case, and set different weights for each resulting word segment to form a feature vector based on the first dictionary and the second dictionary. Then, input the formed feature vector into the trained login interface analysis model for calculation to output whether the API corresponding to the original data is a login interface.
[0043] Step 160: When the API corresponding to the original data output by the login interface analysis model is the login interface: Obtain and record the user credentials, username, and user login status of the login interface from the original data (in specific implementation, the username of the login interface can be obtained from the parameters of the url or the requestBody in the original data, the user credential information of the login interface can be obtained from the responseBody of the original data, and the user login status can be judged according to whether the login interface call is successful. If the call is successful, it is determined that the user is in the logged-in state; otherwise, the user is in the non-logged-in state); then match the user credentials of the login interface obtained with the user credentials output by the session analysis model. If the match is successful, cache the API corresponding to the user credentials, that is, collect the API permission information of the user; if the match is unsuccessful, do not cache the API corresponding to the user credentials.
[0044] In specific implementation, each time the API corresponding to the user credentials is cached, session matching and parameter parsing are also performed to generate the user's API permissions and corresponding data permissions. Specifically: Each time the API corresponding to the user credentials is cached, the API permissions corresponding to the user are also generated according to the cached API. At the same time, the business labels and behavior labels of the API output by the API business analysis model and the API behavior analysis model are used to analyze the recorded user credentials to obtain the API parameters of the corresponding user.
[0045] Step 170: Real-time statistics of the users and APIs of the stored business systems; real-time statistics of the API permissions and API parameters of the users within the stored business systems.
[0046] Step 180: Based on the data output by the session analysis model, the API business analysis model, and the API behavior analysis model, as well as the basic information of the stored APIs and the various statistics, generate a logical vulnerability detection task.
[0047] Step 190: Execute the logical vulnerability detection task to detect logical vulnerabilities and obtain detection response data.
[0048] Step 200: Perform network traffic logical vulnerability analysis based on the response data.
[0049] Exemplarily, based on the data output by the session analysis model, the API business analysis model, and the API behavior analysis model, as well as the basic information of the stored APIs and the various statistics, generating a logical vulnerability detection task includes: Judge in real time whether the user credentials output by the session analysis model are empty. If not, collect the url, requestHeader request headers, and requestBody request body in the original data to generate an unauthorized detection task. If it is empty, do not generate an unauthorized detection task.
[0050] The content of the generated unauthorized detection task includes the url, requestHeader request headers, and requestBody request body in the collected original data, and also includes the request method.
[0051] Among them, the request method can be obtained from the basic information of the corresponding API in the original data.
[0052] The request method can be any one of GET, POST, PUT, DELETE, HEAD, OPTIONS, TRACE, CONNECT.
[0053] Exemplarily, based on the data output by the session analysis model, the API business analysis model, and the API behavior analysis model, as well as the basic information of the stored API and the statistics of each data, generate a logical vulnerability detection task, which also includes: Regularly judge whether the basic information of the currently stored API meets the conditions for generating a horizontal privilege escalation task. If not, do not generate a horizontal privilege escalation task. If so, on the one hand, based on the API permissions and API parameters of the users in the business system currently counted, obtain the data permissions of the users corresponding to the basic information of each currently stored API; on the other hand, according to the API business labels and API behavior labels output by the trained API business analysis model and the trained API behavior analysis model, obtain the data permissions of the users corresponding to the basic information of each currently stored API; then collect the data permissions obtained from the above two aspects to obtain all the data permissions of the users corresponding to the basic information of each currently stored API; then, according to the differences in data permissions between different users corresponding to the basic information of the currently stored API, generate a horizontal privilege escalation task for each API corresponding to the basic information of the currently stored API; Regularly cluster the users in the business system currently counted according to the API permissions of the users in the business system currently counted to obtain several user clusters; obtain the API permissions of the users in each user cluster to obtain the API permissions corresponding to each user cluster; regard each user cluster as a role to obtain the role corresponding to each user cluster, and then obtain the API permissions corresponding to each role; according to the differences in API permissions between each role, generate a vertical privilege escalation task corresponding to each role.
[0054] In this embodiment, the condition for generating a horizontal privilege escalation task is that, among the basic information of the currently stored APIs, the time length between the generation time of the basic information of the earliest stored API and the current time is greater than a preset time length.
[0055] The basic information of the API includes the URL corresponding to the API, the request header of the requestHeader, and the request body of the requestBody.
[0056] Optionally, according to the differences in data permissions among different users corresponding to the basic information of the currently stored APIs, generate horizontal privilege escalation tasks for the APIs corresponding to the basic information of each of the currently stored APIs. The specific implementation steps are as follows: For each API corresponding to the basic information of the currently stored APIs, perform the following steps: Randomly select two users who have the API permissions corresponding to the basic information of the current API from the API permissions of the users in the currently statistically business system, and record them as the first user and the second user; Obtain the data permissions that the first user has but the second user does not have among the data permissions of the first user as the first difference part; Obtain the data permissions that the second user has but the first user does not have among the data permissions of the second user as the second difference part; Based on the first difference part, generate a horizontal privilege escalation task for the second user; Based on the second difference part, generate a horizontal privilege escalation task for the first user.
[0057] For example, among the APIs corresponding to the basic information of the currently stored APIs, there is an API called the first API. User One and User Two are two randomly selected users who have the same API permissions as those corresponding to the basic information of the first API. Both User One and User Two have the interface access permission to the first API. The data permissions of User One are A, B, C (that is, when User One accesses the first API, they can obtain data A, B, and C), and the data permissions of User Two are B, C, D (that is, when User Two accesses the first API, they can obtain data B, C, and D). On this basis, the data permission that User One has but User Two does not have is A, and the data permission that User Two has but User One does not have is D. That is, the first difference part is the data permission A, and the second difference part is the data permission D. Therefore, based on the first difference part, generate a horizontal privilege escalation task for User Two to access data A, and based on the second difference part, generate a horizontal privilege escalation task for User One to access data D.
[0058] The content of the generated horizontal privilege escalation detection task includes the corresponding user, including the URL of the corresponding API, the request header information of the requsetHeader, and the requestBody, and also includes the request method obtained from the basic information of the corresponding API.
[0059] Taking the horizontal privilege escalation task of generating user B's access data A based on the first difference part as an example, the horizontal privilege escalation task of the generated user B's access data A includes the user name of user B, the URL of the corresponding API (i.e., the first API), the request header of the requestHeader, the request body of the requestBody, and also includes the request method obtained from the basic information of the corresponding API (i.e., the first API).
[0060] The task content of other horizontal privilege escalation tasks can refer to the horizontal privilege escalation task of generating user B's access data A based on the first difference part.
[0061] It can be understood that data permission is the permission to access data. Taking the above data permission A as an example, it means having the permission to access data A.
[0062] Exemplarily, the user login status can be obtained according to the login API and its corresponding session information. If the user calls the login API successfully, it is considered that the user is in the logged-in state. The default user online time is 30 minutes. If it is detected through the session information that the user calls other API interfaces, the online time of the user for the successfully logged-in other API interfaces is refreshed.
[0063] Optionally, according to the different API permissions between each role, generate the corresponding vertical privilege escalation task for each role. The method steps are as follows: Traverse each role in the clustering that has not generated a vertical privilege escalation task; Take the currently traversed role as the first role; Randomly select another different role from the roles in the clustering that have not generated a vertical privilege escalation task, and denote it as the second role; Obtain the API permissions that the first role has but the second role does not have, as the first difference permission; Obtain the API permissions that the second role has but the first role does not have, as the second difference permission; Based on the first difference permission, generate the vertical privilege escalation task of the second role; Based on the second difference permission, generate the vertical privilege escalation task of the first role.
[0064] Assume that the currently traversed role is Role 1, and Role 2 is another different role randomly selected from the roles obtained by clustering that have not generated vertical privilege escalation tasks. The APIs that Role 1 has permissions for are API1, API2, and API3, and the APIs that Role 2 has permissions for are API3, API4, and API5. The APIs that Role 1 has but Role 2 does not have permissions for are API1 and API2, and the APIs that Role 2 has but Role 1 does not have permissions for are API4 and API5. That is, API1 and API2 are the APIs of the first differential permissions, and API4 and API5 are the APIs of the second differential permissions. Therefore, the present invention can generate two vertical privilege escalation tasks for the second role based on the first differential permission APIs API1 and API2 (i.e., the vertical privilege escalation task of Role 2 accessing API1 and the vertical privilege escalation task of Role 2 accessing API2), and can generate two vertical privilege escalation tasks for Role 1 to access API4 and API5 based on the second differential permission APIs API4 and API5 (i.e., the vertical privilege escalation task of Role 1 accessing API4 and the vertical privilege escalation task of Role 1 accessing API5).
[0065] The content of the vertical privilege escalation detection task includes the corresponding role, including the url of the corresponding API, the requsetHeader request header information and the requestBody, and also includes the request method obtained from the basic information of the corresponding API.
[0066] Taking the vertical privilege escalation task of Role 1 accessing API4 as an example, the task content of this vertical privilege escalation task includes Role 1 and the url, requestHeader request header, and requestBody request body of the corresponding API (i.e., API4) of this vertical privilege escalation task.
[0067] The content of other vertical privilege escalation tasks can refer to the vertical privilege escalation task of Role 1 accessing API4.
[0068] Optionally, execute the logical vulnerability detection task to perform logical vulnerability detection and obtain the detection response data, including:[[]] Judge whether there is a logical vulnerability detection task generated currently:[[]] If not, no logical vulnerability detection is performed;[[]] If so, then: when an unauthorized detection task is generated, execute the generated unauthorized detection task to perform logical vulnerability detection and obtain the detection response data; when a horizontal authorization detection task is generated, execute the generated horizontal authorization detection task to perform logical vulnerability detection and obtain the detection response data; when a vertical privilege escalation detection task is generated, execute the generated vertical privilege escalation detection task to perform logical vulnerability detection and obtain the detection response data.
[0069] Optionally, execute the unauthorized detection task to perform logical vulnerability detection and obtain the detection response data, including:[[]] Generate an HTTP request based on the URL, request header (requestHeader), and request body (requestBody) in the unauthorized detection task; Send the HTTP request to obtain detection response data.
[0070] Optionally, execute the generated horizontal authorization detection task for logical vulnerability detection to obtain detection response data, including: For each generated horizontal authorization detection task, monitor whether the corresponding user is in the user login state; For each horizontal authorization detection task whose corresponding user is not in the user login state, no detection is performed; For each horizontal authorization detection task whose corresponding user is in the user login state, put the user credentials of its corresponding user into its request header (requestHeader), and then generate an HTTP request together with the request header (requestHeader) with the user credentials, the URL of the horizontal authorization detection task, and the request body (requestBody). Then send this HTTP request to obtain detection response data.
[0071] Optionally, execute the generated vertical privilege escalation detection task for logical vulnerability detection to obtain detection response data, including: For each generated vertical privilege escalation detection task, perform the following steps respectively: Monitor whether there is a user in the user login state under the role corresponding to the vertical privilege escalation detection task: If not, no detection is performed; If there is, randomly select one of the users in the user login state, obtain its user credentials, and put the user credentials into the request header (requestHeader) of the vertical privilege escalation detection task. Then generate an HTTP request together with the request header (requestHeader), the URL in the vertical privilege escalation detection task, and the request body (requestBody). Then send this HTTP request to obtain detection response data.
[0072] Optionally, perform network traffic logical vulnerability analysis based on the response data, including: For the detection response data of each obtained unauthorized detection task, input the detection response data into the vulnerability analysis model for analysis, and output whether there is a privilege escalation vulnerability in the API corresponding to the unauthorized detection task; For the detection response data of each obtained horizontal privilege escalation task, input the detection response data into the vulnerability analysis model for analysis, and output whether there is a privilege escalation vulnerability in the API corresponding to the horizontal privilege escalation task; For the detection response data of each obtained vertical privilege escalation task, the detection response data is input into a vulnerability analysis model for analysis, and whether there is a privilege escalation vulnerability in the API corresponding to the vertical privilege escalation task is output.
[0073] Optionally, the training method of the API behavior analysis model includes: Step 1: Construct a first training set.
[0074] The construction method of the first training set includes: S1. Data collection and processing: Collect the historical mirror traffic data of the target network; Preprocess the historical mirror traffic data to obtain the preprocessed historical mirror traffic data as the original historical data; the data format of the original historical data is in json format, and the parameters are url, requestHeader request header, requestBody request body, responseHeader, and responseBody response body respectively; Annotate the original historical data to obtain the annotated original historical data as the first annotated data; the annotation label is a pre-set interface function label.
[0075] S2. Feature engineering: Build a dictionary: Extract requestHeader, requestBody, responseHeader, and responseBody from the original historical data, and perform word extraction on them. Extract all the letters in requestHeader, requestBody, responseHeader, and responseBody in the original historical data to form a string, and then perform word segmentation on the extracted string according to case, obtaining several word segments. Then, remove duplicates from the obtained word segments, convert them to lowercase, sort them alphabetically, and add them to the index to form a dictionary as the first reference dictionary; Extract the url from the original historical data, and perform word extraction on it. Extract all the letters in the url in the original historical data to form a string, and then perform word segmentation on the extracted string according to case, obtaining several word segments. Then, remove duplicates from the obtained word segments, convert them to lowercase, sort them alphabetically, and add them to the index to form a dictionary as the second reference dictionary; Generate historical data feature vectors: Extract the API data from the original historical data. The API data includes URL, request header, request body, and response body data. Tokenize the extracted URL, request header, request body, and response body data by case, and set different weights for each obtained token to form a feature vector based on the first reference dictionary and the first reference dictionary, that is, obtain the historical data feature vector.
[0076] It can be understood that the interface function tags used in the first annotation data include query list, query tree, query details, delete data, start task, insert data, update data, and upload file.
[0077] In specific implementation, the API function can be judged according to the content of the data. In this embodiment, the API function can be judged according to the content of the API, such as URL and response body, and then the corresponding interface function tag can be determined.
[0078] S3. Dataset division: Match the historical data feature vector with the first annotation data corresponding to the original historical data to obtain matching data; collect the matching data to obtain the first dataset; Divide the first dataset into a training set, a validation set, and a test set in the ratio of 7:1.5:1.5. The training set obtained by this division is the first training set.
[0079] The training set is used for model training, the validation set is used for model hyperparameter tuning, and the test set is used to evaluate the final performance of the model.
[0080] Step 2. Model construction.
[0081] Use the method of ensemble learning to integrate logistic regression, random forest, multi-layer perceptron, and gradient boosting tree into a model using the soft voting method as the initial API behavior analysis model.
[0082] Step 3. Model training.
[0083] Input the first training set into the initial API behavior analysis model for training, and then adjust the learning rate and various hyperparameters of the model according to the validation set until the accuracy rate and recall rate reach more than 95%.
[0084] This model is an ensemble learning model, and the parameters of each model need to be tuned.
[0085] For the logistic regression model, its regularization type selects Lasso regularization, and the optimization algorithm uses the saga optimization algorithm.
[0086] Lasso regularization is suitable for high-dimensional data and sparse data. The saga optimization algorithm is suitable for large-scale data, supports Lasso regularization, has a fast convergence rate, and has more advantages in dealing with sparse matrices.
[0087] During the training process of the logistic regression model, the adjusted parameters are the regularization parameter C and the maximum number of iterations max_iter.
[0088] The regularization parameter C is the reciprocal of the regularization coefficient. The smaller C is, the stronger the regularization is, and the larger C is, the weaker the regularization is. Strong regularization can prevent the model from being too complex, but there may be an underfitting situation. Weak regularization may cause an overfitting situation. By adjusting the regularization parameter, the fitting situation of the model can be adjusted. The maximum iteration coefficient is the control of the model convergence. In the case of a large dataset, the maximum iteration coefficient can be increased to ensure the convergence of the model.
[0089] The training parameter tuning of the random forest includes adjusting the number of trees, the maximum depth of the trees, the minimum number of samples per leaf node, the minimum number of samples per split, and the maximum number of leaf nodes.
[0090] Among them, the number of trees refers to the number of trees in the forest. The maximum depth of the trees, the minimum number of samples per leaf node, the minimum number of samples per split, and the maximum number of leaf nodes are used to control the depth and complexity of the generated trees. The method of parameter tuning is the grid search method, which sets a series of candidate values for multiple hyperparameters to form a parameter grid and tries each combination one by one to find the best combination.
[0091] The training parameter tuning of the multi-layer perceptron involves the network architecture, the optimization algorithm, and the learning rate. The activation function of the multi-layer perceptron selects the rectified linear unit ReLU, the optimization algorithm selects the Adam adaptive moment estimation method, and the learning rate uses a fixed learning rate method. During training, the grid search method is used to continuously adjust the initial learning rate, the number of hidden layers, and the number of neurons in each layer, try each combination one by one, and find the best parameters.
[0092] The parameters adjusted during the training of the gradient boosting tree are similar to those of the random forest, including adjusting the number of trees, the learning rate, and the maximum depth of the trees. The method of parameter tuning is to first fix the number of trees and then vary the learning rate and the maximum depth of the trees to gradually find the best combination. Then adjust the number of trees and repeat the adjustment of the learning rate and the maximum depth of the trees until the final parameters are determined.
[0093] Step 4: Model evaluation and optimization.
[0094] Evaluate the model using the test set to ensure that the accuracy and recall rate of the model on the test set reach over 90%. If not satisfied, modify the model parameters and repeat step three until the model meets the requirements on the test set to obtain the trained API behavior analysis model.
[0095] Optionally, the training method of the API business analysis model includes the following steps L1 to L4.
[0096] Step L1, construct the second training set.
[0097] Step L11, data collection and processing: Obtain the original historical data; Annotate the obtained original historical data to obtain the annotated data as the second annotated data; the annotation labels used in the second annotated data are pre-set business labels.
[0098] Business labels are used to divide the business of the API, and the business systems are divided into different business modules through business labels.
[0099] The pre-set business labels in the present invention include user (used to annotate business modules related to users), department (used to annotate business modules related to organizational structures), menu (used to annotate business modules related to business system menus), role (used to annotate business modules related to role configurations), configuration (used to annotate business modules related to system configurations), dictionary (used to annotate business modules related to system dictionaries), notice (used to annotate business modules related to notices, messages, announcements), process (used to annotate business modules related to business processes such as reviews and approvals), log (used to annotate business modules related to system logs), login (used to annotate business modules related to login services), device (used to annotate business modules related to system device ledgers), report (used to annotate business modules related to statistical analysis), work order (used to annotate business modules related to task work order issuance), file (used to annotate business modules related to file saving, uploading and downloading), email (used to annotate business modules related to email issuance), SMS (used to annotate business modules related to SMS sending services), protocol (used to annotate business modules related to system protocol formulation), monitoring (used to annotate business modules related to system status monitoring), real-time communication (used to annotate business modules related to system real-time communication).
[0100] It should be noted that the data permissions of a user in a business module are unified.
[0101] Preferably, the data permissions of a user in a business module can be obtained by collecting the response data of APIs in the same business module. Specifically, the statistical method for the data permissions of a user in an API includes: Step (1): Collect different data of users accessing the API by collecting mirror traffic data of the target network to obtain the data permissions of users for this API. Step (2): Through the collected mirror traffic data of the target network, obtain the business system to which the collected mirror traffic data of the target network belongs. Then, use the response data of each function interface such as the query tree and query list of the corresponding business module of this API of the user in this business system as the data permissions of the user for this API. Step (3): Take the union of the data permissions obtained in Step (1) and Step (2). The obtained data permissions are the statistically data permissions of users for the API.
[0102] Step L12: Feature engineering: Obtain the first reference dictionary and the second reference dictionary; Obtain the original historical data; Segment each piece of the obtained original historical data by url, requestBody, and responseBody. Then, generate feature vectors for the segmented words corresponding to url, requestBody, and responseBody respectively according to the word frequency, based on the obtained first reference dictionary and second reference dictionary, using TF-IDF. Obtain three feature vectors corresponding to url, requestBody, and responseBody. Set weights for these three feature vectors. Then, perform weighted combination of the three vectors with weights set column by column to generate a new feature vector.
[0103] For example, before setting weights, the above three feature vectors are respectively , and . When specifically implemented, set weights of 0.8, 0.1, and 0.1 for these three feature vectors in sequence. Then, the new feature vector generated after weighted combination of the three vectors column by column is .
[0104] Step L13: Divide the dataset: Match the new feature vectors generated in Step L12 and their corresponding original historical data with the labels marked in Step L11 to obtain matching data. Collect this matching data to obtain the second dataset; Divide the second dataset into a training set, a validation set, and a test set in the ratio of 7:1.5:1.5. The obtained training set is the second training set.
[0105] Step L2: Model construction: Using the method of ensemble learning, a model is integrated by soft voting method with logistic regression, random forest, multi-layer perceptron and gradient boosting tree, and this model is used as the initial API business analysis model constructed.
[0106] Step L3, Model training: Input the second training set into the initial API business analysis model for training, and then continuously adjust the learning rate and various hyperparameters of the model according to the validation set until the accuracy and recall rate reach more than 95%.
[0107] Step L4, Model evaluation and optimization: Use the test set to continue evaluating the trained model to ensure that the accuracy and recall rate of the model on the test set reach more than 90%; if the accuracy and recall rate of the model on the test set cannot both reach more than 90%, modify the model parameters and return to step L3 for retraining until the accuracy and recall rate of the model on the test set reach more than 90%.
[0108] Optionally, the training method of the session analysis model includes: ① Data collection and processing: Obtain the original historical data; After that, label the requestHeader data in the obtained original historical data to get the labeled data as the third labeled data; the method of labeling the requestHeader data in the obtained original historical data is that if the requestHeader data contains user credentials, label the user credential fields, and the user credentials include authrization, token, cookie fields, and if the requestHeader data does not contain user credentials, label it as None; ② Dataset division: Collect the third labeled data to form the third dataset; Divide the data in the third dataset into a training set and a test set according to the ratio of 8:2, and this training set is the third training set; ③ Model construction: Construct a distilbert-base-uncased model distilled from the Bert model as the initial session analysis model; ④ Model training: Since distilbert-base-uncased is a pre-trained model that has been trained on general knowledge, the training model of the present invention fine-tunes on this pre-trained model using a third training set. When fine-tuning, the first four layers of the model are frozen, and only the last two layers are fine-tuned. In this way, on the basis of ensuring the general knowledge of the model, the model can learn task-specific representations, and at the same time, the computational cost can be effectively saved.
[0109] The specific training steps include: First, define the optimizer and loss function. The SGD optimizer is selected as the optimizer, and the cross-entropy loss function is selected as the loss function. The initial learning rate of the optimizer is set to 0.001, and the momentum is set to 0.9.
[0110] Then start the fine-tuning of the model. Specifically, continuously adjust the learning rate, momentum, and appropriately adjust parameters such as the number of training epochs, batch size, and weight decay according to the model training results. Continuously iterate the training until the test results of the model reach the expected goal.
[0111] ⑤ Model evaluation and tuning: Use the test set to evaluate the model. If the accuracy and recall rate of the model cannot both reach more than 90%, modify the optimizer parameters and model parameters, and then repeat step ④ of the model training steps until the recall rate and accuracy of the model on the test set both reach 90% to obtain the trained session analysis model.
[0112] Optionally, the training method of the login interface analysis model includes the following steps H1 to step H5.
[0113] Step H1, data collection and processing: Obtain the first reference dictionary and the second reference dictionary; Obtain the original historical data; Annotate the obtained original historical data to obtain the annotated data as the fourth annotated data. The annotation labels used are login interface and non-login interface; Segment each piece of the obtained original historical data by url, requestBody, and responseBody, and then set different weights for each segmented word to form a feature vector based on the first reference dictionary and the second reference dictionary. Then set weights for the formed feature vector, and perform weighted combination of the feature vectors with weights set by columns to generate a new feature vector.
[0114] Step H2, divide the data set: Match the newly generated feature vectors in step H1 with the labels (i.e., login interface or non-login interface) corresponding to the original historical data to obtain matching data, and collect this matching data to obtain the fourth data set; Divide the fourth data set into a training set, a validation set, and a test set in the ratio of 7:1.5:1.5. The training set obtained from this division is denoted as the fourth training set. The training set is used for model training, the validation set is used for model hyperparameter tuning, and the test set is used to evaluate the final performance of the model.
[0115] Step H3, Model construction: Use the method of ensemble learning to integrate logistic regression, random forest, multi-layer perceptron, and gradient boosting trees into a model using the soft voting method as the initial login interface analysis model.
[0116] Step H4, Model training: Input the fourth training set into the initial API behavior analysis model for training, and then continuously adjust the learning rate and various hyperparameters of the model according to the validation set until the accuracy rate and recall rate reach more than 95%.
[0117] Step H5, Model evaluation and tuning: Use the test set to evaluate the model to ensure that the accuracy rate and recall rate of the model on the test set also reach more than 90%; if not satisfied, modify the model parameters and repeat the steps of H4 until the model meets the requirements on the test set to obtain the trained login interface analysis model.
[0118] Optionally, the training method of the vulnerability analysis model includes the following steps M1 to M5.
[0119] Step M1, Data collection and processing: Obtain the original historical data; Extract the responseBody data from the obtained original historical data and label the extracted responseBody data. Among them, for the responseBody of a successful response, the label is marked as successful, and for the responseBody of an unsuccessful response, the label is marked as failed; Collect the labeled responseBody data to construct a data set.
[0120] Step M2, Divide the data set: Divide the data set constructed in M1 into a training set and a test set in the ratio of 8:2. The training set obtained from this division is denoted as the fifth training set.
[0121] Step M3, Model construction: Build a distilbert-base-uncased model distilled from the Bert model, and add a fully connected layer at the end of the distilbert-base-uncased model to obtain an initial privilege escalation analysis model; the fully connected layer is used as the classification layer of the model to convert the distilbert-base-uncased model into a binary classifier.
[0122] Step M4, Model Training: Set two optimizers: Set an optimizer for distilbert-base-uncased as the first optimizer; set an optimizer for the added fully connected layer as the second optimizer.
[0123] Set learning rates for the two optimizers. Among them, set the learning rate of the first optimizer to be less than that of the second optimizer. Specifically, set a relatively small learning rate for distilbert-base-uncased and only slightly adjust a part of distilbert-base-uncased; set a relatively large learning rate for the added fully connected layer to train the randomly initialized parameters.
[0124] Before training, first freeze the first four layers of distilbert-base-uncased, only take out the data of the last two layers and put them into the optimizer, and then take out the added fully connected layer and put it into the two optimizers for optimization.
[0125] Set different learning rates for the two optimizers, and continuously train the model using the fifth training set, adjusting parameters such as the learning rate and optimizer momentum of model training until a model that meets the expectations is obtained.
[0126] Step M5, Model Evaluation and Tuning: Use the test set to evaluate the model. If the accuracy and recall rate of the model do not reach more than 90%, modify the optimizer parameters and model training parameters, and then repeat the model training steps in Step M4 until the recall rate and accuracy of the model on the test set both reach more than 90% to obtain a trained vulnerability analysis model.
[0127] As Figure 2 shown, the system 300 includes: The first module 301 is used to collect the mirror traffic data of the target network in real time, preprocess the collected mirror traffic data to obtain the preprocessed mirror traffic data as the original data. The data format of the original data is in JSON format, and the parameters are url, requestHeader request header, requestBody request body, responseHeader, and responseBody response body respectively. It is also used to extract and store the basic information of the APIs in the original data, and correspondingly extract and store the business system to which the collected mirror traffic data of the target network belongs based on the extracted basic information of the APIs. The second module 302 is used to extract requestHeader, requestBody, responseHeader, and responseBody from the original data, and extract words from the extracted requestHeader, requestBody, responseHeader, and responseBody of the original data to construct the first dictionary. It extracts the url from the original data and extracts words from the extracted url of the original data to construct the second dictionary. The third module 303 is used to extract url, requestHeader, requestBody, and responseBody from the original data, then segment the extracted url, requestHeader, requestBody, and responseBody by case, and set different weights for each resulting segment to form a feature vector based on the first dictionary and the second dictionary. Then, the feature vector is input into the trained API business analysis model for calculation to output the business labels of the APIs in the original data, and the feature vector is input into the trained API behavior analysis model for calculation to output the behavior labels of the APIs in the original data. The fourth module 304 is used to extract the requestHeader data from the original data, input the extracted requestHeader data into the trained session analysis model for calculation, and output the user credentials in the original data. The fifth module 305 is used to extract url, requestBody, and responseBody from the original data, then segment the extracted url, requestBody, and responseBody by case, and set different weights for each resulting segment to form a feature vector based on the first dictionary and the second dictionary. Then, the formed feature vector is input into the trained login interface analysis model for calculation to output whether the API corresponding to the original data is a login interface. The sixth module 306 is used to, when the API corresponding to the original data output by the login interface analysis model is the login interface: obtain and record the user credentials, username, and user login status of the login interface from the original data; then match the user credentials of the login interface obtained with the user credentials output by the session analysis model. If the match is successful, cache the API corresponding to the user credentials; if the match is unsuccessful, do not cache the API corresponding to the user credentials. The seventh module 307 is used to statistically analyze the users and APIs of the stored business systems in real time; it is also used to statistically analyze the API permissions and API parameters of the users within the stored business systems in real time. The eighth module 308 is used to generate a logical vulnerability detection task based on the data output by the session analysis model, the API service analysis model, and the API behavior analysis model, as well as the basic information of the stored APIs and the various statistically analyzed data. The ninth module 309 is used to perform a logical vulnerability detection task to detect logical vulnerabilities and obtain detection response data. The tenth module 310 is used to perform a logical vulnerability analysis of network traffic based on the response data.
[0128] The sixth module 306 is also used to take no action when the API corresponding to the original data output by the login interface analysis model is not the login interface.
[0129] For the same or similar parts among the various embodiments in this specification, reference can be made to each other.
[0130] Although the present invention has been described in detail by referring to the accompanying drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and all such modifications or substitutions should fall within the scope of the present invention. / Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, and all should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for analyzing logical vulnerabilities in network traffic, characterized in that, The method includes: Collecting mirror traffic data of the target network in real time, and preprocessing the collected mirror traffic data to obtain the preprocessed mirror traffic data as the original data; the data format of the original data is in JSON format, and the parameters are respectively url, requestHeader request header, requestBody request body, responseHeader, and responseBody response body; Extracting and storing the basic information of the API in the original data, and correspondingly extracting and storing the business system to which the collected mirror traffic data of the target network belongs according to the extracted basic information of the API; Extracting requestHeader, requestBody, responseHeader, and responseBody from the original data, and extracting words from the extracted requestHeader, requestBody, responseHeader, and responseBody of the original data to construct a first dictionary; extracting the url from the original data, and extracting words from the extracted url of the original data to construct a second dictionary; Extracting the url, requestHeader, requestBody, and responseBody from the original data, then segmenting the extracted url, requestHeader, requestBody, and responseBody according to upper and lower case, and setting different weights for each obtained segment to form a feature vector based on the first dictionary and the second dictionary, then inputting the feature vector into the trained API business analysis model for calculation to output the business label of the API in the original data, and inputting the feature vector into the trained API behavior analysis model for calculation to output the behavior label of the API in the original data; Extracting the requestHeader data from the original data, and inputting the extracted requestHeader data into the trained session analysis model for calculation to output the user credentials in the original data; Extracting the url, requestBody, and responseBody from the original data, then segmenting the extracted url, requestBody, and responseBody according to upper and lower case, and setting different weights for each obtained segment to form a feature vector based on the first dictionary and the second dictionary, and then inputting the formed feature vector into the trained login interface analysis model for calculation to output whether the API corresponding to the original data is a login interface; The method further includes: When the API corresponding to the original data output by the login interface analysis model is the login interface: obtain and record the user credentials, username, and user login status of the login interface from the original data; then match the user credentials of the login interface obtained with the user credentials output by the session analysis model. If the match is successful, cache the API corresponding to the user credentials; if the match is unsuccessful, do not cache the API corresponding to the user credentials. The method further includes: Real-time statistics of the users and APIs of the stored business systems; Real-time statistics of the API permissions and API parameters of the users within the stored business systems; The method further includes: Generate a logical vulnerability detection task based on the data output by the session analysis model, the API business analysis model, and the API behavior analysis model, as well as the basic information of the stored APIs and the various statistics; Execute the logical vulnerability detection task to perform logical vulnerability detection and obtain detection response data; Perform network traffic logical vulnerability analysis based on the response data.
2. The method according to claim 1, wherein Generating a logical vulnerability detection task based on the data output by the session analysis model, the API business analysis model, and the API behavior analysis model, as well as the basic information of the stored APIs and the various statistics, includes: Real-time determine whether the user credentials output by the session analysis model are empty. If not, collect the url, requestHeader request headers, and requestBody request bodies in the original data to generate an unauthorized detection task; if empty, do not generate an unauthorized detection task; Generating a logical vulnerability detection task based on the data output by the session analysis model, the API business analysis model, and the API behavior analysis model, as well as the basic information of the stored APIs and the various statistics, further includes: Periodically determine whether the basic information of the currently stored APIs meets the conditions for generating a horizontal privilege escalation task. If not, do not generate a horizontal privilege escalation task. If so, on the one hand, based on the API permissions and API parameters of the users within the currently statistical business system, obtain the data permissions of the users corresponding to the basic information of each currently stored API; on the other hand, according to the API business labels and API behavior labels output by the trained API business analysis model and the trained API behavior analysis model, obtain the data permissions of the users corresponding to the basic information of each currently stored API; then gather the data permissions obtained from the above two aspects to obtain all the data permissions of the users corresponding to the basic information of each currently stored API; then generate the horizontal privilege escalation task of the API corresponding to the basic information of each currently stored API according to the differences in data permissions between different users corresponding to the basic information of the currently stored APIs. Timing is based on the API permissions of users in the business system currently being counted, clustering the users in the business system currently being counted to obtain several user clusters; obtaining the API permissions of users within each user cluster to obtain the API permissions corresponding to each user cluster; taking each user cluster as a role to obtain the role corresponding to each user cluster, and then obtaining the API permissions corresponding to each role; generating a vertical privilege escalation task corresponding to each role according to the differences in API permissions between each role.
3. The method according to claim 2, wherein Generating a horizontal privilege escalation task for each API's basic information currently stored according to the differences in data permissions between different users corresponding to the basic information of the currently stored APIs. The specific implementation steps are as follows: For each API corresponding to the basic information of the currently stored APIs, perform the following steps: Randomly select two users with the same API permissions as the users corresponding to the basic information of the current API from the API permissions of users in the business system currently being counted, and record them as the first user and the second user; Obtain the data permissions that the first user has but the second user does not have as the first difference part; Obtain the data permissions that the second user has but the first user does not have as the second difference part; Generate a horizontal privilege escalation task for the second user based on the first difference part; Generate a horizontal privilege escalation task for the first user based on the second difference part.
4. The method according to claim 2, characterized in that Generating a vertical privilege escalation task corresponding to each role according to the differences in API permissions between each role. The method steps are as follows: Traverse each role in the clusters obtained by clustering that has not generated a vertical privilege escalation task; Take the currently traversed role as the first role; Randomly select another different role from the roles in the clusters obtained by clustering that have not generated a vertical privilege escalation task, and record it as the second role; Obtain the API permissions that the first role has but the second role does not have as the first difference permission; Obtain the API permissions that the second role has but the first role does not have as the second difference permission; Generate a vertical privilege escalation task for the second role based on the first difference permission; Generate a vertical privilege escalation task for the first role based on the second difference permission.
5. The method according to claim 2, wherein Execute the logical vulnerability detection task to perform logical vulnerability detection and obtain detection response data, including: Judge whether there is a logical vulnerability detection task generated currently: If not, do not perform logical vulnerability detection; If so, then: when there is an unauthorized detection task generated, execute the generated unauthorized detection task to perform logical vulnerability detection and obtain detection response data; when there is a horizontal authorization detection task generated, execute the generated horizontal authorization detection task to perform logical vulnerability detection and obtain detection response data; when there is a vertical privilege escalation detection task generated, execute the generated vertical privilege escalation detection task to perform logical vulnerability detection and obtain detection response data.
6. The method according to claim 5, wherein Execute the unauthorized detection task to perform logical vulnerability detection and obtain detection response data, including: Generate an HTTP request based on the url, requestHeader request header, requestBody request body, and request method in the unauthorized detection task; Send the HTTP request to obtain probe response data.
7. The method according to claim 5, wherein Execute the generated horizontal privilege escalation probe task for logical vulnerability detection to obtain probe response data, including: For each generated horizontal privilege escalation probe task, monitor whether the corresponding user is in the user logged-in state; For each horizontal privilege escalation probe task whose corresponding user is not in the user logged-in state, no detection is performed; For each horizontal privilege escalation probe task whose corresponding user is in the user logged-in state, put the user credentials of its corresponding user into the requestHeader of its request, and then generate an HTTP request together with the url, requestBody, and request method of the horizontal privilege escalation probe task with the requestHeader containing the user credentials. Then send the HTTP request using the corresponding request method to obtain probe response data.
8. The method according to claim 5, wherein Execute the generated vertical privilege escalation probe task for logical vulnerability detection to obtain probe response data, including: For each generated vertical privilege escalation probe task, perform the following steps respectively: Monitor whether there is a user in the user logged-in state under the role corresponding to the vertical privilege escalation probe task: If not, no detection is performed; If so, randomly select one of the users in the user logged-in state, obtain its user credentials, and put the user credentials into the requestHeader of the vertical privilege escalation probe task. Then generate an HTTP request together with the url, requestBody, and request method of the vertical privilege escalation probe task with the requestHeader, and then send the HTTP request to obtain probe response data.
9. The method according to claim 5, wherein Perform network traffic logical vulnerability analysis based on the response data, including: For the probe response data of each obtained unauthorized probe task, input the probe response data into the vulnerability analysis model for analysis, and output whether there is a privilege escalation vulnerability in the API corresponding to the unauthorized probe task; For the probe response data of each obtained horizontal privilege escalation task, input the probe response data into the vulnerability analysis model for analysis, and output whether there is a privilege escalation vulnerability in the API corresponding to the horizontal privilege escalation task; For the probe response data of each obtained vertical privilege escalation task, input the probe response data into the vulnerability analysis model for analysis, and output whether there is a privilege escalation vulnerability in the API corresponding to the vertical privilege escalation task.
10. A network traffic logic vulnerability analysis system, characterized in that, The system includes: The first module is used to collect the mirror traffic data of the target network in real time, preprocess the collected mirror traffic data to obtain the preprocessed mirror traffic data as the original data; the data format of the original data is json format, and the parameters are url, requestHeader, requestBody, responseHeader, and responseBody respectively; it is also used to extract and store the basic information of the API in the original data, and extract and store the business system to which the collected mirror traffic data of the target network belongs according to the extracted basic information of the API; The second module is used to extract requestHeader, requestBody, responseHeader, and responseBody from the original data, and extract words from the requestHeader, requestBody, responseHeader, and responseBody of the extracted original data to construct the first dictionary; extract the url from the original data, and extract words from the extracted url of the original data to construct the second dictionary; The third module is used to extract the url, requestHeader, requestBody, and responseBody from the original data, then segment the extracted url, requestHeader, requestBody, and responseBody by case, and set different weights for each obtained segment to form a feature vector based on the first dictionary and the second dictionary. Then, input the feature vector into the trained API business analysis model for calculation, output the business label of the API in the original data, and input the feature vector into the trained API behavior analysis model for calculation, output the behavior label of the API in the original data; The fourth module is used to extract the requestHeader data from the original data, and input the extracted requestHeader data into the trained session analysis model for calculation, output the user credentials in the original data; The fifth module is used to extract the url, requestBody, and responseBody from the original data, then segment the extracted url, requestBody, and responseBody by case, and set different weights for each obtained segment to form a feature vector based on the first dictionary and the second dictionary. Then, input the formed feature vector into the trained login interface analysis model for calculation, output whether the API corresponding to the original data is a login interface; The sixth module is used when the login interface analysis model outputs that the API corresponding to the original data is a login interface: obtain and record the user credentials, username, and user login status of the login interface from the original data; then match the user credentials of the obtained login interface with the user credentials output by the session analysis model. If the match is successful, cache the API corresponding to the user credentials; if the match is unsuccessful, do not cache the API corresponding to the user credentials; The seventh module is used to statistically count the users and APIs of the stored business system in real time; It is also used to statistically count the API permissions and API parameters of the users in the stored business system in real time; The eighth module is used to generate a logical vulnerability detection task based on the data output by the session analysis model, the API business analysis model, and the API behavior analysis model, as well as the basic information of the stored APIs and the statistically counted data; The ninth module is used to execute the logical vulnerability detection task to detect logical vulnerabilities and obtain detection response data; The tenth module is used to perform logical vulnerability analysis of network traffic based on the response data.
Citation Information
Cited By
API security vulnerability dynamic detection method based on exception monitoring
CN120880751A