Network state inference and HTTP (Hyper Text Transport Protocol) fuzzy testing method based on large model
Through the network state inference method and fuzz testing technology based on large models, the problem that traditional methods are difficult to detect complex HTTP protocol states is solved, achieving more efficient vulnerability detection and stronger network security.
Patent Information
- Application Number
- CN202510275766.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional fuzz testing methods are difficult to cover complex HTTP protocol states and boundary conditions, resulting in inefficient vulnerability detection.
A large model-based network state inference method is used to perform fuzz testing in combination with the HTTP protocol. The method includes steps such as data collection and preprocessing, large model training and optimization, network state inference and fuzz testing.
Through the intelligent inference capability of the large model, the coverage and efficiency of vulnerability detection can be improved, false alarms and missed reports can be reduced, and the security of the HTTP protocol can be enhanced.
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network security, and in particular relates to a network status inference and HTTP protocol fuzzy testing method based on a large model. Background Art
[0002] With the rapid development of the Internet, HTTP protocol, as one of the most widely used network protocols, is of vital importance for its security. Traditional fuzz testing methods usually rely on predefined rules or randomly generated data, which makes it difficult to cover complex protocol states and boundary conditions, resulting in low vulnerability detection efficiency. In recent years, large models (such as GPT, BERT, etc.) have demonstrated strong capabilities in the fields of natural language processing and data analysis, and can learn complex patterns and states from massive data. Therefore, how to use large models to intelligently infer network states and combine fuzz testing technology to improve the security of the HTTP protocol has become an urgent problem to be solved. Summary of the invention
[0003] 1. Technical issues to be resolved
[0004] The technical problem to be solved by the present invention is: to design a network status inference method based on a large model, and to perform fuzzy testing in combination with the HTTP protocol to improve the coverage and efficiency of vulnerability detection.
[0005] (II) Technical solution
[0006] In order to solve the above technical problems, the present invention provides a network state inference and HTTP protocol fuzzy testing method based on a large model, comprising the following steps:
[0007] Step 1. Data collection and preprocessing
[0008] 1.1 Data Source
[0009] Network traffic data: capture HTTP request and response packets, including request headers, response headers, URL parameters, cookies, session IDs, status codes, and timestamps;
[0010] Log data: collects server logs, application logs, and audit logs to record user behavior, resource access, and abnormal events;
[0011] Contextual information: including user identity, IP address, geographic location, and device information;
[0012] 1.2 Data Preprocessing
[0013] Structured processing: converting raw data into a structured format for easy model input;
[0014] Feature extraction: Extract key features, including:
[0015] Session status: Session ID, Cookie, request time interval;
[0016] Authentication status: Authorization header, Token, login request;
[0017] Resource access permissions: URL path, HTTP method, response status code;
[0018] Data annotation: Manually or semi-automatically annotate data and label session status, authentication status, and resource access permissions;
[0019] Step 2. Large model training and optimization
[0020] 2.1 Model selection
[0021] Use a pre-trained large model as the base model;
[0022] Fine-tune the model based on the characteristics of network data so that it can understand protocol semantics and contextual relationships;
[0023] 2.2 Multi-task Learning
[0024] Task design, including:
[0025] Session state inference: predict whether the current request belongs to the same session;
[0026] Authentication status inference: determine whether the user has passed the authentication;
[0027] Resource access permission inference: predict whether a user has permission to access a specific resource;
[0028] Loss function design: Design an independent loss function for each task and perform joint optimization by combining weighted summation;
[0029] 2.3 Reinforcement Learning Optimization
[0030] Dynamically adjust model parameters through reinforcement learning to enable it to adapt to changes in the network environment;
[0031] Design a reward function with the following design principles: +1 reward for correctly inferring the session state; -1 penalty for incorrectly inferring the session state; +2 reward for detecting abnormal behavior;
[0032] Step 3. Use the large model obtained in step 2 to infer the network status
[0033] 3.1 Session State Inference
[0034] Input data: Session ID, Cookie, timestamp, and historical request sequence of the current request;
[0035] Model inference:
[0036] Use the sequence modeling capabilities of the big model to analyze the time intervals and contextual relationships between requests;
[0037] Determine whether the current request belongs to the same session, or whether there is a risk of session hijacking;
[0038] Output: session status;
[0039] 3.2 Authentication Status Inference
[0040] Input data: Authorization header, Token, login request, historical authentication records;
[0041] Model inference:
[0042] Analyze whether the request contains valid authentication information;
[0043] Combine historical data to determine whether the user has completed the authentication process;
[0044] Output result: authentication status;
[0045] 3.3 Resource Access Permission Inference
[0046] Input data: URL path, HTTP method, user role, historical access records;
[0047] Model inference:
[0048] Analyze the matching relationship between user roles and resource permissions;
[0049] Combined with the context information, determine whether the current request exceeds the authorized access;
[0050] Output: resource access permissions;
[0051] Step 4. Fuzz testing
[0052] 4.1 Test Case Generation
[0053] Based on the inference results of the large model obtained in step 3, generate test cases for specific network states, including:
[0054] For the "unauthenticated" state, generate a request containing an invalid Token;
[0055] For the "session abnormality" state, generate a request containing a forged Session ID;
[0056] For the "abnormal authority" status, a request for unauthorized access is generated;
[0057] 4.1 Test Case Mutation
[0058] Combine mutation strategies to generate diverse test cases;
[0059] 4.3 Test Execution and Monitoring
[0060] Send the generated test cases to the target server and monitor its response behavior;
[0061] Record key indicators such as response status code, response time, and error information;
[0062] 4.4 Vulnerability Analysis
[0063] Analyze response data using big models to identify potential vulnerabilities;
[0064] ·Use graph embedding technology to associate HTTP anomalies with underlying features such as TCP retransmission rate and TLS protocol version, and apply hypergraph neural network to perform association reasoning of multi-hop vulnerabilities;
[0065] 4.5 Vulnerability Detection and Report Generation
[0066] Combine the inference results of the large model with the test data to identify security vulnerabilities in the HTTP protocol;
[0067] Generate vulnerability reports, including vulnerability types, triggering conditions, and repair suggestions.
[0068] The present invention also provides a system for implementing the method.
[0069] The invention also provides a network security analysis method implemented based on the method.
[0070] The invention also provides a network security analysis method implemented based on the system.
[0071] (III) Beneficial effects
[0072] 1. Through intelligent inference of large models, the network status can be identified more accurately and the targetedness of fuzz testing can be improved.
[0073] 2. Combined with the generation capability of large models, it is possible to construct more complex test cases and cover more protocol boundary conditions.
[0074] 3. Improve vulnerability detection efficiency, reduce false positives and missed positives, and enhance the security of network protocols. DETAILED DESCRIPTION
[0075] In order to make the purpose, content and advantages of the present invention more clear, the specific implementation methods of the present invention are further described in detail below in conjunction with embodiments.
[0076] The present invention provides a network status inference method based on a large model, and combines the HTTP protocol to perform fuzzy testing to improve the coverage and efficiency of vulnerability detection. The main ideas of the method are as follows: deploy a large model training environment, collect and preprocess HTTP protocol data; train and optimize the large model so that it can accurately infer the network status; develop a fuzzy testing tool that integrates the inference and generation capabilities of the large model; perform fuzzy testing on the target server, record and analyze the test results; repair vulnerabilities based on the test results, and optimize the network protocol configuration.
[0077] The specific technical solutions are as follows:
[0078] 1. Network state inference method based on large model
[0079] 1.1 Data collection and preprocessing
[0080] Collect a large amount of HTTP protocol network traffic data, including request headers, response headers, status codes, timestamps, etc.
[0081] Clean and annotate the data and extract key features (such as request method, URL parameters, Cookie, Session, etc.).
[0082] Convert the data into a format suitable for large model input (such as text sequences or structured data).
[0083] 1.2 Large Model Training and Optimization
[0084] Use a pre-trained large model (such as GPT-4 or BERT) as the base model and fine-tune it based on the characteristics of the HTTP protocol.
[0085] Design a multi-task learning framework that enables the model to simultaneously learn protocol state inference, anomaly detection, and vulnerability prediction.
[0086] Optimize the model through reinforcement learning so that it can dynamically adapt to changes in the network environment.
[0087] 1.3 Network State Inference
[0088] Input current network traffic data and use the big model to infer network status (such as session status, authentication status, resource access rights, etc.).
[0089] Combine historical data and contextual information to predict likely protocol behaviors (such as redirection, timeout, error response, etc.).
[0090] Output inference results to provide guidance for fuzz testing.
[0091] Inference of network status (such as session status, authentication status, resource access rights, etc.) is the key to understanding system behavior, detecting anomalies, and discovering vulnerabilities. Traditional methods usually rely on rule engines or static analysis, which are difficult to cope with complex and changing network environments. Large models (such as GPT, BERT, etc.) can infer network status more intelligently with their powerful context understanding and pattern recognition capabilities. The following is a detailed implementation method:
[0092] 1. Data collection and preprocessing
[0093] 1.1 Data Source
[0094] Network traffic data: Capture HTTP request and response packets, including request headers, response headers, URL parameters, cookies, session IDs, status codes, timestamps, etc.
[0095] Log data: Collect server logs, application logs, and audit logs to record user behavior, resource access, and abnormal events.
[0096] Contextual information: including user identity, IP address, geographic location, device information, etc.
[0097] 1.2 Data Preprocessing
[0098] Structural processing: Convert raw data into a structured format (such as JSON or a table) for easy model input.
[0099] Feature extraction: Extract key features, such as:
[0100] οSession state: Session ID, Cookie, request time interval.
[0101] οAuthentication status: Authorization header, Token, login request.
[0102] ο Resource access permissions: URL path, HTTP method (GET / POST), response status code.
[0103] Data annotation: Manually or semi-automatically annotate data and label session status, authentication status, and resource access permissions.
[0104] 2. Large model training and optimization
[0105] 2.1 Model selection
[0106] Use pre-trained large models (such as GPT-4, BERT, or T5) as base models, which excel in natural language processing and sequence modeling.
[0107] Fine-tune the model based on the characteristics of network data so that it can understand protocol semantics and contextual relationships.
[0108] 2.2 Multi-task Learning
[0109] Task design:
[0110] oSession state inference: predict whether the current request belongs to the same session.
[0111] oAuthentication status inference: Determine whether the user has been authenticated (Authenticated).
[0112] oResource access permission inference: Predict whether a user has permission to access a specific resource.
[0113] Loss function: Design an independent loss function for each task and combine it with weighted sum for joint optimization.
[0114] 2.3 Reinforcement Learning Optimization
[0115] Dynamically adjust model parameters through reinforcement learning so that it can adapt to changes in the network environment.
[0116] Design reward functions, for example:
[0117] oReward +1 for correctly inferring the session state.
[0118] o Penalty -1 for incorrectly inferred session state.
[0119] oReward +2 for detecting unusual behavior.
[0120] 3. Network state inference
[0121] 3.1 Session State Inference
[0122] Input data: Session ID, Cookie, timestamp, and historical request sequence of the current request.
[0123] Model inference:
[0124] oUse the sequence modeling capability of the big model to analyze the time intervals and contextual relationships between requests.
[0125] oDetermine whether the current request belongs to the same session, or whether there is a risk of session hijacking.
[0126] Output: Session status (such as "new session", "continuous session", "session exception").
[0127] 3.2 Authentication Status Inference
[0128] Input data: Authorization header, Token, login request, historical authentication records.
[0129] Model inference:
[0130] oAnalyze whether the request contains valid authentication information (such as Token or Cookie).
[0131] o Combined with historical data, determine whether the user has completed the authentication process.
[0132] Output result: authentication status (such as "authenticated", "unauthenticated", "authentication expired").
[0133] 3.3 Resource Access Permission Inference
[0134] Input data: URL path, HTTP method, user role, historical access records.
[0135] Model inference:
[0136] oAnalyze the matching relationship between user roles and resource permissions.
[0137] oCombined with the context information, determine whether the current request exceeds the authorized access authority.
[0138] Output results: resource access permissions (such as "access allowed", "access denied", "permission exception").
[0139] 4. Application in fuzz testing
[0140] 4.1 Test Case Generation
[0141] Generate test cases for specific network states based on the inference results of the large model.
[0142] For example:
[0143] oFor the "unauthenticated" state, generate a request containing an invalid token.
[0144] oFor the "Session Abnormal" state, generate a request containing a forged Session ID.
[0145] oFor the "permission exception" status, a request for unauthorized access is generated.
[0146] Leverage the generation capabilities of large models to construct complex request parameters, header fields, and payload data.
[0147] 4.1 Test Case Mutation
[0148] Combine mutation strategies (such as boundary value analysis, random replacement, format destruction, etc.) to generate diverse test data.
[0149] Syntax mutation: node replacement based on the protocol syntax tree (AST Mutation);
[0150] Semantic variation: Generate contextual outliers using CodeBERT;
[0151] Timing mutation: Inject time race test cases.
[0152] ο
[0153] 4.3 Test Execution and Monitoring
[0154] Send the generated test cases to the target server and monitor its response behavior.
[0155] Record key indicators such as response status code, response time, and error information.
[0156] 4.4 Vulnerability Analysis
[0157] Utilize large models to analyze response data and identify potential vulnerabilities (such as session fixation, unauthorized access, authentication bypass, etc.).
[0158] ·Use graph embedding technology to associate HTTP anomalies with underlying features such as TCP retransmission rate and TLS protocol version, and apply HyperGNN to perform association reasoning of multi-hop vulnerabilities.
[0159] 4.5 Vulnerability Detection and Reporting
[0160] Combine the inference results of the large model with the test data to identify security vulnerabilities in the HTTP protocol (such as SQL injection, XSS, CSRF, etc.).
[0161] Generate detailed vulnerability reports, including vulnerability types, triggering conditions, and repair suggestions.
[0162] Provides visual test results to help security analysts quickly locate problems.
[0163] Application scenario: Detecting session fixation vulnerabilities
[0164] 1. Data input: Capture the Session ID of the user login request and subsequent requests.
[0165] 2. State inference: Use a large model to analyze changes in Session IDs to determine whether there is a risk of session fixation.
[0166] 3. Test case generation: Generate a request containing a fixed Session ID and try to hijack other users' sessions.
[0167] 4. Vulnerability mining: monitor server responses and identify session fixation vulnerabilities.
[0168] It can be seen that the present invention provides a network status inference method based on a large model, including steps such as data collection and preprocessing, large model training and optimization, and network status inference; using the large model to perform fuzzy testing on the HTTP protocol, including test case generation, test execution and monitoring, vulnerability detection and reporting, etc.; constructing complex test cases through the generation capability of the large model, and combining mutation strategies to improve test coverage. This method improves the security and vulnerability detection efficiency of the HTTP protocol through the intelligent inference capability of the large model combined with fuzz testing technology. The present invention has broad application prospects and is suitable for fields such as network security detection and protocol vulnerability mining.
[0169] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A network state inference and HTTP protocol fuzz testing method based on a large model, characterized in that: The following steps are involved: Step 1. Data collection and preprocessing 1.1 Data Source ●Network traffic data: capture HTTP request and response data packets, including request headers, response headers, URL parameters, cookies, session ID, status code, and timestamp; ●Log data: collect server logs, application logs and audit logs to record user behavior, resource access and abnormal events; ●Contextual information: including user identity, IP address, geographic location, and device information; 1.2 Data Preprocessing Structured processing: converting raw data into a structured format for easy model input; Feature extraction: Extract key features, including: Session status: SessionID, Cookie, request time interval; Authentication status: Authorization header, Token, login request; Resource access permissions: URL path, HTTP method, response status code; Data annotation: Manually or semi-automatically annotate data and label session status, authentication status, and resource access permissions; Step 2. Large model training and optimization 2.1 Model selection Use a pre-trained large model as the base model; Fine-tune the model based on the characteristics of network data so that it can understand protocol semantics and contextual relationships; 2.2 Multi-task Learning Task design, including: Session state inference: predict whether the current request belongs to the same session; Authentication status inference: determine whether the user has passed the authentication; Resource access permission inference: predict whether a user has permission to access a specific resource; Loss function design: Design an independent loss function for each task and perform joint optimization by combining weighted summation; 2.3 Reinforcement Learning Optimization Dynamically adjust model parameters through reinforcement learning to enable it to adapt to changes in the network environment; ●Design a reward function with the following design principles: +1 reward for correctly inferring the session state; -1 penalty for incorrectly inferring the session state; +2 reward for detecting abnormal behavior; Step 3. Use the large model obtained in step 2 to infer the network status 3.1 Session State Inference ●Input data: Session ID, Cookie, timestamp, and historical request sequence of the current request; ●Model inference: Use the sequence modeling capabilities of the big model to analyze the time intervals and contextual relationships between requests; Determine whether the current request belongs to the same session, or whether there is a risk of session hijacking; ●Output result: session status; 3.2 Authentication Status Inference ●Input data: Authorization header, Token, login request, historical authentication records; ●Model inference: Analyze whether the request contains valid authentication information; Combine historical data to determine whether the user has completed the authentication process; ●Output result: authentication status; 3.3 Resource Access Permission Inference ●Input data: URL path, HTTP method, user role, historical access records; ●Model inference: Analyze the matching relationship between user roles and resource permissions; Combined with the context information, determine whether the current request exceeds the authorized access; Output: resource access permissions; Step 4. Fuzz testing 4.1 Test Case Generation Based on the inference results of the large model obtained in step 3, generate test cases for specific network states, including: For the "unauthenticated" state, generate a request containing an invalid Token; For the "session abnormality" state, generate a request containing a forged SessionID; For the "abnormal permissions" state, a request for unauthorized access is generated; 4.1 Test Case Mutation ● Combine mutation strategies to generate diverse test cases; 4.3 Test Execution and Monitoring ●Send the generated test cases to the target server and monitor its response behavior; ●Record key indicators such as response status code, response time, and error information; 4.4 Vulnerability Analysis Analyze response data using big models to identify potential vulnerabilities; ●Use graph embedding technology to associate HTTP anomalies with underlying features such as TCP retransmission rate and TLS protocol version, and apply hypergraph neural network to perform association reasoning of multi-hop vulnerabilities; 4.5 Vulnerability Detection and Report Generation ● Combine the inference results of the large model with the test data to identify security vulnerabilities in the HTTP protocol; ●Generate vulnerability reports, including vulnerability types, triggering conditions, and repair suggestions.
2. The method according to claim 1, characterized in that In step 1.2, convert the raw data into JSON or table format.
3. The method according to claim 1, characterized in that The pre-trained large models are GPT-4, BERT or T5.
4. The method according to claim 1, characterized in that The valid authentication information contained in the request is a token or a cookie.
5. The method according to claim 1, characterized in that In step 4.1, the generation capability of the large model is also used to construct request parameters, header fields, and payload data.
6. The method according to claim 1, characterized in that The potential vulnerabilities include session fixation, unauthorized access, and authentication bypass.
7. The method according to claim 1, characterized in that The mutation strategy includes: Syntax variation: node replacement based on the protocol syntax tree; ●Semantic variation: Use CodeBERT to generate contextual outliers; ● Timing variation: Inject timing race conditions into test cases.
8. A system for implementing the method according to any one of claims 1 to 7.
9. A network security analysis method implemented based on the method according to any one of claims 1 to 7.
10. A network security analysis method implemented based on the system as claimed in claim 8.
Citation Information
Cited By
Hypertext transfer protocol analysis ambiguity vulnerability detection method and device
CN120434059A