Webpage attack detection method and device and computer equipment

By integrating the characteristics of web page code and log data, and using pre-trained detection models, traditional protection measures cannot cope with complex hybrid XSS attacks, achieving more efficient web attack detection and protection.

CN119945707APending Publication Date: 2025-05-06CHINA TELECOM CLOUD TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411765765.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Traditional XSS protection measures cannot effectively deal with new and complex hybrid XSS attack methods, resulting in increased attack concealment and harm.

Method used

By obtaining web code data and web log data, extracting and fusing web code features and web log features, inputting a pre-trained web attack detection model to output attack results.

Benefits of technology

Improve the accuracy and robustness of attack detection, reduce false alarms and missed reports, and improve web page protection capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119945707A_ABST
    Figure CN119945707A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of network security, and discloses a webpage attack detection method and device and computer equipment. The method comprises the following steps: obtaining webpage code data to be detected and webpage log data corresponding to the webpage code data; extracting corresponding webpage code features from the webpage code data, and extracting corresponding webpage log features from the webpage log data; fusing the webpage code features and the webpage log features to obtain a fused feature vector; and inputting the fusion feature vector into a pre-trained webpage attack detection model, and outputting a webpage attack result through the webpage attack detection model. By implementing the technical scheme of the invention, the potential attack behavior in the webpage can be accurately detected and identified by fusing the webpage code characteristics and the webpage log characteristics and utilizing the pre-trained webpage attack detection model, and the webpage security protection capability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network security, and in particular to a detection method, device and computer equipment for web page attacks. Background Art

[0002] With the development of Internet technology, Cross-Site Scripting (XSS) has gradually evolved into a workflow-based attack method. Attackers first collect vulnerable users and World Wide Web (Web) service information, and then develop targeted attack strategies to exploit the target's weaknesses to launch attacks. Hybrid XSS attacks can cleverly bypass traditional detection methods by combining different attack vectors such as the client and server, as well as other network attack technologies, thereby increasing the concealment and harmfulness of the attack. Summary of the invention

[0003] In view of this, the present invention provides a web page attack detection method, device and computer equipment to solve the problem that traditional XSS protection measures cannot effectively deal with new and complex hybrid XSS attack methods.

[0004] In a first aspect, the present invention provides a method for detecting web page attacks, comprising: obtaining web page code data to be detected and web page log data corresponding to the web page code data; extracting corresponding web page code features from the web page code data, and extracting corresponding web page log features from the web page log data; fusing the web page code features and the web page log features to obtain a fused feature vector; inputting the fused feature vector into a pre-trained web page attack detection model, and outputting the web page attack result through the web page attack detection model.

[0005] The webpage attack detection method provided by the embodiment of the present invention can fully mine the multi-dimensional information of webpage security risks and improve the accuracy and robustness of attack detection by comprehensively utilizing the features of webpage code data and webpage log data to perform feature fusion. Combined with the pre-trained attack detection model, it can more efficiently identify potential webpage attacks, reduce false positives and false negatives, and improve webpage protection capabilities.

[0006] In an optional embodiment, obtaining web page log data corresponding to web page code data includes: obtaining original web page log data, performing data cleaning on the original web page log data to obtain first log data; if there is an attack behavior record in the first log data, determining a first malicious keyword corresponding to the attack behavior record; extracting attack feature keywords from the first log data according to a preset extraction rule; and determining the first malicious keyword and the attack feature keyword as web page log data.

[0007] The webpage attack detection method provided by the embodiment of the present invention can effectively screen out malicious keywords and attack feature keywords related to attack behaviors by cleaning and analyzing the original webpage log data, thereby reducing the interference of irrelevant data and improving the accuracy and pertinence of log data. Therefore, the method can accurately extract key features related to attack behaviors, enhance the sensitivity and recognition ability of webpage attack detection, and help locate and respond to potential security threats more quickly.

[0008] In an optional implementation, extracting attack feature keywords from the first log data according to preset extraction rules includes: performing word segmentation processing on each first log data to obtain corresponding multiple log phrases; extracting attack feature keywords from the multiple log phrases according to the preset extraction rules.

[0009] The webpage attack detection method provided by the embodiment of the present invention can convert the log content into structured phrase information by performing word segmentation processing on the first log data, and extract potential attack feature keywords therefrom. Therefore, the method can flexibly apply preset extraction rules, effectively identify attack-related words or patterns, improve the accuracy and automation level of feature extraction, and thus detect webpage attack behaviors more quickly and accurately.

[0010] In an optional embodiment, obtaining web page code data to be detected includes: obtaining original web page code data of the web page to be detected, performing data cleaning on the original web page code data to obtain first code data; if there is a network attack behavior in the first code data, determining a second malicious keyword corresponding to the network attack behavior; based on attribute information of the first code data, establishing a control flow graph corresponding to the first code data; and determining the second malicious keyword and the control flow graph as the web page code data.

[0011] The webpage attack detection method provided by the embodiment of the present invention can effectively remove irrelevant information and improve data quality by cleaning the original webpage code data. Further, by analyzing the network attack behavior in the code and establishing a control flow graph in combination with the attribute information, the execution path and potential vulnerabilities of the webpage code can be deeply mined. By determining the combination of malicious keywords and control flow graphs, not only the ability to identify attack behaviors is enhanced, but also the detection accuracy of complex attack patterns is improved, thereby improving the overall effect of webpage security protection.

[0012] In an optional embodiment, based on the attribute information of the first code data, a control flow graph corresponding to the first code data is established, including: based on the grammatical structure of the first code data, converting the first code data into an abstract syntax tree; combining the abstract syntax trees based on the execution relationship between the first code data to generate a control flow graph.

[0013] The webpage attack detection method provided by the embodiment of the present invention can clearly represent the logical structure and execution path of the code by converting the first code data into an abstract syntax tree and generating a control flow graph based on the code execution relationship. Therefore, the method provides visualization of code behavior, making potential attack modes and vulnerabilities easier to identify, and helps to deeply analyze the operation logic of the code, thereby improving the detection and prevention capabilities of complex network attacks.

[0014] In an optional implementation, corresponding web page log features are extracted from web page log data, including: performing word vector training on the web page log data to obtain a first feature word vector; inputting the first feature word vector into a pre-trained text processing model, and generating web page log features through the text processing model.

[0015] The webpage attack detection method provided by the embodiment of the present invention can convert the words in the log into efficient numerical feature representations by training the word vectors of the webpage log data, thereby better capturing the semantic information in the log data. The generated word vectors are input into the pre-trained text processing model, and the powerful ability of the deep learning model can be used to further extract more accurate and efficient webpage log features. Therefore, the method improves the automation and accuracy of feature extraction, and enhances the ability to identify potential attack patterns and abnormal behaviors in webpage logs.

[0016] In an optional embodiment, corresponding web page code features are extracted from web page code data, including: performing word vector training on the web page code data to obtain a second feature word vector; inputting the second feature word vector into a pre-trained flow graph processing model, and generating web page code features through the flow graph processing model.

[0017] The web page attack detection method provided by the embodiment of the present invention converts the semantic information in the code into numerical features by training the word vector of the web page code data, thereby improving the expressiveness of the feature representation. By inputting the trained feature word vector into the pre-trained flow graph processing model, the control flow and execution path of the code can be analyzed more deeply, and the key features of the web page code can be accurately extracted. Therefore, the method can make full use of the structured information of the flow graph model, enhance the recognition and attack detection capabilities of complex code behaviors, and improve the efficiency and accuracy of web page security protection.

[0018] In an optional embodiment, the training steps of the web page attack detection model include: obtaining attack data of the web page that has been attacked and normal data of the web page that has not been attacked; performing data processing on the attack data and the normal data to obtain a sample feature vector; performing forward time series processing and backward time series processing on the sample feature vector respectively to obtain a bidirectional time series feature vector; performing feature regularization processing on the bidirectional time series feature vector to obtain a regularized feature vector; performing feature mapping on the regularized feature vector to obtain a predicted classification result of whether the web page has been attacked; determining an actual classification result of whether the web page has been attacked based on the label information of the attack data and the normal data; calculating a loss function based on the predicted classification result and the actual classification result, and when the loss function converges, a web page attack detection model is obtained.

[0019] The web page attack detection method provided by the embodiment of the present invention fully considers the time series characteristics of attack data and normal data, and can more accurately capture the dynamic changes and patterns of web page attacks. Feature regularization processing helps to reduce overfitting and improve the generalization ability of the model, while feature mapping enhances the model's ability to predict whether a web page is under attack. Through the optimization of the training loss function, the method can gradually improve the accuracy and stability of the model, and ultimately achieve efficient and accurate web page attack detection, ensuring that the system can effectively identify and respond to different types of attack behaviors.

[0020] In a second aspect, the present invention provides a detection device for web page attacks, comprising: an acquisition module, used to acquire web page code data to be detected and web page log data corresponding to the web page code data; an extraction module, used to extract corresponding web page code features from the web page code data, and to extract corresponding web page log features from the web page log data; a fusion module, used to fuse the web page code features and the web page log features to obtain a fused feature vector; and a detection module, used to input the fused feature vector into a pre-trained web page attack detection model, and output the web page attack result through the web page attack detection model.

[0021] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the web page attack detection method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0022] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to cause a computer to execute the web page attack detection method of the first aspect or any corresponding embodiment thereof.

[0023] In a fifth aspect, the present invention provides a computer program product, including computer instructions, where the computer instructions are used to enable a computer to execute the web page attack detection method of the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0025] Figure 1 is a flow chart of a method for detecting web page attacks according to an embodiment of the present invention;

[0026] Figure 2 is a flow chart of another webpage attack detection method according to an embodiment of the present invention;

[0027] Figure 3 is a flow chart of another webpage attack detection method according to an embodiment of the present invention;

[0028] Figure 4 is a schematic diagram of the structure of a BiLSTM network according to an embodiment of the present invention;

[0029] Figure 5 is a structural block diagram of a webpage attack detection device according to an embodiment of the present invention;

[0030] Figure 6 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0031] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0032] With the development of Internet technology, Web services have brought convenience to people, but also provided many opportunities for network attackers. Due to the complexity of Web application development and maintenance and insufficient security considerations, Web applications have become the main target of attackers, bringing risks such as information leakage to users. Network attacks can be divided into many types according to the attack methods and targets. As an authoritative reference in the field of Web security, the Open Web Application Security Project (OWASP) points out that XSS attacks are one of the most serious Web application security risks and have widespread harm.

[0033] The main purposes of XSS attacks include stealing cookies, hijacking sessions, phishing, and stealing access history. According to different attack methods, XSS attacks can be divided into three types: reflected XSS (non-persistent), stored XSS (persistent), and DOM XSS. Reflected XSS usually appears in advertising links, phishing emails, etc. Attackers steal user privacy by inducing users to click malicious links; stored XSS stores malicious scripts on the server side and triggers attacks when other users visit the page; DOM XSS is executed through front-end JavaScript rendering to bypass server-side detection. In order to evade detection, attackers often use obfuscation, encoding and transcoding to maliciously deform the attack payload, increasing the difficulty of detection.

[0034] At present, there are three main methods for detecting XSS attacks: static analysis, dynamic analysis, and deep learning-based detection. Static analysis detects vulnerabilities by analyzing the source code of Web applications. However, due to the variety and flexibility of XSS attack methods, it is easy to cause false positives and cannot detect DOM-type XSS attacks. Dynamic analysis simulates real attacks and analyzes server responses. Although it does not rely on source code, its effect is affected by attack vectors and crawler strategies, and it is easy to produce false negatives. Deep learning-based methods automatically extract features by training neural networks, which has high accuracy and efficiency, especially in big data environments. However, with the evolution of XSS attack technology, more complex hybrid XSS attack methods have emerged. Attackers usually collect target user and Web service information, design targeted attack strategies, and use different attack media (such as clients and servers) to launch attacks. Hybrid XSS attacks combine multiple XSS types and other network attack technologies, can bypass traditional detection methods, and have stronger concealment and greater harm.

[0035] At present, XSS detection methods based on deep learning face two major problems. First, hybrid XSS attacks can bypass existing detection systems through malicious obfuscation and combination with other attack techniques. Attackers may use external links or other network attack methods to distract the attention of detection tools, resulting in a decrease in the recognition accuracy of deep learning models, or even failure to identify attacks. Second, existing deep learning detection methods usually only focus on a certain attack medium (such as web log information or program source code), while hybrid XSS attacks often span multiple media, combined with user behavior analysis, malicious obfuscation technology and other attack methods, which increases the flexibility and concealment of attacks, making it difficult for existing detection methods to effectively deal with complex multi-level attacks.

[0036] In view of this, the technical solution of the present invention collects information by acquiring the code data of the web page to be detected and the corresponding web page log data. Features related to the web page code are extracted from the web page code data, and features related to the log are extracted from the web page log data. These two features are fused to form a fused feature vector. Finally, the fused feature vector is input into a pre-trained web page attack detection model, and the model outputs the detection result of whether there is an attack on the web page based on the input feature vector, thereby improving the accuracy and comprehensiveness of the detection.

[0037] According to an embodiment of the present invention, an embodiment of a method for detecting web page attacks is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0038] In this embodiment, a method for detecting web page attacks is provided, which can be used for computer devices, such as desktop computers, notebook computers, etc. Figure 1 is a flow chart of a method for detecting web page attacks according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0039] Step S101, obtaining web page code data to be detected and web page log data corresponding to the web page code data.

[0040] The web page code data to be detected refers to the source code of the web page to be analyzed, which may include, for example, HTML, JavaScript, CSS, etc., which are used to describe the structure, style, and function of the web page. The web page log data corresponding to the web page code data refers to the log data that records the user behavior, request response, and other information interacting with the web page, which may include, for example, user access time, request parameters, response status, access path, etc. Specifically, the page source code can be obtained by accessing the web page to be detected, and the relevant web page log data can be collected through the server-side log system or Web analysis tools.

[0041] Step S102: extracting corresponding web page code features from the web page code data, and extracting corresponding web page log features from the web page log data.

[0042] Web page code features refer to key attributes extracted from web page code data that help identify web page attacks, such as HTML tags, JavaScript code segments, CSS styles, script calls, DOM structures, input validation, etc. These features can reflect the structure and behavior of web pages. Web page log features refer to key indicators extracted from web page log data, such as access time, request method, user behavior (clicks, form submissions, etc.), IP address, access path, response status code, etc. These features help capture user interaction patterns and potential abnormal behaviors. Specifically, web page code features can be extracted by parsing the web page source code to extract the type, attributes, script content, etc. of HTML tags; while web page log features can be extracted by analyzing log files to extract information such as access frequency, time, and request parameters.

[0043] Step S103: fuse the web page code feature and the web page log feature to obtain a fused feature vector.

[0044] The fused feature vector refers to combining the web page code features and web page log features to form a unified feature representation. Specifically, the web page code features and web page log features are converted into numerical vectors after extraction, and can be fused by concatenating these two vectors or by some mathematical operation (such as weighted average, element-by-element addition, etc.) to obtain a fused feature vector. This fused feature vector contains the structural information of the web page (code features) and the user interaction behavior information (log features).

[0045] Step S104: input the fused feature vector into the pre-trained webpage attack detection model, and output the webpage attack result through the webpage attack detection model.

[0046] A pre-trained web attack detection model refers to a machine learning model that has been trained with a large amount of labeled data (such as feature samples of normal web pages and attack web pages). The model has learned how to identify web attacks based on the input feature vector. The web attack result refers to the judgment result output by the model after processing the input fused feature vector. For example, it can be a classification result indicating whether a web page is attacked, such as "attack" or "normal". Specifically, the fused feature vector is used as input data and inference is performed through the web attack detection model. The web attack detection model will analyze whether the feature vector is consistent with the known web attack pattern based on the rules it has learned, and finally output the web attack result.

[0047] The webpage attack detection method provided by the embodiment of the present invention can fully mine the multi-dimensional information of webpage security risks and improve the accuracy and robustness of attack detection by comprehensively utilizing the features of webpage code data and webpage log data to perform feature fusion. Combined with the pre-trained attack detection model, it can more efficiently identify potential webpage attacks, reduce false positives and false negatives, and improve webpage protection capabilities.

[0048] In this embodiment, a method for detecting web page attacks is provided, which can be used for computer devices, such as desktop computers, notebook computers, etc. Figure 2 is a flow chart of a method for detecting web page attacks according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:

[0049] Step S201, obtaining web page code data to be detected and web page log data corresponding to the web page code data.

[0050] Specifically, the above step S201 includes:

[0051] Step S2011, obtaining original web page code data of the web page to be detected, and performing data cleaning on the original web page code data to obtain first code data.

[0052] The original web page code data of the web page to be detected refers to the source code file of the web page, which defines the structure, style and behavior of the web page. The first code data refers to the web page code data after preliminary processing and cleaning, which may include removing irrelevant information, formatting code, deleting duplicate content, etc., so that the data is more standardized and neat, which is convenient for subsequent analysis. Specifically, after obtaining the original web page code data of the web page to be detected, the data is cleaned, including removing spaces, comments, redundant codes or irrelevant content, so as to obtain the first code data, ensuring that the subsequent feature extraction and attack detection can be carried out on a more accurate and efficient basis.

[0053] Step S2012: If there is a network attack behavior in the first code data, determine a second malicious keyword corresponding to the network attack behavior.

[0054] Network attack behavior refers to the behavior of attempting to destroy, steal, tamper with, or disrupt normal network services through malicious web page code or operations, such as cross-site scripting (XSS), SQL injection, malicious redirection, malicious code injection, etc. The second malicious keyword refers to a specific malicious identifier or keyword related to the identified network attack behavior. These keywords can be identified as signs of attack in the web page code, such as certain specific strings, commands, script tags, etc., which will be used to perform attack operations. Specifically, by analyzing the first code data, it is detected whether it contains certain known malicious code fragments, abnormal scripts or attack patterns. Based on the nature and characteristics of the attack, combined with the preset attack pattern or known malicious keyword library, the corresponding malicious keywords are automatically identified and marked. For example, if the code contains specific malicious scripts or SQL injection strings, these strings will be determined as second malicious keywords and used for further attack analysis and identification.

[0055] Step S2013: Based on the attribute information of the first code data, a control flow graph corresponding to the first code data is established; and the second malicious keyword and the control flow graph are determined as web page code data.

[0056] The attribute information of the first code data refers to various structured information contained in the web page code, such as HTML tags, script types, function calls, input and output paths, request methods (GET / POST), variable definitions, function dependencies, etc. These attribute information reflects the specific implementation logic and process of the web page code. The control flow graph constructs a graphical structure by analyzing the structure and logic of the first code data to represent the control path and data flow process of program execution. Each node of the control flow graph represents a basic block in the code (such as a function, loop body or conditional judgment), and the edge represents the order of execution or the relationship of data transfer. Through this graphical representation, the execution path of the code and potential security vulnerabilities or attack surfaces can be intuitively identified. Specifically, information such as function calls, branch structures, loop structures, etc. in the code are extracted, and a control flow graph is generated based on the execution order and data dependency of the code.

[0057] Furthermore, the malicious keyword (ie, the second malicious keyword) identified by analyzing the first code data is combined with a control flow graph constructed based on the code data as the web page code data.

[0058] In some optional implementations, the above step S2013 includes:

[0059] Step a1: convert the first code data into an abstract syntax tree based on the syntax structure of the first code data.

[0060] The grammatical structure of the first code data refers to the grammatical rules and organizational form in the first code data, including the arrangement and hierarchical relationship of labels, attributes, functions, control structures, etc. The abstract syntax tree (AST) is the result of representing the grammatical structure of the first code data in a tree structure. Each node represents a grammatical unit in the code (such as expression, variable, statement, etc.), and the hierarchy and branch structure of the tree reflect the logic and execution order of the code. Specifically, the first code data is restored using the maintained existing coding rule set, and then the variable names and function names in the first code data are uniformly named in the form of par_num and fun_num to avoid the impact of different development specifications on model performance. Then, the parser performs lexical analysis according to the grammatical structure of the first code data, and the code is disassembled into recognizable lexical units (such as labels, keywords, etc.). Through grammatical analysis, these lexical units are organized into a tree structure that conforms to language rules, namely an abstract syntax tree.

[0061] Step a2: Combine the abstract syntax trees based on the execution relationship between the first code data to generate a control flow graph.

[0062] The execution relationship between the first code data refers to the execution order, dependency and control flow of each code fragment (such as function, statement, conditional judgment, loop, etc.) in the web page code. For example, some code statements may be executed when a certain condition is met, while some code statements will be executed only under other conditions. These relationships determine the execution path of the code. Specifically, by analyzing the execution order and conditional judgment in the code, the call or execution relationship between different code blocks is determined and reflected on the basis of the abstract syntax tree. By analyzing the nodes in the abstract syntax tree and combining the code execution path, the nodes are connected to form a control flow graph, which shows how the various parts of the code in the program are executed in a certain logical order, thereby helping to identify potential malicious code behavior.

[0063] In the above implementation, by converting the first code data into an abstract syntax tree and generating a control flow graph based on the code execution relationship, the logical structure and execution path of the code can be clearly represented. Therefore, the method provides visualization of code behavior, making potential attack modes and vulnerabilities easier to identify, and helps to deeply analyze the operation logic of the code, thereby improving the detection and prevention capabilities of complex network attacks.

[0064] Step S2014, obtaining original web page log data, performing data cleaning on the original web page log data, and obtaining first log data.

[0065] The original web page log data refers to the log records corresponding to the original web page code data of the web page to be detected extracted from the log system or other monitoring tools. These log records contain the user's access request to the web page and the server response information. The first log data is the original web page log data after data cleaning. The cleaning process includes removing irrelevant information, formatting log data, filling missing values, etc., so as to extract effective features and potential attack behavior records. Specifically, the original web page log data can be obtained by accessing log files and monitoring tools. After obtaining the original web page log data, since the original web page log data may contain encoded content (such as URL encoding, Base64 encoding, hexadecimal encoding, Unicode encoding, ASCII code, etc.), it is necessary to restore these encodings in the original web page log data by maintaining the existing encoding rule set. For example, decode the Base64 encoded content back to the original data, or convert the hexadecimal encoding into readable text. If the log data contains complex nested encoding (such as multi-layer encoding) or special encoding formats (such as Jsfuck, Aaencode, etc.), it is necessary to decode and restore these special formats layer by layer to ensure that the information in the log data can be correctly interpreted. Furthermore, the restored web page log data is subjected to data denoising to remove redundant or irrelevant log information, such as invalid access records, repeated requests, blank fields or damaged data, to reduce interference with subsequent analysis. After these cleaning steps, the first log data is finally obtained.

[0066] Step S2015: If there is an attack behavior record in the first log data, determine a first malicious keyword corresponding to the attack behavior record.

[0067] Attack behavior records refer to records related to potential attack behaviors detected from the first log data, which reflect abnormal or malicious access patterns. The first malicious keyword refers to specific malicious features or keywords (such as malicious IP, specific attack symbols, etc.) extracted by analyzing these abnormal records. Specifically, if there are attack behavior records in the first log data, malicious keywords related to the attack behavior are identified therefrom according to preset rules or feature matching algorithms.

[0068] Step S2016, extracting attack feature keywords from the first log data according to a preset extraction rule; determining the first malicious keyword and the attack feature keyword as web page log data.

[0069] Preset extraction rules refer to rules for automatically identifying and extracting characteristic information related to attack behaviors from web page log data based on pre-set patterns or algorithms. These rules may include specific keyword matching, pattern recognition, abnormal access patterns, etc., which are used to screen out potential attack traces. Attack characteristic keywords refer to specific information that reflects attack behaviors in the first log data, such as malicious IP, abnormal request type, illegal parameters, etc. Specifically, when attack characteristic keywords are extracted from the first log data according to the preset extraction rules, the logs are scanned according to these rules to identify and extract keywords or data fragments that match the attack patterns. These keywords can be used to further analyze and confirm whether the web page is under attack.

[0070] Furthermore, the first malicious keyword is combined with the attack feature keyword to form web page log data, which provides a basis for subsequent web page attack detection.

[0071] In some optional implementations, the above step S2016 includes:

[0072] Step b1, performing word segmentation processing on each first log data to obtain a corresponding plurality of log phrases.

[0073] Multiple log phrases refer to separate words or phrases extracted from web page log data and formed after word segmentation. These phrases may include keywords, behavior records, IP addresses, timestamps, request types, and other information in web page access records. Specifically, the first log data (such as front-end language, back-end language, server system, etc.) will be split into multiple meaningful units, and keywords related to attacks, such as SQL injection, XSS, etc., will be retained. At the same time, special symbols will be separated by adding spaces (such as converting "&" to "&") to ensure that keywords and symbols are not confused.

[0074] Step b2: extract attack feature keywords from multiple log phrases according to preset extraction rules.

[0075] From these segmented log phrases, keywords related to attack features are extracted according to preset rules. For example, attack-related patterns such as "SQL injection" or "XSS" and other possible malicious behaviors (such as abnormal IP addresses or abnormal request paths, etc.) are identified. Through this rule screening, attack feature keywords can be extracted from the segmented log data, thus providing an effective basis for subsequent XSS or other attack detection.

[0076] In the above implementation, by performing word segmentation on the first log data, the log content can be converted into structured phrase information, from which potential attack feature keywords are extracted. Therefore, the method can flexibly apply preset extraction rules, effectively identify attack-related words or patterns, improve the accuracy and automation level of feature extraction, and thus detect web page attack behaviors more quickly and accurately.

[0077] Step S202: extract the corresponding web page code features from the web page code data, and extract the corresponding web page log features from the web page log data. Figure 1 Step S102 of the illustrated embodiment will not be described in detail here.

[0078] Step S203: fuse the web page code features and web page log features to obtain a fused feature vector. Figure 1 Step S103 of the illustrated embodiment will not be described in detail here.

[0079] Step S204: Input the fused feature vector into the pre-trained web attack detection model, and output the web attack result through the web attack detection model. Figure 1 Step S104 of the illustrated embodiment will not be described in detail here.

[0080] The web page attack detection method provided by the embodiment of the present invention can effectively screen out malicious keywords and attack feature keywords related to attack behaviors by cleaning and analyzing the original web page log data, thereby reducing interference from irrelevant data and improving the accuracy and pertinence of log data. Therefore, the method can accurately extract key features related to attack behaviors, enhance the sensitivity and recognition ability of web page attack detection, and help to locate and respond to potential security threats more quickly. By cleaning the original web page code data, irrelevant information can be effectively removed and data quality can be improved. Further, by analyzing the network attack behaviors in the code and establishing a control flow graph in combination with attribute information, the execution path and potential vulnerabilities of the web page code can be deeply excavated. By determining the combination of malicious keywords and control flow graphs, not only the recognition ability of attack behaviors is enhanced, but also the detection accuracy of complex attack patterns is improved, thereby improving the overall effect of web page security protection.

[0081] In this embodiment, a method for detecting web page attacks is provided, which can be used for computer devices, such as desktop computers, notebook computers, etc. Figure 3 is a flow chart of a method for detecting web page attacks according to an embodiment of the present invention. Figure 3 As shown, the process includes the following steps:

[0082] Step S301, obtain the web page code data to be detected and the web page log data corresponding to the web page code data. Figure 2 Step S201 of the illustrated embodiment will not be described in detail here.

[0083] Step S302: extracting corresponding web page code features from the web page code data, and extracting corresponding web page log features from the web page log data.

[0084] Specifically, the above step S302 includes:

[0085] Step S3021, perform word vector training on the web page code data to obtain a second feature word vector.

[0086] The second feature word vector is a vector representation obtained by training the web page code data with word vectors. Word vector training is the process of converting each element in the web page code (such as HTML tags, JavaScript code snippets, attribute names, etc.) into a digital vector, so that these code elements can be effectively represented in a high-dimensional space, thereby retaining their semantic information. Specifically, the web page code data is processed by a word vector training model (such as Word2Vec or other embedding models), and the text information such as keywords, function names, tags, etc. in the code is converted into the corresponding second feature word vector.

[0087] Step S3022: input the second feature word vector into the pre-trained flow graph processing model, and generate web page code features through the flow graph processing model.

[0088] A pre-trained graph processing model is a model that has been pre-trained on a large amount of data and is designed to process and analyze graph data. The graph processing model is used to process elements in web page code and the relationships between them. Specifically, the graph processing model will establish the relationship between nodes based on the input two-feature word vectors, and through graph convolutional neural networks (GNNs) or other graph processing technologies, combined with the connection information of nodes and edges, generate a high-dimensional feature representation of the web page code (i.e., web page code features), which can better describe the structure and potential attack characteristics of the web page code.

[0089] For example, if a graph convolutional neural network (GCN) is used to construct the topological relationship between flow graph nodes, then by increasing the number of GCN network layers, the scope of information propagation can be expanded, thereby learning the features of more distant neighboring nodes. Through weighted aggregation, GCN can effectively improve the learning ability of important features while reducing the impact of unimportant features. After each layer of convolution, the state of the current node will be updated according to the aggregated information until the convolution process converges and reaches equilibrium. Among them, the convolution operation of GCN can be expressed as the following formula:

[0090]

[0091] Among them, H(l) represents the output of the lth layer, A represents the adjacency matrix, is the sum of the adjacency matrix and the self-loop matrix, is the degree matrix, W is the weight matrix of weighted aggregation, and σ is the nonlinear activation function (ReLU function is generally used here). After the above convolution operation, the current node not only retains its own feature information, but also completes the aggregation of the feature information of adjacent nodes. In the program code segment, the effect of the current code segment is not only related to the context code, but also may have an effect on the code segment farther away, so GCN can well fit the feature extraction function of the program code segment.

[0092] Step S3023, perform word vector training on the web page log data to obtain a first feature word vector.

[0093] The first feature word vector refers to the digital vector extracted from the web page log data that represents the log content. These vectors reflect the semantic information of various elements in the web page log (such as IP address, request path, timestamp, etc.) and are used for subsequent web page attack detection. Specifically, the web page log data is preprocessed and segmented, and a word vector model (such as Word2Vec) is used to learn the semantic relationship between each word or element in the log, and each element is mapped to a high-dimensional vector, and finally the feature representation of the web page log is obtained, that is, the first feature word vector.

[0094] For example, web page code data and web page log data are input into the word2vec neural network for word vector training. The model framework uses the skip-gram model, which can predict the previous and next words of the word based on the current word. The optimal solution output is:

[0095]

[0096] Where θ represents the hyperparameter, w ij represents the i-th word of the j-th sentence, C ij Indicates w ij The formula indicates that in the current context, the model predicts that the current word is w ij After the training, the feature vector of each word segment can be obtained, and the feature vector of the node in the code graph can be obtained by integrating it into the control flow graph.

[0097] Step S3024: input the first feature word vector into the pre-trained text processing model, and generate web page log features through the text processing model.

[0098] A pre-trained text processing model refers to a model that has been trained with a large amount of text data and can effectively process and extract text features. The text processing model can convert the input text data into a feature representation suitable for the task requirements by learning the semantics, syntax and other information in the text data. Specifically, the first feature word vector obtained by word vector training of web log data is used as input. These vectors will be fed into the pre-trained model, and the model will generate web log features through forward propagation based on its existing training experience and semantic understanding. These features are highly abstract and representation of log data after model processing, and can be used for subsequent attack detection tasks.

[0099] Among them, the pre-trained text processing model can be, for example, a TextCNN model. The main structure of the TextCNN model includes: a word vector embedding layer, a convolution layer, a pooling layer, a fully connected layer and a softmax layer. The size of the word vector embedding layer is n×m, where n is the number of word vectors and m is the dimension of the word vector; the convolution layer is used to extract local features of the text, and by adjusting the size of the convolution kernel, features of texts of different sizes can be extracted; the pooling layer is used to reduce the dimensionality of feature data, thereby improving the learning rate and preventing the model from falling into an overfitting state.

[0100] Step S303: fuse the web page code feature and the web page log feature to obtain a fused feature vector.

[0101] The multi-head attention mechanism is used to perform weighted fusion operations on web page code features and web page log features. Specifically, the multi-head attention mechanism can guide the model to learn different semantic information of feature data from different angles. Through multi-dimensional analysis, it can better learn the representation of a certain feature in different dimensions. By fusing multi-dimensional information, the weight matrix of the input feature data can be obtained.

[0102] Step S304: input the fused feature vector into the pre-trained webpage attack detection model, and output the webpage attack result through the webpage attack detection model.

[0103] The web attack detection model can be, for example, a BiLSTM network, in which the fused feature vector is input into the BiLSTM network for bidirectional temporal feature analysis. Specifically, the BiLSTM network consists of a forward LSTM network and a backward LSTM network, and the LSTM network adds a forget gate and memory cells to the traditional RNN. The output of the forget gate indicates which useless information can be forgotten, thereby reducing the possibility of gradient explosion or gradient disappearance; the memory cell can learn past dependent information at the current moment and maintain memory of long-term dependent states, so that the BiLSTM network not only retains memory of past information, but also has dependence on future states. Among them, the structure of the BiLSTM network is as follows: Figure 4 As shown, the transfer function of each gate function is shown as follows:

[0104]

[0105] f t =σ(W f ·[h t-1 , x t ]+b f )

[0106] h t =o t *tanh(C t )

[0107] i t =σ(W i ·[h t-1 , x t ]+b i )

[0108] o t =σ(W o [h t-1 , x t ]+b o )

[0109] Among them C t is the current state of the cell, To indicate the temporary state of the cell, Indicates the part of the input information that is retained, f t is the output of the forget gate, h t is the final output of this layer, σ is the sigmoid activation function, W is the weight, and b is the bias.

[0110] Furthermore, since the fully connected layer in the BiLSTM network has many training parameters, the training rate will be slow. At the same time, due to the lack of a complete data set, overfitting is prone to occur when the amount of training data is small and the number of training parameters is large. Therefore, a Dropout layer can be added to combine with the fully connected layer to ignore some useless features so that the model will not be overly dependent on certain local features, thereby improving the generalization ability of the model. Finally, the output of the fully connected layer is input to the softmax layer and mapped to a probability value between (0, 1). The classification layer then completes the final XSS and non-XSS classification, and the final webpage attack result is obtained.

[0111] The detection method of web page attacks provided by the embodiment of the present invention can convert the vocabulary in the log into an efficient numerical feature representation by training the word vector of the web page log data, so as to better capture the semantic information in the log data. The generated word vector is input into the pre-trained text processing model, and the powerful ability of the deep learning model can be used to further extract more accurate and efficient web page log features. Therefore, the method improves the automation and accuracy of feature extraction, and enhances the ability to identify potential attack patterns and abnormal behaviors in web page logs. By training the word vector of the web page code data, the semantic information in the code is converted into numerical features, thereby improving the expressiveness of the feature representation. By inputting the feature word vector obtained by training into the pre-trained flow graph processing model, the control flow and execution path of the code can be analyzed more deeply, and the key features of the web page code can be accurately extracted. Therefore, the method can make full use of the structured information of the flow graph model, enhance the ability to identify and detect attacks of complex code behaviors, and improve the efficiency and accuracy of web page security protection.

[0112] In some optional implementations, the step of training the webpage attack detection model includes:

[0113] Step c1, acquiring attack data of a web page that has been attacked and normal data of a web page that has not been attacked.

[0114] Attack data includes instances of web page attacks, such as web pages affected by malicious behaviors such as malicious scripts, malicious HTML tags, XSS attacks (cross-site scripting attacks), etc. These attack data include inputs manipulated by attackers, such as malicious JavaScript code, malicious URLs, tampered web page content, etc. Normal data refers to web page data that has not been attacked in any way, and is web page content provided, managed, and maintained by legitimate users and developers, and does not contain malicious code and scripts.

[0115] The dataset is constructed based on the attack data of the web pages that have been attacked and the normal data of the web pages that have not been attacked. Specifically, the attack data can be obtained by crawling the XSSed website (a platform containing a large number of XSS attack records), which contains information such as the source code and URL of the web pages attacked by XSS. In addition, the legitimate website data provided by websites such as Dmoz and Alexa can be used as a source of normal data because of their high credibility and the characteristics of not being attacked maliciously. GitHub and open source projects also contain some malicious HTML tags, JavaScript scripts, etc., which can be used to simulate XSS attacks. At the same time, attack data can also be generated by simulating attacks, such as constructing malicious URLs or combining other attack methods (such as phishing, social engineering, etc.). In the process of constructing the dataset, each piece of data needs to be clearly marked whether it has been attacked. The attack data is marked as the "attack" class, and the normal data is marked as the "normal" class to ensure effective supervised learning in the subsequent training process.

[0116] Step c2: Process the attack data and normal data to obtain a sample feature vector.

[0117] Perform feature extraction on attack data and normal data, for example, extract key features from the source code, URL, Domain and other information of the web page. Specifically, you can extract structured and unstructured data such as HTML tags, JavaScript scripts, URL parameters, etc. in the web page, and build a feature vector containing tag type, tag position, parameter features, script content, etc. For attack data, you can further extract malicious script patterns, abnormal behaviors (such as XSS attack payloads), and potential signs of attacks. Then, merge the features corresponding to the attack data and the features corresponding to the normal data to obtain a sample feature vector.

[0118] Step c3, performing forward time series processing and backward time series processing on the sample feature vector respectively to obtain a bidirectional time series feature vector.

[0119] In forward time series processing, the sample feature vector is input into the network step by step in time, and the processing starts from the first time step, which depends on the output of the previous time step. Backward time series processing starts from the last time step and the processing of each time step depends on the state of the next time step. In this way, the forward and backward dependencies in the input data are captured, and finally bidirectional time series feature vectors are obtained. These feature vectors combine the information of the forward and backward directions in the time series, which improves the model's understanding and prediction capabilities of time series data.

[0120] Step c4, performing feature regularization processing on the bidirectional time series feature vector to obtain a regularized feature vector.

[0121] Use standardization or normalization methods to process each feature dimension to ensure that the scale of the feature values ​​is consistent and to avoid excessive influence of certain features on model training. Specifically, calculate the mean and standard deviation of each feature dimension, and use the mean and standard deviation to standardize the features (z-score standardization); or normalize by mapping the feature values ​​to a fixed interval (such as 0 to 1) to obtain a regularized feature vector.

[0122] Step c5, performing feature mapping on the regularized feature vector to obtain a prediction classification result of whether the web page is attacked.

[0123] The regularized feature vectors are feature mapped and processed by inputting them into a classification model. Specifically, the regularized feature vectors are transformed nonlinearly through one or more fully connected layers to obtain a prediction result for classification. The classification model learns the characteristics of whether the web page is attacked based on the pattern of the feature vector, and finally outputs a binary prediction result, that is, whether the web page is attacked by XSS. This process can generate a predicted probability value through an activation function (such as Sigmoid or Softmax), indicating the possibility of whether the web page is attacked, and then make a classification judgment.

[0124] Step c6, based on the tag information of the attack data and the normal data, determine the actual classification result of whether the web page is attacked.

[0125] Attack data should be labeled with an "attack" label (e.g., 1), while normal data should be labeled with a "normal" label (e.g., 0). By comparing the label of each sample with its feature vector, the model can identify which web pages in the training data are attacked and which are not, thus providing a true classification basis for model training.

[0126] Step c7, based on the predicted classification result and the actual classification result, the loss function is calculated, and when the loss function converges, a webpage attack detection model is obtained.

[0127] The loss function is calculated by comparing the difference between the model prediction result and the actual label (i.e., the actual classification of whether the web page is attacked). The loss function can be, for example, cross entropy loss, mean square error, etc. Specifically, for each sample, the loss function is calculated based on the difference between the prediction result (such as "attack" or "normal" predicted by the model) and the true label ("attack" or "normal"). The value of the loss function reflects the prediction performance of the model. The smaller the loss, the more accurate the model. Through the iterative optimization process, the parameters of the model are continuously adjusted until the loss function value converges to the minimum value, that is, the training of the model reaches the optimal state, thereby obtaining the final web attack detection model.

[0128] The web page attack detection method provided by the embodiment of the present invention fully considers the time series characteristics of attack data and normal data, and can more accurately capture the dynamic changes and patterns of web page attacks. Feature regularization processing helps to reduce overfitting and improve the generalization ability of the model, while feature mapping enhances the model's ability to predict whether a web page is under attack. Through the optimization of the training loss function, the method can gradually improve the accuracy and stability of the model, and ultimately achieve efficient and accurate web page attack detection, ensuring that the system can effectively identify and respond to different types of attack behaviors.

[0129] In this embodiment, a detection device for web page attacks is also provided, which is used to implement the above embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0130] This embodiment provides a webpage attack detection device. Figure 5 As shown, including:

[0131] The acquisition module 501 is used to acquire web page code data to be detected and web page log data corresponding to the web page code data;

[0132] An extraction module 502, for extracting corresponding web page code features from web page code data, and extracting corresponding web page log features from web page log data;

[0133] A fusion module 503 is used to fuse the web page code feature and the web page log feature to obtain a fusion feature vector;

[0134] The detection module 504 is used to input the fused feature vector into the pre-trained web attack detection model, and output the web attack result through the web attack detection model.

[0135] In some optional implementations, the acquisition module 501 includes:

[0136] A first acquisition submodule is used to acquire original webpage log data, perform data cleaning on the original webpage log data, and obtain first log data;

[0137] A first determination submodule, configured to determine a first malicious keyword corresponding to the attack behavior record if there is an attack behavior record in the first log data;

[0138] An extraction submodule, used to extract attack feature keywords from the first log data according to a preset extraction rule;

[0139] The second determination submodule is used to determine the first malicious keyword and the attack feature keyword as web page log data.

[0140] In some optional implementations, the extraction submodule includes:

[0141] A processing unit, configured to perform word segmentation processing on each first log data to obtain a plurality of corresponding log phrases;

[0142] The extraction unit is used to extract attack feature keywords from multiple log phrases according to preset extraction rules.

[0143] In some optional implementations, the acquisition module 501 further includes:

[0144] The second acquisition submodule is used to acquire original web page code data of the web page to be detected, and perform data cleaning on the original web page code data to obtain first code data;

[0145] A third determination submodule, configured to determine a second malicious keyword corresponding to the network attack behavior if there is a network attack behavior in the first code data;

[0146] Establishing a submodule for establishing a control flow graph corresponding to the first code data based on the attribute information of the first code data;

[0147] The fourth determination submodule is used to determine the second malicious keyword and the control flow graph as web page code data.

[0148] In some optional implementations, establishing a submodule includes:

[0149] a conversion unit, configured to convert the first code data into an abstract syntax tree based on a syntax structure of the first code data;

[0150] A generating unit is used to combine the abstract syntax trees based on the execution relationship between the first code data to generate a control flow graph.

[0151] In some optional implementations, the extraction module 502 includes:

[0152] A first training submodule is used to perform word vector training on web page log data to obtain a first feature word vector;

[0153] The first generating submodule is used to input the first feature word vector into the pre-trained text processing model, and generate web page log features through the text processing model.

[0154] In some optional implementations, the extraction module 502 further includes:

[0155] The second training submodule is used to perform word vector training on the webpage code data to obtain a second feature word vector;

[0156] The second generating submodule is used to input the second feature word vector into the pre-trained flow graph processing model, and generate web page code features through the flow graph processing model.

[0157] In some optional implementations, the detection module 504 includes:

[0158] The third acquisition submodule is used to acquire attack data of a web page that has been attacked and normal data of a web page that has not been attacked;

[0159] The first processing submodule is used to process the attack data and normal data to obtain a sample feature vector;

[0160] The second processing submodule is used to perform forward time series processing and backward time series processing on the sample feature vector respectively to obtain a bidirectional time series feature vector;

[0161] The third processing submodule is used to perform feature regularization processing on the bidirectional time series feature vector to obtain a regularized feature vector;

[0162] A mapping submodule is used to perform feature mapping on the regularized feature vector to obtain a prediction classification result of whether the web page is attacked;

[0163] Based on the label information of attack data and normal data, determine the actual classification result of whether the web page is attacked;

[0164] The detection submodule is used to calculate the loss function based on the predicted classification results and the actual classification results. When the loss function converges, a web page attack detection model is obtained.

[0165] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0166] The web attack detection device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0167] The webpage attack detection device provided by the embodiment of the present invention can fully mine the multi-dimensional information of webpage security risks and improve the accuracy and robustness of attack detection by comprehensively utilizing the features of webpage code data and webpage log data to perform feature fusion. Combined with the pre-trained attack detection model, it can more efficiently identify potential webpage attacks, reduce false positives and false negatives, and improve webpage protection capabilities.

[0168] The embodiment of the present invention also provides a computer device having the above Figure 5 A detection device for web page attacks is shown.

[0169] See also Figure 6 , Figure 6 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 6 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 6 A processor 10 is taken as an example.

[0170] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0171] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0172] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0173] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0174] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 6 The example of connecting through bus is taken in the following.

[0175] The input device 30 can receive input digital or character information, and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.

[0176] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0177] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.

[0178] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for detecting web page attacks, characterized in that: The method comprises: Acquire web page code data to be detected and web page log data corresponding to the web page code data; Extracting corresponding web page code features from the web page code data, and extracting corresponding web page log features from the web page log data; Fusing the web page code feature and the web page log feature to obtain a fused feature vector; The fused feature vector is input into a pre-trained webpage attack detection model, and the webpage attack result is output through the webpage attack detection model.

2. The method according to claim 1, characterized in that Obtaining web page log data corresponding to the web page code data includes: Acquire original webpage log data, and perform data cleaning on the original webpage log data to obtain first log data; If there is an attack behavior record in the first log data, determining a first malicious keyword corresponding to the attack behavior record; Extracting attack feature keywords from the first log data according to a preset extraction rule; The first malicious keyword and the attack feature keyword are determined as the web page log data.

3. The method according to claim 2, characterized in that The step of extracting attack feature keywords from the first log data according to a preset extraction rule includes: Performing word segmentation processing on each of the first log data to obtain a corresponding plurality of log phrases; The attack feature keywords are extracted from the multiple log phrases according to the preset extraction rule.

4. The method according to claim 1, characterized in that: Get the web page code data to be tested, including: Acquire original webpage code data of the webpage to be detected, and perform data cleaning on the original webpage code data to obtain first code data; If there is a network attack behavior in the first code data, determining a second malicious keyword corresponding to the network attack behavior; Based on the attribute information of the first code data, establishing a control flow graph corresponding to the first code data; The second malicious keyword and the control flow graph are determined as the web page code data.

5. The method according to claim 4, characterized in that The step of establishing a control flow graph corresponding to the first code data based on the attribute information of the first code data includes: Based on the grammatical structure of the first code data, converting the first code data into an abstract syntax tree; The abstract syntax trees are combined based on the execution relationship between the first code data to generate the control flow graph.

6. The method according to claim 1, characterized in that Extracting corresponding web page log features from the web page log data includes: Performing word vector training on the webpage log data to obtain a first feature word vector; The first feature word vector is input into a pre-trained text processing model, and the web page log feature is generated through the text processing model.

7. The method according to claim 1, characterized in that Extracting corresponding web page code features from the web page code data includes: Performing word vector training on the webpage code data to obtain a second feature word vector; The second feature word vector is input into a pre-trained flow graph processing model, and the web page code feature is generated by the flow graph processing model.

8. The method according to claim 1, characterized in that The training steps of the webpage attack detection model include: Obtain attack data of a web page that has been attacked and normal data of a web page that has not been attacked; Performing data processing on the attack data and the normal data to obtain a sample feature vector; Performing forward time series processing and backward time series processing on the sample feature vector respectively to obtain a bidirectional time series feature vector; Performing feature regularization processing on the bidirectional time series feature vector to obtain a regularized feature vector; Perform feature mapping on the regularized feature vector to obtain the predicted classification result of whether the web page is attacked; Determine an actual classification result of whether a webpage is attacked based on the tag information of the attack data and the normal data; Based on the predicted classification result and the actual classification result, a loss function is calculated, and when the loss function converges, the webpage attack detection model is obtained.

9. A webpage attack detection device, characterized in that: The device comprises: An acquisition module, used to acquire web page code data to be detected and web page log data corresponding to the web page code data; An extraction module, used to extract corresponding web page code features from the web page code data, and to extract corresponding web page log features from the web page log data; A fusion module, used for fusing the web page code feature and the web page log feature to obtain a fusion feature vector; The detection module is used to input the fused feature vector into a pre-trained webpage attack detection model, and output the webpage attack result through the webpage attack detection model.

10. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the web page attack detection method according to any one of claims 1 to 8 by executing the computer instructions.

Citation Information

Cited By

  • Traffic analysis and attack detection method and device, terminal equipment and storage medium

    CN121530710A