Intelligent XSS attack detection and defense method and equipment based on knowledge graph
By constructing an intelligent XSS attack detection method based on knowledge graphs, using a large language model to extract multi-dimensional attack construction elements and construct a graph structure, the problems of weak semantic association and low degree of automation of XSS attack detection in the existing technology are solved, and efficient and intelligent XSS attack detection and defense are achieved.
Patent Information
- Application Number
- CN202510762036.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the XSS attack payload detection method has problems such as weak semantic correlation, difficulty in adapting to complex and diverse attack payloads, lack of effective reasoning and prediction capabilities, and low degree of automation.
Using a knowledge graph-based method, the XSS attack payload is preprocessed and multi-dimensional structural factor extraction is performed through a large language model, entity-relationship-entity data in the form of triple-tuple, and imported it into the graph database to form a graph structure of nodes and edges, and integrated into the Web application firewall for real-time monitoring and defense.
It has realized the intelligent and automated upgrade of XSS attack detection, which can efficiently handle complex attack payloads, adapt to diverse attack methods, improve the automation and intelligence level of detection, and adapt to the ever-changing threat environment through continuous update mechanisms, improving the security and credibility of web applications.
Smart Images

Figure CN120342773A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of network security and artificial intelligence, and in particular to an intelligent detection and defense method and device for XSS attacks based on a knowledge graph. Background Art
[0002] With the wide popularization of Internet applications and the continuous evolution of Web technologies, Web application security issues have become increasingly severe. Among them, cross-site scripting attack (Cross-Site Scripting, abbreviated as XSS), as one of the most common and highly harmful attack methods in Web attacks, has ranked among the top of the security risk list for many consecutive years.
[0003] In recent years, with the wide application of Web 2.0 technologies and front-end frameworks, the interactivity and dynamics of Web pages have been enhanced, and XSS attack techniques have also continuously evolved. There are currently three types of XSS attacks, namely reflected XSS, stored XSS, and DOM-based XSS. The essence of these three types of XSS is that the XSS attack payload is used as the core carrier of the attack. Therefore, the detection and analysis of XSS attack payloads have become an important direction in current network security research, and its accurate identification is crucial for improving the security of Web applications and constructing a highly robust defense mechanism. However, in the face of the increasingly complex combination patterns and diverse obfuscation strategies of XSS attack payloads, how to construct a structured knowledge system, extract high-value attack features, and achieve intelligent identification of complex payloads remains a major problem in the current technical field.
[0004] Currently, there are several key defects in the methods for constructing XSS attack payload knowledge graphs: First, the semantic association is weak. Existing methods usually associate payloads with keywords or simple rules, without deeply capturing the context semantics, attack intentions, or multi-stage attack logics. Second, it is difficult to adapt to complex and diverse attack payloads. Rule-based methods rely on predefined rules or templates and are difficult to adapt to complex texts and diverse relationship expressions, with poor scalability. Third, there is a lack of effective reasoning and prediction capabilities. Traditional knowledge graph reasoning technologies mainly mine the laws of most data, but attack data only accounts for a small proportion of normal system behaviors. Using these methods for reasoning and prediction often results in the behaviors of normal users, which is contrary to the goal of threat detection. Fourth, the degree of automation is low. It relies on manual annotation and limited data sets and is difficult to achieve efficient expansion through NLP or dynamic sandboxes. Summary of the Invention
[0005] In order to solve the problems of weak semantic association and difficulty in adapting to complex and diverse attack payloads in the prior art, the primary object of the present invention is to provide an intelligent detection and defense method for XSS attacks based on a knowledge graph with strong semantic association, which can more efficiently process complex attack payloads and adapt to diverse attack methods.
[0006] To achieve the above object, the present invention adopts the following technical solutions: An intelligent detection and defense method for XSS attacks based on a knowledge graph, the method comprising the following steps in sequence:
[0007] (1) Collect original XSS attack payloads from a publicly available XSS attack sample library; at the same time, use an automated crawler and penetration testing tools to test a Web vulnerability target range to obtain XSS attack payloads that can trigger XSS vulnerabilities, and jointly form an XSS attack payload sample data set from the original XSS attack payloads and the XSS attack payloads that can trigger XSS vulnerabilities;
[0008] (2) Preprocess the XSS attack payload sample data in the XSS attack payload sample data set to obtain a preprocessed XSS attack payload sample data set;
[0009] (3) Use a large language model to extract the elements of the XSS attack payloads in the preprocessed XSS attack payload sample data set to obtain multi-dimensional attack construction elements;
[0010] (4) According to the multi-dimensional attack construction elements, structure the abstract payload information into a graph form, establish edges for the extracted multi-dimensional attack construction elements, and obtain data in the form of triples, i.e., entity-relationship-entity data;
[0011] (5) Import the data in the form of triples into a graph database to form a graph structure of nodes and edges, and complete the construction of the entire XSS attack payload knowledge graph;
[0012] (6) Integrate the XSS attack payload knowledge graph into a Web application firewall to achieve real-time monitoring and intelligent defense of incoming requests in a Web system.
[0013] Step (1) specifically includes the following steps in sequence:
[0014] (1a) Collect original XSS attack payloads from a publicly available XSS attack sample library, and the publicly available XSS attack sample library includes the XSSed database and the OWASP XSS Attack and Defense Cheat Sheet;
[0015] (1b) Use an automated crawler and penetration testing tools to test several Web vulnerability target ranges, namely Firing-Range, WAVSEP, and OWASPBenchmark, and collect XSS attack payloads that can trigger XSS vulnerabilities;
[0016] (1c) The original XSS attack payloads and the XSS attack payloads that can trigger XSS vulnerabilities jointly form an XSS attack payload sample data set.
[0017] Step (2) specifically refers to: The preprocessing includes verification and annotation. The verification means constructing a test page in the browser, inserting the collected XSS attack payloads, and observing whether malicious code is triggered. The annotation means labeling the verified data into two categories: triggerable attack payloads and non-triggerable attack payloads. Among them, the triggerable attack payloads refer to XSS attack payloads that can successfully trigger malicious code on the browser side, and the non-triggerable attack payloads refer to XSS attack payloads that fail to trigger malicious code.
[0018] Step (3) specifically includes the following steps in sequence:
[0019] (3a) Use a large language model to identify XSS attack payloads, disassemble, annotate, and classify them step by step, so that a complex XSS injection statement is structured into a clear set of nodes. The XSS injection statement is the XSS attack payload statement in the preprocessed XSS attack payload sample dataset.
[0020] (3b) Extract multi-dimensional attack construction elements from the XSS injection statement. The multi-dimensional attack construction elements include the original payload, malicious code, HTML tags, events, attributes, encoding techniques, and obfuscation techniques. The original payload is the complete string injected by the attacker into the input point of the Web page. The encoding technique is used to hide the malicious code in an unreadable form to bypass the security detection mechanism of the WAF. The obfuscation technique refers to the attacker's interference processing on the structure or semantics of the XSS malicious payload.
[0021] (3c) Assign a unique identifier to each multi-dimensional attack construction element to provide semantic support for graph reasoning and graph representation learning.
[0022] Step (4) specifically includes the following steps in sequence:
[0023] (4a) Name the multi-dimensional attack construction elements extracted by the large language model as Payload, Code, HTMLTag, Event, Attribute, Encoding, and Obfuscation respectively. Payload, Code, HTMLTag, Event, Attribute, Encoding, and Obfuscation correspond to the original payload, malicious code, HTML tags, events, attributes, encoding techniques, and obfuscation techniques one by one. Each multi-dimensional attack construction element represents a type of entity, and each type of entity is assigned a unique ID in the graph and has an extensible attribute structure.
[0024] (4b) Define the core relationship types between entities, which are used to describe the logical relationships between key elements; there are 6 types of the core relationship types. Among them, the first type is Payload - contains - Code, indicating that the XSS attack payload contains malicious code; the second type is Payload - uses - HTMLTag, indicating that the XSS attack payload uses HTML tags; the third type is HTMLTag - triggers - Event, indicating that the HTML tag of the XSS attack payload can trigger an event; the fourth type is HTMLTag - with - Attribute, indicating the structural relationship between the HTML tag of the XSS attack payload and the XSS attack payload attribute; the fifth type is Payload - encoded_with - Encoding, indicating that the XSS attack payload uses an encoding method; the sixth type is Payload - obfuscated_with - Obfuscation, indicating that the XSS malicious attack payload has been obfuscated.
[0025] (4c) Automatically establish edges for each pair of entities according to the core relationship types, and express them in the form of triples to obtain the triple - form data, that is, entity - relationship - entity data, and finally form a knowledge graph structure composed of entity nodes and semantic edges.
[0026] Step (5) specifically includes the following steps in sequence:
[0027] (5a) Import the triple - form data into the graph database to form a graph structure of nodes and edges; each entity type is modeled as a node in the graph, and each relationship is modeled as an edge connecting the nodes, thus completing the construction of the entire XSS attack payload knowledge graph.
[0028] (5b) Visualize and display the constructed XSS knowledge graph.
[0029] Step (6) specifically includes the following steps in sequence:
[0030] (6a) Analyze the HTTP request parameters, extract the fields that may carry the XSS attack payload, and send the extracted fields into the XSS attack payload knowledge graph for analysis.
[0031] (6b) Perform entity - level parsing on the extracted fields, identify the attack construction elements in the extracted fields, and extract the context relationship and core relationship type of the fields.
[0032] (6c) Compare the core relationship types extracted by the graph matching algorithm with the high-risk combinations in the XSS attack payload knowledge graph, where the high-risk combinations are the core relationship types in the XSS attack payload knowledge graph that are known to bypass the WAF; if the extracted core relationship types are the same as the high-risk combinations in the knowledge graph, it is determined that the extracted core relationship types have a high attack possibility, and corresponding response strategies are taken; the corresponding response strategies include intercepting requests, recording log information, sending the attack samples to the back-end audit system or the alarm module, and dynamically updating the XSS attack payload knowledge graph.
[0033] Another object of the present invention is to provide an electronic device, including:
[0034] A processor; and
[0035] A memory, in which computer program instructions are stored, and when the computer program instructions are run by the processor, the processor executes the above-mentioned intelligent detection and defense method for XSS attacks based on the knowledge graph.
[0036] The present invention also provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are run by a processor, the processor executes the above-mentioned intelligent detection and defense method for XSS attacks based on the knowledge graph.
[0037] As can be seen from the above technical solutions, the beneficial effects of the present invention are as follows: First, the present invention realizes the intelligent and automated upgrade of XSS attack detection. Through the knowledge graph construction technology driven by the large language model, the present invention breaks through the limitation of the traditional rule base relying on manual maintenance, and can automatically analyze the syntax structure, semantic intention and behavioral characteristics of XSS attack payloads, and transform them into computable graph nodes and relationships, with strong semantic associations, so as to process complex attack payloads more efficiently, adapt to diverse attack methods, and greatly improve the automation and intelligence level of XSS detection; Second, the present invention designs an automated data update mechanism, which can automatically perform parsing, entity extraction and relationship annotation as new XSS attack payloads are added, and update the data into the existing knowledge graph. This continuous update mechanism ensures the dynamic growth of the knowledge graph and the continuous enrichment of knowledge, and can continuously adapt to new attack methods and the changing threat environment; Third, the present invention constructs an XSS attack payload knowledge graph to organize key entities and their semantic relationships in a structured manner, realizing the expression and modeling of the essence of attack behaviors; this structured knowledge representation system not only helps to systematically accumulate security knowledge, but also lays a solid foundation for building a new generation of interpretable and traceable Web security defense systems. Through the knowledge graph, security personnel can more intuitively understand and analyze attack behaviors, so as to better formulate defense strategies and further improve the security and credibility of Web applications. Description of the Drawings
[0038] Figure 1 is the flowchart of the method of the present invention;
[0039] Figure 2 is the schematic diagram for extracting multi-dimensional attack construction elements;
[0040] Figure 3 is the schematic diagram for parsing the structure of XSS attack payloads;
[0041] Figure 4 is the schematic diagram of integrating the knowledge graph into the WAF module. Detailed Implementation Manner
[0042] As Figure 1 shown, an intelligent detection and defense method for XSS attacks based on a knowledge graph, the method includes the following steps in sequence:
[0043] (1) Collect original XSS attack payloads from a publicly available XSS attack sample library; at the same time, use automated crawlers and penetration testing tools to test a Web vulnerability test bed to obtain XSS attack payloads that can trigger XSS vulnerabilities. The XSS attack payload sample data set is composed of the original XSS attack payloads and the XSS attack payloads that can trigger XSS vulnerabilities;
[0044] (2) Preprocess the XSS attack payload sample data in the XSS attack payload sample data set to obtain a preprocessed XSS attack payload sample data set;
[0045] (3) Use a large language model to extract the elements of the XSS attack payloads in the preprocessed XSS attack payload sample data set to obtain multi-dimensional attack construction elements;
[0046] (4) According to the multi-dimensional attack construction elements, structure the abstract payload information into a graph form, establish edges for the extracted multi-dimensional attack construction elements, and obtain triple-form data, i.e., entity-relationship-entity data;
[0047] (5) Import the triple-form data into a graph database to form a graph structure of nodes and edges, and complete the construction of the entire XSS attack payload knowledge graph;
[0048] (6) Integrate the XSS attack payload knowledge graph into a Web application firewall to achieve real-time monitoring and intelligent defense of incoming requests in a Web system, so as to make up for the detection blind spots of traditional rule matching, keyword filtering, and blacklist mechanisms in the face of obfuscated and variant attacks.
[0049] Step (1) specifically includes the following steps in sequence:
[0050] (1a)Collect the original XSS attack payloads from the publicly available XSS attack sample libraries, where the publicly available XSS attack sample libraries include the XSSed database and the OWASP XSS Attack Prevention Cheat Sheet; these data samples are rich and diverse in form, and can be used to help the model understand basic features such as typical XSS attack structures, tags and attributes used.
[0051] (1b)Use automated crawlers and penetration testing tools to test several web vulnerability test beds, namely Firing-Range, WAVSEP, and OWASP Benchmark, and collect XSS attack payloads that can trigger XSS vulnerabilities.
[0052] (1c)The original XSS attack payloads and the XSS attack payloads that can trigger XSS vulnerabilities together constitute the XSS attack payload sample data set.
[0053] Step (2) specifically refers to: The preprocessing includes verification and annotation. The verification means constructing a test page in the browser, inserting the collected XSS attack payloads, and observing whether malicious code is triggered; the annotation means labeling the verified data into two categories: attack payloads that can be triggered and attack payloads that cannot be triggered. Among them, the attack payloads that can be triggered refer to the XSS attack payloads that can successfully trigger malicious code on the browser side, and the attack payloads that cannot be triggered refer to the XSS attack payloads that fail to trigger malicious code. The preprocessed XSS attack payload sample data set is characterized by strong representativeness and high sample quality, laying a solid data foundation for the subsequent structure extraction and knowledge graph construction of the large language model.
[0054] Step (3) specifically includes the following steps in sequence:
[0055] (3a)Use the large language model to identify XSS attack payloads, decompose, annotate, and classify them step by step, so that a complex XSS injection statement is structured into a clear set of nodes; the XSS injection statement is the XSS attack payload statement in the preprocessed XSS attack payload sample data set.
[0056] (3b)Extract multi-dimensional attack construction elements from the XSS injection statement, as Figure 2 shown. The multi-dimensional attack construction elements include the original payload, malicious code, HTML tags, events, attributes, encoding techniques, and obfuscation techniques; the original payload is the complete string injected by the attacker into the input point of the web page; the encoding technique is used to hide the malicious code in an unreadable form to bypass the security detection mechanism of the WAF; the obfuscation technique refers to the attacker's interference processing of the XSS malicious payload in terms of structure or semantics.
[0057] Malicious Code refers to JavaScript statements that can be executed on the browser side and pose security threats, such as alert(1), document.cookie, etc. Such code is usually embedded in the event responses of tags or <script>标签内。大语言模型通过上下文分析判断哪些片段为可执行JavaScript语句,排除无害文本或参数干扰,从而提取出恶意代码。
[0058] HTML标签(HTML Tag)是XSS攻击中常用的承载结构,例如 <script>、、<iframe>、<svg> 等,不同标签对应不同的注入媒介和触发机制。JavaScript结合HTML语法规则识别出攻击载荷中使用的具体标签。
[0059] 事件(Event)是HTML标签中用于绑定用户行为响应的属性,如 onerror、onclick、onload、onmouseover等,事件是触发恶意脚本执行的关键机制,大语言模型基于HTML事件词表识别标签中的事件。
[0060] 属性(Attribute)是HTML标签内部用于配置和传参的键值对,如 <img src=x> 中的 src。XSS攻击中属性既可被滥用注入脚本,也可作为诱导事件触发的关键点,大语言模型可通过语法树分析识别每个标签包含的属性。
[0061] 编码技术(Encoding)用于将恶意代码以不可读形式隐藏,以绕过WAF、IDS等安全检测机制。常见技术包括URL编码、HTML编码、Base64编码与Unicode编码等。大语言模型通过对比解码前后文本特征与已知模式匹配,识别其所使用的编码方式。
[0062] 混淆技术(Obfuscation)是指攻击者对XSS攻击载荷进行结构或语义上的干扰处理,以逃避过滤器检测。常见混淆技术包括事件名称替换、空白字符插入、大小写转换、标签嵌套等,一个XSS攻击载荷也会使用多种混淆技术。大语言模型利用自监督提示学习与大量混淆样本训练,在解析语法结构时可准确检测出XSS攻击负载所使用的混淆技术。
[0063] (3c)为每个多维度的攻击构造要素赋予唯一标识符,为图谱推理与图表示学习提供语义支撑。大语言模型具备跨编码、跨语义的泛化能力,能够有效应对不同攻击者风格与新型变种攻击样式。
[0064] 步骤(4)具体包括以下顺序的步骤:
[0065] (4a) 将大语言模型提取的多维度的攻击构造要素分别命名为Payload、Code、HTMLTag、Event、Attribute、Encoding以及Obfuscation,Payload、Code、HTMLTag、Event、Attribute、Encoding以及Obfuscation一一对应为原始载荷、恶意代码、HTML标签、事件、属性、编码技术与混淆技术;每个多维度的攻击构造要素代表一类实体,每类实体在图谱中赋予唯一ID,具备可扩展的属性结构;
[0066] (4b)定义实体间的核心关系类型,用于描述各关键要素之间的逻辑关系;所述核心关系类型有6类,其中,第一类为Payload-contains-Code,表示XSS攻击载荷中包含恶意代码;第二类为Payload-uses-HTMLTag,表示XSS攻击载荷使用HTML标签;第三类为HTMLTag-triggers-Event,表示XSS攻击载荷的HTML标签可触发某事件;第四类为HTMLTag-with-Attribute,表示XSS攻击载荷的HTML标签与XSS攻击载荷属性之间的结构性关系;第五类为Payload-encoded_with-Encoding,表示XSS攻击载荷使用编码方式;第六类为Payload-obfuscated_with-Obfuscation,表示XSS恶意攻击载荷经过混淆处理。这些关系不仅反映出XSS攻击负载的结构特征与行为逻辑,还揭示了不同攻击样式之间的潜在共性与演化路径。在图谱构建过程中,系统会自动为每对实体建立边,并通过三元组形式(实体-关系-实体)表达,最终形成由实体节点和语义边组成的知识图谱结构,为XSS攻击建模与风险评估提供图化支撑。
[0067] (4c)根据核心关系类型自动为每对实体建立边,并通过三元组形式表达,得到三元组形式数据即实体-关系-实体数据,最终形成由实体节点和语义边组成的知识图谱结构。
[0068] 实体是知识图谱中的基本组成单位,代表系统所需表达的"概念”或"对象”。本发明中,实体即为大语言模型提取的多维度的攻击构造要素。关系描述实体之间的语义联系,是知识图谱的语义核心。
[0069] 步骤(5)具体包括以下顺序的步骤:
[0070] (5a)将三元组形式的数据导入到图数据库中,形成节点与边的图结构;每个实体类型都被建模成图中的节点,每种关系都被建模为连接节点的边,从而完成整个XSS攻击载荷知识图谱的构建;
[0071] (5b)将构建完成的XSS知识图谱进行可视化展示。用户可以直观观看某一类XSS攻击载荷的攻击路径、混淆方法、常见标签组合等。例如,安全分析人员可追踪某个事件被哪些XSS攻击负载频繁利用、在那些HTML标签上出现最多,以及是否经常与某类属性相关联。这种图结构的信息表示方式大幅提升了XSS攻击模式的可解释性。
[0072] 步骤(6)具体包括以下顺序的步骤:
[0073] (6a)解析HTTP请求参数,提取可能携带XSS攻击载荷的字段,并将提取的字段送入XSS攻击载荷知识图谱中进行分析;
[0074] (6b)对提取的字段进行实体级解析,识别提取的字段中的攻击构造要素,并提取字段的上下文关系和核心关系类型;
[0075] (6c)通过图谱匹配算法对比提取的核心关系类型与XSS攻击载荷知识图谱中中的高风险组合,所述高风险组合为XSS攻击载荷知识图谱中已知能绕过WAF的核心关系类型;若提取的核心关系类型与知识图谱中的高风险组合相同,则判定提取的核心关系类型有高攻击可能性,并采取相应响应策略;所述相应响应策略包括拦截请求、记录日志信息、将攻击样本送入后端审计系统或告警模块,同时动态更新XSS攻击载荷知识图谱。
[0076] 同时,本发明设计了自动化的数据更新机制。随着新的XSS攻击负载的加入,本发明会自动对其进行解析、实体提取与关系标注,并将数据更新到已有知识图谱中,保证了图谱的持续成长与知识的不断丰富。
[0077] 图3展示XSS攻击载荷的结构,包含多个关键要素及其相互关系,这些要素通过特定的关系类型相互联系,共同构成了描述XSS攻击各要素之间逻辑关系的语义核心,是构建知识图谱的基础。
[0078] 图4中,Web流量首先被引入WAF即Web应用防火墙,WAF负责拦截请求并记录相关信息后告警,进行进一步的分析和处理。同时,WAF还负责将结构化数据发送至知识图谱,知识图谱中包含了良性和攻击两类数据,知识图谱利用这些数据构建图谱驱动的XSS攻击识别机制,以实现对Web系统中传入请求的实时监测与智能防御。
[0079] 综上所述,本发明实现了XSS攻击检测的智能化与自动化升级,通过大语言模型驱动的知识图谱构建技术,本发明突破了传统规则库依赖人工维护的局限性,能够自动解析XSS攻击载荷的语法结构、语义意图及行为特征,并转化为可计算的图谱节点与关系,从而更高效地处理复杂的攻击载荷,适应多样化的攻击方式,极大地提高了XSS检测的自动化和智能化水平;本发明设计了自动化的数据更新机制,能够随着新的XSS攻击载荷的加入,自动进行解析、实体提取与关系标注,并将数据更新到已有知识图谱中,这种持续更新机制保证了知识图谱的动态成长和知识的不断丰富,能够持续适应新的攻击方式和不断变化的威胁环境;本发明通过构建XSS攻击载荷知识图谱,以结构化的方式组织关键实体及其语义关系,实现了对攻击行为本质的表达与建模;这种结构化的知识表示体系不仅有助于系统性地积累安全知识,还为构建新一代可解释、可溯源的Web安全防御体系奠定了坚实基础,通过知识图谱,安全人员可以更直观地理解和分析攻击行为,从而更好地制定防御策略,进一步提升了Web应用的安全性和可信度。
[0080] 以上显示和描述了本发明的基本原理、主要特征和本发明的优点。本行业的技术人员应该了解,本发明不受上述实施例的限制,上述实施例和说明书中描述的只是本发明的原理,在不脱离本发明精神和范围的前提下本发明还会有各种变化和改进,这些变化和改进都落入要求保护的本发明的范围内。本发明要求的保护范围由所附的权利要求书及其等同物界定。< / script>
Claims
1. An intelligent detection and defense method for XSS attacks based on a knowledge graph, characterized in that: The method includes the following steps in sequence: (1) Collect original XSS attack payloads from a publicly available XSS attack sample library; at the same time, use an automated crawler and penetration testing tools to test a Web vulnerability testing ground to obtain XSS attack payloads that can trigger XSS vulnerabilities. The original XSS attack payloads and the XSS attack payloads that can trigger XSS vulnerabilities together form an XSS attack payload sample data set; (2) Preprocess the XSS attack payload sample data in the XSS attack payload sample data set to obtain a preprocessed XSS attack payload sample data set; (3) Use a large language model to extract the elements of the XSS attack payloads in the preprocessed XSS attack payload sample data set to obtain multi-dimensional attack construction elements; (4) According to the multi-dimensional attack construction elements, structure the abstract payload information into a graph form, establish edges for the extracted multi-dimensional attack construction elements, and obtain data in the form of triples, i.e., entity-relationship-entity data; (5) Import the data in the form of triples into a graph database to form a graph structure of nodes and edges, and complete the construction of the entire XSS attack payload knowledge graph; (6) Integrate the XSS attack payload knowledge graph into a Web application firewall to achieve real-time monitoring and intelligent defense of incoming requests in a Web system.
2. The intelligent detection and defense method for XSS attacks based on a knowledge graph according to claim 1, characterized in that: Step (1) specifically includes the following steps in sequence: (1a) Collect original XSS attack payloads from a publicly available XSS attack sample library, where the publicly available XSS attack sample library includes the XSSed database and the OWASP XSS Attack Prevention Cheat Sheet; (1b) Use an automated crawler and penetration testing tools to test several Web vulnerability testing grounds, such as Firing-Range, WAVSEP, and OWASP Benchmark, and collect XSS attack payloads that can trigger XSS vulnerabilities; (1c) The original XSS attack payloads and the XSS attack payloads that can trigger XSS vulnerabilities together form an XSS attack payload sample data set.
3. The intelligent detection and defense method for XSS attacks based on a knowledge graph according to claim 1, characterized in that: Step (2) specifically means that: the preprocessing includes verification and annotation. The verification means constructing a test page in a browser, inserting the collected XSS attack payloads, and observing whether malicious code is triggered; the annotation means annotating the verified data into two categories: attack payloads that can be triggered and attack payloads that cannot be triggered. Among them, attack payloads that can be triggered refer to XSS attack payloads that can successfully trigger malicious code on the browser side, and attack payloads that cannot be triggered refer to XSS attack payloads that fail to trigger malicious code.
4. The intelligent detection and defense method for XSS attacks based on a knowledge graph according to claim 1, characterized in that: Step (3) specifically includes the following steps in sequence: (3a) Use a large language model to identify XSS attack payloads, disassemble, annotate, and classify them step by step, so that a complex XSS injection statement is structured into a clear set of nodes; the XSS injection statement is the XSS attack payload statement in the preprocessed XSS attack payload sample data set; Extract multi-dimensional attack construction elements from the XSS injection statement, where the multi-dimensional attack construction elements include the original payload, malicious code, HTML tags, events, attributes, encoding techniques, and obfuscation techniques; the original payload is the complete string injected by the attacker into the input point of the web page; the encoding technique is used to hide the malicious code in an unreadable form to bypass the security detection mechanism of the WAF; the obfuscation technique refers to the attacker's interference processing on the structure or semantics of the XSS malicious payload. Assign unique identifiers to each multi-dimensional attack construction element to provide semantic support for graph reasoning and graph representation learning.
5. The intelligent detection and defense method for XSS attacks based on a knowledge graph according to claim 1, characterized in that: Step (4) specifically includes the following steps in sequence: (4a) Name the multi-dimensional attack construction elements extracted by the large language model as Payload, Code, HTMLTag, Event, Attribute, Encoding, and Obfuscation respectively. Payload, Code, HTMLTag, Event, Attribute, Encoding, and Obfuscation correspond to the original payload, malicious code, HTML tags, events, attributes, encoding techniques, and obfuscation techniques one by one; each multi-dimensional attack construction element represents a type of entity, and each type of entity is assigned a unique ID in the graph and has an extensible attribute structure. (4b) Define the core relationship types between entities to describe the logical relationships between key elements; there are 6 types of core relationship types. Among them, the first type is Payload-contains-Code, indicating that the XSS attack payload contains malicious code; the second type is Payload-uses-HTMLTag, indicating that the XSS attack payload uses HTML tags; the third type is HTMLTag-triggers-Event, indicating that the HTML tag of the XSS attack payload can trigger an event; the fourth type is HTMLTag-with-Attribute, indicating the structural relationship between the HTML tag of the XSS attack payload and the XSS attack payload attribute; the fifth type is Payload-encoded_with-Encoding, indicating that the XSS attack payload uses an encoding method; the sixth type is Payload-obfuscated_with-Obfuscation, indicating that the XSS malicious attack payload has been obfuscated. (4c) Automatically establish edges for each pair of entities according to the core relationship type and express them in the form of triples to obtain the triple-form data, that is, entity-relationship-entity data, and finally form a knowledge graph structure composed of entity nodes and semantic edges.
6. The intelligent detection and defense method for XSS attacks based on a knowledge graph according to claim 1, characterized in that: Step (5) specifically includes the following steps in sequence: (5a) Import the triple-form data into the graph database to form a graph structure of nodes and edges; each entity type is modeled as a node in the graph, and each relationship is modeled as an edge connecting the nodes, thus completing the construction of the entire XSS attack payload knowledge graph. (5b) Visualize the constructed XSS knowledge graph.
7. The intelligent detection and defense method for XSS attacks based on a knowledge graph according to claim 1, characterized in that: Step (6) specifically includes the following steps in sequence: (6a) Parse the HTTP request parameters, extract the fields that may carry XSS attack payloads, and send the extracted fields into the XSS attack payload knowledge graph for analysis; (6b) Perform entity-level parsing on the extracted fields, identify the attack construction elements in the extracted fields, and extract the context relationships and core relationship types of the fields; (6c) Compare the extracted core relationship types with the high-risk combinations in the XSS attack payload knowledge graph through a graph matching algorithm, where the high-risk combinations are the core relationship types known to be able to bypass the WAF in the XSS attack payload knowledge graph; If the extracted core relationship type is the same as the high-risk combination in the knowledge graph, it is determined that the extracted core relationship type has a high attack possibility, and corresponding response strategies are taken; the corresponding response strategies include intercepting requests, recording log information, sending the attack sample into the backend audit system or the alarm module, and dynamically updating the XSS attack payload knowledge graph.
8. An electronic device, comprising: A processor; And A memory in which computer program instructions are stored, and when the computer program instructions are run by the processor, the processor is caused to execute the knowledge graph-based XSS attack intelligent detection and defense method according to any one of claims 1-7.
9. A computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are run by a processor, the processor is caused to execute the knowledge graph-based XSS attack intelligent detection and defense method according to any one of claims 1-7.