Attack detection method, apparatus and electronic device
Patent Information
- Application Number
- CN202410968301.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-18
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2044-07-18
AI Technical Summary
[0007]本申请提供了一种攻击检测方法、装置及电子设备,用以解决现有技术难以准确、高效地检测到XSS注入攻击的问题
[0072]第四方面,本申请提供了一种计算机可读存储介质,计算机可读存储介质内存储有计算机程序,计算机程序被处理器执行时实现上述的一种攻击检测方法步骤。
Smart Images

Figure CN118740479B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cybersecurity technology, and in particular to an attack detection method, apparatus, and electronic device. Background Technology
[0002] Cross-site scripting (XSS) attacks essentially involve attackers injecting malicious script code into the Hypertext Markup Language (HTML) of a webpage. This injection modifies the original data packets sent from the server to the client, inserting malicious script code into the webpage. When a user browses the page, the embedded malicious script code is executed, achieving the attacker's goal of harming the user. Therefore, detecting XSS injection attacks is of great importance.
[0003] Currently, XSS injection attacks are primarily detected using regular expression matching or machine learning. When using regular expression matching or machine learning to detect XSS injection attacks, it's possible to detect similar... <script>标签或alert等函数来确定XSS注入攻击。
[0004] 然而,在采用正则匹配来检测XSS注入攻击时,特定业务的接口参数需要提交描述或提交带一些标签(如带有<script>标签)的内容,导致该请求被误认为XSS注入攻击,从而导致用户的正常业务无法使用,造成XSS注入攻击的误报。并且,正则匹配对XSS注入攻击的变形难以识别,从而难以准确地确定出XSS注入攻击。此外,正则匹配需要进行复杂的规则调整才能达到降低误报率的目的,使得XSS注入攻击的检测效率低。
[0005] 在采用机器学习来检测XSS注入攻击时,机器学习仅仅依靠内部训练样本,且仅仅采用分词方式来实现XSS注入攻击的检测,从而无法理解HTML或Javascript的上下文语句含义,导致XSS注入攻击的误报率高。
[0006] 因此,通过正则匹配或者机器学习来检测XSS注入攻击时,均难以准确、高效地检测到XSS注入攻击。发明内容
[0007] 本申请提供了一种攻击检测方法、装置及电子设备,用以解决现有技术难以准确、高效地检测到XSS注入攻击的问题。具体实现方案如下:
[0008] 第一方面,本申请提供了一种攻击检测方法,所述方法包括:
[0009] 获取待检测原始数据,并对所述待检测原始数据进行解码,得到待检测数据;
[0010] 针对所述待检测数据进行分词处理,并在每得到一个分词后,对所述分词进行语义分析,生成包含目标特征的语法树;
[0011] 根据所述语法树中的所述目标特征与攻击特征集合的匹配结果,确定所述待检测原始数据的威胁得分;
[0012] 根据所述威胁得分与告警阈值的比较结果,确定所述待检测原始数据是否为攻击数据。
[0013] 通过上述申请实施例,对待检测原始数据进行解码,从而处理编码绕过的恶意攻击,然后,对待检测数据进行分词处理,以便于提升编码、混淆、变形的绕过检测能力,同时通过对分词进行的语义分析,消除不符合语法规范的误报,并且理解待检测数据中的代码的上下文信息(如变量赋值、函数调用、DOM操作等),从而更好地识别潜在的XSS注入攻击,最终基于语法树中的目标特征与内置攻击特征集合的匹配结果,确定出了待检测原始数据的威胁得分,以确定待检测数据的危害级别,再通过威胁得分与告警阈值的比较结果,来确定待检测原始数据是否为攻击数据,从而确定是否拦截该待检测原始数据。本申请实施例所提供的上述方案,相比于现有技术中的正则匹配方案与机器学习方案,误报和漏报率更低,且配置更加简单,从而可以更加准确性、高效地识别出XSS注入攻击。
[0014] 在一种可能的实施方式中,所述对所述待检测原始数据进行解码,得到待检测数据,包括:
[0015] 在所述待检测原始数据中,确定经过了编码处理的待解码数据;
[0016] 根据所述待解码数据对应的解码类型,对所述待解码数据进行解码,得到还原数据;
[0017] 将所述还原数据与所述待检测数据中未经过所述编码处理的原始数据作为所述待检测数据。
[0018] 通过上述申请实施例,根据待检测原始数据中经过了编码处理的待解码数据所对应的解码类型,对经过了编码处理的待解码数据进行解码,再将解码后的还原数据与待检测原始数据中未进行编码处理的原始数据作为待检测数据,使得得到的待检测数据中的数据均为解码后的数据或者未经过编码处理的数据,从而可以识别出编码绕过的恶意攻击,保证了编码绕过的恶意攻击的检测能力,避免了编码绕过导致的XSS注入攻击的漏报,进而有利于进一步地提高XSS注入攻击的检测准确性。
[0019] 在一种可能的实施方式中,所述针对所述待检测数据进行分词处理,包括:
[0020] 根据包括分词闭合条件的词法分析方法,对所述待检测数据进行分词处理;其中,所述分词闭合条件为第一闭合条件、第二闭合条件、第三闭合条件、第四闭合条件、第五闭合条件、第六闭合条件中的任一闭合条件;
[0021] 所述第一闭合条件为,若初始分词模式为HTML词法模式,则在第一拆分字符处进行拆分,得到第一个分词;所述第一拆分字符为空白字符或者大于符号;
[0022] 所述第二闭合条件为,若所述初始分词模式为所述HTML词法模式,则在第二拆分字符处进行拆分,得到第一个分词;所述第二拆分字符为单引号;
[0023] 所述第三闭合条件为,若所述初始分词模式为所述HTML词法模式,则在第三拆分字符处进行拆分,得到第一个分词;所述第三拆分字符为双引号;
[0024] 所述第四闭合条件为,若所述初始分词模式为Javascript词法模式,则在第四拆分字符处进行拆分,得到第一个分词;所述第四拆分字符为所述空白字符或JS边界字符;
[0025] 所述第五闭合条件为,若所述初始分词模式为所述Javascript词法模式,则在第五拆分字符处进行拆分,得到第一个分词;所述第五拆分字符为引号;
[0026] 所述第六闭合条件为,若所述初始分词模式为所述Javascript词法模式,则在第六拆分字符处进行拆分,得到第一个分词;所述第六拆分字符为星号字符与斜杠字符。
[0027] 通过上述申请实施例,根据上述第一闭合条件、第二闭合条件、第三闭合条件、第四闭合条件、第五闭合条件、第六闭合条件来对待检测数据进行分词,基本可以确定出待检测数据中的XSS注入点,大幅度地提高了确定XSS攻击的准确性与效率。
[0028] 在一种可能的实施方式中,所述针对所述待检测数据进行分词处理,包括:
[0029] 若当前词法模式为HTML词法模式,且触发第一切换条件,则将所述当前词法模式切换为Javascript词法模式,并按照所述Javascript词法模式,对所述待检测数据进行所述分词处理;
[0030] 在初始词法模式为所述HTML词法模式时,若所述当前词法模式为所述Javascript模式,且触发第二切换条件,则将所述当前词法模式切换为所述HTML词法模式,并按照所述HTML词法模式,对所述待检测数据进行所述分词处理。
[0031] 通过上述申请实施例,根据当前词法模式、初始词法模式、以及第一切换条件与第二切换条件,实现了HTML词法模式与Javascript词法模式的相互切换,以便于对待检测数据进行更加全面的分词。
[0032] 在一种可能的实施方式中,所述根据所述语法树中的所述目标特征与攻击特征集合的匹配结果,确定所述待检测原始数据的威胁得分,包括:
[0033] 针对所述语法树中的每一个所述目标特征,均执行以下匹配打分操作:
[0034] 将所述目标特征与所述攻击特征集合中的每一个攻击特征进行匹配,确定出符合所述目标特征的目标攻击特征;
[0035] 根据所述目标攻击特征对应的权重,确定所述目标特征对应的分值;
[0036] 直至所述语法树中的每一个所述目标特征均执行完所述匹配打分操作,根据所述语法树中的每一个所述目标特征各自对应的分值,确定所述待检测原始数据在对应的分词闭合条件下的所述威胁得分。
[0037] 通过上述申请实施例,根据目标特征与攻击特征集合中的每一个攻击特征的匹配以及分值确定的匹配打分操作,确定出了具有攻击特征特性的目标特征,再通过对语法树中每一个目标特征对应的分值的累加,可以确定出待检测原始数据在对应的分词闭合条件下的威胁得分,从而综合考虑了待检测原始数据中的各个目标特征,避免个别特征对攻击检测结果的影响,有利于进一步地提高攻击检测的准确性。
[0038] 在一种可能的实施方式中,所述根据所述威胁得分与告警阈值的比较结果,确定所述待检测原始数据是否为攻击数据,包括:
[0039] 获取每一种分词闭合条件各自对应的所述威胁得分;
[0040] 判断所有所述威胁得分中是否存在高于所述告警阈值的目标威胁得分;
[0041] 若否,则确定所述待检测原始数据非攻击数据;
[0042] 若是,则确定所述待检测原始数据为所述攻击数据。
[0043] 通过上述申请实施例,每一种分词闭合条件对应的各个威胁得分中,只要有一个威胁得分高于告警阈值,就可以确定待检测原始数据为攻击数据。由于不同分词条件下的分词结果不同,从而导致威胁得分也不同。因此,根据每一种闭合条件对应的威胁得分来确定待检测原始数据是否为攻击数据,考虑到了不同分词结果的情形,避免了分词方式对XSS攻击检测的绕过,进一步地提高了待检测原始数据的攻击检测的准确性,进而根据每一种分词闭合条件对应的威胁得分与告警阈值的比较结果,更加准确地确定出了待检测原始数据是否为攻击数据。
[0044] 第二方面,本申请还提供了一种攻击检测装置,所述装置包括:
[0045] 解码模块,用于获取待检测原始数据,并对所述待检测原始数据进行解码,得到待检测数据;
[0046] 分词与语义分析模块,用于针对所述待检测数据进行分词处理,并在每得到一个分词后,对所述分词进行语义分析,生成包含目标特征的语法树;
[0047] 匹配打分模块,用于根据所述语法树中的所述目标特征与攻击特征集合的匹配结果,确定所述待检测原始数据的威胁得分;
[0048] 处理模块,用于根据所述威胁得分与告警阈值的比较结果,确定所述待检测原始数据是否为攻击数据。
[0049] 在一种可能的实施方式中,所述解码模块,具体用于在所述待检测原始数据中,确定经过了编码处理的待解码数据;
[0050] 根据所述待解码数据对应的解码类型,对所述待解码数据进行解码,得到还原数据;
[0051] 将所述还原数据与所述待检测数据中未经过所述编码处理的原始数据作为所述待检测数据。
[0052] 在一种可能的实施方式中,所述分词与语义分析模块,具体用于根据包括分词闭合条件的词法分析方法,对所述待检测数据进行分词处理;其中,所述分词闭合条件为第一闭合条件、第二闭合条件、第三闭合条件、第四闭合条件、第五闭合条件、第六闭合条件中的任一闭合条件;
[0053] 所述第一闭合条件为,若初始分词模式为HTML词法模式,则在第一拆分字符处进行拆分,得到第一个分词;所述第一拆分字符为空白字符或者大于符号;
[0054] 所述第二闭合条件为,若所述初始分词模式为所述HTML词法模式,则在第二拆分字符处进行拆分,得到第一个分词;所述第二拆分字符为单引号;
[0055] 所述第三闭合条件为,若所述初始分词模式为所述HTML词法模式,则在第三拆分字符处进行拆分,得到第一个分词;所述第三拆分字符为双引号;
[0056] 所述第四闭合条件为,若所述初始分词模式为Javascript词法模式,则在第四拆分字符处进行拆分,得到第一个分词;所述第四拆分字符为所述空白字符或JS边界字符;
[0057] 所述第五闭合条件为,若所述初始分词模式为所述Javascript词法模式,则在第五拆分字符处进行拆分,得到第一个分词;所述第五拆分字符为引号;
[0058] 所述第六闭合条件为,若所述初始分词模式为所述Javascript词法模式,则在第六拆分字符处进行拆分,得到第一个分词;所述第六拆分字符为星号字符与斜杠字符。
[0059] 在一种可能的实施方式中,所述分词与语义分析模块,具体用于若当前词法模式为HTML词法模式,且触发第一切换条件,则将所述当前词法模式切换为Javascript词法模式,并按照所述Javascript词法模式,对所述待检测数据进行所述分词处理;
[0060] 在初始词法模式为所述HTML词法模式时,若所述当前词法模式为所述Javascript模式,且触发第二切换条件,则将所述当前词法模式切换为所述HTML词法模式,并按照所述HTML词法模式,对所述待检测数据进行所述分词处理。
[0061] 在一种可能的实施方式中,所述匹配打分模块,具体用于针对所述语法树中的每一个所述目标特征,均执行以下匹配打分操作:
[0062] 将所述目标特征与所述攻击特征集合中的每一个攻击特征进行匹配,确定出符合所述目标特征的目标攻击特征;
[0063] 根据所述目标攻击特征对应的权重,确定所述目标特征对应的分值;
[0064] 直至所述语法树中的每一个所述目标特征均执行完所述匹配打分操作,根据所述语法树中的每一个所述目标特征各自对应的分值,确定所述待检测原始数据在对应的分词闭合条件下的所述威胁得分。
[0065] 在一种可能的实施方式中,所述处理模块,具体用于获取每一种分词闭合条件各自对应的所述威胁得分;
[0066] 判断所有所述威胁得分中是否存在高于所述告警阈值的目标威胁得分;
[0067] 若否,则确定所述待检测原始数据非攻击数据;
[0068] 若是,则确定所述待检测原始数据为所述攻击数据。
[0069] 第三方面,本申请提供了一种电子设备,包括:
[0070] 存储器,用于存放计算机程序;
[0071] 处理器,用于执行所述存储器上所存放的计算机程序时,实现上述的一种攻击检测方法步骤。
[0072] 第四方面,本申请提供了一种计算机可读存储介质,计算机可读存储介质内存储有计算机程序,计算机程序被处理器执行时实现上述的一种攻击检测方法步骤。
[0073] 上述第二方面至第四方面中的各个方面以及各个方面可能达到的技术效果请参照上述针对第一方面或第一方面中的各种可能方案可以达到的技术效果说明,这里不再重复赘述。附图说明
[0074] 图1为本申请实施例提供的一种攻击检测方法的流程示意图;
[0075] 图2为本申请实施例提供的分词处理与语义分析流程示意图;
[0076] 图3a为本申请实施例提供的代码片段示意图一;
[0077] 图3b为本申请实施例提供的代码片段示意图二;
[0078] 图4为本申请实施例提供的攻击检测方法的处理过程示意图;
[0079] 图5为本申请实施例提供的一种攻击检测装置的结构示意图;
[0080] 图6为本申请提供的一种电子设备示意图。具体实施方式
[0081] 为了使本申请的目的、技术方案和优点更加清楚,下面将结合附图对本申请作进一步地详细描述。方法实施例中的具体操作方法也可以应用于装置实施例或系统实施例中。需要说明的是,在本申请的描述中"多个”理解为"至少两个”。"和 / 或”,描述关联对象的关联关系,表示可以存在三种关系,例如,A和 / 或B,可以表示:单独存在A,同时存在A和B,单独存在B这三种情况。A与B连接,可以表示:A与B直接连接和A与B通过C连接这两种情况。另外,在本申请的描述中,"第一”、"第二”等词汇,仅用于区分描述的目的,而不能理解为指示或暗示相对重要性,也不能理解为指示或暗示顺序。
[0082] 下面结合附图,对本申请实施例进行详细描述。
[0083] 当前,在采用正则匹配或者机器学习来检测XSS注入攻击时,特定业务的接口参数需要提交描述或提交带一些标签(如带有<script>标签)的内容,导致该请求被误认为XSS注入攻击,从而导致用户的正常业务无法使用,造成XSS注入攻击的误报。并且,正则匹配对XSS注入攻击的变形难以识别,从而难以准确地确定出XSS注入攻击。此外,正则匹配需要进行复杂的规则调整才能达到降低误报率的目的,使得XSS注入攻击的检测效率低。而机器学习仅仅依靠内部训练样本,且仅仅采用分词方式来实现XSS注入攻击的检测,从而无法理解HTML或Javascript的上下文语句含义,导致XSS注入攻击的误报率高。
[0084] 因此,通过正则匹配或者机器学习来检测XSS注入攻击时,均存在难以准确且高效地检测到XSS注入攻击的问题。
[0085] 故,本申请提出了一种攻击检测方法,通过对待检测原始数据进行解码,来处理编码绕过的恶意攻击,然后通过对待检测数据进行分词处理,以便于提升编码、混淆、变形的绕过检测能力,并将语义分析技术应用于XSS注入攻击检测中,每拆分出一个分词,就对该分词进行语义分析,生成包含目标特征的语法树,从而消除不符合语法规范的误报,并能更好地理解待检测原始数据中的代码的上下文(如变量赋值、函数调用、DOM操作等),从而更好地识别潜在的XSS注入攻击,最终根据语法树中的目标特征与攻击特征集合的匹配结果,确定待检测原始数据的威胁得分,再根据威胁得分与告警阈值的比较结果,确定待检测原始数据是否为攻击数据,显著降低了攻击检测的误报率和漏报率,大幅度地提高了对攻击检测的准确性。并且,还可以根据业务灵活调整告警阈值,从而实现对策略的灵活调整,省去了正则匹配中频繁进行的复杂的规则调整,使得攻击检测的效率更高。
[0086] 参照图1所示为本申请实施例提供的一种攻击的检测方法流程图,该方法包括:
[0087] S101,获取待检测原始数据,并对待检测原始数据进行解码,得到待检测数据。
[0088] S102,针对待检测数据进行分词处理,并在每得到一个分词后,对分词进行语义分析,生成包含目标特征的语法树。
[0089] S103,根据语法树中的目标特征与攻击特征集合的匹配结果,确定待检测原始数据的威胁得分。
[0090] S104,根据威胁得分与告警阈值的比较结果,确定待检测原始数据是否为攻击数据。
[0091] 为了识别出编码绕过的恶意攻击,本申请实施例在对待检测原始数据进行攻击检测前,首先对待检测原始数据进行解码,得到待检测数据。
[0092] 具体地,首先获取待检测原始数据。该待检测原始数据可以为一段程序,也可以为一句代码,可以根据具体应用场景来确定待检测原始数据。
[0093] 然后,在待检测原始数据中,确定出经过了编码处理的待解码数据。
[0094] 接着,根据待解码数据对应的解码类型,对该待解码数据进行解码,得到还原数据。
[0095] 再将还原数据与待检测原始数据中未经过编码处理的原始数据,作为待检测数据。
[0096] 在本申请实施例中,上述解码类型可以包括HTML实体解码、Unicode解码,但并不仅限于此。可以通过待解码数据的编码特征,来确定待解码数据对应的解码类型,但并不仅限于此。
[0097] 示例性的,待解码数据中的连续字符串为开头,则将待解码数据中的后续数据按HTML实体编码进行解析,即,该待解码数据对应的解码类型为HTML实体解码。再比如待解码数据中的连续字符串为%u后接2个16进制字符,则按Unicode编码对待解码数据进行解析,即,该待解码数据对应的解码类型为Unicode编码。
[0098] 上述通过待解码数据的编码特征确定待解码数据对应的解码类型,可以通过解码函数来确定。
[0099] 通过上述方式,将待检测原始数据中经过了编码处理的待解码数据所对应的解码类型,对该经过了编码处理的待解码数据进行了解码,从而对该待解码数据进行了编码还原,以便于识别出XSS注入攻击的绕过场景,从而解决了混合编码绕过的场景。由于在进行XSS攻击时,特征经过HTML实体编码(包括10进制、16进制等格式)或unicode编码后依然能够被浏览器解析执行,对于此种编码层面的绕过,在经过编码还原后可以对真实的攻击特征进行检测,保证了编码绕过的检测能力,而传统的正则匹配方式如果需要检测编码绕过的场景需要编写大量复杂的正则表达式。因此,通过上述方式,使得攻击检测具备高效性。
[0100] 在另一种可能的实时方式中,首先获取待检测原始数据。若在待检测原始数据中,确定不存在经过了编码处理的待解码数据,则直接将待检测原始数据作为待检测数据。
[0101] 在步骤S101得到待检测数据后,针对该待检测数据进行分词处理,并在每得到一个分词后,对该分词进行语义分析,生成包含目标特征的语法树(即执行步骤S102)。
[0102] 在本申请实施例中,上述针对该待检测数据进行分词处理,可以根据包括分词闭合条件的词法分析方法,来对该待检测数据进行分词处理。
[0103] 上述分词闭合条件可以为第一闭合条件、第二闭合条件、第三闭合条件、第四闭合条件、第五闭合条件、第六闭合条件中的任一闭合条件,但并不仅限于此。
[0104] 其中,第一闭合条件为HTML词法模式下的无引号闭合。具体地,若初始分词模式为HTML词法模式,则在第一拆分字符处进行拆分,得到第一个分词(即返回第一个TOKEN);第一拆分字符可以为空白字符或者大于符号(即>)。
[0105] 第二闭合条件为HTML词法模式下的单引号闭合。具体地,若初始分词模式为HTML词法模式,则在第二拆分字符处进行拆分,得到第一个分词(即返回第一个TOKEN);第二拆分字符可以为单引号(即’)。
[0106] 第三闭合条件为HTML词法模式下的双引号闭合。具体地,若初始分词模式为HTML词法模式,则在第三拆分字符处进行拆分,得到第一个分词(即返回第一个TOKEN);第三拆分字符可以为双引号(即”)。
[0107] 第四闭合条件为Javascript词法模式下的无引号闭合。具体地,若初始分词模式为Javascript词法模式,则在第四拆分字符处进行拆分,得到第一个分词(即返回第一个TOKEN);第四拆分字符可以为空白字符或JS边界字符。
[0108] 第五闭合条件为Javascript词法模式下的字符串闭合。具体地,若初始分词模式为Javascript词法模式,则在第五拆分字符处进行拆分,得到第一个分词(即返回第一个TOKEN);第五拆分字符为可以引号。
[0109] 第六闭合条件为Javascript词法模式下的注释闭合。具体地,若初始分词模式为Javascript词法模式,则在第六拆分字符处进行拆分,得到第一个分词(即返回第一个TOKEN);第六拆分字符可以为星号字符与斜杠字符(即* / )。
[0110] 另外,上述词法分析方法还包括状态记录。具体地,每拆分出一个分词后,记录当前状态(如记录当前状态为文本状态),以便于根据当前状态,确定下一个拆分字符,从而根据拆分字符,继续对待检测数据进行分词处理。
[0111] 此外,为了区分按照HTML词法模式或者Javascript词法模式进行分词处理后的TOKEN类型(包括HTML类型与JS类型),在TOKEN前加上类型标记来进行区分。
[0112] 当类型标记为HTML(如[HTML])时,表示TOKEN类型为HTML类型,即按照HTML词法模式进行分词处理,也就是说,基于标准HTML标记语言模型进行解析。
[0113] 当类型标记为JS(如[JS])时,表示TOKEN类型为JS类型,即按照Javascript词法模式进行的分词处理,也就是说,基于标准Javascript语法进行解析。
[0114] 示例性的,待检测数据如下所示:
[0115] 1”onclick=”alert(1)” / >
[0116] 假设分词闭合条件为第二闭合条件,即在HTML语法下的单引号闭合语境中进行拆分。由于上述待检测数据中没有单引号,则在第二闭合条件下,将上述待检测数据进行拆分后只得到一个分词(即TOKEN),该分词的内容为1”onclick=”alert(1)” / >。得到的TOKEN为[HTML]标签属性值。
[0117] 假设分词模式为第五闭合条件,即在Javascript语法下的字符串闭合语境中进行拆分。根据该第五闭合条件,按照Javascript词法模式对上述待检测数据进行拆分后,得到第一个分词的内容为1,即第一个TOKEN为[JS]字符串;第二个分词的内容为onclick,即第二个TOKEN为[JS]标识符;第三个分词的内容为=,即第三个TOKEN为[JS]等号;第四个分词的内容为alert(1),即第四个TOKEN为[JS]字符串;第五个分词的内容为 / ,即第五个TOKEN为[JS]斜杠;第六个分词的内容为>,即第六个TOKEN为[JS]大于符号。
[0118] 在本申请实施例中,上述HTML词法模式与Javascript词法模式还可以根据第一切换条件或第二切换条件进行切换。
[0119] 具体地,若当前词法模式为HTML词法模式,则在触发第一切换条件时,将当前词法模式切换为Javascript词法模式,然后按照Javascript词法模式,继续对后续数据进行分词处理。
[0120] 若初始词法模式为HTML词法模式,在当前词法模式为Javascript词法模式时,当触发第二切换条件后,将当前词法模式切换为HTML词法模式,然后按照HTML词法模式,继续对后续数据进行分词处理。
[0121] 上述第一切换条件可以为如下所示的第一条件、第二条件、第三条件中的任一条件:
[0122] 第一条件:属性为事件类型(如onerror、onclick)。
[0123] 第二条件:属性类型为src、href等链接属性,且使用Javascript:开头的伪协议。
[0124] 第三条件:当解析到script标签的内容部分。
[0125] 上述第二切换条件可以为如下所示的第四条件、第五条件、第六条件中的任一条件:
[0126] 第四条件:当前属性为事件类型(如onerror、onclick),且遇到属性名结束(如单引号、双引号)。
[0127] 第五条件:当前属性类型为src、href等链接属性,且遇到属性名结束。
[0128] 第六条件:当前位于script标签内,且遇到< / script> End tag.
[0129] For example, the data to be detected is shown below:
[0130] 1”onclick="alert(1)" / >
[0131] Assume the word segmentation closure condition is the third closure condition, which is double quote closure in HTML lexical mode. Then, when encountering the first double quote, we split it to get the content of the first word segment as 1, that is, the first TOKEN is [HTML]<tag attribute value>; the content of the second word segment is onclick=, that is, the second TOKEN is [HTML]<tag attribute name>.
[0132] After reading the second double quote, the first condition is triggered, at which point the lexical mode is switched from HTML lexical mode to JavaScript lexical mode.
[0133] According to the JavaScript lexical pattern, the content of the third token is alert, that is, the third token is the [JS] identifier; the content of the fourth token is (, that is, the fourth token is [JS] brackets; the content of the fifth token is 1, that is, the fifth token is the [JS] identifier; the content of the sixth token is ), that is, the sixth token is [JS] brackets.
[0134] After reading the third double quote, it is determined that the third double quote is the end of the attribute name. If the fourth condition is met, the fourth condition is triggered, and the lexical mode is switched from JavaScript lexical mode to HTML lexical mode.
[0135] According to the HTML lexical pattern, the content of the seventh token is / >, that is, the seventh TOKEN is the [HTML] closing tag.
[0136] The event type, attribute type, and tag of the data attributes in the first, second, third, fourth, fifth, and sixth conditions mentioned above can be determined according to the HTML standard syntax rules and the Javascript standard syntax rules.
[0137] Furthermore, when performing word segmentation on the data to be detected, in order to identify the various variations of HTML and thus better perform word segmentation on the data to be detected, the lexical analysis method also needs to include the first rule.
[0138] Unlike Extensible Markup Language (XML) or Extensible HyperText Markup Language (XHTML), HTML has very loose syntax requirements. Even in highly ambiguous or potentially risky scenarios, HTML can guess the webpage author's intent. Therefore, the first rule of lexical analysis methods needs to be able to understand the behavior of the browser's HTML parser, thus enabling it to parse various loose HTML fragments like a browser.
[0139] For example, code snippets such as Figure 3a As shown. At position ①, Internet Explorer allows the insertion of the null byte 0x00.
[0140] At position ③, most browsers can insert the 0x00 character and various types of space characters.
[0141] For the spaces at positions ② and ④, all browsers can replace the space with a vertical tab (0x0B) or a form feed (0x0C); Opera may replace the space with a non-breaking space in UTF-8 format (0xA0); Firefox can replace the space at position ② with a single ordinary forward slash (i.e., / ), but cannot replace the space at position ④.
[0142] For single and double quotes, strings containing spaces or angle brackets can be placed in HTML parameters. At position ⑤, Internet Explorer also accepts backticks ('). Additionally, a space character is implicitly followed after the quoted parameter, so the space at position ⑥ can be identified.
[0143] Therefore, the first rule needs to include content that can identify the information in the above example, so that browsers such as IE, Opera, and Firefox can identify whether the characters at each position of the code snippet are correct and whether the characters contain special information, in order to determine whether it is necessary to split at that position, and thus perform word segmentation processing on the data to be detected more accurately.
[0144] To understand the characteristics of script character encoding and bypassing in JavaScript, lexical analysis methods also need to include a second rule. This second rule has the ability to parse the characteristics of script character encoding and bypassing.
[0145] For example, code snippets such as Figure 3b As shown. At position ①, you can insert a whitespace character between \x01 and \x20. At position ②, you can insert a tab character (\x09), a newline character (\x0a, \x0d), etc. At position ③ (i.e., after the colon following the javascript protocol name), you can insert a tab character (\x09), a newline character (\x0a, \x0d), etc.
[0146] Therefore, the second rule needs to include content that can identify the information in the above examples, thereby understanding the characteristics of script character encoding and bypassing in JavaScript, in order to determine whether it is necessary to bypass certain rules at a certain location (as mentioned above). Figure 3b The positions shown (①, ②, and ③) are split to more accurately segment the data to be detected.
[0147] Furthermore, HTML tags, attribute names, and pseudo-protocol names (javascript:) are case-insensitive, while JavaScript code is case-sensitive. Therefore, lexical analysis methods also need to include a third rule. This third rule has the ability to recognize the case of data.
[0148] Therefore, in this embodiment of the application, the above-mentioned word segmentation processing of the data to be detected can specifically be as follows:
[0149] When the word segmentation closure condition is the first closure condition, or the second closure condition, or the third closure condition, the data to be detected is segmented according to the word segmentation closure condition, the HTML lexical pattern (i.e., the standard specification of HTML), and based on the first and third rules, as well as the state record.
[0150] When the word segmentation closure condition is the fourth, fifth, or sixth closure condition, the data to be detected is segmented according to the word segmentation closure condition, following the JavaScript lexical pattern (i.e., the standard JavaScript specification), and based on the second and third rules, as well as the state record. This allows for better word segmentation of the data to be detected, thereby improving the completeness and accuracy of word segmentation and better identifying XSS injection points.
[0151] Furthermore, after segmenting the data to be detected, semantic analysis is performed on each segment obtained. In this embodiment, semantic analysis can be performed using a semantic analysis method.
[0152] Then, during the semantic analysis of the segmented word, after confirming that the segmented word is grammatically correct, the remaining data to be tested is segmented. This prevents the waste of resources caused by continuing semantic analysis on other segments after a grammatical error is identified in the semantic analysis of a segmented word. This process can be as follows: Figure 2 As shown:
[0153] S201. Based on the word segmentation closure condition, according to the lexical pattern corresponding to the word segmentation closure condition, and based on the first rule, the second rule, the third rule, and the state record, the data to be detected is segmented to obtain the i-th word.
[0154] Where i is a positive integer.
[0155] S202, Based on the semantic analysis method, perform syntactic analysis on the i-th word to determine whether there is a syntactic error in the i-th word.
[0156] If it is determined that the i-th word segment does not have a grammatical error, then execute S203.
[0157] If it is determined that there is a grammatical error in the i-th word segment, then proceed to step S205.
[0158] S203, determine whether the i-th word segment is the end symbol.
[0159] If it is determined that the i-th word segment is not the end character, then proceed to step S204.
[0160] If the i-th word is determined to be the end character, then proceed to step S205.
[0161] S204, determine that i equals i plus 1.
[0162] After executing step S204, continue executing step S201.
[0163] S205, End attack detection on the raw data to be tested.
[0164] Optionally, after determining that there is a syntax error in the data to be detected and executing step S205 to end the attack detection of the original data to be detected, the location of the syntax error in the original data to be detected can be modified, and then steps S101-S104 can be performed on the modified data to continue detecting whether the original data to be detected is attack data.
[0165] Furthermore, since HTML is only a markup language and lacks a complex syntactic structure, and is only parsed but not executed in a browser, no syntactic analysis is performed on tokens with the TOKEN type "HTML".
[0166] Furthermore, in the process of performing semantic analysis on the word segmentation and generating a syntax tree containing the target features based on the semantic analysis method mentioned above, it is necessary to first determine whether the TOKEN type of the word segmentation is JS type or HTML type.
[0167] When the TOKEN type of the word segment is HTML, the semantic analysis of the word segment is not performed, and the word segmentation process continues for the data to be detected.
[0168] When the token type is JS, the token is analyzed using the syntax analysis method in the semantic analysis method to determine whether the syntax of the token is correct.
[0169] Therefore, according to the grammar rules in the semantic analysis method, the generated syntax tree containing the target features is the target feature corresponding to the JS-type tokens contained therein.
[0170] The above-mentioned syntax analysis method can be based on the standard JS syntax rules, and then perform syntax analysis on the word segmentation according to the standard JS syntax rules.
[0171] In this embodiment of the application, the above-mentioned syntactic analysis of word segmentation according to the JS standard syntax rules can be performed by the bottom-up LALR(1) analysis (LookAhead-LR) method in the context-free text processing of the JS standard syntax rules, so as to maintain linear time complexity and maximum similarity to Javascript syntax.
[0172] The aforementioned JavaScript standard syntax rules also include compatibility handling to ensure compatibility across different statement segments. For example, for locations where XSS injection points exist, these points might be within an assignment statement, a function parameter, or a map structure. This can be addressed by writing corresponding Backus-Naur Form (BNF) syntax rules to describe the reduction mechanism of the syntax rules as comprehensively as possible, thereby enabling semantic analysis of tokenization.
[0173] The JS standard syntax rules include the first symbol and the second symbol.
[0174] The first symbol is "expression". This expression represents a syntax and is the basic unit of syntax, such as variables, operators, constants, and function calls.
[0175] The second symbol is "statement". "Statement" represents a statement, which is the unit of execution in the language, such as basic statements like import, switch, break, try catch, and function.
[0176] In this embodiment, the above-mentioned JS standard syntax rules can handle fragmented JS code. When processing the syntax of fragmented JS code, the first symbol (i.e., expression) is used as the starting point, and stepwise reduction is performed based on LALR(1) parsing. The syntax rules list any closing symbols that can be processed in the current state. For example, when closing at the positions of function parameters or actual parameters, array members, map members, return values, catch statements, etc., the closing symbols that can be processed include ), ],}, ;, etc. After a reduction with the closing symbol, the reduced value will attempt a second reduction to satisfy the scenario of continuous closing statements. At the same time, a reduction is also required after encountering the second symbol (i.e., statement).
[0177] For example, a JavaScript injection snippet is shown below:
[0178] 1),1){alert( / XSS / )}),1); / /
[0179] The above XSS injection into Payload may result in the following injection scenarios:
[0180] ;a(b(function c(d(1,XSS_PAYLOAD)){}));
[0181] The complete statement after replacing XSS_PAYLOAD using the above JavaScript injection snippet is shown below:
[0182] ;a(b(function c(d(1,1),1){alert( / XSS / )}),1); / / )){}));
[0183] In the JavaScript injection example above, the first expression 1 exists in the function call parameters. After being closed by the ')' symbol, it is closed again by ',1)'. ',1)' is the closing form of the function parameter, followed by the function body. The function body is represented by a statement, and after being closed by the statement, subsequent '),', '1', ';', etc., are similarly closed. Furthermore, the content after the comment symbol / / is ignored during tokenization.
[0184] Furthermore, semantic analysis of JavaScript syntax fragments is performed through closed-form conversion. After the semantic analysis is completed, the syntax tree generated during the conversion process is subjected to feature matching in step S103. The generated syntax tree mainly contains expression nodes, function call nodes, etc.
[0185] Furthermore, since the word segmentation closure condition can be any of the first, second, third, fourth, fifth, or sixth closure conditions, to avoid missing XSS injection points and improve the accuracy of XSS injection point detection, the above-mentioned word segmentation processing and semantic analysis are performed under each closure condition, thereby generating a syntax tree for each closure condition. That is, the first syntax tree under the first closure condition, the second syntax tree under the second closure condition, the third syntax tree under the third closure condition, the fourth syntax tree under the fourth closure condition, the fifth syntax tree under the fifth closure condition, and the sixth syntax tree under the sixth closure condition.
[0186] After step S102, when the word segmentation and syntax analysis of the data to be detected are completed and a syntax tree containing the target features is generated, the threat score of the original data to be detected is determined based on the matching results of the target features and the attack feature set in the syntax tree (i.e., step S103 is executed).
[0187] Specifically, for each target feature in the syntax tree, the following matching and scoring operation is performed:
[0188] The target feature is matched with each attack feature in the attack feature set to determine the target attack feature that matches the target feature.
[0189] Then, based on the weights corresponding to the target attack features, the scores corresponding to the target features are determined.
[0190] However, it should be noted that if there is no target attack feature that matches the target feature in the target attack feature, then the score corresponding to the target feature is determined to be zero.
[0191] The matching and scoring operation is performed on each target feature in the syntax tree. Based on the score corresponding to each target feature in the syntax tree, the threat score of the original data to be detected under the corresponding word segmentation closure condition is determined.
[0192] The threat score can range from 0 to 10, and it represents the confidence level of the raw data to be detected.
[0193] In this embodiment of the application, the aforementioned attack feature set may include, but is not limited to, a first attack feature, a second attack feature, a third attack feature, a fourth attack feature, a fifth attack feature, a seventh attack feature, an eighth attack feature, a ninth attack feature, a tenth attack feature, and an eleventh attack feature. Furthermore, each attack feature has its own corresponding weight, which can be adjusted according to the specific application scenario.
[0194] The first characteristic of the attack mentioned above is the presence of tags from an active blacklist within the HTML tag structure.
[0195] The second attack characteristic mentioned above is the blacklist attribute name in the feature closing tag.
[0196] The third attack characteristic mentioned above is the use of JavaScript pseudo-protocol access based on src, href, etc.
[0197] The fourth attack characteristic mentioned above is based on accessing JS event attributes within tags.
[0198] The fifth attack feature mentioned above is the injection of link reference boxes in Frame tags.
[0199] The sixth attack characteristic mentioned above is the introduction of external links based on Object and Script tags.
[0200] The seventh attack characteristic mentioned above is the execution of CSS expressions using the style property.
[0201] The eighth attack feature mentioned above is introduced through HTML comments.
[0202] The ninth attack characteristic mentioned above is the invocation of sensitive JavaScript functions or objects.
[0203] The tenth attack characteristic mentioned above is access to sensitive JavaScript attribute values.
[0204] The eleventh attack characteristic mentioned above is JavaScript's array-based function call behavior.
[0205] In this application embodiment, the matching method between the target feature and each of the above attack features is different, and the matching method needs to be determined according to the specific content of the attack feature.
[0206] The threat score of the original data to be detected under the corresponding word segmentation closure condition is determined based on the score corresponding to each target feature in the syntax tree. Specifically, it can be:
[0207] First, the sum of the scores corresponding to each target feature in the syntax tree is calculated to obtain the JS score.
[0208] Since the syntax tree contains target features corresponding to JS-type segmentations, the determined JS score is the score of the JS information corresponding to the JS-type segmentation. Therefore, it is also necessary to determine the HTML score of the HTML information corresponding to the HTML-type segmentation, so that the final threat score fully considers all segmentations in the data to be detected, making the information contained in the threat score more comprehensive, thereby improving the accuracy of attack detection.
[0209] Then calculate the sum between the JS score and the HTML score, and use this sum as the threat score of the original data to be detected under the corresponding word segmentation closure condition.
[0210] The HTML score of the HTML information (including tag names, attribute names, attribute values, etc.) corresponding to the HTML-type tokens can be determined by combining the HTML-type tokens (i.e., TOKENs) with the Document Object Model (DOM) structural context and the attack feature set.
[0211] For example, for HTML-type tokenization, such as <script src= brutelogic.com.br 1.js>,在DOM上下文中记录了当前为script标签且属性为src,属性值为JS文件路径。然后,根据该信息,与攻击特征集合中的每一个攻击特征进行匹配,以确定对应的分值。该匹配确定分值的过程同语法树中目标特征与攻击特征集合的匹配确定分值的过程一致,在此不再赘述。
[0212] 此外,在本申请实施例中,若语法树中的某一目标特征(即单一目标特征)出现多次,在确定单一目标特征对应的分值时,每出现一次该单一目标特征,就会计算一次该单一目标特征对应的分值,然后对每次计算的单一目标特征的分值进行累加,再将该累加后的分值(即累加分值)作为单一目标特征的最终分值。但此种方法会存在如下问题:
[0213] 若单一目标特征为弱特征,且出现频次加高,则会导致威胁得分偏高,从而导致攻击检测容易出现误报。
[0214] 因此,为了避免弱攻击性的特征的出现频次较高,使得威胁得分较高,导致攻击数据的误报的问题出现,针对语法树中的每一种目标特征均会设置最大值。具体地,在计算单一目标特征对应的分值时,若该单一目标特征对应的累加分值大于为该第一目标特征设置的最大值,则将该最大值作为该单一目标特征的最终分值。若该单一目标特征对应的累加分值小于等于最大值,则将累加分值作为该单一目标特征的最终分值,从而进一步地提高了待检测原始数据的检测准确性。
[0215] 另外,由于步骤S102中不同分词闭合条件下生成的语法树不同,则通过不同语法树得到的JS分值也不相同。并且,不同闭合条件下的HTML类型的分词也会不相同,则HTML类型的分词所对应的HTML信息的HTML分值也会不相同。因此,在不同分词闭合条件下,得到的威胁得分也不相同。
[0216] 所以,针对每一种闭合条件下生成的语法树,均采用上述方式来计算每一种语法树中的目标特征所对应的分值,从而得到每一种语法树对应的JS分值。并且,针对每一种分词闭合条件下的HTML类型的分词,采用上述方式来计算每一种分词闭合条件下的HTML类型的分词所对应的HTML分值,从而根据每一种分词闭合条件下的JS分值与HTML分值,得到每一种分词闭合条件下的威胁得分。也就是说,得到待检测数据在第一分词闭合条件下的第一威胁得分,在第二分词闭合条件下的第二威胁得分,在第三分词闭合条件下的第三威胁得分,在第四分词闭合条件下的第四威胁得分,在第五分词闭合条件下的第五威胁得分,在第六分词闭合条件下的第六威胁得分。
[0217] 在步骤S103得到待检测原始数据的威胁得分后,根据威胁得分与告警阈值的比较结果,确定待检测原始数据是否为攻击(即执行步骤S104)。
[0218] 由于待检测原始数据在不同的分词闭合条件下,得到的是不同威胁得分。因此,根据威胁得分与告警阈值的比较结果,确定待检测原始数据是否为攻击,具体可以为:
[0219] 获取待检测原始数据在每一种分词闭合条件下的威胁得分。
[0220] 然后,将所有威胁得分与告警阈值进行比较,判断所有威胁得分中是否存在高于告警阈值的目标威胁得分。
[0221] 若确定所有威胁得分中均不存在高于告警阈值的目标威胁得分,则确定待检测原始数据非攻击数据。
[0222] 若确定所有威胁得分中存在高于告警阈值的目标威胁得分,即只要所有威胁得分中存在任一个威胁得分高于告警阈值,则确定待检测原始数据为攻击数据。
[0223] 并且,还可以根据威胁得分的高低,确定待检测原始数据的危害性。示例性的,威胁得分越高,待检测原始数据的危害性越高,且该待检测原始数据为XSS注入攻击的概率越大。
[0224] 在本申请实施例中,告警阈值可以为2,可以根据具体的应用场景来进行灵活调整。
[0225] 可选地,在确定待检测原始数据为攻击数据时,可以拦截或阻断该待检测原始数据。在确定待检测原始非攻击数据时,可以放行该待检测原始数据。
[0226] 综上来说,本申请所提出的攻击检测方法,在XSS注入攻击检测中,相比使用正则匹配或者机器学习的方式,基于解码、分词处理、语义分析、特征打分、告警阈值比较等一系列的处理操作,更能将误报率大幅度降低,同时也大幅度提升了对编码变形绕过的检测能力。基于HTML和JS两个维度的分词处理、语法分析、特征匹配,能更灵活有效的对输入数据(即待检测原始数据)进行威胁评分。
[0227] 并且,根据不同的分词闭合条件,对待检测数据进行分词处理,以便于解析出待检测数据中的标签和属性等内容或JS关键词等,使得能够更好的还原对标签的具体属性行为,从而有效应对插入混淆字符后无法识别关键标签或函数的绕过场景,防止在词法层面的变形绕过,大大提升了变形攻击的检测能力。
[0228] 同时,基于语义分析方法,可以进行容错性处理,对于分词中的简单错误进行兼容,从而大幅度的降低了XSS注入攻击的误报与漏报概率。
[0229] 此外,通过攻击特征的权重来描述XSS注入攻击中各类特征的强度,攻击特征越明显权重越高。对于待检测原始数据是否为XSS注入攻击,是依据众多的攻击特征的打分并进行加权计算后确定的,客观上反映了该待检测原始数据的危害级别。相比传统的正则特征匹配输出一个是否有害的布尔值,本申请实施例所提供的方案提高了XSS注入攻击检测的灵活度。
[0230] 下面结合具体的应用过程对本申请技术方案做进一步的说明。
[0231] 如图4所示为攻击检测方法的处理过程示意图,首先在解码模块中,获取待检测原始数据。然后,在该待检测原始数据中,确定出经过了编码处理的待解码数据。接着,根据待解码数据对应的解码类型,对该待解码数据进行解码,得到还原数据。再将该还原数据与待检测原始数据中未经过编码处理的原始数据,作为待检测数据。并将该待检测数据传输至分词处理模块,以使分词处理模块对待检测数据进行分词处理。
[0232] 在分词处理模块中,接收解码模块传输来的待检测数据,或者接收语义分析模块传输来的指示分词语法正常的第一语法分析结果。然后,在第一分词子模块中,根据第一分词闭合条件,按照第一分词闭合条件对应的HTML词法模式,并根据第一规则与第三规则,以及状态记录,对待检测数据进行拆分,并在拆分出一个分词后,将该分词传输至语义分析模块,以使语义分词模块对该分词进行语义分析。
[0233] 同样地,在第二分词子模块中,根据第二分词闭合条件,按照第二分词闭合条件对应的HTML词法模式,并根据第一规则与第三规则,以及状态记录,对待检测数据进行拆分,并在拆分出一个分词后,将该分词传输至语义分析模块,以使语义分词模块对该分词进行语义分析。
[0234] 在第三分词子模块中,根据第三分词闭合条件,按照第三分词闭合条件对应的HTML词法模式,并根据第一规则与第三规则,以及状态记录,对待检测数据进行拆分,并在分出一个拆分后,将该分词传输至语义分析模块,以对该分词进行语义分析。
[0235] 在第四分词子模块中,根据第四分词闭合条件,按照第四分词闭合条件对应的Javascript词法模式,并根据第二规则与第三规则,以及状态记录,对待检测数据进行拆分,并在拆分出一个分词后,将该分词传输至语义分析模块,以使语义分词模块对该分词进行语义分析。
[0236] 在第五分词子模块中,根据第五分词闭合条件,按照第五分词闭合条件对应的Javascript词法模式,并根据第二规则与第三规则,以及状态记录,对待检测数据进行拆分,并在拆分出一个分词后,将该分词传输至语义分析模块,以使语义分词模块对该分词进行语义分析。
[0237] 在第六分词子模块中,根据第六分词闭合条件,按照第六分词闭合条件对应的Javascript词法模式,并根据第二规则与第三规则,以及状态记录,对待检测数据进行拆分,并在拆分出一个分词后,将该分词传输至语义分析模块,以使语义分词模块对该分词进行语义分析。
[0238] 在语义分析模块中,当接收到分词处理模块中的第一分词子模块传输来的分词后,确定该分词对应的分词类型(即TOKEN类型)是否为JS类型。若确定该分词类型为JS类型,则按照JS标准语法规则对该分词进行语法分析。并在语法分析的过程中,生成包含目标特征(即关键表达式与关键语句符号)的第一语法树。若在对分词进行语法分析中,确定该分词存在语法错误,则停止分词处理与语法分析处理,结束待检测数据的攻击检测。若在对分析进行语法分析中,确定该分词不存在语法错误,则将第一语法分析结果传输至分词处理模块,以使分词处理模块继续对待检测数据进行分词。若确定该分词类型为HTML类型,则结合DOM结构上下文,得到HTML类型的分词所对应的HTML信息。直至接收的分词对应的TOKEN为结束符时,进入特征匹配打分模块,以使特征匹配打分模块根据第一语法树、HTML信息与攻击特征集合,确定出待检测原始数据在第一分词闭合条件下对应的第一威胁得分。
[0239] 同样地,在语义分析模块中,当接收到分词处理模块中的第二分词子模块传输来的分词,或者第三分词子模块传输来的分词,或者第四分词子模块传输来的分词,或者第五分词子模块传输来的分词,或者第六分词子模块传输来的分词后,同接收到第一分词子模块传输来的分词后的操作一致,从而得到每一个分词子模块各自对应的语法树,即第一分词子模块对应的第一语法树,第二分词子模块对应的第二语法树,第三分词子模块对应的第三语法树,第四分词子模块对应的第四语法树,第五分词子模块对应的第五语法树,第六分词子模块对应的第六语法树,以及HTML信息。
[0240] 在特征匹配打分模块中,获取语义分析模块中在语义分析过程中生成的第一语法树。然后,遍历第一语法树,针对第一语法树中的每一个目标特征,均执行以下匹配打分操作:
[0241] 将目标特征与攻击特征集合中的每一个攻击特征进行匹配,确定出符合目标特征的目标攻击特征;再根据目标攻击特征对应的权重,确定目标特征对应的分值;若攻击特征集合中不存在符合目标特征的目标攻击特征,则确定目标特征对应的分值为零。
[0242] 直至第一语法树中的每一个目标特征均执行完上述匹配打分操作后,计算每一个目标特征各自对应的分值之间的累加和,得到JS分值。
[0243] 同时,获取语义分析模块中得到的HTML信息。再根据该HTML信息与攻击特征集合中每一个攻击特征的匹配结果,得到HTML信息对应的HTML分值。然后,计算JS分值与HTML分值之间的和,并将该和作为待检测原始数据在第一分词闭合条件下的第一威胁得分。
[0244] 同样地,在特征匹配打分模块中,针对第二闭合条件下的第二语法树与得到的HTML信息、第三闭合条件下的第三语法树与得到的HTML信息、第四闭合条件下的第四语法树与得到的HTML信息、第五闭合条件下的第五语法树与得到的HTML信息、第六闭合条件下的第六语法树与得到的HTML信息,同上述针对第一闭合条件下的第一语法树与得到的HTML信息所执行的操作一致,从而得到待检测原始数据在每一种分词闭合条件下的威胁得分,即得到第一威胁得分、第二威胁得分、第三威胁得分、第四威胁得分、第五威胁得分与第六威胁得分。并将每一个威胁得分均传输至告警检测模块,以使告警检测模块确定待检测原始数据是否为攻击数据。
[0245] 在告警检测模块中,接收特征匹配模块传输来的待检测原始数据在每一种分词闭合条件下的威胁得分。然后,判断所有威胁得分中是否存在高于告警阈值的目标威胁得分。若不存在,则确定待检测原始数据非攻击数据。若存在,则确定待检测原始数据为攻击数据。
[0246] 基于同一发明构思,本申请实施例中还提供了一种攻击检测装置,如图5所示为本申请提供的一种攻击检测装置的结构示意图,该装置包括:
[0247] 解码模块501,用于获取待检测原始数据,并对所述待检测原始数据进行解码,得到待检测数据;
[0248] 分词与语义分析模块502,用于针对所述待检测数据进行分词处理,并在每得到一个分词后,对所述分词进行语义分析,生成包含目标特征的语法树;
[0249] 匹配打分模块503,用于根据所述语法树中的所述目标特征与攻击特征集合的匹配结果,确定所述待检测原始数据的威胁得分;
[0250] 处理模块504,用于根据所述威胁得分与告警阈值的比较结果,确定所述待检测原始数据是否为攻击数据。
[0251] 在一种可能的实施方式中,解码模块501,具体用于在所述待检测原始数据中,确定经过了编码处理的待解码数据;
[0252] 根据所述待解码数据对应的解码类型,对所述待解码数据进行解码,得到还原数据;
[0253] 将所述还原数据与所述待检测数据中未经过所述编码处理的原始数据作为所述待检测数据。
[0254] 在一种可能的实施方式中,分词与语义分析模块502,具体用于根据包括分词闭合条件的词法分析方法,对所述待检测数据进行分词处理;其中,所述分词闭合条件为第一闭合条件、第二闭合条件、第三闭合条件、第四闭合条件、第五闭合条件、第六闭合条件中的任一闭合条件;
[0255] 所述第一闭合条件为,若初始分词模式为HTML词法模式,则在第一拆分字符处进行拆分,得到第一个分词;所述第一拆分字符为空白字符或者大于符号;
[0256] 所述第二闭合条件为,若所述初始分词模式为所述HTML词法模式,则在第二拆分字符处进行拆分,得到第一个分词;所述第二拆分字符为单引号;
[0257] 所述第三闭合条件为,若所述初始分词模式为所述HTML词法模式,则在第三拆分字符处进行拆分,得到第一个分词;所述第三拆分字符为双引号;
[0258] 所述第四闭合条件为,若所述初始分词模式为Javascript词法模式,则在第四拆分字符处进行拆分,得到第一个分词;所述第四拆分字符为所述空白字符或JS边界字符;
[0259] 所述第五闭合条件为,若所述初始分词模式为所述Javascript词法模式,则在第五拆分字符处进行拆分,得到第一个分词;所述第五拆分字符为引号;
[0260] 所述第六闭合条件为,若所述初始分词模式为所述Javascript词法模式,则在第六拆分字符处进行拆分,得到第一个分词;所述第六拆分字符为星号字符与斜杠字符。
[0261] 在一种可能的实施方式中,分词与语义分析模块502,具体用于若当前词法模式为HTML词法模式,且触发第一切换条件,则将所述当前词法模式切换为Javascript词法模式,并按照所述Javascript词法模式,对所述待检测数据进行所述分词处理;
[0262] 在初始词法模式为所述HTML词法模式时,若所述当前词法模式为所述Javascript模式,且触发第二切换条件,则将所述当前词法模式切换为所述HTML词法模式,并按照所述HTML词法模式,对所述待检测数据进行所述分词处理。
[0263] 在一种可能的实施方式中,匹配打分模块503,具体用于针对所述语法树中的每一个所述目标特征,均执行以下匹配打分操作:
[0264] 将所述目标特征与所述攻击特征集合中的每一个攻击特征进行匹配,确定出符合所述目标特征的目标攻击特征;
[0265] 根据所述目标攻击特征对应的权重,确定所述目标特征对应的分值;
[0266] 直至所述语法树中的每一个所述目标特征均执行完所述匹配打分操作,根据所述语法树中的每一个所述目标特征各自对应的分值,确定所述待检测原始数据在对应的分词闭合条件下的所述威胁得分。
[0267] 在一种可能的实施方式中,处理模块504,具体用于获取每一种分词闭合条件各自对应的所述威胁得分;
[0268] 判断所有所述威胁得分中是否存在高于所述告警阈值的目标威胁得分;
[0269] 若否,则确定所述待检测原始数据非攻击数据;
[0270] 若是,则确定所述待检测原始数据为所述攻击数据。
[0271] 基于同一发明构思,本申请实施例中还提供了一种电子设备,上述电子设备可以实现前述攻击检测装置的功能,参考图6,上述电子设备包括:
[0272] 至少一个处理器601,以及与至少一个处理器601连接的存储器602,本申请实施例中不限定处理器601与存储器602之间的具体连接介质,图6中是以处理器601和存储器602之间通过总线600连接为例。总线600在图6中以粗线表示,其它部件之间的连接方式,仅是进行示意性说明,并不引以为限。总线600可以分为地址总线、数据总线、控制总线等,为便于表示,图6中仅用一条粗线表示,但并不表示仅有一根总线或一种类型的总线。或者,处理器601也可以称为控制器,对于名称不做限制。
[0273] 在本申请实施例中,存储器602存储有可被至少一个处理器601执行的指令,至少一个处理器601通过执行存储器602存储的指令,可以执行前文论述的攻击检测方法。处理器601可以实现图5所示的装置中各个模块的功能。
[0274] 其中,处理器601是该装置的控制中心,可以利用各种接口和线路连接整个该控制设备的各个部分,通过运行或执行存储在存储器602内的指令以及调用存储在存储器602内的数据,该装置的各种功能和处理数据,从而对该装置进行整体监控。
[0275] 在一种可能的设计中,处理器601可包括一个或多个处理单元,处理器601可集成应用处理器和调制解调处理器,其中,应用处理器主要处理操作系统、用户界面和应用程序等,调制解调处理器主要处理无线通信。可以理解的是,上述调制解调处理器也可以不集成到处理器601中。在一些实施例中,处理器601和存储器602可以在同一芯片上实现,在一些实施例中,它们也可以在独立的芯片上分别实现。
[0276] 处理器601可以是通用处理器,例如中央处理器(英文:Central ProcessingUnit,缩写为CPU)、数字信号处理器、专用集成电路、现场可编程门阵列或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件,可以实现或者执行本申请实施例中公开的各方法、步骤及逻辑框图。通用处理器可以是微处理器或者任何常规的处理器等。结合本申请实施例所公开的攻击检测方法的步骤可以直接体现为硬件处理器执行完成,或者用处理器中的硬件及软件模块组合执行完成。
[0277] 存储器602作为一种非易失性计算机可读存储介质,可用于存储非易失性软件程序、非易失性计算机可执行程序以及模块。存储器602可以包括至少一种类型的存储介质,例如可以包括闪存、硬盘、多媒体卡、卡型存储器、随机访问存储器(英文:Random AccessMemory,缩写为RAM)、静态随机访问存储器(英文:Static Random Access Memory,缩写为SRAM)、可编程只读存储器(英文:Programmable Read Only Memory,缩写为PROM)、只读存储器(英文:Read Only Memory,缩写为ROM)、带电可擦除可编程只读存储器(英文:Electrically Erasable Programmable Read-Only Memory,缩写为EEPROM)、磁性存储器、磁盘、光盘等等。存储器602是能够用于携带或存储具有指令或数据结构形式的期望的程序代码并能够由计算机存取的任何其他介质,但不限于此。本申请实施例中的存储器602还可以是电路或者其它任意能够实现存储功能的装置,用于存储程序指令和 / 或数据。
[0278] 通过对处理器601进行设计编程,可以将前述实施例中介绍的攻击检测方法所对应的代码固化到芯片内,从而使芯片在运行时能够执行图1所示的实施例的攻击检测方法的步骤。如何对处理器601进行设计编程为本领域技术人员所公知的技术,这里不再赘述。
[0279] 基于同一发明构思,本申请实施例还提供一种存储介质,该存储介质存储有计算机指令,当该计算机指令在计算机上运行时,使得计算机执行前文论述的攻击检测方法。
[0280] 在一些可能的实施方式中,本申请提供的攻击检测方法的各个方面还可以实现为一种程序产品的形式,其包括程序代码,当程序产品在装置上运行时,程序代码用于使该控制设备执行本说明书上述描述的根据本申请各种示例性实施方式的攻击检测方法中的步骤。
[0281] 本领域内的技术人员应明白,本申请的实施例可提供为方法、系统、或计算机程序产品。因此,本申请可采用完全硬件实施例、完全软件实施例、或结合软件和硬件方面的实施例的形式。而且,本申请可采用在一个或多个其中包含有计算机可用程序代码的计算机可用存储介质(包括但不限于磁盘存储器、CD-ROM、光学存储器等)上实施的计算机程序产品的形式。
[0282] 本申请是参照根据本申请实施例的方法、设备(系统)、和计算机程序产品的流程图和 / 或方框图来描述的。应理解可由计算机程序指令实现流程图和 / 或方框图中的每一流程和 / 或方框、以及流程图和 / 或方框图中的流程和 / 或方框的结合。可提供这些计算机程序指令到通用计算机、专用计算机、嵌入式处理机或其他可编程数据处理设备的处理器以产生一个机器,使得通过计算机或其他可编程数据处理设备的处理器执行的指令产生用于实现在流程图一个流程或多个流程和 / 或方框图一个方框或多个方框中指定的功能的装置。
[0283] 这些计算机程序指令也可存储在能引导计算机或其他可编程数据处理设备以特定方式工作的计算机可读存储器中,使得存储在该计算机可读存储器中的指令产生包括指令装置的制造品,该指令装置实现在流程图一个流程或多个流程和 / 或方框图一个方框或多个方框中指定的功能。
[0284] 这些计算机程序指令也可装载到计算机或其他可编程数据处理设备上,使得在计算机或其他可编程设备上执行一系列操作步骤以产生计算机实现的处理,从而在计算机或其他可编程设备上执行的指令提供用于实现在流程图一个流程或多个流程和 / 或方框图一个方框或多个方框中指定的功能的步骤。
[0285] 显然,本领域的技术人员可以对本申请进行各种改动和变型而不脱离本申请的精神和范围。这样,倘若本申请的这些修改和变型属于本申请权利要求及其等同技术的范围之内,则本申请也意图包含这些改动和变型在内。< / script>
Claims
1. An attack detection method, characterized in that, include: Obtain the original data to be detected, and decode the original data to be detected to obtain the data to be detected; A lexical analysis method with different word segmentation closure conditions is used to segment the data to be detected. After each word segmentation is obtained, semantic analysis is performed on the word segmentation. After confirming that the word segmentation is lexical correct, the remaining data of the data to be detected is segmented to generate a syntax tree containing the target features. Each word segmentation closure condition corresponds to a syntax tree. For each target feature in the syntax tree, the following matching and scoring operations are performed: matching the target feature with each attack feature in the attack feature set to determine the target attack feature that matches the target feature; determining the score corresponding to the target feature based on the weight corresponding to the target attack feature; until each target feature in the syntax tree has completed the matching and scoring operation, calculating the sum of the scores corresponding to each target feature in the syntax tree to obtain the JS score; determining the HTML score of the HTML information corresponding to the HTML type segmentation; calculating the sum of the JS score and the HTML score, and using the sum of the JS score and the HTML score as the threat score of the original data to be detected under the corresponding segmentation closure condition; wherein, a maximum value is set for each target feature in the syntax tree, and when calculating the score corresponding to a single target feature, if the cumulative score corresponding to the single target feature is greater than the maximum value set for the single target feature, the maximum value of the single target feature is used as the final score of the single target feature; Obtain the threat score corresponding to each word segmentation closure condition; determine whether there is a target threat score higher than the alarm threshold among all the threat scores; if not, determine that the original data to be detected is not attack data; if so, determine that the original data to be detected is attack data.
2. The method as described in claim 1, characterized in that, Decoding the original data to be detected to obtain the data to be detected includes: In the original data to be detected, the data to be decoded that has undergone encoding processing is identified; According to the decoding type corresponding to the data to be decoded, the data to be decoded is decoded to obtain the restored data; The restored data and the original data in the data to be detected that has not undergone the encoding process are used as the data to be detected.
3. The method as described in claim 1, characterized in that, The word segmentation closure condition is any one of the following: first closure condition, second closure condition, third closure condition, fourth closure condition, fifth closure condition, and sixth closure condition. The first closure condition is that if the initial word segmentation mode is HTML lexical mode, then the word segmentation is performed at the first splitting character to obtain the first word; the first splitting character is a whitespace character or a greater than sign. The second closure condition is that if the initial word segmentation mode is the HTML lexical mode, then the word segmentation is performed at the second splitting character to obtain the first word; the second splitting character is a single quote. The third closure condition is that if the initial word segmentation mode is the HTML lexical mode, then the word segmentation is performed at the third splitting character to obtain the first word; the third splitting character is a double quotation mark. The fourth closure condition is that if the initial word segmentation mode is a Javascript lexical mode, then the word segmentation is performed at the fourth splitting character to obtain the first word; the fourth splitting character is the whitespace character or the JS boundary character. The fifth closure condition is that if the initial word segmentation mode is the Javascript lexical mode, then the word segmentation is performed at the fifth splitting character to obtain the first word; the fifth splitting character is a quotation mark. The sixth closure condition is that if the initial word segmentation mode is the Javascript lexical mode, then the word segmentation is performed at the sixth splitting character to obtain the first word; the sixth splitting character is an asterisk character and a forward slash character.
4. The method as described in claim 1, characterized in that, The word segmentation process for the data to be detected includes: If the current lexical mode is HTML lexical mode and the first switching condition is triggered, then the current lexical mode is switched to Javascript lexical mode, and the word segmentation processing is performed on the data to be detected according to the Javascript lexical mode; When the initial lexical mode is the HTML lexical mode, if the current lexical mode is the Javascript mode and the second switching condition is triggered, the current lexical mode is switched to the HTML lexical mode, and the word segmentation processing is performed on the data to be detected according to the HTML lexical mode.
5. An attack detection device, characterized in that, include: The decoding module is used to acquire the original data to be detected and decode the original data to be detected to obtain the data to be detected. The word segmentation and semantic analysis module is used to perform word segmentation on the data to be detected using lexical analysis methods with different word segmentation closure conditions. After obtaining each word segment, semantic analysis is performed on the word segment. After confirming that the word segmentation is lexical correct, word segmentation is performed on the remaining data of the data to be detected to generate a syntax tree containing target features. Each word segmentation closure condition corresponds to a syntax tree. The matching and scoring module is used to perform the following matching and scoring operations for each target feature in the syntax tree: matching the target feature with each attack feature in the attack feature set to determine the target attack feature that matches the target feature; determining the score corresponding to the target feature according to the weight corresponding to the target attack feature; until the matching and scoring operations have been performed for each target feature in the syntax tree, calculating the sum of the scores corresponding to each target feature in the syntax tree to obtain the JS score; determining the HTML score of the HTML information corresponding to the HTML type segmentation; calculating the sum of the JS score and the HTML score, and using the sum of the JS score and the HTML score as the threat score of the original data to be detected under the corresponding segmentation closure condition; wherein, a maximum value is set for each target feature in the syntax tree, and when calculating the score corresponding to a single target feature, if the cumulative score corresponding to the single target feature is greater than the maximum value set for the single target feature, the maximum value of the single target feature is used as the final score of the single target feature; The processing module is used to obtain the threat score corresponding to each word segmentation closure condition; determine whether there is a target threat score higher than the alarm threshold among all the threat scores; if not, determine that the original data to be detected is not attack data; if so, determine that the original data to be detected is attack data.
6. The apparatus as claimed in claim 5, characterized in that, The word segmentation closure condition is any one of the following: first closure condition, second closure condition, third closure condition, fourth closure condition, fifth closure condition, and sixth closure condition. The first closure condition is that if the initial word segmentation mode is HTML lexical mode, then word segmentation is performed at the first splitting character to obtain the first word; the first splitting character is a whitespace character or a greater than sign. The second closure condition is that if the initial word segmentation mode is the HTML lexical mode, then word segmentation is performed at the second splitting character to obtain the first word; the second splitting character is a single quote. The third closure condition is that if the initial word segmentation mode is the HTML lexical mode, then word segmentation is performed at the third splitting character to obtain the first word; the third splitting character is a double quotation mark. The fourth closure condition is that if the initial word segmentation mode is Javascript lexical mode, then word segmentation is performed at the fourth splitting character to obtain the first word; the fourth splitting character is the whitespace character or the JS boundary character. The fifth closure condition is that if the initial word segmentation mode is the Javascript lexical mode, then word segmentation is performed at the fifth splitting character to obtain the first word; the fifth splitting character is a quotation mark. The sixth closure condition is that if the initial word segmentation mode is the Javascript lexical mode, then word segmentation is performed at the sixth splitting character to obtain the first word; the sixth splitting character is an asterisk character and a forward slash character.
7. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, when executing a computer program stored in the memory, implements the method of any one of claims 1-4.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1-4.
Citation Information
Patent Citations
Detection method and device for cross site scripting and firewall with device
CN102833269A
Cross-site scripting attack detection method and device
CN112883372A