A website AI intelligent risk assessment real-time avoidance system and its avoidance method

By combining VPN with artificial intelligence technology, web page content can be captured and evaluated in real time, solving the problem of traditional security protection's inadequate recognition of dynamic content and advanced threats, and achieving real-time risk assessment and security avoidance in complex network environments.

CN120512308BActive Publication Date: 2025-09-12SHANGHAI ZUOQI NETWORK TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510999685.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-09-12
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

Traditional web security protection methods are unable to effectively identify dynamic web content, encrypted and obfuscated code, and advanced persistent threat (APT) attacks. VPNs lack real-time risk analysis and protection capabilities, resulting in weak user network access security.

Method used

Combining VPN technology with artificial intelligence, it uses multimodal fusion methods to capture web content in real time, and uses deep learning models to perform multi-dimensional feature extraction and risk assessment, including domain name, web page structure, media resources, script code and network behavior analysis, to provide graded response strategies to avoid risks.

Benefits of technology

It achieves real-time risk assessment and accurate identification of complex network attacks, reduces missed reports and false alarms, provides flexible access policies to ensure user security, and adapts to complex and changing network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120512308B_ABST
    Figure CN120512308B_ABST
Patent Text Reader

Abstract

The present invention discloses a real-time website AI risk assessment and avoidance system and its avoidance method. The system includes a VPN connection module to ensure secure and concealed data capture; a data capture module to collect multiple types of information from the target website in real time; a risk feature extraction module to perform multi-dimensional feature extraction; an AI risk analysis and assessment model to identify and rate risks from multiple aspects such as images and scripts; a risk avoidance and warning module to implement a graded response strategy based on the results; and a feedback and model optimization module to collect feedback and sample optimization models. The method implements website risk avoidance through the steps of VPN connection, data collection, feature extraction, risk assessment, strategy execution, and model optimization. Combining VPN with artificial intelligence technology, the present invention can comprehensively and real-timely assess website risks, dynamically adjust access policies, effectively respond to complex network attacks, and ensure user network access security.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and specifically to a website AI intelligent risk assessment real-time avoidance system and avoidance method. The system aims to combine VPN technology with artificial intelligence technology to conduct real-time risk assessment of target websites and take corresponding avoidance measures according to the risk level to ensure user network access security. Background Art

[0002] With the rapid development of the Internet, network security issues are becoming increasingly prominent. Especially in the context of the popularization of remote work, cloud computing and mobile Internet, the network environment that users access is becoming increasingly complex.

[0003] Traditional web security protection methods, mostly based on static feature matching and rule engines, suffer from high rates of missed detections, false positives, and long detection response times, making them difficult to adapt to complex attack scenarios. This is particularly true for dynamic web content, encrypted and obfuscated code, and advanced persistent threat (APT) attacks, which traditional solutions often fail to effectively identify.

[0004] On the other hand, virtual private network (VPN) technology is widely used to protect user privacy and communication security due to its ability to encrypt user network communications and hide real IP addresses. However, VPNs lack real-time risk analysis and protection capabilities for accessed content, leaving users with weak links in the security chain unaddressed.

[0005] Therefore, how to combine VPN security connections with advanced artificial intelligence technology to capture and access web content in real time, use multimodal fusion methods such as deep learning to conduct comprehensive risk assessments of websites, and dynamically adjust access strategies based on risks has become a technical problem that urgently needs to be solved. Summary of the Invention

[0006] The purpose of this invention is to provide a website AI intelligent risk assessment real-time avoidance system and method. By combining VPN technology with artificial intelligence technology, real-time risk assessment of target websites can be achieved, and corresponding avoidance measures can be taken according to the assessment results to reduce the risk of users visiting malicious websites, improve the security and reliability of network access, and effectively respond to the current complex and changeable network attack environment.

[0007] In a first aspect, the technical solution of the present invention provides a website AI intelligent risk assessment and real-time avoidance system, comprising:

[0008] VPN connection module, used to establish a virtual private network connection to ensure the security and confidentiality of the data capture process;

[0009] The data capture module is used to collect the target website's web page structure, media resources, script code, dynamic content, and network request behavior in real time after the VPN connection is established;

[0010] A risk feature extraction module is used to extract multi-dimensional features of the target website content, including domain name information features, hypertext structure features, image content features, script static behavior features and dynamic behavior features, network communication behavior features, and historical reputation information;

[0011] An AI risk analysis and assessment model is used to identify and assess the risk level of a target website based on the multi-dimensional features. The AI ​​risk analysis and assessment model includes:

[0012] Image analysis module, used to identify potential embedded malicious content in images;

[0013] Script analysis module, used to detect obfuscation, encryption or malicious logic in JavaScript scripts;

[0014] Network behavior analysis module, used to analyze suspicious patterns in API calls, data interactions, and network requests;

[0015] Dynamic content analysis module, used to identify deceptive structures or social engineering attack behaviors dynamically generated by scripts;

[0016] a behavior pattern analysis module that uses a deep learning model to identify potentially malicious behavior patterns, wherein the deep learning model is a bidirectional encoding representation text model, a convolutional neural network model, a graph neural network model, or a combination thereof; and

[0017] A risk assessment module, which is used to integrate the analysis results of each module in the AI ​​risk analysis and assessment model and output the risk type and risk level of the target website;

[0018] A risk avoidance and warning module, which is used to automatically execute a graded response strategy based on the results provided by the risk assessment module. The strategy includes access blocking, script interception, content degradation, permission restriction, and user warning prompts;

[0019] The feedback and model optimization module is used to continuously collect user feedback, record security incidents, and collect real attack samples. It is used to adjust parameters and optimize the structure of the AI ​​risk analysis and assessment model to improve the overall detection accuracy and adaptability of the system.

[0020] Wherein, the risk feature extraction module includes:

[0021] A domain name information extraction unit, used to obtain the domain name and related registration information of the target website;

[0022] HTML structural feature extraction unit, used to parse and extract the hypertext tag hierarchy and attributes;

[0023] Media feature extraction unit, used to analyze the content features of image and video media resources;

[0024] Script feature extraction unit, used to extract the static syntax structure and calling behavior of the script;

[0025] Communication behavior feature extraction unit, used to count and analyze network request headers, parameters and frequencies;

[0026] The reputation information extraction unit is used to query and record historical reputation scores.

[0027] The risk levels include:

[0028] Safety, the corresponding value range is 0≤R<0.1;

[0029] Low risk, corresponding to the value range of 0.1≤R<0.5;

[0030] Medium risk, corresponding to the value range of 0.5≤R<0.8;

[0031] High risk, the corresponding numerical range is 0.8≤R≤1.0.

[0032] The risk avoidance and warning module includes:

[0033] A risk warning unit is used to display corresponding warning information and safety suggestions to the user according to the risk level;

[0034] The interception decision unit is used to execute access blocking, forced redirection to a secure page, or user confirmation access strategies when the risk level is high.

[0035] The feedback and model optimization module includes:

[0036] The active learning unit is used to prioritize the misjudgment samples reported by users and the successfully identified attack samples into the training set to enhance the model adaptability and detection accuracy.

[0037] In a second aspect, the technical solution of the present invention provides a method for real-time avoidance of website AI intelligent risk assessment, comprising the following steps:

[0038] Step 1: Establish a secure network connection through the VPN connection module, and use the data capture module to collect the target website's web page structure, media resources, script code, and network request behavior in real time;

[0039] Step 2: Input the data collected in step 1 into the risk feature extraction module to extract multi-dimensional features such as domain name information, HTML structure, image content, script static and dynamic behavior, network communication behavior, and historical reputation;

[0040] Step 3: The features extracted in Step 2 are input into the AI ​​risk analysis and assessment model. The risk is assessed in parallel by the image analysis module, script analysis module, network behavior analysis module, dynamic content analysis module, and behavior pattern analysis module.

[0041] Step 4: The risk assessment module integrates the assessment results of each sub-module in step 3 to determine the risk type and risk level of the target website;

[0042] Step 5: The risk avoidance and warning module executes a graded response strategy of access blocking, script interception, content degradation, permission restriction or user warning prompt according to the risk type and risk level;

[0043] Step 6: The feedback and model optimization module collects user interaction feedback and attack samples generated during actual detection, and continuously learns and optimizes the structure of the AI ​​risk analysis and assessment model.

[0044] The process of the image analysis module detecting malicious content hidden in image pixels includes:

[0045] Step 7.1: The data crawling module extracts the image resources of the target website in real time;

[0046] Step 7.2: The image analysis module uses a convolutional neural network to extract pixel differences, color distribution, and frequency domain features;

[0047] Step 7.3: The image analysis module calculates the image histogram based on statistical methods and determines the steganographic features of the abnormal color distribution areas;

[0048] Step 7.4: The image analysis module compares the extracted features with the malicious image feature library and outputs the risk score for each image area. and the corresponding timestamp ;

[0049] Step 7.5: Risk Assessment Module Based on Formula Calculate the overall image risk score,

[0050] in:

[0051] represents the dynamic risk score of the image calculated by the risk assessment module;

[0052] Indicates the total number of regions divided by the analyzed image;

[0053] Indicates the The risk score of each image region is calculated by the image analysis module;

[0054] Indicates the Timestamp of the risk score of each image region;

[0055] Indicates the current time when the risk assessment is performed;

[0056] Represents the time decay coefficient, which is used to dynamically reduce the impact of outdated risk scores;

[0057] Step 7.6: Risk avoidance and warning module according to the The score determines whether to block access or prompt a security warning.

[0058] The process of detecting JavaScript code variants by the script analysis module includes:

[0059] Step 8.1: The data scraping module extracts the JavaScript script embedded in the web page;

[0060] Step 8.2: The static analysis unit identifies obfuscated, encrypted, and disguised instructions in the script based on the abstract syntax tree technology;

[0061] Step 8.3: The behavior pattern analysis unit uses the pre-trained language model to encode the script semantics and compares it with known malicious patterns to obtain the probability value p_{mal};

[0062] Step 8.4: The dynamic analysis unit executes the script in the sandbox environment to monitor and count abnormal behavior events. and total number of events ;

[0063] Step 8.5: Risk Assessment Module Based on Formula Calculate the script risk score,

[0064] in,

[0065] : Comprehensive risk score of JavaScript scripts, ranging from [0,1];

[0066] : Risk normalization factor, a positive constant used to limit the upper limit of the score;

[0067] : Risk score weight coefficient, satisfying ;

[0068] : The number of suspicious syntax structures detected during static analysis, including obfuscation, encryption, and renamed instructions;

[0069] : The number of abnormal behavior events detected in dynamic execution, including cross-site scripting, redirect hijacking, and data leakage;

[0070] : The total number of behavioral events identified during the dynamic analysis process, used to standardize the abnormality ratio;

[0071] : The probability of the script being malicious, obtained through semantic model analysis, ranges from [0,1];

[0072] Step 8.6: Risk avoidance and warning module according to the Scoring executes script blocking or user prompting strategies.

[0073] The process of analyzing the API request by the network behavior analysis module includes:

[0074] Step 9.1: The data capture module collects all API request records during the loading process;

[0075] Step 9.2: The network request parsing unit parses the URL, request headers and parameters of each request and counts the call frequency ;

[0076] Step 9.3: The network request parsing unit identifies cross-domain, high-frequency, and forged identity behaviors and outputs an identity forgery score. ;

[0077] Step 9.4: The API behavior analysis unit tracks the connection path between the request and the sensitive data node based on the graph neural network and outputs the graph structure risk score. ;

[0078] Step 9.5: Risk Assessment Module Based on Formula Calculate the API risk score, where:

[0079] Indicates the web API behavior risk score;

[0080] Indicates the total number of API requests triggered during the loading process of the target web page;

[0081] Indicates the The calling frequency of each API request is obtained by the network request analysis module;

[0082] Indicates the The graph structure risk score of each API request connecting to a sensitive node is generated by the API behavior analysis module;

[0083] Indicates the The identity forgery risk score of each API request is output by the network request analysis module;

[0084] a, b, and c are preset or trained weight coefficients used to adjust the relative contributions of frequency, structure, and identity factors in the overall risk score;

[0085] Step 9.6: Risk avoidance and warning module according to the The scoring execution request interception or demotion strategy is implemented.

[0086] The process of calculating the overall risk of a webpage by the risk assessment module includes:

[0087] Step 10.1: The dynamic content monitoring unit monitors the advertisement frames, pop-ups, and prompt bars rendered by the script in real time;

[0088] Step 10.2: The dynamic content monitoring unit identifies phishing links, fake login boxes, and misleading icons;

[0089] Step 10.3: The behavior pattern analysis unit extracts the abnormal intensity of the behavior based on the user interaction trajectory ;

[0090] Step 10.4: Dynamic content monitoring unit counts dynamic loading frequency and static risk factors ;

[0091] Step 10.5: Risk Assessment Module Based on Formula Calculate the overall risk score of the webpage, where:

[0092] Indicates the overall risk level of the webpage;

[0093] represents the number of risk factors;

[0094] Normalize the total weight of all factors;

[0095] For the The weighting coefficients of the static risk factors;

[0096] The first Static risk factor scores, including but not limited to script source credibility and ad source domain reputation;

[0097] For the The weight coefficient of each behavior-related factor;

[0098] The abnormal intensity value of user behavior identified by the behavior pattern analysis module, including abnormal mouse trajectory and multiple interactions in a short period of time;

[0099] Dynamic content loading frequency factor, indicating the number of newly added DOM nodes, scripts, or resources per unit time;

[0100] is the frequency sensitivity adjustment factor, which is used to control the nonlinear effect of high-frequency content loading on the score;

[0101] Step 10.6: Risk avoidance and warning module according to the The access policy is adjusted dynamically based on the score.

[0102] The beneficial effects of the technical solution of the present invention are:

[0103] 1. This invention comprehensively assesses website risks through multi-dimensional feature extraction and multi-module parallel analysis. It extracts features from multiple perspectives, including domain name information, web page structure, media resources, script code, and network communication behavior. It utilizes multiple modules, including image analysis, script analysis, and network behavior analysis, to identify different types of risks. Compared to traditional assessment methods based on single features or rules, this method can more accurately identify various potential risks and reduce the rates of missed and false positives.

[0104] 2. This invention securely captures data by integrating VPN connections, collecting various information about target websites in real time and promptly conducting risk assessments and responses. It can monitor changes in website content in real time while users are accessing the website, promptly identifying emerging risks such as dynamically generated malicious scripts and real-time changes in network request behavior. It can then quickly implement appropriate mitigation measures based on the risk level, ensuring real-time user access security.

[0105] 3. This invention utilizes deep learning models for risk analysis, such as convolutional neural networks, pre-trained language models, and graph neural networks, to automatically learn the patterns and characteristics of malicious behavior and adapt to complex and diverse cyberattacks. By learning from large amounts of sample data, the model continuously optimizes its risk identification capabilities, particularly for novel, encrypted, and obfuscated malicious code and attack behaviors, providing enhanced detection capabilities.

[0106] 4. This invention provides a variety of tiered response strategies based on risk assessment results, including access blocking, script interception, content degradation, permission restrictions, and user warnings. This allows for flexible selection of the most appropriate mitigation strategy based on risk levels, ensuring user safety while minimizing the impact on normal user access. For example, for low-risk websites, users are simply notified of the risk without impacting normal access; for high-risk websites, access is decisively blocked to prevent losses.

[0107] 5. This invention continuously optimizes the AI ​​risk analysis and assessment model through the feedback and model optimization module, collecting user feedback and real attack samples. The system can timely update the model, improving detection accuracy and adaptability. User feedback of false positives and successfully identified attack samples provides valuable data for model optimization, enabling the model to continuously learn and improve, maintaining its ability to detect the latest risks. BRIEF DESCRIPTION OF THE DRAWINGS

[0108] Figure 1 This is a structural block diagram of the website AI intelligent risk assessment and real-time avoidance system in an embodiment of the present invention;

[0109] Figure 2 This is a structural block diagram of a risk feature extraction module in an embodiment of the present invention;

[0110] Figure 3 This is a structural block diagram of the risk avoidance and warning module in an embodiment of the present invention;

[0111] Figure 4 This is a structural block diagram of the feedback and model optimization module in an embodiment of the present invention;

[0112] Figure 5 This is a method flow chart of the real-time avoidance method for website AI intelligent risk assessment in an embodiment of the present invention. DETAILED DESCRIPTION

[0113] In order to better understand the above technical solutions, the technical solutions of the embodiments of this specification are described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of this specification and the specific features in the embodiments are detailed descriptions of the technical solutions of the embodiments of this specification, rather than limitations on the technical solutions of this specification. In the absence of conflict, the embodiments of this specification and the technical features in the embodiments can be combined with each other.

[0114] 1. System composition

[0115] like Figure 1 As shown, the website AI intelligent risk assessment and real-time avoidance system of the present invention includes the following modules:

[0116] VPN Connection Module: This module establishes a virtual private network connection to ensure the security and confidentiality of the data capture process. By encrypting network communications between the user and the target website, the user's true IP address is hidden, protecting the user from attacks such as network monitoring and IP tracking during the data capture process, and providing a secure network channel for subsequent data collection. For example, the VPN Connection Module can utilize common VPN protocols such as IPsec and OpenVPN, automatically selecting the appropriate protocol and configuration parameters for connection establishment based on the user's network environment and security requirements.

[0117] Data Capture Module: After the VPN connection is established, this module collects the target website's webpage structure, media resources, script code, dynamic content, and network request behavior in real time. It simulates a browser loading the target website and captures various data and information during the website's rendering process. For example, it parses HTML documents to obtain the webpage's hierarchical structure and identify various HTML tags and their attributes; extracts links to media resources such as images and videos from the webpage and further obtains the media resource content; captures JavaScript script code embedded in the webpage; and records the network requests generated by the website during the loading process, including the request URL, request headers, and parameters.

[0118] Risk feature extraction module: extracts multi-dimensional features from the target website content. The extracted features include domain name information features, hypertext structure features, image content features, script static and dynamic behavior features, network communication behavior features, and historical reputation information. Figure 2 As shown, it specifically includes the following units:

[0119] Domain Name Information Extraction Unit: This unit is responsible for obtaining the target website's domain name and related registration information, such as the domain name's registration date, registrar, and domain owner information. This information can be obtained through methods such as WHOIS queries. A recent domain name registration, a poorly reputable registrar, or unclear domain owner information may indicate a website is risky.

[0120] HTML structural feature extraction unit: parses and extracts the hypertext tag hierarchy and attributes, such as the nested relationship of tags, specific tags (such as <script>、等)的属性值。异常的HTML结构,如大量嵌套不合理的标签、标签指向不明的链接等,可能与恶意网页相关。

[0121] 媒体特征提取单元:分析图片、视频等媒体资源的内容特征,对于图片,可提取其像素分布、颜色直方图、纹理特征等;对于视频,可提取关键帧特征、视频编码格式等信息。隐藏在媒体资源中的恶意代码或异常内容,可通过这些特征进行识别。

[0122] 脚本特征提取单元:提取脚本的静态语法结构和调用行为,例如函数定义、变量声明、函数调用关系等。通过分析脚本的静态结构,可发现混淆、加密或异常的脚本逻辑。

[0123] 通信行为特征提取单元:统计并分析网络请求头、参数及频率,例如请求头中的User-Agent、Referer等字段信息,请求参数的类型和值,以及特定请求的发起频率。异常的请求头信息、参数内容或过高的请求频率,都可能是恶意行为的表现。

[0124] 信誉信息提取单元:查询并记录历史信誉评分,可通过与第三方信誉评级机构合作,获取目标网站的历史信誉数据,如是否曾经被标记为恶意网站、信誉评级的变化趋势等。

[0125] AI风险分析评估模型:基于多维特征对目标网站进行风险识别与等级评估,该模型包括以下模块:

[0126] 图像分析模块:识别图像中潜在的嵌入式恶意内容。它使用卷积神经网络(CNN)提取像素差异、颜色分布及频域特征,基于统计方法计算图像直方图,并对色彩分布异常区域进行隐写特征判断,将提取特征与恶意图像特征库比对,输出各图像区域的风险评分及对应时间戳。例如,CNN模型通过卷积层和池化层对图像进行特征提取,学习正常图像与隐藏恶意内容图像之间的特征差异,从而识别出图像中可能隐藏的恶意代码、恶意链接等内容。

[0127] 脚本分析模块:检测JavaScript脚本中的混淆、加密或恶意逻辑。通过基于抽象语法树(AST)技术识别脚本中的混淆、加密及伪装指令,调用预训练语言模型对脚本语义进行编码,并与已知恶意模式进行比对获得概率值,在沙盒环境执行脚本监控并统计异常行为事件数及总事件数,最终计算脚本风险评分。例如,AST技术能够将JavaScript脚本解析为树状结构,便于分析脚本的语法和逻辑,从而发现混淆、加密等隐藏恶意逻辑的手段。

[0128] 网络行为分析模块:分析API调用、数据交互及网络请求中的可疑模式。通过采集加载过程中的所有API请求记录,解析请求的URL、请求头及参数,统计调用频率,识别跨域、高频及伪造身份行为,基于图神经网络(GNN)追踪请求与敏感数据节点的连接路径,输出图结构风险评分,进而计算API风险评分。例如,GNN模型可以将API请求和数据交互构建为图结构,通过分析图中节点和边的关系,发现潜在的数据泄露、异常数据交互等风险。

[0129] 动态内容分析模块:识别由脚本动态生成的诱导性结构或社会工程学攻击行为。实时监听由脚本渲染的广告框、弹窗及提示条,识别钓鱼链接、伪造登录框及诱导性图标,并结合用户交互轨迹提取行为异常强度,统计动态加载频率及静态风险因子,计算网页整体风险评分。例如,通过分析用户与动态生成内容的交互行为,如点击链接的频率、鼠标移动轨迹等,判断是否存在社会工程学攻击的迹象。

[0130] 行为模式分析模块:使用深度学习模型对潜在恶意行为模式进行识别,深度学习模型包括双向编码表示的文本模型(如BERT)、卷积神经网络模型(CNN)、图神经网络模型(GNN)或其组合。不同的模型适用于不同类型的数据和行为模式分析,例如BERT可用于分析文本内容中的语义信息,识别潜在的恶意文本;CNN适用于图像和视频数据的特征提取;GNN适用于分析具有图结构的数据,如网络拓扑、数据交互关系等。通过这些模型的组合使用,可以更全面地识别各种潜在恶意行为模式。

[0131] 风险评估模块:综合AI风险分析评估模型中各模块的分析结果,输出目标网站的风险类型与风险等级。根据不同模块输出的风险评分,按照预设的权重和算法进行综合计算,确定网站的风险类型(如恶意软件感染、钓鱼攻击、数据泄露风险等)和风险等级。

[0132] 如图3所示,风险规避与警告模块:根据风险评估模块提供的结果自动执行分级响应策略,策略包括访问阻断、脚本拦截、内容降级、权限限制以及用户警告提示。具体包括以下单元:

[0133] 风险提示单元:根据风险等级向用户展示相应警告信息及安全建议。例如,对于低风险等级,提示用户"该网站可能存在轻微风险,建议谨慎操作”,并提供一些通用的安全建议,如不随意点击不明链接;对于高风险等级,强烈警告用户"该网站存在严重风险,可能导致信息泄露或设备感染病毒,请立即停止访问”。

[0134] 拦截决策单元:当风险等级为高风险时执行访问阻断、强制跳转安全页面或用户确认访问的策略。访问阻断可通过修改本地主机文件、设置防火墙规则等方式阻止用户访问目标网站;强制跳转安全页面则将用户引导至系统预设的安全页面,告知用户风险情况;用户确认访问策略则在提示用户风险后,由用户自行决定是否继续访问,但会对后续操作进行更严格的监控。

[0135] 如图4所示,反馈与模型优化模块:持续收集用户反馈、记录安全事件并采集真实攻击样本,用以对AI风险分析评估模型进行参数调整与结构优化,提升系统整体检测准确率与适应性。该模块包括主动学习单元,其用于将用户反馈的误判样本及成功识别的攻击样本优先加入训练集,以增强模型适应性和检测准确率。例如,当用户反馈某个被判定为高风险的网站实际上是正常网站时,主动学习单元将该网站的相关数据及用户反馈信息整理为误判样本,加入训练集,对AI风险分析评估模型进行微调,使其能够更准确地识别类似网站。

[0136] 二、如图5所示,本发明的网站AI智能风险评估实时规避方法包括以下步骤:

[0137] 步骤1:通过VPN连接模块建立安全网络连接,并由数据抓取模块实时采集目标网站的网页结构、媒体资源、脚本代码及网络请求行为。首先,VPN连接模块根据用户的网络环境和安全需求,选择合适的VPN协议和配置参数,与VPN服务器建立加密连接。数据抓取模块在VPN连接成功后,模拟浏览器行为加载目标网站,获取网站的各种数据信息,为后续的风险评估提供数据基础。

[0138] 步骤2:将步骤1中采集的数据输入至风险特征提取模块,提取域名信息、HTML结构、图像内容、脚本静态与动态行为、网络通信行为及历史信誉等多维度特征。风险特征提取模块的各个单元分别对输入数据进行处理,域名信息提取单元获取域名及注册信息,HTML结构特征提取单元解析超文本标签结构和属性,媒体特征提取单元分析媒体资源内容特征,脚本特征提取单元提取脚本的静态和动态行为特征,通信行为特征提取单元统计网络通信行为特征,信誉信息提取单元查询并记录历史信誉信息。

[0139] 步骤3:将步骤2中提取的特征输入至AI风险分析评估模型,分别由图像分析模块、脚本分析模块、网络行为分析模块、动态内容分析模块及行为模式分析模块这几个子模块并行评估风险。各个子模块根据自身的功能和算法对输入特征进行分析,图像分析模块识别图像中的恶意内容,脚本分析模块检测脚本中的恶意逻辑,网络行为分析模块分析网络行为中的可疑模式,动态内容分析模块识别动态生成内容中的诱导性结构和攻击行为,行为模式分析模块使用深度学习模型识别潜在恶意行为模式。

[0140] 步骤4:由风险评估模块综合步骤3各子模块的评估结果,判断目标网站的风险类型及风险等级。风险评估模块根据预设的权重和算法,对各子模块输出的风险评分进行综合计算,确定目标网站的风险类型和风险等级。风险等级分为安全(0≤R<0.1)、低风险(0.1≤R<0.5)、中风险(0.5≤R<0.8)、高风险(0.8≤R≤1.0)。

[0141] 步骤5:由风险规避与警告模块根据风险类型及风险等级执行访问阻断、脚本拦截、内容降级、权限限制或用户警告提示等分级响应策略。风险规避与警告模块的风险提示单元根据风险等级向用户展示相应警告信息及安全建议,拦截决策单元在风险等级为高风险时执行相应的拦截或跳转策略。

[0142] 步骤6:由反馈与模型优化模块收集用户交互反馈及实际检测中产生的攻击样本,对AI风险分析评估模型进行持续学习与结构优化。反馈与模型优化模块的主动学习单元将用户反馈的误判样本及成功识别的攻击样本优先加入训练集,对AI风险分析评估模型进行参数调整和结构优化,提升模型的检测准确率和适应性。

[0143] 图像分析模块检测隐藏在图片像素中的恶意内容的过程:

[0144] 步骤7.1:数据抓取模块实时提取目标网站的图片资源,确保获取到网站中所有的图片内容,包括嵌入在网页中的普通图片、背景图片、图标等。

[0145] 步骤7.2:图像分析模块使用卷积神经网络提取像素差异、颜色分布及频域特征。通过卷积层和池化层对图片进行处理,学习图片的局部特征和全局特征,识别像素值的异常变化、颜色分布的偏离以及频域中的异常信号。

[0146] 步骤7.3:图像分析模块基于统计方法计算图像直方图,并对色彩分布异常区域进行隐写特征判断。通过分析直方图的形状、峰值和谷值,发现颜色分布异常的区域,这些区域可能隐藏有恶意内容,进一步对这些区域进行隐写特征分析,判断是否存在隐藏的恶意代码或信息。

[0147] 步骤7.4:图像分析模块将提取特征与恶意图像特征库比对,输出各图像区域的风险评分及对应时间戳。恶意图像特征库中存储了已知的恶意图像特征,通过比对计算各图像区域与恶意特征的相似度,得出风险评分,并记录评分的时间戳。

[0148] 步骤7.5:风险评估模块根据公式计算整体图像风险评分。其中:表示由风险评估模块计算得到的图像动态风险评分;表示被分析图像划分的区域总数;表示第个图像区域的风险评分,由图像分析模块计算得出;表示第个图像区域风险评分的时间戳;表示风险评估时的当前时间;表示时间衰减系数,用于动态降低过时风险评分的影响。通过这个公式,综合考虑各图像区域的风险评分以及评分的时效性,得出整体图像的风险评分。

[0149] 步骤7.6:风险规避与警告模块根据的评分判定是否阻断访问或提示安全警告。如果的评分超过预设的阈值,表明图像存在较高风险,风险规避与警告模块执行阻断访问操作,阻止用户查看该图片或相关网页;如果的评分处于较低范围,则向用户提示相应的安全警告,告知用户图片可能存在一定风险。

[0150] 脚本分析模块检测JavaScript代码变种的过程:

[0151] 步骤8.1:数据抓取模块提取网页中嵌入的JavaScript脚本,确保获取到网页运行所需的所有脚本代码,包括外部引用的脚本和内联脚本。

[0152] 步骤8.2:静态分析单元基于抽象语法树技术识别脚本中的混淆、加密及伪装指令。将JavaScript脚本解析为抽象语法树,通过分析树的节点结构和关系,识别出使用混淆、加密手段隐藏恶意逻辑的指令,统计检测到的可疑语法结构数量。

[0153] 步骤8.3:行为模式分析单元调用预训练语言模型对脚本语义进行编码,并与已知恶意模式进行比对,获得概率值。预训练语言模型能够学习脚本的语义信息,通过与已知恶意脚本模式的比对,计算该脚本为恶意脚本的概率。

[0154] 步骤8.4:动态分析单元在沙盒环境执行脚本,监控并统计异常行为事件数及总事件数。沙盒环境提供了一个安全的执行空间,在其中运行脚本并监控其行为,记录异常行为事件(如跨站脚本攻击、数据外泄等)的数量以及总的行为事件数。

[0155] 步骤8.5:风险评估模块根据公式计算脚本风险评分。其中,:JavaScript脚本的综合风险评分,范围为[0,1];:风险归一化因子,为正数常量,用于限制评分上限;:风险评分权重系数,满足;:在静态分析中检测到的可疑语法结构数量,包括混淆、加密、重命名指令等;:在动态执行中检测到的异常行为事件数,包括跨站脚本、重定向劫持、数据外泄等;:动态分析过程中总共识别的行为总事件数,用于标准化异常比例;:语义模型分析得出的该脚本为恶意脚本的概率,取值范围为[0,1]。通过这个公式,综合考虑静态分析、动态分析以及语义分析的结果,得出脚本的风险评分。

[0156] 步骤8.6:风险规避与警告模块根据的评分执行脚本阻断或用户提示策略。如果的评分超过一定阈值,表明脚本存在较高风险,风险规避与警告模块执行脚本阻断操作,阻止脚本的执行;如果的评分处于较低范围,则向用户提示脚本可能存在风险,由用户决定是否继续执行脚本。

[0157] 网络行为分析模块分析API请求的过程:

[0158] 步骤9.1:数据抓取模块采集加载过程中的所有API请求记录,确保获取到网页在加载过程中发起的所有API请求信息,包括请求的URL、请求头、参数以及请求的时间等。

[0159] 步骤9.2:网络请求解析单元解析每请求的URL、请求头及参数,统计调用频率。分析请求的URL结构,提取请求头中的关键信息(如User-Agent、Referer等),解析请求参数的内容和类型,并统计每个API请求的调用频率。

[0160] 步骤9.3:网络请求解析单元识别跨域、高频及伪造身份行为,输出身份伪造评分。通过检查请求的源域和目标域是否相同来判断跨域行为,设定合理的频率阈值来识别高频请求行为,同时通过分析请求头中的身份标识信息(如Cookie、Token等)以及与已知合法身份模式的比对,评估请求的身份伪造可能性,从而输出身份伪造评分。

[0161] 步骤9.4:API行为分析单元基于图神经网络追踪请求与敏感数据节点的连接路径,输出图结构风险评分。将API请求、数据交互以及相关的数据节点构建成图结构,利用图神经网络分析节点之间的连接关系和数据流向。例如,若发现某个API请求频繁访问敏感数据节点(如用户账户信息、财务数据等)且连接路径异常,表明存在数据泄露风险,进而输出图结构风险评分。

[0162] 步骤9.5:风险评估模块根据公式计算API风险评分。其中,表示网页API行为风险评分;表示目标网页在加载过程中共触发的API请求总数;表示第个API请求的调用频率,由网络请求分析模块统计获得;表示第API请求连接到敏感节点的图结构风险评分,由API行为分析模块生成;表示第个API请求的身份伪造风险评分,由网络请求分析模块输出;a、b、c为预设或训练得到的权重系数,用于调节频率、结构和身份因素在整体风险评分中的相对贡献。通过该公式,综合考虑各个API请求的调用频率、连接敏感节点的风险以及身份伪造风险,得出API的整体风险评分。

[0163] 步骤9.6:风险规避与警告模块根据的评分执行请求拦截或降权策略。若的评分超过设定的阈值,表明该API请求存在较高风险,风险规避与警告模块执行请求拦截操作,阻止该API请求的进一步执行;若的评分处于相对较低但仍有一定风险的范围,则对该API请求进行降权处理,例如降低其网络请求优先级,减少其占用的网络资源,以降低潜在风险。

[0164] 风险评估模块计算网页整体风险的过程:

[0165] 步骤10.1:动态内容监控单元实时监听由脚本渲染的广告框、弹窗及提示条。通过监测网页DOM树的变化,及时捕捉由脚本动态生成的这些元素,因为这些元素常被用于诱导用户进行危险操作或实施社会工程学攻击。

[0166] 步骤10.2:动态内容监控单元识别钓鱼链接、伪造登录框及诱导性图标。利用模式匹配、特征提取等技术,将监测到的元素与已知的钓鱼链接模式、伪造登录框特征以及诱导性图标样式进行比对,识别出潜在的恶意元素。

[0167] 步骤10.3:行为模式分析单元结合用户交互轨迹提取行为异常强度。通过记录用户的鼠标移动轨迹、点击行为、输入内容等交互信息,分析用户行为是否符合正常模式。例如,如果用户在短时间内频繁点击可疑链接,或者鼠标移动轨迹呈现异常的急促或不规则,表明用户行为存在异常,进而提取行为异常强度。

[0168] 步骤10.4:动态内容监控单元统计动态加载频率及静态风险因子。动态加载频率指单位时间内网页新增DOM节点、脚本或资源的数量,反映了网页动态内容的更新活跃程度。静态风险因子包括脚本来源可信度、广告来源域名声誉等信息,通过查询相关的信誉数据库或利用预设的规则进行评估。

[0169] 步骤10.5:风险评估模块根据公式计算网页整体风险评分。其中:表示网页整体风险等级;表示风险因子的数量;为所有因子归一化总权重;为第个静态风险因子的加权系数;为由动态内容监控模块提取的第个静态风险因子得分,包括但不限于脚本来源可信度、广告来源域名声誉等;为第个行为相关因子的权重系数;为由行为模式分析模块识别出的用户行为异常强度值,包括鼠标轨迹反常、短时间内多次交互等行为特征;为动态内容加载频率因子,表示单位时间内新增DOM节点、脚本或资源数量;为频率敏感度调节因子,用于控制高频内容加载对评分的非线性影响程度。该公式综合考虑了静态风险因子、用户行为异常强度以及动态加载频率等因素,全面评估网页的整体风险。

[0170] 步骤10.6:风险规避与警告模块根据的评分动态调整访问策略。若的评分较高,表明网页风险较大,风险规避与警告模块可能执行访问阻断操作,阻止用户继续访问该网页;若的评分处于中等风险范围,可能采取内容降级策略,例如限制网页部分功能的加载,或对高风险的动态内容进行屏蔽;若的评分较低,向用户提示潜在风险,并提供相关的安全建议。

[0171] 以下结合具体场景对本发明的网站AI智能风险评估实时规避系统及方法进行详细说明。

[0172] 假设用户尝试访问一个未知网站,系统开始工作。

[0173] VPN连接与数据抓取

[0174] VPN连接模块:用户设备上的系统检测到用户发起网页访问请求,VPN连接模块启动。根据用户设备当前的网络环境和预设的安全策略,VPN连接模块选择OpenVPN协议与VPN服务器建立连接。经过身份验证和密钥协商等过程,成功建立起加密的虚拟专用网络连接,确保后续数据传输的安全性和隐蔽性。

[0175] 数据抓取模块:在VPN连接建立后,数据抓取模块模拟浏览器加载目标网站。它首先获取网页的HTML文档,解析其中的标签结构和属性,提取网页的层次结构信息。例如,识别出页面中的导航栏、内容区域、广告区域等不同部分的HTML标签及其嵌套关系。同时,提取网页中的媒体资源,如图片、视频等。假设网页中有一张产品展示图片和一个介绍视频,数据抓取模块获取它们的链接,并进一步下载图片和视频内容。此外,数据抓取模块还抓取嵌入网页的JavaScript脚本代码,以及记录网页在加载过程中产生的网络请求,如向服务器请求用户数据的API请求,记录请求的URL、请求头(如User-Agent:Mozilla / 5.0(WindowsNT10.0;Win64;x64)AppleWebKit / 537.36(KHTML,likeGecko)Chrome / 91.0.4472.124Safari / 537.36)和参数(如user_id=12345)。

[0176] 风险特征提取

[0177] 域名信息提取单元:通过WHOIS查询获取目标网站的域名注册信息,假设该网站域名为"example.com”,注册时间为2023年1月1日,注册商为一家不太知名的公司,且域名所有者信息显示为匿名。这些信息被记录作为域名信息特征。

[0178] HTML结构特征提取单元:对获取的HTML文档进行深入解析,发现网页中存在大量嵌套过深的标签,且部分标签的href属性指向一些可疑的URL,如"http: / / abcpqr.com / redirect.php”,这些异常的HTML结构特征被提取出来。

[0179] 媒体特征提取单元:对于之前抓取的产品展示图片,媒体特征提取单元使用图像处理算法提取其像素分布、颜色直方图等特征。发现图片的颜色分布在某些区域与正常图片有明显差异,可能存在隐藏信息的迹象。对于视频,提取关键帧特征和编码格式等信息,未发现明显异常。

[0180] 脚本特征提取单元:对抓取的JavaScript脚本进行分析,基于抽象语法树技术,识别出脚本中存在一些变量名混淆的情况,例如使用无意义的短字符串作为变量名,同时发现一些函数调用关系较为复杂,可能存在加密或伪装的恶意逻辑。这些脚本的静态语法结构和调用行为特征被提取。

[0181] 通信行为特征提取单元:统计网络请求的相关信息,发现某个API请求的频率明显高于正常水平,每分钟达到50次,同时请求头中的Referer字段为空,这两个通信行为特征被记录。

[0182] 信誉信息提取单元:通过与第三方信誉评级机构合作的接口,查询"example.com”的历史信誉评分。发现该网站在过去一个月内信誉评分呈下降趋势,且曾被标记为存在潜在风险,这些历史信誉信息被提取。

[0183] AI风险分析评估模型

[0184] 图像分析模块:将产品展示图片输入图像分析模块。首先,使用卷积神经网络提取像素差异、颜色分布及频域特征,通过卷积层和池化层的多次运算,学习图片的特征表示。接着,基于统计方法计算图像直方图,发现直方图在某些颜色区间出现异常峰值。然后,将提取的特征与恶意图像特征库比对,输出各图像区域的风险评分及对应时间戳。假设将图片划分为10个区域,其中3个区域的风险评分较高,分别为,对应的时间戳分别为。风险评估模块根据公式计算整体图像风险评分,假设,其他区域风险评分为0,则=1 / 10[(1-0.1×10)×0.8+(1-0.1×15)×0.7+(1-0.1×8)×0.85+0+0+0+0+0+0+0]=0.45。

[0185] 脚本分析模块:数据抓取模块提取的JavaScript脚本被送入脚本分析模块。静态分析单元基于抽象语法树技术识别出脚本中存在5个可疑的混淆变量名和2个可能的加密函数调用,即。行为模式分析单元调用预训练的语言模型(如BERT)对脚本语义进行编码,并与已知恶意模式进行比对,获得该脚本为恶意脚本的概率值。动态分析单元在沙盒环境执行脚本,监控到异常行为事件数(如检测到一次跨站脚本攻击尝试和两次异常的数据请求),总事件数。风险评估模块根据公式计算脚本风险评分,假设,则。

[0186] 网络行为分析模块:数据抓取模块采集的API请求记录被输入网络行为分析模块。网络请求解析单元解析每个请求的URL、请求头及参数,统计调用频率。例如,某个API请求"http: / / example.com / api / user_data”的调用频率。同时,识别出该请求存在跨域行为且请求头中的身份标识信息存在伪造嫌疑,输出身份伪造评分。API行为分析单元基于图神经网络追踪该请求与敏感数据节点的连接路径,发现该请求试图访问用户的账户余额信息节点,输出图结构风险评分。风险评估模块根据公式计算API风险评分,假设(目标网页加载过程中共触发10个API请求),,则=1 / 10×(0.2×50+0.4×0.8+0.4×0.7)=1.06(这里由于公式计算结果可能因假设数据而超出范围,实际应用中会通过归一化等方式处理)。

[0187] 动态内容分析模块:动态内容监控单元实时监听由脚本渲染的广告框、弹窗及提示条,发现一个弹窗中包含一个看似官方的登录框,但仔细分析发现其URL指向一个可疑域名,识别为伪造登录框。行为模式分析单元结合用户交互轨迹,发现用户在看到该弹窗后,鼠标快速移动并点击了登录框,行为异常强度。动态内容监控单元统计动态加载频率(单位时间内新增DOM节点数量),并评估脚本来源可信度等静态风险因子。风险评估模块根据公式计算网页整体风险评分,假设(这里简化为两个风险因子),,则。

[0188] 行为模式分析模块:结合上述各模块的分析结果,行为模式分析模块使用深度学习模型(如将BERT用于文本分析部分,CNN用于图像相关分析部分,GNN用于网络行为相关分析部分进行组合)对潜在恶意行为模式进行综合识别。通过对多维度特征的融合分析,进一步确认网站存在的风险模式。

[0189] 风险评估模块:综合图像分析模块的、脚本分析模块的、网络行为分析模块的(经归一化处理后假设为0.7)以及动态内容分析模块的等各模块的评估结果,按照预设的权重(假设图像分析权重0.2,脚本分析权重0.3,网络行为分析权重0.3,动态内容分析权重0.2)进行综合计算,得出目标网站的风险类型为可能存在钓鱼攻击和数据泄露风险,风险等级计算为0.2×0.45+0.3×0.32+0.3×0.7+0.2×0.54=0.492,属于低风险等级。

[0190] 风险规避与警告

[0191] 风险提示单元:风险规避与警告模块的风险提示单元根据风险等级向用户展示相应警告信息及安全建议。由于风险等级为低风险,提示用户"该网站可能存在一定风险,建议您谨慎操作,避免点击不明链接和输入敏感信息”。

[0192] 拦截决策单元:因为风险等级未达到高风险标准,拦截决策单元不执行访问阻断或强制跳转安全页面操作,用户可以继续访问该网站,但系统会对后续用户与网站的交互进行更密切的监控。

[0193] 反馈与模型优化

[0194] 假设用户在访问该网站后,发现网站并无恶意行为,向系统反馈这是一个误判。反馈与模型优化模块的主动学习单元收集用户反馈信息,将该网站的相关数据(包括之前提取的各种特征数据、各模块的分析结果等)整理为误判样本,优先加入AI风险分析评估模型的训练集。模型根据这些新样本进行参数调整和结构优化,例如调整各模块风险评分计算中的权重系数,或者对深度学习模型的网络结构进行微调,使得模型在后续对类似网站的风险评估中能够更加准确,减少误判情况的发生。随着系统不断运行,收集到更多的用户反馈和真实攻击样本,AI风险分析评估模型将持续优化,不断提升系统整体检测准确率与适应性,更好地保障用户的网络访问安全。

[0195] 通过以上具体实施方式,本发明的网站AI智能风险评估实时规避系统及方法能够在用户访问网站过程中,全面、实时地评估网站风险,并根据风险等级采取合理的规避措施,同时借助用户反馈不断优化风险评估模型,有效应对复杂多变的网络安全威胁,为用户提供更加安全可靠的网络访问环境。本发明的系统和方法凭借其持续优化的能力,能够不断适应新的安全挑战,保持对各类恶意网站的有效防范。

[0196] 以上仅为本发明较佳的实施例,并非因此限制本发明的实施方式及保护范围,对于本领域技术人员而言,应当能够意识到凡运用本发明说明书及图示内容所作出的等同替换和显而易见的变化所得到的方案,均应当包含在本发明的保护范围内。< / script>

Claims

1. A website AI intelligent risk assessment and real-time avoidance system, characterized by: include: VPN connection module, used to establish a virtual private network connection to ensure the security and confidentiality of the data capture process; The data capture module is used to collect the target website's web page structure, media resources, script code, dynamic content, and network request behavior in real time after the VPN connection is established; A risk feature extraction module is used to extract multi-dimensional features of the target website content, including domain name information features, hypertext structure features, image content features, script static behavior features and dynamic behavior features, network communication behavior features, and historical reputation information; An AI risk analysis and assessment model is used to identify and assess the risk level of a target website based on the multi-dimensional features. The AI ​​risk analysis and assessment model includes: Image analysis module, used to identify potential embedded malicious content in images; Script analysis module, used to detect obfuscation, encryption or malicious logic in JavaScript scripts; Network behavior analysis module, used to analyze suspicious patterns in API calls, data interactions, and network requests; Dynamic content analysis module, used to identify deceptive structures or social engineering attack behaviors dynamically generated by scripts; A behavior pattern analysis module that uses a deep learning model to identify potentially malicious behavior patterns, wherein the deep learning model is a text model represented by bidirectional encoding, a convolutional neural network model, a graph neural network model, or a combination thereof; A risk assessment module, which is used to integrate the analysis results of each module in the AI ​​risk analysis and assessment model and output the risk type and risk level of the target website; A risk avoidance and warning module, which is used to automatically execute a graded response strategy based on the results provided by the risk assessment module. The strategy includes access blocking, script interception, content degradation, permission restriction, and user warning prompts; The feedback and model optimization module is used to continuously collect user feedback, record security incidents, and collect real attack samples to adjust parameters and optimize the structure of the AI ​​risk analysis and assessment model, thereby improving the overall detection accuracy and adaptability of the system; The risk feature extraction module includes: A domain name information extraction unit, used to obtain the domain name and related registration information of the target website; HTML structural feature extraction unit, used to parse and extract the hypertext tag hierarchy and attributes; Media feature extraction unit, used to analyze the content features of image and video media resources; Script feature extraction unit, used to extract the static syntax structure and calling behavior of the script; Communication behavior feature extraction unit, used to count and analyze network request headers, parameters and frequencies; The reputation information extraction unit is used to query and record historical reputation scores.

2. The system according to claim 1, wherein: The risk levels include: Safety, the corresponding value range is 0≤R<0.1; Low risk, corresponding to the value range of 0.1≤R<0.5; Medium risk, corresponding to the value range of 0.5≤R<0.8; High risk, the corresponding numerical range is 0.8≤R≤1.

0.

3. The system according to claim 1, wherein: The risk avoidance and warning module includes: A risk warning unit is used to display corresponding warning information and safety suggestions to the user according to the risk level; The interception decision unit is used to execute access blocking, forced redirection to a secure page, or user confirmation access strategies when the risk level is high.

4. The system according to claim 1, wherein: The feedback and model optimization module includes: The active learning unit is used to prioritize the misjudgment samples reported by users and the successfully identified attack samples into the training set to enhance the model adaptability and detection accuracy.

5. A real-time avoidance method for website AI intelligent risk assessment, characterized in that: The following steps are involved: Step 1: Establish a secure network connection through the VPN connection module, and use the data capture module to collect the target website's web page structure, media resources, script code, and network request behavior in real time; Step 2: Input the data collected in step 1 into the risk feature extraction module to extract multi-dimensional features such as domain name information, HTML structure, image content, script static and dynamic behavior, network communication behavior, and historical reputation; Step 3: The features extracted in Step 2 are input into the AI ​​risk analysis and assessment model. The risk is assessed in parallel by the image analysis module, script analysis module, network behavior analysis module, dynamic content analysis module, and behavior pattern analysis module. Step 4: The risk assessment module integrates the assessment results of each sub-module in step 3 to determine the risk type and risk level of the target website; Step 5: The risk avoidance and warning module executes a graded response strategy of access blocking, script interception, content degradation, permission restriction or user warning prompt according to the risk type and risk level; Step 6: The feedback and model optimization module collects user interaction feedback and attack samples generated during actual detection, and continuously learns and optimizes the AI ​​risk analysis and assessment model. The process of the image analysis module detecting malicious content hidden in image pixels includes: Step 7.1: The data crawling module extracts the image resources of the target website in real time; Step 7.2: The image analysis module uses a convolutional neural network to extract pixel differences, color distribution, and frequency domain features; Step 7.3: The image analysis module calculates the image histogram based on statistical methods and determines the steganographic features of the abnormal color distribution areas; Step 7.4: The image analysis module compares the extracted features with the malicious image feature library and outputs the risk score for each image area. and the corresponding timestamp ; Step 7.5: Risk Assessment Module Based on Formula Calculate the overall image risk score, in: represents the dynamic risk score of the image calculated by the risk assessment module; Indicates the total number of regions divided by the analyzed image; Indicates the The risk score of each image region is calculated by the image analysis module; Indicates the Timestamp of the risk score of each image region; Indicates the current time when the risk assessment is performed; Represents the time decay coefficient, which is used to dynamically reduce the impact of outdated risk scores; Step 7.6: Risk avoidance and warning module according to the The score determines whether to block access or prompt a security warning.

6. The method according to claim 5, characterized in that The process of detecting JavaScript code variants by the script analysis module includes: Step 8.1: The data scraping module extracts the JavaScript script embedded in the web page; Step 8.2: The static analysis unit identifies obfuscated, encrypted, and disguised instructions in the script based on the abstract syntax tree technology; Step 8.3: The behavior pattern analysis unit uses the pre-trained language model to encode the script semantics and compares it with known malicious patterns to obtain the probability value p_{mal}; Step 8.4: The dynamic analysis unit executes the script in the sandbox environment to monitor and count abnormal behavior events. and total number of events ; Step 8.5: Risk Assessment Module Based on Formula Calculate the script risk score, in, : Comprehensive risk score of JavaScript scripts, ranging from [0,1]; : Risk normalization factor, a positive constant used to limit the upper limit of the score; : Risk score weight coefficient, satisfying ; : The number of suspicious syntax structures detected during static analysis, including obfuscation, encryption, and renamed instructions; : The number of abnormal behavior events detected in dynamic execution, including cross-site scripting, redirect hijacking, and data leakage; : The total number of behavioral events identified during the dynamic analysis process, used to standardize the abnormality ratio; : The probability of the script being malicious, obtained through semantic model analysis, ranges from [0,1]; Step 8.6: Risk avoidance and warning module according to the Scoring executes script blocking or user prompting strategies.

7. The method according to claim 5, characterized in that The process of analyzing API requests by the network behavior analysis module includes: Step 9.1: The data capture module collects all API request records during the loading process; Step 9.2: The network request parsing unit parses the URL, request headers and parameters of each request and counts the call frequency ; Step 9.3: The network request parsing unit identifies cross-domain, high-frequency, and forged identity behaviors and outputs an identity forgery score. ; Step 9.4: The API behavior analysis unit tracks the connection path between the request and the sensitive data node based on the graph neural network and outputs the graph structure risk score. ; Step 9.5: Risk Assessment Module Based on Formula Calculate the API risk score, where: Indicates the web API behavior risk score; Indicates the total number of API requests triggered during the loading process of the target web page; Indicates the The calling frequency of each API request is obtained by the network request analysis module; Indicates the The graph structure risk score of each API request connecting to a sensitive node is generated by the API behavior analysis module; Indicates the The identity forgery risk score of each API request is output by the network request analysis module; a, b, and c are preset or trained weight coefficients used to adjust the relative contributions of frequency, structure, and identity factors in the overall risk score; Step 9.6: Risk avoidance and warning module according to the The scoring execution request interception or demotion strategy is implemented.

8. The method according to claim 5, characterized in that The process of calculating the overall risk of a web page by the risk assessment module includes: Step 10.1: The dynamic content monitoring unit monitors the advertisement frames, pop-ups, and prompt bars rendered by the script in real time; Step 10.2: The dynamic content monitoring unit identifies phishing links, fake login boxes, and misleading icons; Step 10.3: The behavior pattern analysis unit extracts the abnormal intensity of the behavior based on the user interaction trajectory ; Step 10.4: Dynamic content monitoring unit counts dynamic loading frequency and static risk factors ; Step 10.5: Risk Assessment Module Based on Formula Calculate the overall risk score of the webpage, where: Indicates the overall risk level of the webpage; represents the number of risk factors; Normalize the total weight of all factors; For the The weighting coefficients of the static risk factors; The first Static risk factor scores, including but not limited to script source credibility and ad source domain reputation; For the The weight coefficient of each behavior-related factor; The abnormal intensity value of user behavior identified by the behavior pattern analysis module, including abnormal mouse trajectory and multiple interactions in a short period of time; Dynamic content loading frequency factor, indicating the number of newly added DOM nodes, scripts, or resources per unit time; is the frequency sensitivity adjustment factor, which is used to control the nonlinear effect of high-frequency content loading on the score; Step 10.6: Risk avoidance and warning module according to the The access policy is adjusted dynamically based on the score.

Citation Information

Patent Citations

  • Comprehensive urban network space governance system

    CN107958322A

  • Multi-source software supply chain intelligent analysis method and system

    CN119720225A