Network attack tracing method, system and equipment based on user portrait, and medium
By collecting and analyzing user interaction behavior data and terminal fingerprint data, and using dynamic hashing algorithms to generate user unique identifiers, it solves the problem that traditional technology is difficult to identify real users and attackers, realizes accurate identification and traceability of network attacks, and improves network security protection capabilities.
Patent Information
- Application Number
- CN202510695810.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-28
AI Technical Summary
Traditional user information collection and image construction technology is difficult to effectively identify real users and attackers in modern network environments, especially when facing complex data and hidden means, it is difficult to accurately identify the attack source, and the defense effect is relatively weak.
By collecting user interaction behavior data on the web page, standardizing user terminal fingerprint data, and using dynamic hashing algorithms to generate user unique identifiers containing risk level tags, combining multiple data sources to analyze abnormal behaviors, the identification and traceability of network attacks can be achieved.
It realizes accurate identification and traceability of network attacks, improves network security protection capabilities, provides effective attack warning and traceability tracking methods, and improves the response speed and accuracy of the defense system.
Smart Images

Figure CN120238369A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of network security technology, and in particular, to a network attack tracing method, system, device and medium based on user portraits. Background Art
[0002] In today's enterprise Web business scenarios, with the increasing complexity of user requirements and the continuous change of the network environment, traditional user information collection and portrait construction technologies are facing more and more challenges. To improve the user experience and enhance network security protection, many enterprises have adopted a variety of technical means to collect user data and construct user portraits. These technologies usually involve multiple data sources such as Cookies, local storage, server logs, JavaScript tracking codes, API interface call records, and user purchase history. By analyzing these data, enterprises can deeply understand the interest preferences and personal information of users, and thereby identify real users and automated attackers (such as crawlers, malicious scripts, etc.).
[0003] However, with the continuous evolution of network attack means, traditional technical means are gradually showing their limitations. First of all, these technologies are difficult to effectively meet the high requirements of the modern network environment. Especially when facing a large amount of complex data, the capture and parsing capabilities of traditional means are insufficient, and they cannot accurately identify the differences between real users and attackers. Secondly, since most traditional data collection methods rely on static data sources (such as IP addresses, browser Cookies, etc.), they are unable to accurately identify the attack source when attackers use hiding means such as proxies, VPNs, and dynamic IPs, and the defense effect is relatively weak. Moreover, traditional user portrait construction methods have the problem of limited data dimensions and are difficult to comprehensively consider various complex factors, such as the dynamic behavior of users and device characteristics, which greatly reduces the accuracy and practicality of the portrait.
[0004] Therefore, how to accurately construct user portraits on the basis of efficiently collecting user behavior data, effectively identify and trace network attack behaviors, and improve the security and defense capabilities of Web services is an urgent problem to be solved at present. Summary of the Invention
[0005] In order to improve the efficiency and accuracy of network security protection, the present application provides a network attack tracing method, system, device and medium based on user portraits.
[0006] In a first aspect, the present application provides a network attack tracing method based on user portraits, adopting the following technical solution: A network attack tracing method based on user portraits, the attack tracing method includes: Collect the interaction behavior data of users in the Web page; the interaction behavior data includes DOM node change records, mouse track coordinates, page scroll positions, input content, and timestamps; Compress and serialize the interaction behavior data to generate a structured interaction log file; Collect user terminal fingerprint data through the browser environment interface; the user terminal fingerprint data includes device models, screen resolutions, GPU rendering features, IP addresses, time zone information, and WebGL fingerprints; Standardize the user terminal fingerprint data to generate fingerprint feature vectors; Based on the structured interaction log file and the fingerprint feature vectors, generate a user unique identifier containing a risk level label through a dynamic hashing algorithm; Perform cluster analysis on the interaction behavior data according to the user unique identifier, identify the abnormal behavior patterns of abnormal users and mark the attack types; Obtain the public network IP address of the abnormal user and locate the geographical location in combination with the IP database; Build an attack traceability report based on the user unique identifier, abnormal behavior pattern, public network IP address, and geographical location of the abnormal user.
[0007] By adopting the above technical solutions, combining user behavior data, terminal fingerprint information, and dynamic hashing algorithms, a comprehensive and accurate user portrait is constructed, and based on this portrait, network attack behaviors can be effectively identified and traced. By collecting the interaction behaviors of users on the Web page, standardizing the terminal fingerprint data, generating unique identifiers, and combining multiple data sources for abnormal behavior analysis, various potential attack types such as crawlers, credential stuffing attacks, DDoS, etc. can be detected in real time. At the same time, using technologies such as WebRTC to break through traditional IP protection means, accurately locate the real IP and geographical location of the attacker, generate a detailed attack traceability report, greatly enhance the network security protection ability, provide effective attack early warning and traceability tracking means, and improve the response speed and accuracy of the defense system.
[0008] Optionally, the step of generating a user unique identifier containing a risk level label through a dynamic hashing algorithm based on the structured interaction log file and the fingerprint feature vectors includes: Collect the time series of the interaction behavior data in the structured interaction log file; the time series includes operation interval times and behavior pattern features; Based on a pre-configured weight coefficient matrix, assign high weight values to unconventional operations according to the behavior pattern features, perform weighted summation on the operation interval times, and calculate the behavior consistency score; Collect the hardware features in the fingerprint data of the user terminal; the hardware features include device model, screen resolution, and GPU rendering parameters; Perform hash encryption on the hardware features, and calculate the device stability score based on the matching degree between the hash value and the preset trusted device library; Based on the behavior consistency score and the device stability score, generate a risk level score through a linear combination formula, and determine the corresponding risk level label; Encode the hashed hardware features and the risk level label to obtain a user unique identifier.
[0009] By adopting the above technical solution, comprehensively collect the interaction behavior data and device fingerprint features of the user, and generate a user unique identifier (UID) including a risk level label based on the dynamic hash algorithm. Through the combination of the behavior consistency score and the device stability score, the system can accurately identify the differences between normal users and potential attackers, and improve the defense ability against complex network attacks (such as crawlers, credential stuffing, DDoS attacks, etc.). This technical solution not only enhances the system's ability to identify abnormal behaviors and trace the source of attacks, but also provides accurate risk assessment and attack detection under the premise of user privacy protection, thus greatly improving the security and protection ability of Web applications.
[0010] Optionally, the steps of performing clustering analysis on the interaction behavior data according to the user unique identifier, identifying the abnormal behavior patterns of abnormal users, and marking the attack types include: Group the interaction behavior data according to the user unique identifier, and extract the operation path sequence, time series features, and interaction intensity features in the interaction behavior data; Perform sliding window segmentation on the time series features, and calculate the click frequency and scrolling speed within the window; Perform operation path word segmentation processing on the operation path sequence to generate operation path segments; Perform normalization processing on the click frequency, scrolling speed, operation path segments, and interaction intensity features to obtain the user's behavior features; Input the user's behavior features into a clustering algorithm to generate a user behavior clustering result; Based on the user behavior clustering result, calculate the Mahalanobis distance between the user's behavior features and the clustering center, and mark the users whose Mahalanobis distance exceeds the preset abnormal threshold as abnormal users; Perform pattern analysis on the behavior features of the abnormal users to identify the abnormal behavior patterns; Input the behavior features and abnormal behavior patterns of the abnormal users into a pre-trained classification model to obtain attack type labels.
[0011] By adopting the above technical solutions, it is possible to accurately identify the abnormal behavior patterns of users and mark their attack types. By utilizing multi-dimensional user behavior data and adopting advanced unsupervised learning and supervised learning methods, the detection and defense capabilities against network attacks have been comprehensively improved. At the same time, through real-time monitoring and dynamic updates, it is possible to flexibly respond to complex and changeable attack patterns, thereby providing efficient and accurate technical support for network security protection.
[0012] Optionally, the steps of performing pattern analysis on the behavior characteristics of the abnormal user and identifying the abnormal behavior pattern include: Based on the user behavior clustering results, construct a user behavior benchmark template; wherein, the user behavior benchmark template includes a library of regular operation path segments, the mean and standard deviation of time series features. Calculate the difference degrees between the behavior characteristics of the abnormal user and the user behavior benchmark template, including path difference degree, click frequency deviation degree, and scrolling speed deviation degree. Compare the difference degrees with preset difference thresholds respectively and mark abnormal behavior labels. Perform weighted fusion based on the difference degrees, generate a comprehensive anomaly score, and determine the comprehensive anomaly level. Combine the abnormal behavior labels and the comprehensive anomaly level to obtain the abnormal behavior pattern.
[0013] By adopting the above technical solutions, through the difference degree analysis based on user behavior clustering results, abnormal behaviors are accurately identified and label outputs are provided. By combining weighted fusion calculations to generate a comprehensive anomaly score, through a comprehensive analysis of the comprehensive anomaly level of abnormal users, and marking abnormal labels such as high-frequency clicks and unconventional paths, an intuitive and effective means of attack identification and traceability is provided for network security protection. The system can flexibly respond to different types of attack behaviors, improve the accuracy and response speed of abnormal behavior detection, and enhance network protection capabilities.
[0014] Optionally, after the step of constructing the attack traceability report, the following steps are further included: Perform threat intelligence association on the attack traceability report based on a preset threat intelligence library to obtain a threat intelligence association result. Calculate the attack risk score based on the threat intelligence association result. Generate a list of disposal priorities according to the attack risk score and call a preset policy library to match countermeasures. Based on the list of disposal priorities and countermeasures, execute policy instructions through a security orchestration platform, and record the policy execution status and result logs. Monitor the changes in attack metrics after policy execution, evaluate the effectiveness of the policy, and update the preset threat intelligence library and preset policy library.
[0015] By adopting the above technical solution, based on the detailed data of the attack traceability report, combined with the correlation and risk scoring of threat intelligence, a disposal priority list is automatically generated and corresponding countermeasures are executed. Through the security orchestration platform, the policies can be quickly executed and the status can be recorded. At the same time, the system monitors the changes in attack metrics, evaluates the effectiveness of the policies, and dynamically updates the threat intelligence library and policy library. This technical solution ensures an efficient response to network attacks, improves the intelligence and response speed of defense, and enhances the protection ability against complex attack scenarios through continuous optimization and update.
[0016] Optionally, the attack traceability method further includes: Obtain the social network data associated with the abnormal user; wherein, the social network data includes public metadata and social graph data; Perform sensitive information filtering and correlation analysis on the social network data to generate a social behavior feature vector; Fuse the social behavior feature vector with the fingerprint feature vector and interaction behavior data to generate a multi-dimensional user profile; Based on the multi-dimensional user profile, integrate the risk attributes of the multi-dimensional user profile in the attack traceability report to form an enhanced attack traceability report; According to the enhanced attack traceability report, dynamically adjust the matching logic of the countermeasures in the preset policy library.
[0017] By adopting the above technical solution, based on the fusion of multi-source data and dynamic policy adjustment, the intelligence level of attack recognition and defense is significantly improved. By obtaining social network data, the system can reveal the social behavior background and attack propagation path of the attacker; through the generation of multi-dimensional user profiles, the system can comprehensively evaluate the potential risks of users; through the enhanced attack traceability report, the system provides security personnel with more detailed attack context to help quickly locate key nodes; finally, based on these data and reports, the system can dynamically adjust the protection policy to improve the accuracy and response speed of defense. While improving the attack traceability ability, this solution also reduces the mis-blocking rate, enhances the personalization and adaptability of the protection policy, and provides a powerful means to deal with complex network attacks.
[0018] Optionally, the step of performing sensitive information filtering and correlation analysis on the social network data to generate a social behavior feature vector includes: Perform hash encryption on the user identification information in the social network data, delete the original user identifiers in the social relationship chain, and perform regional aggregation on the geographical location information to obtain the filtered social network data; Construct a social graph based on the filtered social network data, apply graph algorithms to identify key nodes and risk association paths, and obtain the graph algorithm analysis result; Perform time series correlation analysis on the filtered social network data, align the social behavior timestamps with the abnormal login event times, and count the high-risk behaviors within a specified time window after the abnormal login events to obtain the time correlation analysis results; Generate social behavior feature vectors based on the graph algorithm analysis results and the time correlation analysis results.
[0019] By adopting the above technical solutions, social behavior feature vectors of users are generated from multiple dimensions (including social graphs, time series, and behavior characteristics, etc.), which can accurately identify potential attack behaviors. Through the encryption and regional aggregation processing of sensitive information, the protection of user privacy is ensured, while key social behavior characteristics are retained. The combination of graph algorithms and time correlation analysis enables the system to extract high-value security information from users' social interactions and behavior patterns, providing solid data support for attack identification and protection strategies.
[0020] In a second aspect, the present application provides a network attack traceability system based on user portraits, adopting the following technical solutions: A network attack traceability system based on user portraits, the attack traceability system includes: An interaction behavior collection module, used to collect interaction behavior data of users in Web pages; the interaction behavior data includes DOM node change records, mouse trajectory coordinates, page scrolling positions, input content, and timestamps; A data processing module, used to compress and serialize the interaction behavior data to generate a structured interaction log file; A fingerprint data collection module, used to collect user terminal fingerprint data through browser environment interfaces; the user terminal fingerprint data includes device models, screen resolutions, GPU rendering characteristics, IP addresses, time zone information, and WebGL fingerprints; A normalization processing module, used to normalize the user terminal fingerprint data to generate fingerprint feature vectors; An identifier generation module, used to generate a user-unique identifier containing a risk level label based on the structured interaction log file and the fingerprint feature vectors through a dynamic hashing algorithm; A clustering analysis module, used to perform clustering analysis on the interaction behavior data according to the user-unique identifier, identify abnormal behavior patterns of abnormal users and mark attack types; A positioning module, used to obtain the public network IP address of the abnormal user and locate the geographical location in combination with an IP database; An attack traceability module, used to construct an attack traceability report based on the user-unique identifier, abnormal behavior pattern, public network IP address, and geographical location of the abnormal user.
[0021] In a third aspect, the present application provides a computer device, adopting the following technical solution: A computer device includes a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the method described in the first aspect.
[0022] In a fourth aspect, the present application provides a computer-readable storage medium, adopting the following technical solution: A computer-readable storage medium stores a computer program that can be loaded and executed by a processor to implement any one of the methods in the first aspect.
[0023] In summary, the present application includes at least one of the following beneficial technical effects: comprehensively collecting user interaction behavior data, user terminal fingerprint data, and dynamically generated user unique identifiers to construct an accurate network attack tracing system. The present application can record user operation behaviors more efficiently, generate user action replays with small volume, lossless image quality, and higher fineness, can better track the specific operation paths of users, identify abnormal behaviors, and effectively cluster and analyze user data through a dynamic hashing algorithm to accurately identify automated attacks and real users, and then automatically merge attack behaviors from the same attack source. At the same time, the WebRTC method is used to break through the protection of proxies and VPNs to obtain the real public network IP address, which helps to locate the detailed network information of the attacker. Finally, by combining this information to generate an attack tracing report, it not only improves the accuracy and efficiency of attack detection and analysis, but also provides data support for the optimization of defense strategies. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 FIG. 1 is a first flowchart of a network attack tracing method based on user portraits according to one embodiment of the present application.
[0025] Figure 2 FIG. 2 is a second flowchart of a network attack tracing method based on user portraits according to one embodiment of the present application.
[0026] Figure 3 FIG. 3 is a third flowchart of a network attack tracing method based on user portraits according to one embodiment of the present application.
[0027] Figure 4 FIG. 4 is a fourth flowchart of a network attack tracing method based on user portraits according to one embodiment of the present application.
[0028] Figure 5 FIG. 5 is a fifth flowchart of a network attack tracing method based on user portraits according to one embodiment of the present application.
[0029] Figure 6It is the sixth process schematic diagram of a network attack traceability method based on user portraits according to one embodiment of the present application.
[0030] Figure 7 It is the seventh process schematic diagram of a network attack traceability method based on user portraits according to one embodiment of the present application. Detailed implementation manners
[0031] In order to make the purpose, technical solutions and advantages of the present application clearer, the following further describes the present application in detail with reference to the Figures 1-7 accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0032] The embodiments of the present application disclose a network attack traceability method based on user portraits.
[0033] Referring to Figure 1 , a network attack traceability method based on user portraits, the attack traceability method includes: Step S101, collecting the interaction behavior data of the user in the Web page; Among them, the interaction behavior data includes DOM node change records, mouse trajectory coordinates, page scroll positions, input content, and timestamps; It can be understood that on the Web page, the user's interaction behavior not only reflects the user's browsing preferences, but also reveals potential security risks. The collection of interaction behavior data on the Web page refers to recording all interaction operations of the user with the page through a front-end script tool. This operation includes but is not limited to the user's mouse clicks, mouse movement trajectories, page scroll positions, keyboard inputs, and interactions with other DOM elements. These behaviors can help the system build a dynamic portrait based on the user's real behavior.
[0034] Specifically, the rrweb.js library can be introduced into the web site. The rrweb.js library is a front-end interaction behavior capture tool that can be seamlessly integrated into the Web page. By listening to DOM changes, mouse events, input content, page scrolling, etc., it can capture all the behaviors of the user on the page in real time. The record of each behavior is timestamped and bound to the user's operation path to form a traceable behavior chain. In order to ensure the efficient transmission and storage of data, the collected interaction data will be compressed and serialized, usually in JSON format. This can not only effectively reduce the data volume, but also ensure the clear data structure for subsequent analysis.
[0035] It should be noted that the importance of these interaction behavior data lies in that it can provide key clues for subsequent behavior analysis. For example, if a user repeatedly accesses multiple sensitive pages and performs high-frequency click operations, such behavior may be marked as a potential attack behavior. Therefore, the collection of interaction behavior data lays a foundation for further user portrait construction and attack tracing.
[0036] Step S102, compress and serialize the interaction behavior data to generate a structured interaction log file; Among them, by converting the original behavior data into a data format that is easy to store, transmit, and structured, a large amount of data can be transmitted over the network and stored in the database without causing excessive performance pressure.
[0037] Specifically, data compression refers to reducing the redundant part of the data through an algorithm, thereby reducing the storage space and network transmission bandwidth requirements. Common compression methods include text-based data compression algorithms, such as JSON format compression. This compression method can significantly reduce the data volume while retaining the data integrity. Serialization refers to the process of converting an object into a specific format, usually using the JSON format to express these interaction behavior data. Through the JSON format, different types of data (such as strings, numbers, arrays, boolean values, etc.) can be encapsulated in a unified structure, which is easy to parse and process subsequently.
[0038] For example, after capturing a user's click on a button, the system will package this interaction behavior and its attached metadata (such as timestamp, button ID, click position, etc.) into a log data. After compression and serialization, this data can be efficiently stored and transmitted to the background. Finally, the processed structured interaction log will form a data file with a standardized format, which is convenient for aggregating and analyzing user behavior.
[0039] Step S103, collect user terminal fingerprint data through the browser environment interface; Among them, the user terminal fingerprint data includes device model, screen resolution, GPU rendering characteristics, IP address, time zone information, and WebGL fingerprint; It can be understood that the terminal fingerprint data is the unique identifier of the user device, which includes multi-dimensional information at the hardware, software, and network levels. The purpose of collecting terminal fingerprint data is to uniquely identify the user by recognizing device characteristics and maintain the uniqueness of the device between different sessions. Compared with the traditional IP-based identity recognition method, the terminal fingerprint can still effectively identify the user's identity in cases where the user frequently changes the network, uses a proxy, etc.
[0040] Specifically, the collection of terminal fingerprint data is mainly carried out through browser environment interfaces. Common collected information includes device model, operating system version, screen resolution, GPU rendering characteristics (extracted through WebGL), IP address, time zone information, and WebGL fingerprints, etc. These information constitute a unique "fingerprint" of the user device, and its uniqueness and stability are much higher than ordinary IP addresses or Cookies. Through the navigator interface of JavaScript, WebGL API, and WebRTC technology, the system can obtain this information without directly requesting user privacy data.
[0041] For example, WebGL fingerprints are generated by obtaining the rendering characteristics of the browser for the hardware. It is based on the browser's graphics rendering context, including but not limited to GPU model, WebGL version, anti-aliasing support, etc. These characteristics can provide a high degree of uniqueness because the GPU rendering capabilities of different devices vary greatly. By collecting these fingerprint data, the system can construct a highly reliable and stable user identity identifier and effectively distinguish the behavior characteristics of different users in the same network environment.
[0042] Step S104, standardize the user terminal fingerprint data to generate a fingerprint feature vector; Among them, since the collected terminal fingerprint data has a high degree of diversity, and the data such as device characteristics and network environment information often have inconsistent formats, it is necessary to standardize this data. The standardization process includes converting and encrypting different types of data to ensure that all collected data can be compared and used under the same standard.
[0043] In this step, standardization not only refers to the unification of formats but also includes the processing of sensitive data. For example, IP addresses, device IDs, etc. may contain sensitive information. To protect user privacy, it is necessary to perform hash encryption on this data. Hash encryption technology converts the original data into an irreversible encrypted value through a specific algorithm to ensure that the original data cannot be restored and avoid disclosing user privacy.
[0044] Finally, the standardized data will generate a fingerprint feature vector. This feature vector is an array composed of multi-dimensional data, and each dimension represents a certain attribute of the terminal device (such as device model, screen resolution, GPU characteristics, etc.). This feature vector can uniquely identify a device and has a certain degree of stability, and can still be valid when the device is restarted or the user logs in again. By unifying data of different formats and sources into a comparable feature vector, subsequent user analysis, clustering analysis, and attack recognition can all be carried out under the same standard.
[0045] Step S105: Based on the structured interaction log file and the fingerprint feature vector, generate a user - unique identifier containing a risk - level label through a dynamic hashing algorithm; Among them, generating a user - unique identifier (UID) based on user behavior and terminal fingerprint data is the key to subsequent analysis. After combining the user's behavior and device fingerprint data, a comprehensive user profile is formed. Through the dynamic hashing algorithm, these behavior data and device fingerprint features are mapped to a unique identifier. This identifier is not only unique but also can be classified according to the user's behavior risk and is given a risk - level label.
[0046] Specifically, the dynamic hashing algorithm generates a behavior consistency score by weighting the time series in the interaction log and the user's behavior pattern; then, it generates a device stability score based on the hardware features in the device fingerprint data. Finally, these scores will form a UID through a linear combination. The UID contains the user's behavior characteristics, device stability, and potential risk information. Therefore, it can not only uniquely identify the user but also reflect the user's risk level.
[0047] Exemplarily, assume that a certain user has a high behavior consistency score and a low device stability score. The system may mark this user as "high - risk". On the contrary, if the user's behavior consistency is low and the device stability is high, the system may mark this user as "low - risk". This way of generating UID based on behavior and device characteristics can help the system more accurately identify and mark potential attackers or abnormal users.
[0048] Step S106: Perform clustering analysis on the interaction behavior data according to the user - unique identifier, identify the abnormal behavior patterns of abnormal users, and mark the attack types; Among them, clustering analysis is an unsupervised learning method for analyzing the user's behavior pattern to identify different types of behaviors. Based on the generated UID, the system can group users with similar behaviors into one category through a clustering algorithm, and then identify abnormal behavior patterns. Abnormal behavior patterns usually show interaction methods that are quite different from the normal user behavior. For example, a user who frequently accesses sensitive information or makes abnormal login attempts within a short period may be marked as a potential attack behavior.
[0049] Specifically, during the behavior analysis process, the system will extract features such as the user's operation path, click frequency, and scrolling speed, and classify these data in combination with a machine - learning model (such as XGBoost) to output the attack type. For example, a user's frequent login failures and high - frequency API calls within a short period may be classified as a "brute - force attack". If it is found that the user's behavior is significantly different from that of normal users, the system will mark them as potential attackers.
[0050] Step S107, obtain the public IP address of the abnormal user and locate the geographical location in combination with the IP database; It can be understood that although a user may hide their true IP address through a proxy or VPN, through technologies such as WebRTC, the system can effectively penetrate these protective measures and obtain the user's true public IP address. In combination with the IP database, the system can locate the geographical location of the user, providing additional basis for attack traceability.
[0051] Specifically, WebRTC is a technology that can directly obtain the user's true IP. It is usually used in real-time communication scenarios, but can also be used to break through traditional proxy detection, obtain the user's unencrypted true IP address, and locate the actual geographical location of the attacker in combination with the IP database. Through this technology, the system can discover and track attackers hiding behind proxies.
[0052] It should be noted that in some embodiments, the geographical location information of the user may involve privacy data, so sensitive information filtering and encryption processing may be required. To ensure compliance with privacy protection regulations (such as GDPR), these data can be processed through hash encryption and regional aggregation, etc. For example, only the provincial administrative division information is retained to avoid disclosing the user's personal information while retaining the key features related to the attack.
[0053] Step S108, construct an attack traceability report based on the user unique identifier, abnormal behavior pattern, public IP address, and geographical location of the abnormal user.
[0054] Among them, the attack traceability report is one of the important tools for network security defense. By combining the user unique identifier, abnormal behavior pattern, public IP address, and its geographical location, the system can accurately construct the attacker's attack path and identity association information. The attack traceability report not only includes the attacker's identity information, but also can reveal the source of the attack, the attack path, and the possible attack purpose.
[0055] Exemplarily, if an attacker conducts a DDoS attack through an automated script, the system will generate a complete attack traceability report based on the attacker's IP address, device characteristics, behavior pattern, etc., indicating its attack path and related accounts.
[0056] In the above embodiments, by combining user behavior data, terminal fingerprint information, and dynamic hashing algorithms, a comprehensive and accurate user profile is constructed, and based on this profile, network attack behaviors can be effectively identified and traced. By collecting the interaction behaviors of users on web pages, standardizing terminal fingerprint data, generating unique identifiers, and combining multiple data sources for abnormal behavior analysis, various potential attack types such as crawlers, credential stuffing attacks, DDoS, etc. can be detected in real time. At the same time, technologies such as WebRTC are used to break through traditional IP protection means, accurately locate the real IP and geographical location of the attacker, generate a detailed attack traceability report, greatly enhance the network security protection ability, provide effective attack warning and traceability tracking means, and improve the response speed and accuracy of the defense system.
[0057] Referring to Figure 2 , as an implementation manner of step S105, the step of generating a user unique identifier including a risk level label through a dynamic hashing algorithm based on a structured interaction log file and a fingerprint feature vector includes: Step S201, collecting the time series of interaction behavior data in the structured interaction log file; Among them, the time series includes the operation interval time and the behavior pattern features; It can be understood that during the interaction process between the user and the web page, the interval time between each operation and the characteristics of the operation (i.e., the behavior pattern) reflect the user's behavior habits. The operation interval time represents the time difference between different interaction operations of the user, and the behavior pattern features include the user's click frequency, operation path, scrolling speed, input rhythm, etc. These behavior pattern features help the system analyze the user's operation method and evaluate whether there are irregular or abnormal behavior patterns.
[0058] Specifically, by analyzing the interaction log file, the system extracts the time interval between each operation of the user and the relevant behavior pattern features. For example, assuming that the user continuously clicks multiple buttons on a certain page, the system will record the timestamp of each click and the time interval between them. In addition, the behavior pattern features may include the text content input by the user or the sliding speed of the scroll bar, etc.
[0059] Exemplarily, when a user accesses a website, quickly clicks multiple links and the time interval between each click is very short, this may be an unusual operation mode, and the system will record these behavior characteristics as the basis for subsequent analysis.
[0060] Step S202, based on a pre-configured weight coefficient matrix, assigning high weight values to unconventional operations according to the behavior pattern features, performing weighted summation on the operation interval time, and calculating the behavior consistency score; Among them, in order to accurately evaluate whether the user's behavior is normal, the system needs to perform weighted processing according to the "regularity" of each operation. Unconventional operations usually refer to those operations that do not conform to the behavior of ordinary users, such as frequent clicking, rapid page switching, or excessive repeated input, etc. By assigning higher weight values to these unconventional operations, the impact of these abnormal behaviors on the overall behavior score can be highlighted.
[0061] Specifically, a weight coefficient matrix can be defined according to the behavior pattern characteristics, and different weight values are assigned to each behavior characteristic. For example, if a user's click speed is abnormally fast, a higher weight value can be assigned to this operation; then, the system calculates the behavior consistency score through weighted summation. The behavior consistency score reflects the degree of conformity of the user's operations with the normal behavior pattern. The higher the score, the more consistent the user's behavior.
[0062] Exemplarily, if a user continuously clicks multiple buttons quickly, and the interval between each click is less than 1 second, the system will mark this operation as an "unconventional" operation according to the predefined weight coefficient matrix and assign a higher weight to it. Finally, the system calculates the behavior consistency score of this user through weighted summation. If the score is low, it may indicate that the user's behavior has potential risks.
[0063] Step S203, collect the hardware features in the user terminal fingerprint data; among them, the hardware features include device model, screen resolution, and GPU rendering parameters; It can be understood that the characteristics of the user's terminal device are an important basis for identifying the user's identity. Hardware features such as device model, screen resolution, and GPU rendering parameters are usually inherent and relatively stable information of the device, with high uniqueness. By collecting these hardware features, the system can effectively distinguish different users. Especially when facing concealment means such as proxy or VPN, the hardware features can still provide strong support for the user's identity.
[0064] Specifically, through the interfaces provided by the browser (such as the navigator object, WebGL API, etc.), the system can collect information such as the device model, screen resolution, and GPU rendering characteristics of the device. For example, the GPU rendering parameters can be extracted through the WebGL interface. These parameters are based on the device's graphics hardware and can provide a highly unique identifier.
[0065] Exemplarily, assume that a user uses a MacBook Pro with a resolution of 2880x1800, and its GPU rendering characteristics show a specific model of AMD graphics card. Through these hardware features, the system can uniquely identify the user's device.
[0066] Step S204, perform hash encryption on the hardware features, and calculate the device stability score based on the matching degree between the hash value and the preset trusted device library; Among them, in order to protect user privacy and ensure the security of device identification, the hardware feature data needs to be processed by hash encryption. Hash encryption converts the hardware features into irreversible encrypted values to avoid leaking the real information of the device. At the same time, the hash value can be matched with the preset trusted device library to calculate the device stability score. The device stability score reflects the consistency of the device in different sessions, and users with higher device stability are considered more trustworthy.
[0067] Specifically, a hash algorithm (such as SHA-256) can be used to encrypt the hardware features of the device (such as device model, screen resolution, GPU rendering parameters). Then, the system compares the encrypted hash value with the preset trusted device library. If the matching degree is high, it means that the device is a trusted device, and thus a higher device stability score is generated.
[0068] Exemplarily, assume that a user uses a MacBook Pro to access a website. The system encrypts the hardware features (such as model, GPU, etc.) of the device to generate an encrypted value. Subsequently, the encrypted value is compared with the value in the device library. If the matching degree is high, the stability score of the device is higher.
[0069] Step S205, based on the behavior consistency score and the device stability score, generate a risk level score through a linear combination formula, and determine the corresponding risk level label; Specifically, combine the behavior consistency score and the device stability score, and generate a risk level score through a linear combination. In the linear combination formula, the weights (α and β) of the behavior consistency score and the device stability score are adjusted according to the actual situation to ensure the balance between the two.
[0070] In some embodiments, the linear combination formula is: RiskScore = α·(1 - Score_behavior) + β·(1 - Score_hardware); In the above formula, α and β are preset weight parameters, Score_behavior is the behavior consistency score, Score_hardware is the device stability score, and RiskScore is the risk level score.
[0071] As an implementation method of the risk level label, if the risk level score RiskScore ≥ high risk threshold, the corresponding risk level label is "high risk"; if the risk level score RiskScore < high risk threshold and ≥ medium risk threshold, the corresponding risk level label is "medium risk"; if the risk level score RiskScore < medium risk threshold, the corresponding risk level label is "low risk".
[0072] Step S206, encode the hashed hardware features and risk level labels to obtain a unique user identifier.
[0073] Among them, the user's unique identifier (UID) can be generated through Base64 encoding. The encoded content contains hashed hardware characteristics and risk level labels, which can not only uniquely identify the user, but also reflect the user's risk level.
[0074] It can be understood that the calculated risk level score (RiskScore) is combined with the user's behavior and device characteristics to ultimately generate a UID to reflect the user's behavior and device stability, and to mark the corresponding risk level label.
[0075] In the above implementation, the user's interactive behavior data and device fingerprint features are comprehensively collected, and a user unique identifier (UID) containing a risk level label is generated based on a dynamic hash algorithm. Through the combination of behavior consistency scoring and device stability scoring, the system can accurately identify the difference between normal users and potential attackers, and improve the defense capabilities against complex network attacks (such as crawlers, database collisions, DDoS attacks, etc.). This technical solution not only enhances the system's ability to identify abnormal behaviors and trace the source of attacks, but also provides accurate risk assessment and attack detection under the premise of protecting user privacy, thereby greatly improving the security and protection capabilities of Web applications.
[0076] Reference Figure 3 As an implementation method of step S106, the steps of performing cluster analysis on the interactive behavior data according to the user unique identifier, identifying the abnormal behavior pattern of the abnormal user and marking the attack type include: Step S301, grouping the interactive behavior data according to the user unique identifier, and extracting the operation path sequence, time series features and interactive intensity features in the interactive behavior data; Among them, the user unique identifier (UID) is the key element to identify the user's identity. By grouping the interactive behavior data based on the UID, the behavior data of the same user in multiple sessions can be grouped together. The core goal of this step is to extract the user's operation path, time interval, and interaction intensity, which provide an important basis for subsequent behavioral pattern recognition.
[0077] Specifically, the interaction log files are grouped by UID to ensure that all behavior data of the same user can be centralized. Such data includes, but is not limited to, the click order of the user on the page (operation path sequence), the time interval between each operation (time series feature), and the interaction intensity feature (number of clicks, scrolling distance, and number of input characters per unit time).
[0078] Exemplarily, assume that user A performs the following operations in sequence during a session: "Login → Personal Center → Payment Page". The system will record this path as the operation path sequence; if the user stays on the "Payment Page" for 30 seconds and makes 2 clicks, then the time series feature of this operation is the stay time of "30 seconds", and the interaction intensity features are "number of clicks = 2" and "scrolling distance = 200px".
[0079] Step S302: Perform sliding window segmentation on the time series features, and calculate the click frequency and scrolling speed within the window; Among them, the sliding window method is an effective technique for processing time series data, especially suitable for dynamically changing behavior data. By defining the window size and step size, the behavior data can be segmented and analyzed to calculate the click frequency and scrolling speed within the window. These dynamic features can help the system detect anomalies in behavior, such as abnormally frequent clicks or overly fast scrolling speeds.
[0080] Specifically, define the sliding window parameters, such as the window size (e.g., 5 seconds) and the step size (e.g., 1 second), and then segment the time series data. For each window, the system calculates the click frequency (i.e., the number of clicks within the window divided by the window duration) and the scrolling speed (i.e., the scrolling distance within the window divided by the window duration).
[0081] Exemplarily, assume that the user clicks 3 times within 5 seconds and the page scrolls 500px within these 5 seconds. The system calculates the click frequency as 3 / 5 = 0.6 times / second, and the scrolling speed as 500 / 5 = 100px / second. These calculation results can be used as the basis for further analysis.
[0082] Step S303: Perform operation path word segmentation on the operation path sequence to generate operation path segments; Among them, operation path word segmentation is a method of disassembling the continuous operation sequence of the user on the page into small segments, which can effectively analyze the user's behavior trajectory and identify whether there are abnormal operation sequences or abnormal jump patterns.
[0083] Specifically, perform n-gram tokenization on the user's operation path in the page, that is, decompose the continuous operation path into several operation segments. For example, "login → payment → logout" can be split into two segments: "login payment" and "payment logout". In this way, the system can capture the details of the user's operations more precisely, providing information for subsequent clustering and anomaly detection.
[0084] Exemplarily, assume the user's operation path is: "login → personal center → payment page → payment success". After tokenization, the system can decompose it into operation path segments such as "login personal center", "personal center payment page", "payment page payment success", etc. These segments can help the system identify whether the operation process meets the expectations.
[0085] Step S304, perform standardization processing on the click frequency, scrolling speed, operation path segments, and interaction intensity features to obtain the user's behavior features; Among them, the purpose of standardization processing is to convert features with different dimensions into a unified standard to ensure that all features are compared and analyzed on the same scale. Z-score normalization is a commonly used standardization method that can convert feature data into a distribution form with a mean of 0 and a standard deviation of 1, thus avoiding the adverse effects of the dimensional differences of different features on the clustering and classification results.
[0086] Specifically, perform Z-score normalization processing on the behavior features of each user. The formula is . Where X is the original data, μ is the mean, and σ is the standard deviation. In this way, all behavior features (such as click frequency, scrolling speed, etc.) are standardized so that they can be compared on the same scale.
[0087] Exemplarily, assume that a certain user's click frequency is 0.8 times per second and the scrolling speed is 100 px per second. After standardization processing, the values of the click frequency and scrolling speed will be converted into Z-score values, thus eliminating the influence of the dimension between them.
[0088] Step S305, input the user's behavior features into a clustering algorithm to generate a user behavior clustering result; Among them, clustering analysis is an unsupervised learning method used to group users according to the similarity of their behaviors. Through clustering analysis, the system can group users with similar behavior features into one category, thereby identifying normal behaviors and abnormal behaviors.
[0089] Specifically, input the standardized user behavior features into the clustering algorithm to group the user behaviors. The DBSCAN clustering algorithm can automatically identify abnormal clusters with low density, while K-means requires presetting the number of clusters and relies on optimizing the cluster centers to group users.
[0090] Exemplarily, assume that the system uses DBSCAN to cluster the user's behaviors. The system may classify normal users into a high-density cluster and users with frequent clicks into a low-density cluster. The low-density cluster may represent potential attackers or abnormal users.
[0091] Step S306: Based on the user behavior clustering result, calculate the Mahalanobis distance between the user's behavior features and the cluster center, and mark users whose Mahalanobis distance exceeds the preset anomaly threshold as abnormal users; Among them, the Mahalanobis distance is a method to measure the similarity between the user's behavior features and the cluster center, which can consider the correlation between features. Based on the clustering result of the behavior features, calculate the Mahalanobis distance between the behavior features of each user and the cluster center, so as to determine whether the user belongs to the normal behavior range. Users whose Mahalanobis distance exceeds the preset anomaly threshold will be marked as abnormal users, and those who do not exceed the preset anomaly threshold are normal users.
[0092] Specifically, use the Mahalanobis distance formula to calculate the distance between each user and its cluster center: ; In the above formula, X is the user behavior feature, C is the cluster center, and S is the covariance matrix of the features. In the embodiment of the present application, if the Mahalanobis distance exceeds the preset anomaly threshold (such as 3 times the standard deviation), the user will be marked as an abnormal user.
[0093] Step S307: Perform pattern analysis on the behavior features of abnormal users to identify abnormal behavior patterns; Among them, by further analyzing the behavior features of abnormal users, the differences between their behaviors and those of other normal users are extracted, and these differences are transformed into specific abnormal behavior patterns.
[0094] Step S308: Input the behavior features and abnormal behavior patterns of abnormal users into a pre-trained classification model to obtain attack type labels.
[0095] Among them, the behavior features of abnormal users (such as click frequency, scrolling speed, operation path differences) are used as inputs and provided to the pre-trained classification model. The classification model can adopt the XGBoost model, which is an ensemble learning algorithm that can learn the features of different attack types through training data.
[0096] In one embodiment of the present application, the attack type classification model based on XGBoost adopts a gradient boosting decision tree (GBDT) ensemble learning architecture, which mainly consists of four modules: data input and feature processing, gradient boosting tree ensemble, model training and tuning, and prediction and confidence output. At the data processing layer, the model processes multi-dimensional user behavior features through one-hot encoding and Z-score standardization, and optimizes the input data through feature screening. The gradient boosting tree optimizes the classification error through residual learning, and uses the logarithmic loss function and regularization terms (L1 and L2) to prevent overfitting. The model training uses cross-validation and grid search to adjust hyperparameters, and evaluates feature contributions through feature importance. In the prediction stage, the model outputs confidence based on Softmax and supports incremental training to cope with new types of attacks. Through non-linear modeling, regularization, and dynamic update mechanisms, the model achieves accurate attack classification, strong generalization ability, and high robustness.
[0097] Specifically, the model can predict the attack type label of the user based on the input features, such as "credential stuffing attack", "web crawler behavior", or "DDoS simulation". In addition, the model also outputs a confidence score for the attack type, indicating the degree of trust of the model in the prediction result. For example, assume that user A exhibits an abnormal behavior pattern of "high-frequency clicks + unconventional path". The system inputs the user's behavior features and abnormal behavior pattern into the XGBoost model, and the model can output the "credential stuffing attack" label and give a confidence score of 80%.
[0098] In the above implementation, it is possible to accurately identify the abnormal behavior pattern of the user and mark its attack type. By using multi-dimensional user behavior data and adopting advanced unsupervised learning and supervised learning methods, the detection and defense capabilities against network attacks are comprehensively improved. At the same time, through real-time monitoring and dynamic update, it is possible to flexibly cope with complex and changeable attack patterns, thus providing efficient and accurate technical support for network security protection.
[0099] Referring to Figure 4 , as an implementation of step S307, the steps of performing pattern analysis on the behavior characteristics of abnormal users and identifying the abnormal behavior pattern include: Step S401, constructing a user behavior benchmark template based on the user behavior clustering result; Among them, the user behavior benchmark template includes a library of regular operation path segments, the mean and standard deviation of time series features; Specifically, the system first performs clustering analysis on a large number of user behaviors to extract the behavior patterns of regular users. The benchmark template represents the typical behavior characteristics of regular users and is a reference standard used to compare with the behaviors of abnormal users. The template includes two core parts: a library of regular operation path segments and the mean and standard deviation of time series features (such as click frequency and scrolling speed).
[0100] Specifically, by analyzing the clustering results, the system extracts and constructs common operation path segments. Each regular operation path segment represents the operation sequence of users in a specific scenario and reflects the regular pattern of user behavior. Moreover, the system calculates the mean and standard deviation of the time series features by statistically analyzing the behavior data of regular users. For example, data such as click frequency (number of clicks / unit time), page stay time, and scrolling speed are statistically analyzed to obtain their mean and standard deviation, which are used for subsequent analysis of behavior deviation.
[0101] Exemplarily, assume that the system discovers that "login → personal center → payment page" is a typical path accessed by most users, and the system records it as a regular operation path segment. When statistically analyzing the click frequency, the system finds that most normal users click 5 times per minute, and the standard deviation of the click frequency is 1 time. These data are then used as part of the benchmark template.
[0102] Step S402: Calculate the difference degree between the behavior characteristics of the abnormal user and the user behavior benchmark template, including path difference degree, click frequency deviation degree, and scrolling speed deviation degree; Among them, calculating the difference degree between the behavior characteristics of the abnormal user and the regular behavior benchmark template is the core of abnormal behavior recognition. The greater the difference degree, the more abnormal the behavior of the user, and the more likely it is to be inclined to attack behavior. This step evaluates the degree of abnormality by comparing the operation path, click frequency, and scrolling speed of the abnormal user with the regular user behavior benchmark template.
[0103] Specifically, by using an edit distance algorithm (such as Levenshtein distance), calculate the minimum matching distance between the operation path of the abnormal user and the regular path segment library to obtain the path difference degree, which reflects the difference between the user behavior and the regular path. By calculating the deviation degree of the click frequency of the abnormal user from the mean of the regular users in the benchmark template, obtain the click frequency deviation degree to evaluate whether the user frequently performs abnormal click operations. The calculation formula is: click frequency deviation degree = |abnormal value - regular mean| / regular standard deviation. Similarly, calculate the deviation of the scrolling speed of the abnormal user from the mean and standard deviation of the regular users to obtain the scrolling speed deviation degree.
[0104] Exemplarily, assume that an abnormal user frequently jumps paths and makes a large number of clicks after logging in, and the click frequency reaches 20 times per minute, which is much higher than the 5 times of regular users. The system will calculate the deviation degree of the click frequency of this user. Assume that the calculation result of the deviation degree is 3, indicating that the click behavior of this user is significantly higher than that of normal users.
[0105] Step S403: Compare the difference degrees with the preset difference thresholds respectively and mark the abnormal behavior labels; Specifically, by comparing the operation paths of abnormal users with those of normal users, the path difference degree is calculated. If the path difference degree of the abnormal user exceeds the preset difference threshold, it can be marked as an abnormal behavior. At the same time, calculate the click frequency and scrolling speed of the abnormal user within a unit time, and compare them with the average value of normal users. If the deviation degrees of the click frequency and scrolling speed of the abnormal user exceed the preset difference threshold, it can also be marked as an abnormal behavior. Through the analysis of the above features, specific abnormal behavior labels are identified, such as "high-frequency clicks + unconventional path", indicating that the user clicks frequently and jumps to uncommon pages within a short period of time.
[0106] Exemplarily, assume that a user frequently switches pages after logging in, and their click frequency is abnormally high, and the page jump path is unconventional (such as directly jumping from the "login page" to the "payment page"). By comparing the operation paths and click behaviors of other normal users, the system can identify the abnormal behavior label of this user as "high-frequency clicks + unconventional path".
[0107] Step S404, perform weighted fusion based on the difference degree to generate a comprehensive abnormal score and determine the comprehensive abnormal level; Among them, the weighted fusion of the difference degree is to comprehensively consider the difference degrees of multiple features to generate a comprehensive abnormal score. By assigning weights to different difference degrees (such as path difference degree, click frequency deviation degree, etc.), the system can give higher attention to certain features according to the actual situation. The higher the comprehensive abnormal score, the more abnormal the user's behavior is, and the closer it may be to an attack behavior.
[0108] In some embodiments, the comprehensive abnormal score can be calculated through weighted fusion, that is, weights are respectively assigned to the path difference degree, click frequency deviation degree, and scrolling speed deviation degree, and a final score is comprehensively obtained: Comprehensive abnormal score = w1 ⋅ path difference degree + w2 ⋅ click frequency deviation degree + w3 ⋅ scrolling speed deviation degree; where w1, w2, and w3 are preset weight parameters, indicating the contribution sizes of different features to the abnormal score.
[0109] Furthermore, compare the comprehensive abnormal score with the preset score threshold to determine the comprehensive abnormal level; for example, if the comprehensive abnormal score ≥ 0.8, it indicates "high-risk abnormality"; if 0.5 ≤ comprehensive abnormal score < 0.8, it indicates "medium-risk abnormality"; if the comprehensive abnormal score < 0.5, it indicates "low-risk abnormality".
[0110] Step S405, combine the abnormal behavior label and the comprehensive abnormal level to obtain the abnormal behavior pattern.
[0111] Exemplarily, assume that the comprehensive anomaly score of user A is 0.85, and the label is "high-frequency clicks + unconventional path". The system will output the abnormal behavior pattern as: "high-risk anomaly: unconventional path + high-frequency clicks". If the score is 0.75, it will output "medium-risk anomaly: unconventional path + high-frequency clicks"; if the score is 0.4, it will output "low-risk anomaly: unconventional path + high-frequency clicks".
[0112] In the above embodiments, based on the difference analysis of the user behavior clustering results, abnormal behaviors are accurately identified and label outputs are performed. Combined with weighted fusion calculation, a comprehensive anomaly score is generated. By comprehensively analyzing the comprehensive anomaly level of abnormal users and marking abnormal labels such as high-frequency clicks and unconventional paths, intuitive and effective attack identification and traceability means are provided for network security protection. The system can flexibly respond to different types of attack behaviors, improve the accuracy and response speed of abnormal behavior detection, and enhance the network protection ability.
[0113] Refer to Figure 5 , as a further embodiment of the network attack traceability method, after the step of constructing the attack traceability report, it further includes: Step S501, perform threat intelligence association on the attack traceability report based on a preset threat intelligence library to obtain a threat intelligence association result; Among them, the attack traceability report contains basic information of the attacker, such as the unique identifier of the abnormal user, the abnormal behavior pattern, the public network IP address, and geographical location information, etc. These information provide preliminary data support for attack analysis. By associating this information with the preset threat intelligence library, the understanding of the attacker can be further deepened. The threat intelligence library usually contains information such as historical attack records and malicious IP blacklists. These data sources help to identify the threat intelligence association result, such as the attacker's infrastructure, attack source, and historical attack activities.
[0114] Specifically, match the data in the attack traceability report with the preset threat intelligence library, and call the internal threat intelligence library (such as the historical attack IP blacklist) and the external threat intelligence API (such as the malicious IP scoring interface) to associate the identity, behavior, and infrastructure of the attacker, and generate a more comprehensive threat intelligence association result. This association result contains key information such as the C2 (Command and Control) server IP used by the attacker and different stages of the attack chain (such as the data stealing stage).
[0115] Exemplarily, assume that the IP address in the traceability report matches the historical attack IP blacklist. The system will mark this IP as a malicious source and infer that the attacker may be in the data stealing stage.
[0116] Step S502, calculate the attack risk score based on the threat intelligence association result; Among them, through the threat intelligence correlation results, the system can comprehensively evaluate the potential risks of attacks. When calculating the attack risk score, multiple factors such as the scope of influence of the attack, the attack frequency, and the historical success probability are comprehensively considered. Attacks with a high risk score need to be prioritized for handling to prevent greater damage.
[0117] Specifically, the calculation of the risk score is based on a pre-set formula, combined with multiple influencing factors. The scope of attack influence is the number of users affected by the attack, the degree of damage, etc.; the attack frequency is the frequency of attack occurrence, and frequent attacks usually mean a greater threat; the historical success probability is the proportion of successful attacks in the attacker's historical records. Each factor is assigned a different weight according to its impact on the attack, and finally a comprehensive attack risk score is generated.
[0118] Step S503, generate a disposal priority list based on the attack risk score, and call the pre-set policy library to match the response strategy; Among them, once the risk score of the attack is calculated, the system will prioritize different attacks according to the score and call the predefined policy library to respond according to this priority. Different attack types (such as credential stuffing attacks, crawler behavior, DDoS simulation) may require different response strategies.
[0119] Specifically, the system generates a disposal priority list according to the level of the attack risk score. High-risk attacks will be ranked at the top of the list and immediately matched with the preset response strategies. The policy library contains specific measures for different attack types. For example, for credential stuffing attacks, malicious IP addresses are blocked; for crawler behavior, dynamic verification codes are implemented; for DDoS simulation, traffic cleaning is performed.
[0120] For example, if the risk score of an attack is 0.85 (high risk) and it is identified as a credential stuffing attack, the system will automatically select the "block IP" policy and execute it with priority.
[0121] Step S504, based on the disposal priority list and the response strategy, execute the policy instructions through the security orchestration platform, and record the policy execution status and result log; Among them, the security orchestration platform acts as the command center for policy execution, ensuring that the policy can be executed smoothly and the execution status and results are recorded in real time for subsequent auditing and improvement. By receiving the policy instructions generated in the disposal priority list and performing automatic execution, the policy instructions include policy ID, target IP address, execution actions (such as blocking, traffic limiting) and effective duration. After each policy execution, the platform will record the status and result log of the policy execution, including information such as whether the execution is successful, execution time, execution result, etc.
[0122] Exemplarily, assume that a blocking policy is executed for a certain IP. The platform will record the execution status of the blocked IP, including the time point of blocking, the blocked IP address, and the execution result.
[0123] Step S505: Monitor the changes in attack metrics after the policy is executed, evaluate the effectiveness of the policy, and update the preset threat intelligence library and the preset policy library.
[0124] Among them, monitoring the effect after the policy is executed is an important link in evaluating its effectiveness. The system will evaluate the effectiveness of the policy based on changes in attack metrics (such as attack frequency, traffic, etc.). If the policy is ineffective, the system will roll back and generate alternative policies to ensure continuous improvement of the defense mechanism.
[0125] Specifically, after the policy is executed, the system continuously monitors the changes in relevant attack metrics (such as attack traffic, behavior of malicious IPs, etc.) to evaluate the effectiveness of the policy. If, after the policy is executed, the attack metrics decrease and remain at a low level within a preset time, the policy is considered effective; otherwise, if the attack metrics are not effectively controlled, the policy is considered ineffective. According to the effectiveness of the policy, the system will adjust the threat intelligence library and the policy library to optimize the response strategy.
[0126] Exemplarily, assume that after the IP blocking policy is executed, the monitoring data shows that the attack traffic has decreased significantly, then the policy is evaluated as effective; if the attack traffic has not decreased, the policy rollback is triggered and a new policy is generated for trial.
[0127] In the above embodiments, based on the detailed data of the attack traceability report, combined with the association and risk scoring of threat intelligence, a list of disposal priorities is automatically generated and the corresponding response strategies are executed. Through the security orchestration platform, the policy can be quickly executed and the status can be recorded. At the same time, the system monitors the changes in attack metrics, evaluates the effectiveness of the policy, and dynamically updates the threat intelligence library and the policy library. This technical solution ensures an efficient response to network attacks, improves the intelligence and response speed of the defense, and enhances the protection ability against complex attack scenarios through continuous optimization and update.
[0128] Refer to Figure 6 , as a further embodiment of the network attack traceability method, the attack traceability method further includes: Step S601: Obtain the social website data associated with the abnormal user; Among them, the sources of social website data mainly include public metadata and social graph data. Public metadata generally includes the publicly posted content of users, like records, comment history, etc., and social graph data mainly includes information such as friend relationships, group affiliations, and interaction frequencies. These data can provide a multi-dimensional perspective for subsequent analysis and help reveal potential attack behaviors.
[0129] Specifically, to obtain this data, the system usually calls the open APIs provided by social platforms. Many platforms offer specific API interfaces that allow external applications to obtain users' public information through authentication. For example, the API interfaces of some platforms can provide users' friend lists, like records, and posted content, while other platforms can obtain users' tweet content and interaction records. Through these APIs, the system can extract users' social behavior data, identify potential attack behaviors, and thus provide support for traceability analysis.
[0130] It should be noted that for some social platforms, due to the existence of the same-origin policy or API call restrictions, it may not be possible to directly scrape data through the browser or client. In this case, the system can bypass the restrictions of the same-origin policy through server-side proxy technology and use the server to simulate browser behavior to scrape users' public social data. This method can collect users' social graph data without violating the platform's privacy policy, further enriching the behavioral analysis of attackers.
[0131] It can be understood that by analyzing the data of these social websites, the system can identify users' frequent interactions with malicious accounts, phishing websites, malicious links, etc., thus revealing their possible associations with attack behaviors. For example, attackers may use social engineering methods to spread malicious links by leveraging the influence of a well-known account. These social interaction data can reveal the social propagation path of the attackers, providing important clues for subsequent analysis.
[0132] Step S602, filter sensitive information and perform correlation analysis on the social website data to generate social behavior feature vectors; Among them, the data of social websites contains a large amount of user behavior information, and some of this data may involve users' sensitive privacy information, such as personal nicknames, friend relationships, etc. When processing this data, it is first necessary to filter and encrypt sensitive information to ensure that the data meets the requirements of privacy protection, especially under the premise of compliance with international privacy regulations (such as GDPR).
[0133] Specifically, for sensitive information filtering, an encryption algorithm (such as SHA-256) can be used to hash-encrypt users' sensitive information. In this way, even if the data is leaked, the users' real information cannot be recovered through the hash value. Through this method, the system can ensure the protection of users' privacy while still being able to use the topological structure of social data for attack traceability analysis.
[0134] In addition, the correlation analysis of social data itself is very crucial, especially when the user's behavior poses potential attack risks. The behavioral data of social websites usually presents a complex graph structure, where each node in the social relationship represents a user, and each edge represents the interaction relationship between the two. In this graph structure, the system can apply graph algorithms (such as PageRank, Dijkstra algorithm, etc.) to identify key nodes. By analyzing the social behaviors of these key nodes, users with strong correlations to abnormal behaviors can be effectively identified. For example, when a user forwards content containing a phishing link within 1 hour after a login failure, the system can identify the potential threat of this behavior through correlation analysis, inferring that the user may have been attacked or that it is a malicious behavior spread by an attacker on the social platform.
[0135] Moreover, the correlation analysis is not limited to social relationships and can also be carried out in a multi-dimensional manner by combining time series. For example, when detecting abnormal user login behaviors, the system can analyze the user's social behaviors after a login failure to determine whether they have forwarded malicious links or participated in other abnormal activities. If there is a significant temporal matching between these social behaviors and attack events, the system can further determine that the user has a relatively high risk.
[0136] Exemplarily, the user nickname "Zhang San" is encrypted using SHA-256 hash to obtain the encrypted value "abcd1234"; and through time series analysis, it is found that user A forwarded a phishing email link within 1 hour after a login failure.
[0137] Step S603, fuse the social behavior feature vector with the fingerprint feature vector and interaction behavior data to generate a multi-dimensional user profile; Among them, the user profile not only includes the social behavior characteristics of the user, but also combines the user's device fingerprint and interaction behavior data to comprehensively evaluate the potential risks of the user. In network security protection, the user profile has become an important tool for identifying attackers. Through multi-dimensional feature fusion, the system can more accurately identify the risks of users and make corresponding adjustments to the protection strategies.
[0138] Specifically, the social behavior feature vector usually includes social interaction information related to the attacker, such as the frequency of associating with phishing website accounts, the interaction with malicious accounts, etc. The fingerprint feature vector determines the risk of malicious behavior of the user by identifying device information. For example, an attacker may use methods such as virtual machines and anonymous networks to hide their true identity, and these device fingerprint features can provide a more accurate risk assessment for user profiling. Interaction behavior data includes various behavior data generated by the user during the use of the system or application. If the user is in a normal situation, they will follow a specific operation path, while during an attack, they will deviate from this path, and this behavior difference can be used as a key basis for identifying attacks. By analyzing these behavior data, the system can further determine whether there is abnormal behavior, and combine social and device fingerprint features to generate a comprehensive user profile.
[0139] Exemplarily, the multi-dimensional user profile fields are {device type: PC, social association risk: high, operation path deviation degree: 0.8}. The weight allocation rule is: device type (0.3) + social association (0.4) + operation path (0.3).
[0140] Step S604, based on the multi-dimensional user profile, integrate the risk attributes of the multi-dimensional user profile in the attack traceability report to form an enhanced attack traceability report; Among them, the purpose of generating the enhanced attack traceability report is to combine the multi-dimensional user profile information with the attack traceability data, so as to provide more detailed and accurate attacker information. This enhanced report not only includes the basic behavior patterns of the attacker, but also includes information such as the social background, device fingerprint, and social behavior characteristics related to the attacker, which can provide a more comprehensive context for the security team and help them quickly locate the key nodes of the attack chain.
[0141] Exemplarily, the original attack traceability report is: abnormal user UID: A123, behavior pattern: high-frequency clicks. The enhanced attack traceability report is: UID: A123, behavior pattern: high-frequency clicks, risk attributes: virtual machine device + associated phishing account. The enhanced report further provides the risk attributes of the attacker, such as using a virtual machine device and associating with known phishing accounts.
[0142] It can be understood that in the traceability report, the behavior of the attacker is usually presented in ways such as IP address, login time, abnormal behavior patterns, etc., but these data cannot fully reflect the background information of the attacker. By adding the risk attributes in the user profile to the report, the system can supplement more background information, such as whether the attacker uses a virtual machine device, whether there is frequent social interaction with phishing websites or malicious accounts, etc. These information can help security personnel better understand the attacker's behavior motivation and further infer the attacker's identity and attack path.
[0143] Step S605: Dynamically adjust the matching logic of the response strategies in the preset policy library according to the enhanced attack traceability report.
[0144] Among them, based on the risk attributes in the enhanced attack traceability report, the system can adjust the matching rules in the policy library in real time, thereby improving the accuracy and adaptability of the protection strategy.
[0145] Specifically, the system dynamically adjusts the policy library according to the risk attributes of the attacker (such as whether to use virtual machine devices, whether to frequently click malicious links, etc.). For example, for high-risk users, the system can automatically enable more stringent verification mechanisms, such as two-factor verification, behavioral sandbox isolation, etc. For low-risk users, the system can simplify the verification process to improve the user experience. Through this dynamic adjustment mechanism, the system can effectively reduce the misblocking rate and improve the pertinence and flexibility of the protection strategy.
[0146] Exemplarily, the original response strategy was: "Frequent clicks" corresponded to "Blocking the IP". The response strategy after dynamic adjustment is: "Frequent clicks + virtual machine device" corresponds to "Enabling behavioral sandbox isolation".
[0147] In the above embodiments, based on the fusion of multi-source data and dynamic policy adjustment, the intelligent level of attack recognition and defense has been significantly improved. By obtaining social website data, the system can reveal the social behavior background and attack propagation path of the attacker; through the generation of multi-dimensional user portraits, the system can comprehensively evaluate the potential risks of users; through the enhanced attack traceability report, the system provides more detailed attack context for security personnel to help quickly locate key nodes; finally, based on these data and reports, the system can dynamically adjust the protection strategy to improve the accuracy and response speed of defense. While improving the attack traceability ability, this solution also reduces the misblocking rate, enhances the personalization and adaptability of the protection strategy, and provides a powerful means to deal with complex network attacks.
[0148] Referring to Figure 7 , as an implementation manner of step S602, the steps of filtering sensitive information and performing correlation analysis on social website data to generate social behavior feature vectors include: Step S701: Hash-encrypt the user identification information in the social website data, delete the original user identifiers in the social relationship chain, and perform regional aggregation on the geographical location information to obtain the filtered social website data; It is understandable that when collecting social media data, users' personal information (such as nicknames, email addresses, etc.) and geographical location information may involve privacy data. Therefore, sensitive information filtering and encryption processing are required. To ensure compliance with privacy protection regulations (such as GDPR), this data needs to be processed through hash encryption and regional aggregation, etc., to avoid disclosing users' personal information while retaining key features related to attacks.
[0149] Specifically, an irreversible hash algorithm (such as SHA-256) can be used to encrypt user identifiers (such as nicknames, email addresses, etc.). The hash algorithm can convert the user's original identifier into a hash value of a fixed length, thereby preventing the leakage of sensitive information. Since the hash operation is irreversible, even if the data is leaked, the original information cannot be recovered from the hash value. Also, to avoid exposing the real user identity, the original user identifiers (such as user IDs) in the social relationship chain are deleted, and only the relationship topology structure is retained. This means that the connection of social relationships still exists, but it cannot be directly traced back to a specific user. Additionally, to maintain the effectiveness of geographical information without disclosing the precise geographical location, regional aggregation is performed on the users' geographical location information. For example, only the provincial administrative division information is retained. In this way, the system can obtain the approximate location of the user while avoiding the leakage of the precise location.
[0150] Step S702: Based on the filtered social media data, construct a social graph, apply graph algorithms to identify key nodes and risk association paths, and obtain the graph algorithm analysis result. Among them, the purpose of constructing a social graph is to model the user relationships on the social platform, and through graph algorithms, analyze the key nodes and potential risk association paths in the social network. These key nodes and risk paths can reveal the social influence of users, potential attacker associations, and possible attack propagation paths.
[0151] Specifically, a social graph is a graph structure composed of nodes (representing users) and edges (representing social relationships between users). When constructing the graph, the system generates edges between nodes based on the social relationship chains (such as friend relationships, group memberships, etc.) in the filtered social data. Each edge can represent different types of social relationships (such as friend relationships, group member relationships, etc.). By applying graph algorithms (such as PageRank, Dijkstra algorithm, etc.), the system can identify key nodes (such as users with greater social influence) and risk association paths between users from the social graph.
[0152] In some embodiments, the PageRank algorithm can be used to calculate the importance of each node and identify users with greater influence, who may be potential sources of attack or targets of attack. The Dijkstra algorithm, on the other hand, can calculate the shortest path between nodes to evaluate the risk association degree between users and known attackers.
[0153] Step S703: Conduct a time series correlation analysis on the filtered social network site data, align the social behavior timestamps with the abnormal login event times, and count the high-risk behaviors within a specified time window after the abnormal login event to obtain the time correlation analysis result. Among them, the core purpose of time series analysis is to align social behaviors with abnormal login events in terms of time and analyze the social behavior patterns after abnormal events occur. Aligning the timestamps of social behaviors with the times of abnormal login events can help identify potential attack behaviors or signs of social engineering attacks. By analyzing the behaviors within a certain time window after an abnormal event, the possible behavior trajectories of attackers can be identified, thus providing more accurate early warnings for the defense system.
[0154] Specifically, align the timestamps in the social data with the times of abnormal login events. For example, when a login failure event of a certain user is detected, the system can compare whether the user has performed high-risk social behaviors (such as forwarding malicious links, joining phishing groups, etc.) within 1 hour after the login failure. Moreover, within the time window (such as 1 hour), the system counts the user's social behaviors, such as whether the user has forwarded content containing malicious links or participated in discussions in abnormal groups. These behaviors are usually related to attack activities such as social engineering attacks and malicious dissemination.
[0155] Step S704: Generate a social behavior feature vector based on the graph algorithm analysis result and the time correlation analysis result.
[0156] Among them, the social behavior feature vector includes graph structure features (such as PageRank scores, shortest path lengths), time matching degrees (such as behavior matching degree scores after abnormal events), and social behavior features (such as frequencies of forwarding malicious links, probabilities of being associated with attackers).
[0157] Specifically, graph structure features can describe the importance of users in the social network and the tightness of their relationships with potential attackers; time matching degrees are used to evaluate whether social behaviors are associated with abnormal login events, and a high matching degree indicates that the behavior may be a malicious behavior of the attacker; social behavior features can directly reflect the association strength between users and attack behaviors.
[0158] It can be understood that by combining the graph algorithm analysis results with the time correlation analysis results, the generated social behavior feature vectors can comprehensively reflect the social behavior risks of users, providing more accurate data support for subsequent attack identification and risk assessment, and helping the defense system to identify potential attack behaviors and high-risk users in a complex social environment.
[0159] In the above embodiments, social behavior feature vectors of users are generated from multiple dimensions (including social graphs, time series, and behavior characteristics, etc.), which can accurately identify potential attack behaviors. Through the encryption and regional aggregation processing of sensitive information, the protection of user privacy is ensured, while key social behavior characteristics are retained. The combination of the graph algorithm and time correlation analysis enables the system to extract high-value security information from users' social interactions and behavior patterns, providing solid data support for attack identification and protection strategies.
[0160] The embodiment of the present application also discloses a network attack tracing system based on user portraits.
[0161] A network attack tracing system based on user portraits includes: An interaction behavior collection module, which is used to collect the interaction behavior data of users in Web pages; the interaction behavior data includes DOM node change records, mouse trajectory coordinates, page scrolling positions, input content, and timestamps; A data processing module, which is used to compress and serialize the interaction behavior data to generate a structured interaction log file; A fingerprint data collection module, which is used to collect user terminal fingerprint data through browser environment interfaces; the user terminal fingerprint data includes device models, screen resolutions, GPU rendering characteristics, IP addresses, time zone information, and WebGL fingerprints; A normalization processing module, which is used to normalize the user terminal fingerprint data to generate fingerprint feature vectors; An identifier generation module, which is used to generate a unique user identifier containing a risk level label based on the structured interaction log file and the fingerprint feature vectors through a dynamic hashing algorithm; A clustering analysis module, which is used to perform clustering analysis on the interaction behavior data according to the unique user identifier, identify the abnormal behavior patterns of abnormal users and mark the attack types; A positioning module, which is used to obtain the public IP address of the abnormal user and locate the geographical location in combination with the IP database; An attack tracing module, which is used to construct an attack tracing report based on the unique user identifier, abnormal behavior pattern, public IP address, and geographical location of the abnormal user.
[0162] One network attack traceability system based on user portraits according to an embodiment of the present application can implement any of the above network attack traceability methods, and the specific working processes of each module in the network attack traceability system can refer to the corresponding processes in the above method embodiments.
[0163] In several embodiments provided by the present application, it should be understood that the provided methods and systems can be implemented in other ways. For example, the system embodiments described above are only illustrative; for example, the division of a certain module is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0164] An embodiment of the present application also discloses a computer device.
[0165] The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements a network attack traceability method based on user portraits as described above.
[0166] An embodiment of the present application also discloses a computer-readable storage medium.
[0167] The computer-readable storage medium stores a computer program that can be loaded and executed by a processor to implement any of the above network attack traceability methods based on user portraits.
[0168] Among them, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component; the program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.
[0169] It should be noted that in the above embodiments, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0170] The above are all preferred embodiments of the present application. Without restricting the protection scope of the present application accordingly, any feature disclosed in this specification (including the abstract and drawings), unless specifically described, can be replaced by other equivalent or similar-purpose alternative features. That is, unless specifically described, each feature is only an example of a series of equivalent or similar features.
Claims
1. A network attack traceability method based on user portraits, characterized in that, The attack traceability method includes: Collecting the interaction behavior data of the user in the Web page; the interaction behavior data includes DOM node change records, mouse trajectory coordinates, page scrolling positions, input content, and timestamps; Compressing and serializing the interaction behavior data to generate a structured interaction log file; Collecting user terminal fingerprint data through the browser environment interface; the user terminal fingerprint data includes device models, screen resolutions, GPU rendering features, IP addresses, time zone information, and WebGL fingerprints; Normalizing the user terminal fingerprint data to generate a fingerprint feature vector; Based on the structured interaction log file and the fingerprint feature vector, generating a user unique identifier containing a risk level label through a dynamic hashing algorithm; Performing cluster analysis on the interaction behavior data according to the user unique identifier, identifying the abnormal behavior patterns of abnormal users, and marking the attack types; Obtaining the public IP address of the abnormal user, and locating the geographical location in combination with the IP database; Based on the user unique identifier, abnormal behavior pattern, public IP address, and geographical location of the abnormal user, constructing an attack traceability report.
2. The network attack tracing method based on user portraits according to claim 1, characterized in that The step of generating a user unique identifier containing a risk level label through a dynamic hashing algorithm based on the structured interaction log file and the fingerprint feature vector includes: Collecting the time series of the interaction behavior data in the structured interaction log file; the time series includes operation interval times and behavior pattern features; Based on a pre-configured weight coefficient matrix, assigning high weight values to unconventional operations according to the behavior pattern features, performing weighted summation on the operation interval times, and calculating a behavior consistency score; Collecting the hardware features in the user terminal fingerprint data; the hardware features include device models, screen resolutions, and GPU rendering parameters; Performing hash encryption on the hardware features, and calculating a device stability score based on the matching degree between the hash value and a preset trusted device library; Based on the behavior consistency score and the device stability score, generating a risk level score through a linear combination formula, and determining the corresponding risk level label; Encoding the hashed and encrypted hardware features and the risk level label to obtain a user unique identifier.
3. A method for tracing network attacks based on user portraits according to claim 1, characterized in that, The step of performing cluster analysis on the interaction behavior data according to the user unique identifier, identifying the abnormal behavior patterns of abnormal users, and marking the attack types includes: Grouping the interaction behavior data according to the user unique identifier, and extracting the operation path sequence, time series features, and interaction intensity features in the interaction behavior data; Performing sliding window segmentation on the time series features, and calculating the click frequency and scrolling speed within the window; Performing operation path word segmentation on the operation path sequence to generate operation path segments; Normalizing the click frequency, scrolling speed, operation path segments, and interaction intensity features to obtain the behavior features of the user; Inputting the behavior features of the user into a clustering algorithm to generate a user behavior clustering result; Based on the user behavior clustering results, calculate the Mahalanobis distance between the user's behavior characteristics and the cluster center, and mark the users whose Mahalanobis distance exceeds the preset anomaly threshold as abnormal users; Conduct pattern analysis on the behavior characteristics of the abnormal users to identify abnormal behavior patterns; Input the behavior characteristics and abnormal behavior patterns of the abnormal users into a pre-trained classification model to obtain attack type labels.
4. The network attack tracing method based on user profiling according to claim 3, characterized in that The steps of conducting pattern analysis on the behavior characteristics of the abnormal users to identify abnormal behavior patterns include: Based on the user behavior clustering results, construct a user behavior benchmark template; where the user behavior benchmark template includes a library of regular operation path segments, mean and standard deviation of time series features; Calculate the difference degrees between the behavior characteristics of the abnormal users and the user behavior benchmark template, including path difference degree, click frequency deviation degree, and scrolling speed deviation degree; Compare the difference degrees with the preset difference thresholds respectively to mark abnormal behavior labels; Conduct weighted fusion based on the difference degrees to generate a comprehensive anomaly score and determine the comprehensive anomaly level; Combine the abnormal behavior labels and the comprehensive anomaly level to obtain abnormal behavior patterns.
5. A method for tracing the origin of a cyber attack based on a user profile according to any one of claims 1 to 4, characterized in that, After the step of constructing the attack traceability report, it also includes: Conduct threat intelligence association on the attack traceability report based on a preset threat intelligence library to obtain threat intelligence association results; Calculate the attack risk score based on the threat intelligence association results; Generate a list of handling priorities according to the attack risk score, and call a preset policy library to match response strategies; Based on the list of handling priorities and response strategies, execute policy instructions through a security orchestration platform, and record the policy execution status and result logs; Monitor the changes in attack metrics after policy execution, evaluate the effectiveness of the policies, and update the preset threat intelligence library and preset policy library.
6. The network attack traceability method based on user portraits according to claim 5, characterized in that The attack traceability method also includes: Obtain the social website data associated with the abnormal users; where the social website data includes public metadata and social graph data; Conduct sensitive information filtering and correlation analysis on the social website data to generate social behavior feature vectors; Fuse the social behavior feature vectors with the fingerprint feature vectors and interaction behavior data to generate a multi-dimensional user profile; Based on the multi-dimensional user profile, integrate the risk attributes of the multi-dimensional user profile into the attack traceability report to form an enhanced attack traceability report; According to the enhanced attack traceability report, dynamically adjust the matching logic of the response strategies in the preset policy library.
7. The network attack traceability method based on user portraits according to claim 6, characterized in that, The steps of conducting sensitive information filtering and correlation analysis on the social website data to generate social behavior feature vectors include: Perform hash encryption on the user identification information in the social website data, delete the original user identifiers in the social relationship chain, and perform regional aggregation on the geographical location information to obtain filtered social website data; Construct a social graph based on the filtered social website data, apply graph algorithms to identify key nodes and risk association paths to obtain graph algorithm analysis results; Perform time series correlation analysis on the filtered social network data, align the social behavior timestamps with the abnormal login event times, and count the high-risk behaviors within a specified time window after the abnormal login events to obtain the time correlation analysis results; Generate social behavior feature vectors based on the graph algorithm analysis results and the time correlation analysis results.
8. A network attack tracing system based on user portraits, characterized in that, The attack tracing system includes: An interaction behavior collection module, configured to collect interaction behavior data of a user in a Web page; the interaction behavior data includes DOM node change records, mouse trajectory coordinates, page scroll positions, input content, and timestamps; A data processing module, configured to perform compression and serialization processing on the interaction behavior data to generate a structured interaction log file; A fingerprint data collection module, configured to collect user terminal fingerprint data through a browser environment interface; the user terminal fingerprint data includes device model, screen resolution, GPU rendering characteristics, IP address, time zone information, and WebGL fingerprint; A normalization processing module, configured to perform normalization processing on the user terminal fingerprint data to generate fingerprint feature vectors; An identifier generation module, configured to generate a unique user identifier including a risk level label based on the structured interaction log file and the fingerprint feature vectors through a dynamic hashing algorithm; A clustering analysis module, configured to perform clustering analysis on the interaction behavior data according to the unique user identifier, identify the abnormal behavior patterns of abnormal users, and mark the attack types; A positioning module, configured to obtain the public network IP address of the abnormal user and locate the geographical location in combination with an IP database; An attack tracing module, configured to construct an attack tracing report based on the unique user identifier, abnormal behavior pattern, public network IP address, and geographical location of the abnormal user.
9. A computer device, characterized in that: It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: A computer program is stored that can be loaded and executed by a processor to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and system for constructing large-scale trapping scene based on cloud computing
CN111935185A
Network attack tracing method based on behavior portraits
CN111988285A
Large-scale network attack-oriented tracing system and method
CN114584401A
APT detection method and system based on continuous time dynamic heterogeneous graph neural network
CN115883213A
Deep traceability method and system based on canvas fingerprint identification and electronic equipment
CN117714094A
Cited By
Network attack intelligent identification method and apparatus, and electronic device
CN120710803A
A method, device, and electronic device for intelligent identification of network attacks
CN120710803B
Safety control method and system for cloud inventory management platform
CN120805200A
Dynamic data tracing method and system for network security
CN121037128A
A network security-oriented dynamic data provenance method and system
CN121037128B