Remote Office Network Security Protection Method and System Based on Big Data
By performing multi-dimensional security feature analysis and multi-layer abnormality detection in remote office scenarios, and dynamically adjusting the protection strategy, the problem of inability to flexibly respond to complex network security challenges in the existing technology is solved, and efficient network security protection is achieved.
Patent Information
- Application Number
- CN202510237030.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-01
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-01
AI Technical Summary
The existing remote office network security protection methods cannot be flexibly adjusted according to specific abnormal situations, and the lack of effective feedback mechanisms leads to poor protection effects and inability to adapt to increasingly complex network security challenges.
By obtaining network activity data in remote office scenarios, performing multi-dimensional security feature analysis, generating a dynamic security feature set, using multi-layer anomaly detection model to identify abnormal behavior, and dynamically adjusting protection policies based on the abnormal detection results, and controlling network connections, device permissions and user access paths in real time.
It improves the accuracy and reliability of network security detection, reduces the false alarm rate and omission rate, realizes accurate customization of protection strategies and dynamic self-adjustment, adapts to complex network environment changes, and provides adaptive and intelligent protection.
Smart Images

Figure CN119728311B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of network security technology. Specifically, it relates to a method and system for network security protection of remote work based on big data. Background Art
[0002] With the rapid development of information technology and the increasing transformation of work patterns, the remote work mode has gradually become the choice of many enterprises and workers. Remote work breaks the limitations of traditional work in terms of time and space, greatly improving the flexibility and efficiency of work. However, network security issues in the remote work environment have also become prominent, posing severe challenges to enterprises and users.
[0003] In the traditional field of network security protection, early protection methods mainly focused on protecting the network boundary. For example, external illegal access was blocked through devices such as firewalls. However, for the remote work scenario, this method is too simple and one-sided, and cannot effectively cope with various security risks generated by internal terminal devices in a complex network environment.
[0004] The existing method for formulating protection strategies is relatively rigid. Usually, a fixed set of protection strategies is preset. Regardless of the severity of the actual security threats, the same countermeasures are adopted, and it is impossible to flexibly adjust according to specific abnormal situations, resulting in poor protection effects and inability to meet the diverse security protection needs in the remote work environment.
[0005] Moreover, most traditional protection methods lack an effective feedback mechanism. After implementing protection measures, the change information of the network state cannot be timely fed back to the detection model, so that the detection model cannot optimize its own recognition logic according to the actual protection effect. This leads to the difficulty for the protection system to continuously evolve with the continuous change of the network environment and gradually unable to adapt to the increasingly complex network security challenges of remote work. Summary of the Invention
[0006] In view of the problems mentioned above, in combination with the first aspect of the present application, embodiments of the present application provide a method for network security protection of remote work based on big data. The method includes:
[0007] Obtain network activity data generated by multiple terminal devices in the remote work scenario. The network activity data includes traffic data, device status data, and user behavior data;
[0008] Conduct multi-dimensional security feature analysis on the network activity data to generate a dynamic security feature set corresponding to the network activity data. The dynamic security feature set includes traffic features, device fingerprint features, and behavior pattern features;
[0009] Based on the dynamic security feature set, an abnormal behavior recognition is performed on the network activity data through a multi-layer anomaly detection model, and an anomaly detection result including an anomaly type and an anomaly confidence level is output;
[0010] According to the anomaly detection result, a target protection policy corresponding to the anomaly type is matched from a preset protection policy library, and the target protection policy is dynamically weighted based on the anomaly confidence level to generate a real-time protection policy instruction;
[0011] Execute the real-time protection policy instruction to synchronously regulate the network connection, device permissions, and user access paths in the remote work scenario, and feedback the regulated network status data to the multi-layer anomaly detection model to optimize the anomaly recognition logic.
[0012] On the other hand, an embodiment of the present application further provides a remote work protection system, including a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor. The machine-readable storage medium is used to store programs, instructions, or codes, and the processor is used to execute the programs, instructions, or codes in the machine-readable storage medium to implement the above method.
[0013] Based on the above aspects, the embodiment of the present application comprehensively obtains network activity data including traffic data, device status data, and user behavior data generated by multiple terminal devices in the remote work scenario, and constructs an all-round and multi-dimensional network activity portrait. When performing multi-dimensional security feature analysis on the network activity data and generating a dynamic security feature set, the traffic features, device fingerprint features, and behavior pattern features included in the dynamic security feature set are not simply a list of static features. The dynamic security feature set will be continuously updated with the real-time changes of network activities, and can adapt to the diverse and changeable network behavior patterns in the remote work scenario. Compared with the traditional fixed feature analysis method, it can more accurately depict the real security status of network activities at each moment, and effectively identify the continuously evolving new network attack patterns.
[0014] The multi-layer anomaly detection model performs abnormal behavior recognition on the network activity data based on the dynamic security feature set, and outputs an anomaly detection result including an anomaly type and an anomaly confidence level, which not only improves the accuracy and reliability of the detection, but also through multi-level analysis, can dig out abnormal behaviors from different angles, greatly reducing the false alarm rate and the missed alarm rate. Then, according to the anomaly detection result, a target protection policy is matched from a preset protection policy library, and its dynamic weight is adjusted based on the anomaly confidence level to generate a real-time protection policy instruction, abandoning the limitations of the traditional fixed protection policy, being able to flexibly adjust the intensity of the protection measures according to the severity and credibility of the abnormal situation, and can be accurately customized according to the actual threat situation, significantly improving the pertinence and effectiveness of the protection policy.
[0015] Finally, execute the real-time protection policy instructions to synchronously regulate the network connections, device permissions, and user access paths in the remote work scenario, and feedback the regulated network status data to the multi-layer anomaly detection model to optimize the anomaly recognition logic, realizing the dynamic self-adjustment and continuous optimization of the protection process. It can continuously improve the anomaly recognition logic with the change of the network environment, enabling the protection system to always maintain an efficient response ability to various security threats, providing an adaptive, intelligent, and continuously evolving protection guarantee for the remote work network security. Description of the Drawings
[0016] Figure 1 is a schematic execution flowchart of the remote work network security protection method based on big data provided by an embodiment of the present application.
[0017] Figure 2 is a schematic hardware architecture diagram of the remote work protection system provided by an embodiment of the present application. Detailed Embodiments
[0018] The present application will be specifically described below in conjunction with the accompanying drawings of the specification. Figure 1 is a schematic flowchart of the remote work network security protection method based on big data provided by an embodiment of the present application. The remote work network security protection method based on big data will be introduced in detail below.
[0019] Step S110, obtain network activity data generated by multiple terminal devices in the remote work scenario, where the network activity data includes traffic data, device status data, and user behavior data.
[0020] In this embodiment, in the remote work scenario of an enterprise, employees connect to the company's office network through various terminal devices (such as laptops, tablets, and smartphones). These terminal devices will generate a large amount of network activity data during daily work.
[0021] For traffic data, when employees use office software (such as enterprise email clients, project management tools, etc.) to interact with the company's server, traffic data will be generated for each data transmission. For example, when employee A downloads an important project document from the company's server, the data transmission volume, transmission protocol (such as HTTP or HTTPS), etc. during this download process will be recorded as traffic data. At the same time, when employees use instant messaging tools to communicate with colleagues, the sent and received messages will also generate a certain amount of traffic data, including traffic information involved in chat messages, file transfers, etc.
[0022] Regarding the device status data, taking the laptop of Employee B as an example, the hardware identification features of the device (such as the MAC address, etc.) are the information that uniquely identifies the device. If there are vulnerabilities in the operating system of the laptop, they can be recorded in the device status data. For example, when the operating system detects that a certain security patch is not installed, this constitutes a vulnerability feature of the operating system. In addition, during the office work, the resource occupancy of various processes will also be recorded. For instance, when Employee B opens multiple office software (such as document editing software, spreadsheet software, and presentation software) simultaneously, the information such as the CPU usage rate and memory usage amount occupied by each software process is part of the device status data.
[0023] The user behavior data is also rich and diverse. The time when Employee C logs in to the company's office system every morning is part of the login time series feature. When Employee C operates on files during the office work, such as creating new documents, modifying existing documents, or deleting useless documents in the shared folder, the tracks of these file operations will be recorded as file operation track features. If the role permissions of Employee C change, for example, being promoted from an ordinary employee to a project leader, the corresponding permission change history features will also be recorded. These traffic data, device status data, and user behavior data are collected centrally to provide a basis for subsequent security analysis.
[0024] Step S120, perform multi-dimensional security feature analysis on the network activity data to generate a set of dynamic security features corresponding to the network activity data, and the set of dynamic security features includes traffic features, device fingerprint features, and behavior pattern features.
[0025] When analyzing the collected network activity data, various features are first extracted from the traffic data. Looking at the extraction of the protocol type distribution feature from the previously collected traffic data. For example, in the remote working network of the entire company, most of the normal business traffic uses common protocols such as HTTP or HTTPS. However, if it is found that a certain terminal device frequently uses a relatively rare protocol, such as a certain custom encryption protocol, within a short period of time, this may be an abnormal situation. By capturing the nested relationship between the transport layer protocol and the application layer protocol in the traffic data, this unconventional protocol combination pattern can be identified. At the same time, the proportion of the number of data packets of different protocols under the same source address is statistically calculated to generate a protocol distribution vector. Suppose that among the data packets sent by an employee's terminal device in a day, 80% are HTTP protocols, 15% are HTTPS protocols, and 5% are other protocols. This proportional relationship is part of the protocol distribution vector. By comparing the similarity of this protocol distribution vector with the known malicious protocol library, if it is found that the similarity between a certain protocol used by the device and a malicious protocol in the malicious protocol library is greater than the dynamically set score (for example, when the proportion of normal business traffic decreases, the dynamically set score is reduced according to the rules, so that more protocols are regarded as risk protocols), then this protocol will be marked as a risk protocol.
[0026] The device fingerprint feature is parsed from the device status data. Taking the device of employee D as an example, the hardware identification feature (MAC address) of the device is like the ID card of the device and is an important part of the device fingerprint feature. If there is a vulnerability in the operating system of the device, such as a certain remote execution vulnerability in the Windows system, this operating system vulnerability feature will be recorded. During the work process, when employee D opens multiple office software, the process resource occupancy feature will also be analyzed. For example, if a certain office software process suddenly occupies a large amount of CPU resources, this may be a sign that the software has a malfunction or is under a malware attack, and this process resource occupancy feature is also part of the device fingerprint feature.
[0027] Regarding the behavioral pattern features in user behavior data, take employee E as an example. The time when employee E logs in to the company's office system every day is relatively fixed. If the login time suddenly advances or is postponed by a long time one day, this constitutes an anomaly in the login time series feature. When employee E operates on files in the shared folder, the file operation trace feature will be recorded in detail. Suppose employee E only modifies files in specific folders under normal circumstances, but suddenly starts to frequently access other irrelevant folders and perform file deletion operations. This may be an abnormal behavioral pattern. Additionally, if the permissions of employee E change, the permission change history feature will record information such as the time of the change and the content of the change (e.g., changing from only being able to read certain files to being able to edit these files). Finally, align the traffic feature, device fingerprint feature, and behavioral pattern feature according to a time window. For example, take each hour as a time window, and perform cross-dimensional correlation analysis on the aligned features. For example, based on the target address clustering feature in the traffic feature and the file operation trace feature in the behavioral pattern feature, identify abnormal data transfer paths. If it is found that the file operation trace of employee F shows that they frequently transfer internal company files to an external suspicious address, and this external address matches the malicious address clustering in the target address clustering feature of the traffic feature, then an abnormal data transfer path is identified. Similarly, based on the process resource occupancy feature in the device fingerprint feature and the data packet transmission frequency feature in the traffic feature, potential threats of mismatched resource consumption and traffic can be detected. For example, if a certain device has a high process resource occupancy but a low data packet transmission frequency, this may mean that there is a malicious program on the device that is consuming a large amount of resources in the background.
[0028] Step S130, based on the dynamic security feature set, use a multi-layer anomaly detection model to identify abnormal behaviors in the network activity data, and output an anomaly detection result including the anomaly type and anomaly confidence level.
[0029] The multi-layer anomaly detection model includes a primary detection layer, an intermediate aggregation layer, and a high-level decision layer. First, input the dynamic security feature set into the primary detection layer. For the traffic feature, assume that in previous analyses, it is found that there is a suspicious protocol in the protocol type distribution feature of a certain terminal device, and the data packet transmission frequency is abnormally high. The primary detection layer will independently score the traffic feature according to predefined rules. For the device fingerprint feature, if the operating system vulnerability feature of a certain device indicates the existence of high-risk vulnerabilities, and the process resource occupancy feature shows that resources are abnormally occupied, corresponding anomaly scores will also be obtained. For the behavioral pattern feature, such as the login time series feature of employee G shows an anomaly, and the file operation trace feature shows abnormal file access behavior, an anomaly score will also be obtained, thus obtaining a set of primary anomaly scores.
[0030] Input the set of primary anomaly scores into the intermediate aggregation layer. For example, after receiving the traffic feature anomaly score, device fingerprint feature anomaly score, and behavior pattern feature anomaly score in the set of primary anomaly scores, extract the feature attribute labels and timestamp sequences corresponding to each anomaly score. Suppose the feature attribute label corresponding to the traffic feature anomaly score is "protocol anomaly", and the timestamp sequence shows that the anomaly occurred between 9 am and 10 am. Parse the combined logic rules defined in the threat scenario template, match the corresponding logical condition branches according to the feature attribute labels, and determine the dynamic weight coefficients of each anomaly score in the combined logic rules. If, according to the threat scenario template, when the traffic feature anomaly score is higher than the first set score and the behavior pattern feature anomaly score is lower than the second set score, it is marked as a spoofing attack scenario. In this case, if the traffic feature anomaly score is indeed higher than the first set score and the behavior pattern feature anomaly score is lower than the second set score, then the dynamic weight coefficient of the traffic feature anomaly score in the combined logic rules of this spoofing attack scenario may be higher. Then, based on the timestamp sequence, perform a fluctuation trend analysis of the traffic feature anomaly score, device fingerprint feature anomaly score, and behavior pattern feature anomaly score within a time window to generate the trend intensity factor of each anomaly score. For example, if the traffic feature anomaly score continuously rises within an hour, its trend intensity factor may be higher. According to the product result of the dynamic weight coefficient and the trend intensity factor, calculate the weighted contribution values of the traffic feature anomaly score, device fingerprint feature anomaly score, and behavior pattern feature anomaly score. Accumulate the weighted contribution values in the aggregation direction defined in the threat scenario template to generate the initial aggregation metric. Suppose during the calculation process, the initial aggregation metric exceeds the upper limit of the metric threshold range of the same type of threat scenario in the historical aggregation metric library, then the dynamic compensation mechanism will be triggered. For example, adjust the decay rate of the weighted contribution value according to the overrun ratio of the initial aggregation metric, and apply the adjusted decay rate to the correlation constraint condition between the traffic feature anomaly score and the device fingerprint feature anomaly score, and finally generate the intermediate anomaly aggregation metric.
[0031] Input the intermediate anomaly aggregation metrics into the high-level decision-making layer. Analyze the aggregated feature vectors and time range markers in the intermediate anomaly aggregation metrics, and extract the traffic feature correlation strength, device fingerprint fluctuation amplitude, and behavior pattern deviation degree in the aggregated feature vectors. For example, the traffic feature correlation strength represents the degree of deviation of traffic features from the normal mode, the device fingerprint fluctuation amplitude reflects the instability of the device state, and the behavior pattern deviation degree reflects the difference between user behavior and the normal behavior pattern. Retrieve historical attack case fragments with the same time attributes from the historical anomaly case library according to the time range marker, and generate a set of candidate anomaly types. Assume that the time range marker is from 9 am to 10 am, then find similar attack case fragments that occurred during this time period from the historical anomaly case library. Perform multi-dimensional feature matching between the aggregated feature vectors and each historical attack case fragment in the set of candidate anomaly types. For example, calculate the similarity between the traffic feature correlation strength and the protocol anomaly metric in the historical attack case fragment, the coverage between the device fingerprint fluctuation amplitude and the resource tampering metric in the historical attack case fragment, and the coincidence degree between the behavior pattern deviation degree and the privilege abuse metric in the historical attack case fragment. Based on the number of online devices, user activity level, and network throughput fluctuation value in the real-time network environment parameters, dynamically weight and adjust the similarity, coverage, and coincidence degree to generate a weighted matching degree list. Sort the matching degree values in the weighted matching degree list in descending order, select the top N matching degree values corresponding to the historical attack case fragments as the target candidate cases, and extract the anomaly type labels and confidence correction coefficients marked in the target candidate cases. For example, if the anomaly type label of the target candidate case is "masquerade attack" and the confidence correction coefficient is 0.8. Calculate the initial anomaly confidence corresponding to the intermediate anomaly aggregation metrics according to the occurrence frequency of the anomaly type label in the target candidate case and the confidence correction coefficient. Assume that the label "masquerade attack" appears frequently in the target candidate case, and the initial anomaly confidence is calculated to be 0.7. Perform time decay compensation on the initial anomaly confidence based on the network state baseline data within the time range marker. For example, if the network state baseline data shows large network fluctuations during this time period, according to the rules of time decay compensation (such as adjusting the decay curve slope of the initial anomaly confidence according to the duration of the time range marker and the deviation rate of the network state baseline data, so that the confidence decay rate in the short-term high-deviation scenario is lower than that in the long-term low-deviation scenario), obtain the dynamic confidence threshold. If the initial anomaly confidence continuously exceeds the dynamic confidence threshold, determine the anomaly type label with the highest occurrence frequency in the target candidate case as the final anomaly type, which is "masquerade attack" here, and output "masquerade attack" and the corresponding initial anomaly confidence as the anomaly detection result. At the same time, write the matching path association relationship between the aggregated feature vectors and the target candidate cases into the retrieval index structure of the historical anomaly case library to update the pattern matching logic.
[0032] Step S140: According to the abnormal detection result, match the target protection policy corresponding to the abnormal type from the preset protection policy library, and dynamically adjust the weight of the target protection policy based on the abnormal confidence level to generate a real-time protection policy instruction.
[0033] Suppose the abnormal detection result is "disguised attack". First, according to the threat level corresponding to the abnormal type, screen the candidate policy set that meets the minimum response level from the protection policy library. In the defined mapping relationship table between threat levels and response levels, the "disguised attack" scenario corresponds to a secondary response. Monitor the resource occupancy baseline of the available policies in the current protection policy library (including CPU occupancy rate, memory consumption, and network latency growth coefficient, which are calculated and updated in real time through the sliding window algorithm). If it is found that the resource requirements of a certain available policy are greater than the upper limit of the system idle resources, then raise the minimum response level to the next level (here it is the primary response). Screen out the candidate policy set that meets the secondary response from the protection policy library.
[0034] Rank the policies in the candidate policy set based on the abnormal confidence level. For example, the abnormal confidence level is 0.7. For policies A, B, and C in the candidate policy set, according to the preset priority rules related to the abnormal confidence level (such as the higher the abnormal confidence level, the higher the priority of the policy with stricter network traffic monitoring), assume that policy A has the strictest network traffic monitoring, then policy A has the highest priority, and select policy A as the basic protection policy.
[0035] Then, according to the bandwidth load rate, device online status, and user role permissions in the real-time network environment parameters, dynamically adapt the execution parameters of the basic protection policy to generate the target protection policy. For example, when the current bandwidth load rate is greater than the preset critical value, reduce the traffic monitoring frequency and increase the device fingerprint verification intensity according to the rules. If it is detected that a user with set permissions (such as senior company management) is online, while retaining the access channels for key data, restrict the access of external devices. Through these adjustments, a target protection policy instruction for the "disguised attack" scenario is generated.
[0036] Step S150: Execute the real-time protection policy instruction, synchronously regulate the network connection, device permissions, and user access paths in the remote work scenario, and feedback the regulated network status data to the multi-layer abnormal detection model to optimize the abnormal recognition logic.
[0037] In this embodiment, the real-time protection policy instruction is decomposed into a network connection control instruction, a device permission reset instruction, and a user access interception instruction. For example, the network connection control instruction may include operations such as blocking certain suspicious IP address segments and disabling specific ports. When executing the network connection control instruction, verify its compatibility with the current network topology. Parse the IP address segment blocking list and port disabling list in the network connection control instruction, and compare the blocking list with the current active connection list. If an overlapping address segment is found between the blocking list and the current active connection list, generate an address conflict report. According to the overlapping address segments in the address conflict report, query the corresponding device fingerprint features to determine the department to which the device belongs and the business importance level. For example, it is found that a device corresponding to a certain overlapping address segment is a device used by the marketing department for promoting an important project, and its business importance level is relatively high. Re-divide the blocked address segments based on the business importance level so that the connections of core business devices are not affected by the network connection control instruction. Merge the re-divided blocked address segments into the original network connection control instruction, generate a final control instruction that passes the compatibility verification, and execute it.
[0038] According to the device permission reset instruction, batch modify the file read / write permissions and peripheral interface enable status of the target device (such as a device suspected of being affected by a spoofing attack). For example, change the read / write permissions of some suspicious files from read / write to read-only, disable some unnecessary peripheral interfaces (such as USB interfaces), and record the permission difference logs before and after the modification.
[0039] Based on the user access interception instruction, inject virtual access paths into the user behavior data to induce the attacker to trigger the protection mechanism, and at the same time back up an encrypted copy of the real access path. For example, after an employee logs in to the login page of the office system, create some seemingly valuable but actually virtual file access paths. If the attacker tries to access these paths, the protection mechanism will be triggered. At the same time, encrypt and back up the employee's normal real access path.
[0040] Finally, use the permission difference logs and the encrypted copy as the regulated network status data and input them into the multi-layer anomaly detection model. For example, the primary detection layer in the multi-layer anomaly detection model can re-evaluate the security of the device based on the permission change information in the permission difference logs, and the intermediate aggregation layer can optimize the combination weight calculation rules in the threat scenario template according to the real access path information in the encrypted copy, thereby optimizing the anomaly recognition logic and improving the detection and prevention capabilities of the entire system against network security threats in the remote work scenario.
[0041] Based on the above steps, the embodiment of this application comprehensively obtains network activity data including traffic data, device status data, and user behavior data generated by multiple terminal devices in the remote work scenario, constructs an all-round and multi-dimensional network activity portrait. When performing multi-dimensional security feature analysis on the network activity data and generating a dynamic security feature set, the traffic features, device fingerprint features, and behavior pattern features included in the dynamic security feature set are not simply a list of static features. The dynamic security feature set will be continuously updated with the real-time changes of network activities, and can adapt to the diverse and changeable network behavior patterns in the remote work scenario. Compared with the traditional fixed feature analysis method, it can more accurately depict the real security status of network activities at each moment, and effectively identify the evolving new network attack patterns.
[0042] The multi-layer anomaly detection model uses the dynamic security feature set to identify abnormal behaviors in the network activity data, and outputs anomaly detection results including anomaly types and anomaly confidence levels. It not only improves the accuracy and reliability of detection, but also through multi-level analysis, can dig out abnormal behaviors from different angles, greatly reducing the false alarm rate and missed alarm rate. Then, according to the anomaly detection results, the target protection policy is matched from the preset protection policy library, and its dynamic weight is adjusted based on the anomaly confidence level to generate a real-time protection policy instruction, abandoning the limitations of the traditional fixed protection policy. It can flexibly adjust the intensity of protection measures according to the severity and credibility of abnormal situations, and can be accurately customized according to the actual threat situation, significantly improving the pertinence and effectiveness of the protection policy.
[0043] Finally, execute the real-time protection policy instruction to synchronously regulate the network connection, device permissions, and user access paths in the remote work scenario, and feedback the regulated network state data to the multi-layer anomaly detection model to optimize the anomaly recognition logic, realizing the dynamic self-adjustment and continuous optimization of the protection process. It can continuously improve the anomaly recognition logic with the change of the network environment, making the protection system always maintain an efficient response ability to various security threats, providing an adaptive, intelligent and continuously evolving protection guarantee for remote work network security.
[0044] In a possible implementation manner, step S120 includes:
[0045] Step S121, extract the protocol type distribution feature, packet transmission frequency feature, and target address clustering feature from the traffic data to form the traffic feature.
[0046] In this embodiment, the extraction process of the protocol type distribution feature in traffic data is relatively complex. In the entire enterprise remote office network, data interaction between office software and servers generates traffic, and this traffic is transmitted based on different protocols. For example, the enterprise email client and the server mainly use the standard IMAP or POP3 protocol to obtain email content, and the HTTPS protocol may also be involved during the transmission process to ensure data security. By analyzing a large amount of traffic data and capturing the nested relationship between the transport layer protocol and the application layer protocol, the conventional protocol combination patterns can be identified. The proportion of the number of data packets of different protocols under the same source address is statistically calculated to generate a protocol distribution vector. For example, a certain employee's laptop is used as the source address. Among the data packets sent within a working day, the data packets related to the enterprise office system account for 70%. Among them, the HTTP protocol accounts for 40%, the HTTPS protocol accounts for 30%, and the remaining 30% of the data packets are generated by other applications, which may include a small amount of FTP protocol (5%) for file sharing and other operations. The similarity between this protocol distribution vector and the known malicious protocol library is compared. If the proportion of normal business traffic in the enterprise office network decreases, the dynamic setting score is reduced according to the rules, which may cause some protocols that were originally regarded as normal to be marked as risk protocols under the new scoring criteria. For example, a certain custom encryption protocol is regarded as normal when the normal business traffic is sufficient. However, as the proportion of normal business traffic decreases, due to its certain similarity to a certain malicious encryption protocol in the malicious protocol library, it is marked as a risk protocol under the new scoring.
[0047] The feature of the data packet transmission frequency is also an important part of the traffic feature. In daily office work, employees' usage frequencies of the office network are different at different time periods, which is reflected in the data packet transmission frequency. For example, from 9 am to 11 am is the time when most employees concentrate on processing emails and conducting project communication. At this time, the data packet transmission frequency in the office network is relatively high, and a large amount of email data, instant messaging data, etc. are transmitted in the network. During lunch time, the data packet transmission frequency will decrease significantly. If a certain terminal device suddenly shows an abnormally high data packet transmission frequency during a normal low-traffic period (such as lunch time), this may be a potential security threat signal.
[0048] In terms of the clustering characteristics of target addresses, the terminal devices in an enterprise's office network usually communicate with specific servers or other devices. For example, the terminal devices in the finance department often interact with the finance server, and these frequently communicated target addresses can be clustered together. If a device suddenly starts to perform a large amount of data transmission to an external address that has never had a communication record, and this external address does not conform to the normal clustering characteristics of target addresses, this constitutes a suspicious clustering characteristic of target addresses. These protocol type distribution characteristics, packet transmission frequency characteristics, and target address clustering characteristics together constitute the traffic characteristics.
[0049] Step S122: Parse the device hardware identification characteristics, operating system vulnerability characteristics, and process resource occupancy characteristics from the device status data to form the device fingerprint characteristics.
[0050] The device hardware identification characteristic is the unique identifier of the device. Taking the laptop used by an employee as an example, its MAC address is fixed and unchanged, and this hardware identification characteristic is recorded when the device accesses the network, becoming the basic part of the device fingerprint characteristics. The operating system vulnerability characteristics are crucial for the security of the device. For example, most of the office computers in an enterprise use the Windows operating system. When Microsoft releases a security update patch, if a device does not update in time, its operating system may have known vulnerabilities. For example, for a certain remote code execution vulnerability, hackers may use this vulnerability to execute malicious code on the device without authorization. The process resource occupancy characteristic reflects the resource usage situation of the device during operation. When an employee opens multiple office software at the same time, such as document editing software, project management software, and instant messaging software, each software process will occupy a certain CPU usage rate and memory usage amount. If a certain process suddenly occupies excessive CPU resources, for example, a certain office software process originally only occupied 10% of the CPU resources under normal circumstances but suddenly soars to 80%, and at the same time the memory usage amount also increases significantly, this may be because the software has a malfunction or is attacked by malware, and this process resource occupancy characteristic becomes a part of the device fingerprint characteristics. These device hardware identification characteristics, operating system vulnerability characteristics, and process resource occupancy characteristics together constitute the device fingerprint characteristics.
[0051] Step S123: Identify the login time series characteristics, file operation track characteristics, and permission change history characteristics from the user behavior data to form the behavior pattern characteristics.
[0052] For example, the login time series features can reflect the working patterns of employees. For example, an employee logs in to the company's office system at around 9:00 every morning, and this login time is relatively fixed. If one day this employee logs in to the office system at 3:00 am, which does not match the normal login time series features, it may be a signal of potential security risks. The file operation track features record the operations of employees on files during the work process. In the shared folders of an enterprise, all operations of employees on files, such as creation, modification, and deletion, will be recorded in detail. For example, an employee is mainly responsible for editing project documents and usually operates in a specific project folder. If it is suddenly found that this employee starts to frequently access other folders unrelated to work and modifies or deletes some important confidential files, this constitutes abnormal file operation track features. The permission change history features record the changes in employees' permissions in the enterprise office system. For example, after an employee is promoted from an ordinary employee to a project leader, their permissions change from only being able to read certain project files to being able to edit and manage these files, and information such as the time of this permission change and the specific content of the change will be recorded. These login time series features, file operation track features, and permission change history features together constitute the behavior pattern features.
[0053] Step S124: Align the traffic features, device fingerprint features, and behavior pattern features according to a time window, and perform cross-dimensional correlation analysis on the aligned features to generate the dynamic security feature set.
[0054] Among them, the cross-dimensional correlation analysis includes: identifying abnormal data transmission paths based on the target address clustering features in the traffic features and the file operation track features in the behavior pattern features. Detecting potential threats of mismatched resource consumption and traffic based on the process resource occupancy features in the device fingerprint features and the data packet transmission frequency features in the traffic features.
[0055] For example, taking one hour as a time window, the traffic characteristics, device fingerprint characteristics, and behavior pattern characteristics within the same hour are correlated. When performing cross-dimensional correlation analysis, abnormal data transmission paths are identified based on the target address clustering characteristics in the traffic characteristics and the file operation track characteristics in the behavior pattern characteristics. For example, the department where an employee belongs mainly conducts data interaction with internal servers, and the target address clustering characteristics show that the addresses he often accesses are all internal server addresses of the enterprise. However, from the file operation track characteristics, it is found that this employee copies a large number of internal confidential files of the enterprise to an external address, which is completely different from the addresses in the normal target address clustering characteristics. This identifies an abnormal data transmission path, and there may be a risk of data leakage. Based on the process resource occupancy characteristics in the device fingerprint characteristics and the packet transmission frequency characteristics in the traffic characteristics, potential threats of mismatched resource consumption and traffic are detected. For example, the process resource occupancy of a certain device is very high, such as the CPU usage rate remains above 80% for a long time, but its packet transmission frequency is very low. This indicates that there may be malicious programs on the device that occupy a large amount of resources in the background without normal network data interaction. This is a potential threat situation of mismatched resource consumption and traffic. Through such cross-dimensional correlation analysis, a set of dynamic security characteristics corresponding to network activity data is generated, providing an important basis for subsequent operations such as anomaly detection.
[0056] In a possible implementation manner, the multi-layer anomaly detection model includes a primary detection layer, an intermediate aggregation layer, and a high-level decision layer.
[0057] Step S130 includes:
[0058] Step S131, input the set of dynamic security characteristics into the primary detection layer, and independently perform anomaly scoring on the traffic characteristics, device fingerprint characteristics, and behavior pattern characteristics respectively to obtain a set of primary anomaly scores.
[0059] In this embodiment, for the traffic characteristics, if an abnormality is found in the analysis of traffic data, such as the protocol type distribution characteristics of a certain terminal device showing an abnormality, for example, using a protocol highly similar to a malicious protocol, and the packet transmission frequency characteristics showing that the transmission frequency is too high, exceeding the average transmission frequency of this device during normal working hours, and the target address clustering characteristics showing that this device frequently communicates with external suspicious addresses, then according to the predefined traffic characteristic anomaly scoring rules, the traffic characteristics of this device will obtain a relatively high anomaly score.
[0060] In terms of device fingerprint features, if the operating system vulnerability features of a certain device indicate the existence of multiple unpatched high-risk vulnerabilities, and the process resource occupancy features show that a certain process continuously occupies a large amount of CPU and memory resources, this abnormal resource occupancy situation combined with the operating system vulnerability situation, according to the abnormal scoring rules of device fingerprint features, the device fingerprint features of this device will be given corresponding abnormal scores.
[0061] From the perspective of behavior pattern features, the login time series features of a certain employee show obvious abnormalities, such as frequent logins during non-working hours, and the file operation track features show that he has abnormal access to and operations on a large number of confidential files, and the permission change history features show that there are non-compliances in permission changes. According to the abnormal scoring rules of behavior pattern features, the corresponding behavior pattern features of this employee will also get an abnormal score. These abnormal scores for traffic features, device fingerprint features, and behavior pattern features together constitute the primary abnormal score set.
[0062] Step S132, input the primary abnormal score set into the intermediate aggregation layer, calculate the combined weights of multiple primary abnormal scores according to the preset threat scenario template, and generate intermediate abnormal aggregation indicators.
[0063] For example, in a possible implementation manner, step S132 includes:
[0064] Step S1321, receive the traffic feature abnormal score, device fingerprint feature abnormal score, and behavior pattern feature abnormal score in the primary abnormal score set, and extract the feature attribute tags and timestamp sequences corresponding to each abnormal score.
[0065] Step S1322, parse the combined logic rules defined in the threat scenario template, match the corresponding logical condition branches according to the feature attribute tags, and determine the dynamic weight coefficients of each abnormal score in the combined logic rules.
[0066] Step S1323, perform a fluctuation trend analysis of the traffic feature abnormal score, device fingerprint feature abnormal score, and behavior pattern feature abnormal score within a time window based on the timestamp sequence, and generate the trend intensity factors of each abnormal score.
[0067] Step S1324, calculate the weighted contribution values of the traffic feature abnormal score, device fingerprint feature abnormal score, and behavior pattern feature abnormal score according to the product results of the dynamic weight coefficients and the trend intensity factors.
[0068] Step S1325, accumulate the weighted contribution values in the aggregation direction defined in the threat scenario template to generate an initial aggregation indicator.
[0069] Step S1326: Detect the initial aggregation metric and the metric threshold range of the same type of threat scenario in the historical aggregation metric library. If the initial aggregation metric exceeds the upper limit of the metric threshold range, trigger the dynamic compensation mechanism.
[0070] Step S1327: Non-linearly correct the accumulated result of the weighted contribution value through the dynamic compensation mechanism to generate the intermediate anomaly aggregation metric.
[0071] For example, the feature attribute label corresponding to the traffic feature anomaly score is "protocol and transmission anomaly", and the timestamp sequence shows that the anomaly occurred between 10:00 am and 11:00 am; the feature attribute label corresponding to the device fingerprint feature anomaly score is "resource and vulnerability anomaly", and the timestamp sequence is from 10:30 am to 11:00 am; the feature attribute label corresponding to the behavior pattern feature anomaly score is "login and operation anomaly", and the timestamp sequence is from 10:15 am to 11:00 am.
[0072] Then, according to the threat scenario template, when the traffic feature anomaly score is higher than the first set score and the behavior pattern feature anomaly score is lower than the second set score, it is marked as a disguised attack scenario. Assume that the current traffic feature anomaly score is indeed higher than the first set score and the behavior pattern feature anomaly score is lower than the second set score. Then, the dynamic weight coefficient of the traffic feature anomaly score in the combined logic rules related to this disguised attack scenario will be higher.
[0073] Furthermore, within the above time window (10:00 am to 11:00 am), if the traffic feature anomaly score continuously rises over time, its trend intensity factor will be higher; if the device fingerprint feature anomaly score is relatively stable during this period, its trend intensity factor may be lower; and if the behavior pattern feature anomaly score fluctuates greatly during this period, its trend intensity factor will also be higher.
[0074] Next, assume that the dynamic weight coefficient of the traffic feature anomaly score is 0.6 and the trend intensity factor is 0.8. Then, the weighted contribution value is 0.6 * 0.8 = 0.48; the dynamic weight coefficient of the device fingerprint feature anomaly score is 0.3 and the trend intensity factor is 0.2, and the weighted contribution value is 0.3 * 0.2 = 0.06; the dynamic weight coefficient of the behavior pattern feature anomaly score is 0.1 and the trend intensity factor is 0.5, and the weighted contribution value is 0.1 * 0.5 = 0.05.
[0075] Thus, it can be assumed that the initial aggregation metric obtained by summing up the above three weighted contribution values is 0.59. Detect the metric threshold range of this initial aggregation metric and similar threat scenarios (such as spoofing attack scenarios) in the historical aggregation metric library. If the initial aggregation metric exceeds the upper limit of the metric threshold range, for example, the metric threshold range of the spoofing attack scenario is 0.3 - 0.5 and 0.59 exceeds this upper limit, then the dynamic compensation mechanism is triggered.
[0076] Furthermore, assume that the overrun ratio of the initial aggregation metric is (0.59 - 0.5) / 0.5 = 0.18. According to this overrun ratio, adjust the decay rate of the weighted contribution value. If the original correlation constraint condition between the abnormal score of traffic characteristics and the abnormal score of device fingerprint characteristics is that their sum does not exceed 0.8, after adjusting the decay rate, this correlation constraint condition may become that their sum does not exceed 0.7. Through such non-linear correction, an intermediate abnormal aggregation metric is generated.
[0077] Step S133, input the intermediate abnormal aggregation metric into the high-level decision-making layer, and combine the pattern matching results in the historical abnormal case library and the real-time network environment parameters to calculate the abnormal confidence level and determine the abnormal type.
[0078] Among them, the threat scenario template defines the following combination logic: when the abnormal score of traffic characteristics is higher than the first set score and the abnormal score of behavior pattern characteristics is lower than the second set score, it is marked as a spoofing attack scenario. When the abnormal score of device fingerprint characteristics continuously increases and is negatively correlated with the abnormal score of traffic characteristics, it is marked as a device hijacking scenario.
[0079] In a possible implementation manner, the execution process of the dynamic compensation mechanism includes: adjusting the decay rate of the weighted contribution value according to the overrun ratio of the initial aggregation metric, and applying the adjusted decay rate to the correlation constraint condition between the abnormal score of traffic characteristics and the abnormal score of device fingerprint characteristics.
[0080] For example, in a possible implementation manner, step S133 includes:
[0081] Step S1331, parse the aggregation feature vector and time range marker in the intermediate abnormal aggregation metric, and extract the traffic feature correlation strength, device fingerprint fluctuation amplitude, and behavior pattern deviation degree in the aggregation feature vector.
[0082] For example, the traffic feature correlation strength shows that the traffic deviates greatly from the normal mode, the device fingerprint fluctuation amplitude indicates that the device state is significantly unstable, the behavior pattern deviation degree reflects a large difference between the user behavior and the normal mode, and the time range marker is from 10 am to 11 am.
[0083] Step S1332: Mark historical attack case fragments with the same time attribute retrieved from the historical anomaly case library according to the time range, and generate a candidate anomaly type set.
[0084] For example, mark historical attack case fragments with the same time attribute (from 10 am to 11 am) retrieved from the historical anomaly case library according to the time range, and generate a candidate anomaly type set. Suppose three historical attack case fragments are retrieved, namely a masquerade attack case, a device hijacking attack case, and a data leakage attack case.
[0085] Step S1333: Perform multi-dimensional feature matching between the aggregated feature vector and each historical attack case fragment in the candidate anomaly type set, and calculate the similarity between the traffic feature association strength and the protocol anomaly index in the historical attack case fragment, the coverage of the device fingerprint fluctuation range and the resource tampering index in the historical attack case fragment, and the coincidence degree between the behavior pattern deviation degree and the privilege abuse index in the historical attack case fragment.
[0086] For example, for the masquerade attack case fragment, the similarity between the traffic feature association strength and the protocol anomaly index therein is 0.6, the coverage of the device fingerprint fluctuation range and the resource tampering index is 0.3, and the coincidence degree between the behavior pattern deviation degree and the privilege abuse index is 0.4.
[0087] Step S1334: Based on the number of online devices, user activity level, and network throughput fluctuation value in the real-time network environment parameters, perform dynamic weighted adjustment on the similarity, coverage, and coincidence degree to generate a weighted matching degree list.
[0088] Suppose the current number of online devices is large, the user activity level is high, and the network throughput fluctuation value is large. According to the pre-set weighted rules, perform weighted adjustment on the above similarity, coverage, and coincidence degree. For example, the adjusted weighted matching degree for the masquerade attack case fragment is 0.7 * 0.6 + 0.2 * 0.3 + 0.1 * 0.4 = 0.52.
[0089] Step S1335: Sort the matching degree values in the weighted matching degree list in descending order, select the historical attack case fragments corresponding to the top N matching degree values as target candidate cases, and extract the anomaly type labels and confidence correction coefficients marked in the target candidate cases.
[0090] Assume N = 2. After sorting, the weighted matching degrees of the spoofing attack case segment and the device hijacking attack case segment are relatively high and are selected as target candidate cases. Then, extract the abnormal type labels and confidence correction coefficients marked in the target candidate cases. The abnormal type label of the spoofing attack case segment is "Spoofing Attack" and the confidence correction coefficient is 0.8. The abnormal type label of the device hijacking attack case segment is "Device Hijacking" and the confidence correction coefficient is 0.7.
[0091] Step S1336: Calculate the initial abnormal confidence corresponding to the intermediate abnormal aggregation index according to the occurrence frequency of the abnormal type label of the target candidate case and the confidence correction coefficient.
[0092] Assume that the label "Spoofing Attack" appears with a relatively high frequency in the target candidate cases. According to the calculation rule, the initial abnormal confidence is 0.7.
[0093] Step S1337: Perform time decay compensation on the initial abnormal confidence based on the network state baseline data within the time range marker to generate a dynamic confidence threshold.
[0094] Adjust the decay curve slope of the initial abnormal confidence according to the duration of the time range marker (from 10 am to 11 am) and the deviation rate of the network state baseline data, so that the confidence decay rate in the short-term high-deviation scenario is lower than that in the long-term low-deviation scenario. For example, if the network state baseline data deviates greatly within this time period but the time is short, the adjusted decay curve slope will result in a slower confidence decay rate. After time decay compensation, a dynamic confidence threshold is generated.
[0095] Step S1338: Compare the initial abnormal confidence with the dynamic confidence threshold. If the initial abnormal confidence continuously exceeds the dynamic confidence threshold, determine the abnormal type label with the highest occurrence frequency in the target candidate case as the final abnormal type.
[0096] Step S1339: Output the final abnormal type and the corresponding initial abnormal confidence as the abnormal detection result, and write the association relationship between the aggregated feature vector and the matching path of the target candidate case into the retrieval index structure of the historical abnormal case library to update the pattern matching logic.
[0097] Among them, the execution process of the time decay compensation includes: adjusting the decay curve slope of the initial abnormal confidence according to the duration of the time range marker and the deviation rate of the network state baseline data, so that the confidence decay rate in the short-term high-deviation scenario is lower than that in the long-term low-deviation scenario.
[0098] Assume that the initial anomaly confidence level of 0.7 continuously exceeds the dynamic confidence threshold. Since the label of "masquerade attack" has the highest occurrence frequency among the target candidate cases, the final anomaly type is determined as "masquerade attack", and the "masquerade attack" and the corresponding initial anomaly confidence level are output as the anomaly detection result. At the same time, the matching path association relationship between the aggregated feature vector and the target candidate case is written into the retrieval index structure of the historical anomaly case library to update the pattern matching logic.
[0099] In a possible implementation manner, step S140 includes:
[0100] Step S141, screening a set of candidate policies that meet the minimum response level from the protection policy library according to the threat level corresponding to the anomaly type.
[0101] In this embodiment, in the protection policy library of this enterprise, corresponding protection policies are set for different threat levels. For the anomaly type of masquerade attack, its corresponding threat level is defined as a certain specific level (for example, the second-level threat level). The policies in the protection policy library are divided into different response levels, and each level contains multiple protection measures. According to the mapping relationship between the threat level and the response level, all the policies that meet the second-level response level are screened out, and these policies constitute the set of candidate policies. For example, one of the policies may be to strengthen the protocol analysis of network traffic, and another policy may be to conduct more frequent audits on specific user behaviors, etc.
[0102] Step S142, based on the anomaly confidence level, sorting the policies in the set of candidate policies by priority, and selecting the policy with the highest priority as the basic protection policy.
[0103] For example, the previously calculated anomaly confidence level for the masquerade attack is 0.7, and this confidence value will affect the priority of the policies. Each policy in the set of candidate policies is sorted according to the pre-set priority rules related to the anomaly confidence level. For example, for a higher anomaly confidence level, those policies that can more directly target the possible attack routes of the masquerade attack (such as network traffic masquerade, user identity masquerade, etc.) and have better protection effects will be given higher priorities. Assume that there is a policy in the set of candidate policies that deeply detects the masquerade source in network traffic and is set to have a higher priority when the anomaly confidence level is high, and this policy will be selected as the basic protection policy.
[0104] Step S143, dynamically adapting the execution parameters of the basic protection policy according to the bandwidth load rate, device online status, and user role permissions in the real-time network environment parameters to generate the target protection policy.
[0105] Among them, the dynamic adaptation includes: when the bandwidth load rate is greater than a preset critical value, reducing the traffic monitoring frequency and increasing the device fingerprint verification intensity. When it is detected that a user with set permissions is online, while retaining the access channels for critical data, restricting the access of external devices.
[0106] For example, in the current telecommuting scenario, the network environment is in a dynamic state of change. If it is detected that the bandwidth load rate is greater than the preset critical value, it means that the network resources are relatively tense. According to the dynamic adaptation rules, it is necessary to reduce the traffic monitoring frequency to relieve the network burden, and at the same time increase the device fingerprint verification intensity to ensure the legitimacy of the device. For example, originally, traffic monitoring was performed every 10 seconds, and now it is adjusted to every 30 seconds; in terms of device fingerprint verification, from the previous single hardware identification verification, it is increased to a combined verification of multiple aspects such as hardware identification and operating system characteristics.
[0107] When it is detected that a user with set permissions (such as enterprise senior management or core project leaders, etc.) is online, in order to ensure their normal office operations, it is necessary to restrict the access of external devices while retaining the access channels for critical data. For example, for enterprise senior management, they need to access key data such as the enterprise's core financial data and strategic planning documents, so it is necessary to ensure that the channels for them to access these data normally remain unblocked. At the same time, in order to prevent potential security risks brought by external devices (such as malicious device access to steal data or conduct attacks, etc.), the access of external devices (such as unauthorized external hard drives, USB flash drives, etc.) is restricted. Through these dynamic adaptation operations based on real-time network environment parameters, the execution parameters of the basic protection strategy are adjusted, and finally a target protection strategy for the disguised attack scenario is generated to ensure network security while guaranteeing office efficiency.
[0108] In a possible implementation manner, step S150 includes:
[0109] Step S151, decomposing the real-time protection strategy instruction into a network connection control instruction, a device permission reset instruction, and a user access interception instruction.
[0110] Taking the previously detected disguised attack as an example, the network connection control instruction may include restricting the connection of a specific suspicious IP address segment to the enterprise office network, prohibiting communication on certain ports, etc. The device permission reset instruction involves adjusting the permissions of the device suspected of being affected by the disguised attack. For example, the target device may be an employee's laptop computer, and abnormal network activities have been detected on this computer. The user access interception instruction aims to monitor and restrict the user's access behavior to prevent the attacker from further obtaining the enterprise's sensitive information.
[0111] Step S152: Verify the compatibility between the network connection control instruction and the current network topology. If there is a conflict between the network connection control instruction and the current network topology, reassign the IP address pool according to the device fingerprint features.
[0112] The network topology of an enterprise is relatively complex, including the connection relationships between subnets of multiple departments, server groups, and various network devices (such as routers, switches, etc.). When executing a network connection control instruction, for example, when the instruction requires blocking a certain IP address segment, it is necessary to compare the IP address segment blocking list with the current active connection list. If an overlapping address segment is found between the blocking list and the current active connection list, an address conflict report will be generated. Suppose there is an IP address segment in the blocking list that overlaps with the IP address of a device in the marketing department that is promoting an important project and requires continuous network connection. At this time, reassign the IP address pool according to the device fingerprint features. By querying the device fingerprint features (such as the MAC address of the device, the type of operating system, etc.), it is determined that the device belongs to the marketing department, and since it is promoting an important project, the business importance level is relatively high. Based on this, the blocked address segment is re-divided so that the connection of the devices in the marketing department is not affected, and then the re-divided blocked address segment is merged into the original network connection control instruction to generate a final control instruction with passed compatibility verification and execute it, so as to ensure the effective protection of network security under the premise of not affecting the connection of important business devices.
[0113] Step S153: According to the device permission reset instruction, batch modify the file read / write permissions and the enabled status of peripheral interfaces of the target device, and record the permission difference log before and after the modification.
[0114] According to the device permission reset instruction, batch modify the file read / write permissions and the enabled status of peripheral interfaces of the target device (such as an employee's laptop suspected of having a spoofing attack risk). For example, for confidential files within the enterprise, change the read / write permissions of these files on the target device from readable and writable to read-only to prevent attackers from tampering with or stealing the file content. For peripheral interfaces, such as the USB interface, change its enabled status from the default usable to disabled to prevent external devices (such as USB drives that may be implanted with malicious programs) from being inserted into the device to steal data or spread malware. While performing these permission modification operations, record the permission difference log before and after the modification, detailing which files' permissions have changed, what permissions they have changed from and to, and how the enabled status of the peripheral interfaces has changed, etc.
[0115] Step S154: Based on the user access interception instruction, inject a virtual access path into the user behavior data to induce the attacker to trigger the protection mechanism, and at the same time back up an encrypted copy of the real access path.
[0116] For example, in the operation interface after an employee logs in to the enterprise office system, create some seemingly valuable but actually virtual file access paths. These virtual access paths may seem to lead to the paths of the enterprise's core confidential files, but in fact, they are specially set traps. If an attacker attempts to access these virtual paths by disguising as a legitimate user, it will trigger protection mechanisms, such as issuing an alarm, restricting the attacker's further operations, etc. At the same time, encrypt and back up the employee's normal real access path to ensure that the employee's normal office operations are not affected, and the real access path can be restored according to the encrypted copy when needed.
[0117] Step S155, input the permission difference log and the encrypted copy as the regulated network status data into the multi-layer anomaly detection model to update the combined weight calculation rule in the threat scenario template.
[0118] The primary detection layer in the multi-layer anomaly detection model can re-evaluate the security of the device according to the permission change information in the permission difference log. For example, if the file read and write permissions of a certain device change, and the subsequent network activity data shows an abnormal decrease or increase, this may imply other potential security problems with the device. The intermediate aggregation layer can optimize the combined weight calculation rule in the threat scenario template according to the real access path information in the encrypted copy. For instance, if it is found that the characteristics of a certain real access path in the encrypted copy are similar to the access path characteristics in the previously detected disguised attack, the corresponding combined weight calculation rule can be adjusted so that similar disguised attacks can be more accurately identified in future detections, thereby improving the detection and prevention capabilities of the entire system against network security threats in the remote office scenario.
[0119] In a possible implementation manner, step S121 includes:
[0120] Step S1211, capture the nested relationship between the transport layer protocol and the application layer protocol in the traffic data, and identify the unconventional protocol combination pattern.
[0121] In the enterprise's remote work network, traffic data contains a lot of information about network activities. First, specific network monitoring technologies are used to capture the nested relationship between the transport layer protocol and the application layer protocol in the traffic data, so as to identify unconventional protocol combination patterns. For example, in the enterprise office network, common transport layer protocols such as TCP and UDP have normal nested relationships with application layer protocols such as HTTP, HTTPS, SMTP, etc. Employees use the enterprise email client (sending emails based on the SMTP protocol and receiving emails based on the POP3 or IMAP protocol) to communicate with the enterprise email server, which is a normal application scenario, and the nested relationship between the transport layer protocol and the application layer protocol is clear and conforms to the convention. However, when an uncommon nested relationship appears in the traffic data of a certain terminal device, such as nesting a custom encryption protocol on top of the UDP protocol, and this combination has never appeared in the normal business of the enterprise office network, it is identified as an unconventional protocol combination pattern.
[0122] Step S1212, count the proportion of the number of data packets of different protocols under the same source address, and generate a protocol distribution vector.
[0123] Taking a certain employee's laptop in the enterprise as the source address, within a specific time period (such as within a working day), analyze the traffic data generated by it. This laptop conducts data interaction with the enterprise internal server, other employees' devices, and the external network (such as accessing the Internet to obtain business-related materials). Through statistics, it is found that among all the data packets generated during this period, the proportion of data packets based on the HTTP protocol is 40%, and these data packets are mainly generated when employees use the browser to access the enterprise internal office system web page; the proportion of data packets based on the HTTPS protocol is 30%, which may be used for secure file transfer or logging in to the enterprise encryption service; there are also 15% of the data packets based on the FTP protocol, which is because employees download some large project files from the enterprise internal file server; the remaining 15% are data packets generated by various other protocols, such as the SNMP protocol for network management, etc. In this way, a protocol distribution vector based on this source address is generated, that is, [HTTP: 40%, HTTPS: 30%, FTP: 15%, other: 15%].
[0124] Step S1213, compare the similarity of the protocol distribution vector with the known malicious protocol library. If the similarity is greater than the dynamically set score, mark the corresponding protocol as a risk protocol.
[0125] Among them, the dynamically set score is automatically adjusted according to the proportion of normal business traffic in the current network environment: when the proportion of normal business traffic decreases, reduce the dynamically set score to expand the scope of risk protocol identification.
[0126] For example, the known malicious protocol library contains a series of protocol features that have been confirmed to have malicious intentions or are highly relevant to malicious activities. In a normal enterprise office network environment, when the proportion of normal business traffic is relatively high, the dynamically set score is at a relatively high level. For example, the dynamically set score is set to 0.8. When comparing, if the similarity calculated between a certain protocol in the protocol distribution vector and the protocol features in the known malicious protocol library is lower than 0.8, then this protocol is regarded as a normal protocol. However, when the proportion of normal business traffic in the enterprise office network decreases, in order to expand the scope of risk protocol identification, the dynamically set score will be reduced. Suppose the proportion of normal business traffic drops to a certain extent due to some reason (such as some business services being interrupted due to a network attack), and at this time the dynamically set score is reduced to 0.6. If in this case, a certain protocol that was previously regarded as normal (such as a rarely used custom encryption protocol, whose similarity with the malicious protocol library was calculated as 0.7 when the proportion of normal business traffic was high and was regarded as normal because it was lower than 0.8), due to the reduction of the dynamically set score, its similarity with the malicious protocol library is calculated as 0.65, which is greater than 0.6, then this protocol will be marked as a risk protocol. This mechanism of automatically adjusting the dynamically set score according to the proportion of normal business traffic in the network environment helps to more accurately identify potentially risky protocols in different network states, thereby improving the security of the entire network and ensuring the normal progress of enterprise remote work.
[0127] In a possible implementation manner, the construction process of the historical abnormal case library includes:
[0128] Step S210, collect the full life cycle data of historical network attack events, and extract the features of the initial attack stage, the lateral movement path features, and the data leakage node features.
[0129] In the network environment of an enterprise, there have been many network attack events in history. For each attack event, the full life cycle data covers all relevant information from the start to the end of the attack. Taking an attack on the data of the enterprise's finance department as an example, the features of the initial attack stage may include the IP address of the attack source (suppose it is from a certain external suspicious IP address segment), the time when the attack was launched (such as 2 am, which is a non-working time and does not conform to the normal office network activity pattern), and the initial means used in the attack (for example, taking advantage of an operating system vulnerability on a device in the finance department that was not updated in time, and information such as the number and type of this vulnerability).
[0130] The lateral movement path feature describes the activity trajectory of the attacker after entering the enterprise network. The attacker may first invade an edge device in the finance department, and then use the permissions of this device to access devices in other departments through the enterprise internal network. For example, the attacker uses the permissions shared by the internal network to move laterally from the device in the finance department to the device in the human resources department. In this process, a specific protocol (such as the SMB protocol in the Windows network) may be used for data transmission and permission acquisition.
[0131] The data leakage node feature involves information related to the attacker obtaining enterprise sensitive data and attempting to transmit it out of the enterprise network. For example, the attacker finds a database file containing enterprise employee salary information on the device in the human resources department, and then encrypts and compresses these files and attempts to leak the data through a hidden external channel (such as disguising as normal network traffic and sending it to an overseas server). These attack initial stage features, lateral movement path features, and data leakage node features are all carefully extracted from the full - life - cycle data, providing a basis for subsequent processing.
[0132] Step S220: Perform event slicing on the full - life - cycle data to generate multiple independent attack - stage cases.
[0133] For the above - mentioned attack events targeting the data of the enterprise's finance and human resources departments, take the time point when the attacker first breaks through the network boundary as the starting slicing point. This time point may be the moment when the attacker successfully invades using an operating system vulnerability of the device in the finance department. Take the time point when the attacker obtains key data (such as the database file of enterprise employee salary information) or triggers the protection mechanism (assuming the enterprise's intrusion detection system detects abnormal traffic and issues an alarm) as the ending slicing point.
[0134] Between the starting slicing point and the ending slicing point, perform sub - slicing division at fixed time intervals or key operation events. For example, perform sub - slicing division at intervals of every 10 minutes, or perform sub - slicing division when the attacker performs key operations (such as obtaining new device permissions, encrypting data, etc.). For each sub - slice, extract the following features: the vulnerability number used (such as the vulnerability number CVE - 2021 - 1234), the protocol type used for lateral movement (such as the SMB protocol mentioned above), and the hash value of the injected malicious payload (assuming the malicious payload is hashed to obtain a specific hash value, such as abcdef1234567890).
[0135] In this way, the full - life - cycle data of the entire attack event is divided into multiple independent sub - slices, and each sub - slice represents a specific stage in the attack process. These independent sub - slices generate multiple independent attack - stage cases.
[0136] Step S230, label the corresponding protection strategy effectiveness labels for each attack phase case, where the protection strategy effectiveness labels include strategy response speed, resource consumption ratio, and false interception rate.
[0137] For each independent attack phase case, analyze the effectiveness of the protection strategy adopted for that phase at that time. For example, in the initial stage of the attack, an enterprise deployed a signature-based intrusion detection system as a protection strategy. In terms of the strategy response speed of this strategy, the time interval from the start of the attack to the intrusion detection system issuing an alarm is 5 minutes, and this time is recorded as the strategy response speed.
[0138] In terms of the resource consumption ratio, the intrusion detection system occupies a certain amount of CPU resources and memory resources during operation. Assume that during the detection of this attack phase, the average CPU occupancy rate of the intrusion detection system is 10%, and the memory consumption is 500MB (relative to the total resources of the enterprise server). These resource consumption situations are calculated as the resource consumption ratio.
[0139] In terms of the false interception rate, if during the detection process, the intrusion detection system wrongly intercepts some normal office network traffic (such as employees' normal file download traffic) as attack traffic, count the ratio of the intercepted traffic to the total detected traffic. Assume it is 2%, and this ratio is the false interception rate. Then label these strategy response speed, resource consumption ratio, and false interception rate on the corresponding attack phase case.
[0140] Step S240, cluster the labeled attack phase cases according to the attack type to form the retrieval index structure of the historical abnormal case library.
[0141] Among them, the retrieval index structure adopts multi-dimensional hash mapping, and the dimensions of the multi-dimensional hash mapping include attack tool hash, vulnerability exploitation hash, and data encryption method hash.
[0142] In the historical network attack events of an enterprise, there are different types of attacks, such as the previously mentioned attack for data theft, and there may also be denial-of-service attacks, malicious software implantation attacks, etc. Cluster the labeled attack phase cases according to the attack type. For example, gather all the attack phase cases related to data theft together to form a cluster of the data theft attack type.
[0143] For this clustering, the retrieval index structure adopts a multi-dimensional hash map. Taking the previously mentioned attack on enterprise data theft as an example, the dimensions of the multi-dimensional hash map include the attack tool hash, the exploit hash, and the data encryption method hash. Suppose an attacker uses a specific hacking tool during the attack, and this tool is hashed to obtain a unique hash value, which serves as the attack tool hash. If the attacker exploits a specific vulnerability (such as the previously mentioned CVE - 2021 - 1234 vulnerability), the exploit method for this vulnerability is hashed to obtain the exploit hash. When the attacker encrypts the stolen data (such as using the AES encryption algorithm), the encryption method is hashed to obtain the data encryption method hash. A multi-dimensional hash map is constructed through these hash values to form the retrieval index structure of the historical anomaly case library, so that when querying and matching similar attack cases subsequently, relevant cases can be quickly and accurately located.
[0144] In a possible implementation manner, the determining step of the minimum response level includes:
[0145] Step S310, define a mapping relationship table between threat levels and response levels, where the device hijacking scenario corresponds to a first-level response, and the spoofing attack scenario corresponds to a second-level response.
[0146] In this embodiment, in the enterprise's network security policy, the mapping relationship between different threat levels and response levels is clearly defined. For the device hijacking scenario, this is a very serious threat because the attacker may completely control important enterprise devices, such as core servers or key office devices. In this case, it is defined as a first-level response, which means that the most strict and rapid protection measures need to be taken. For example, the first-level response may include immediately cutting off the network connection of the hijacked device, conducting a comprehensive security check and recovery operation on the device, and at the same time deeply investigating other devices that have communicated with the hijacked device.
[0147] For the spoofing attack scenario, although it is also a threat, its harm level is slightly lower compared to device hijacking. In this scenario, the attacker may disguise as a legitimate user to obtain enterprise sensitive information. The corresponding is the second-level response, such as taking measures to strengthen user authentication, monitoring and auditing suspicious spoofing behaviors, etc.
[0148] Step S320, monitor the resource occupancy baseline of the available policies in the current protection policy library. If the resource requirements of the available policies are greater than the upper limit of the system idle resources, then raise the minimum response level to the next level.
[0149] In the enterprise's network environment, the resource occupancy of available policies in the real-time monitoring protection policy library is monitored. Protection policies will occupy a certain amount of CPU, memory, and network resources during execution. For example, an enterprise deploys an intrusion prevention system (IPS) as a protection policy. When the IPS is running normally, the average CPU occupancy rate is 5%, the memory consumption is 300MB, and the network latency growth factor is 0.1 (indicating the proportion of increased network latency due to the presence of the IPS under normal network traffic), which constitutes the resource occupancy baseline.
[0150] If the upper limit of the system idle resources in the enterprise network is a CPU idle rate of 30%, a memory idle amount of 1GB, and a network latency growth factor not exceeding 0.2. When a certain threat situation (such as a suspected spoofing attack) occurs and a protection policy needs to be activated, if it is found that the resource requirements of the available protection policy (such as an advanced user behavior analysis protection policy) are a CPU occupancy rate of 20% (greater than the CPU idle rate of 30% of the system idle resource upper limit), a memory consumption of 800MB (less than the memory idle amount of 1GB), and a network latency growth factor of 0.3 (greater than the network latency growth factor of 0.2), then the original secondary response level corresponding to the spoofing attack scenario will be upgraded to the primary response level to ensure that threats can be more effectively addressed under limited resources.
[0151] Step S330, when multiple abnormal types are detected concurrently, select the response level corresponding to the abnormal type with the highest threat level as the global minimum response level.
[0152] Among them, the resource occupancy baseline includes the CPU occupancy rate, memory consumption, and network latency growth factor, which are calculated and updated in real time through a sliding window algorithm.
[0153] Specifically, during the network monitoring process of the enterprise, multiple abnormal types may be detected simultaneously. For example, at a certain moment, signs of both spoofing attacks (such as abnormal login behaviors of some users, suspected spoofing) and resource abnormal occupancy of devices are detected, which may be a prelude to device hijacking. Since the threat level of the device hijacking scenario is higher than that of the spoofing attack scenario, according to the rules, select the primary response corresponding to the device hijacking scenario as the global minimum response level. This means that in this concurrent abnormal situation, the entire network security protection system will start the strictest protection measures according to the requirements of the primary response, and comprehensively protect and investigate all devices and network areas that may be threatened.
[0154] In a possible implementation manner, step S152 includes:
[0155] Step S1521, analyze the IP address segment block list and port disable list in the network connection control instruction.
[0156] Specifically, when a security threat is detected, a network connection control instruction can be generated. For example, when it is suspected that there is an external malicious attack source attempting to invade the enterprise network, the IP address segment block list in the network connection control instruction may include an external suspicious IP address segment (such as an IP address segment from a certain specific country or region, which has been considered related to malicious activities in previous security analyses). The port disable list may include some commonly attacked and exploited ports, such as port 8080 (assuming this port is found to have security vulnerabilities and may be exploited by attackers for malicious access). By analyzing these IP address segment block lists and port disable lists, the network connection range to be restricted by the network connection control instruction is clarified.
[0157] Step S1522, compare the block list with the current active connection list. If there is an overlapping address segment between the block list and the current active connection list, generate an address conflict report.
[0158] In this embodiment, the enterprise's network management system maintains the current active connection list, which records the IP addresses and port information of the devices that are currently communicating in the enterprise network. Compare the IP address segment block list and port disable list in the network connection control instruction with the current active connection list. For example, if there is an IP address segment 192.168.1.10 - 192.168.1.20 in the block list, and it is found that the IP address of the device used by the enterprise's marketing department for an important project promotion is 192.168.1.15 in the current active connection list, there is an overlapping address segment. At this time, the system will generate an address conflict report, which details the overlapping IP address segment and the relevant device information.
[0159] Step S1523, according to the overlapping address segment in the address conflict report, query the corresponding device fingerprint features to determine the department to which the device belongs and the business importance level.
[0160] Specifically, according to the overlapping address segment (such as 192.168.1.15) in the address conflict report, query the device fingerprint features. The device fingerprint features include information such as the device's MAC address, operating system type, installed software, etc. By querying these device fingerprint features, it can be determined that the department to which the device belongs is the marketing department. Then, according to the business importance assessment criteria preset by the enterprise, determine the business importance level of the device. For example, since the marketing department is conducting an important project promotion, the business importance level of this device is determined to be a high level.
[0161] Step S1524, re - divide the blocked address segments based on the business importance level so that the connections of core business devices are not affected by the network connection control instruction.
[0162] Specifically, based on the high business importance level of the marketing department devices, re - divide the blocked address segments. For example, modify the original IP address segment block list 192.168.1.10 - 192.168.1.20 to 192.168.1.10 - 192.168.1.14 and 192.168.1.16 - 192.168.1.20. In this way, the IP address 192.168.1.15 of the marketing department devices is excluded, so that the connections of core business devices are not affected by the network connection control instruction, ensuring the normal progress of important project promotion work in the marketing department.
[0163] Step S1525, merge the re - divided blocked address segments into the original network connection control instruction to generate a final control instruction that passes the compatibility verification.
[0164] Specifically, merge the re - divided blocked address segments (192.168.1.10 - 192.168.1.14 and 192.168.1.16 - 192.168.1.20) into the original network connection control instruction and replace the original IP address segment block list part. At the same time, keep the port disable list unchanged to generate a final control instruction that passes the compatibility verification. This final control instruction can effectively control network security (block suspicious IP address segments and disable dangerous ports) without affecting the normal network connections of enterprise core business devices.
[0165] In a possible implementation manner, the specific steps of the event slicing process include:
[0166] Step S410, use the time point when the attacker first breaks through the network boundary as the starting slicing point.
[0167] In this embodiment, taking a network attack on the data of an enterprise's R & D department as an example, the attacker uses an unpatched operating system vulnerability (assumed to be a specific vulnerability number CVE - 2023 - 5678) on a test server in the R & D department and successfully invades at 3:00 am on October 15, 2023. This time point is the starting slicing point.
[0168] Step S420, use the time point when the attacker obtains key data or triggers the protection mechanism as the ending slicing point.
[0169] In this embodiment, the termination slicing point is the time point when the attacker obtains critical data (such as the source code file of a core project being carried out by the R & D department) or triggers the protection mechanism. Suppose the attacker successfully obtained the source code file at 8:00 am on October 15, 2023 and began to transmit it out of the enterprise network. At this time, abnormal traffic was detected by the intrusion detection system within the enterprise, triggering the protection mechanism. This 8:00 am is the termination slicing point.
[0170] Step S430: Between the start slicing point and the termination slicing point, perform sub - slicing according to a fixed time interval or critical operation events.
[0171] For example, between the start slicing point (3:00 am on October 15, 2023) and the termination slicing point (8:00 am on October 15, 2023), perform sub - slicing according to a fixed time interval or critical operation events.
[0172] If divided according to a fixed time interval, assume that each 30 - minute interval is used for sub - slicing. During this time period, a total of 10 sub - slices are divided from 3:00 am to 8:00 am. At the same time, if there are critical operation events during this process, sub - slicing will also be performed separately. For example, at 4:30 am, the attacker successfully elevated their privileges on the invaded server, which is a critical operation event, so sub - slicing will also be performed at this time point.
[0173] Step S440: Extract the following features for each sub - slice: the vulnerability number used, the protocol type adopted for lateral movement, and the hash value of the injected malicious payload.
[0174] For example, in one of the sub - slices (from 4:00 am to 4:30 am), the vulnerability number used by the attacker is CVE - 2023 - 5678 (which may be the same as the vulnerability of the initial intrusion or a new vulnerability discovered by exploiting this initial vulnerability), the protocol type adopted for lateral movement is the SMB protocol (the attacker uses the SMB protocol shared within the enterprise network to move from the invaded test server to other R & D devices), and the hash value of the injected malicious payload is 123456789abcdef (assuming this hash value is obtained by hashing the malicious program injected into the target device).
[0175] Step S450: Perform temporal correlation analysis on the sub - slice features and the full - life - cycle features of the full - life - cycle data to establish a causal relationship chain for the attack - phase cases.
[0176] For example, step S450 includes:
[0177] Step S451: Extract the vulnerability number, protocol type, and malicious payload hash value from the sub-slice features, and obtain the timestamp interval of the start slice point and the end slice point corresponding to each sub-slice.
[0178] For example, for the sub-slice from 4:00 am to 4:30 am mentioned above, the vulnerability number is CVE - 2023 - 5678, the protocol type is SMB protocol, the malicious payload hash value is 123456789abcdef, the timestamp of the start slice point is 2023 - 10 - 15 04:00:00, and the timestamp of the end slice point is 2023 - 10 - 15 04:30:00.
[0179] Step S452: Extract the timestamp start point of the attack initial stage feature, the timestamp sequence of the lateral movement path feature, and the timestamp end point of the data leakage node feature from the full life cycle features of the full life cycle data, and generate a full life cycle timeline.
[0180] Specifically, extract the timestamp start point of the attack initial stage feature (i.e., 3:00 am on October 15, 2023), the timestamp sequence of the lateral movement path feature (such as the recorded time points when the attacker moves between different devices, such as a series of time points starting from 4:00 am and using the SMB protocol to move to other devices), and the timestamp end point of the data leakage node feature (8:00 am on October 15, 2023) from the full life cycle features, and generate a full life cycle timeline.
[0181] Step S453: Map the timestamp interval of the sub-slice to the full life cycle timeline, identify the time overlap area and the interval area between adjacent sub-slices, and generate an event sequence pattern.
[0182] For example, there may be an overlap between the end slice point of a certain sub-slice (such as 4:30 am) and the start slice point of the next sub-slice (such as 4:30 am), which indicates that the attacker's operations are continuous; and if there is an interval area (such as no attacker activity is detected within 10 minutes after a certain sub-slice ends, and then there is activity at the start of the next sub-slice), this is also part of the event sequence pattern.
[0183] Step S454: According to the continuity of the vulnerability numbers in the event sequence pattern, match the vulnerability exploitation paths in the lateral movement path features, and generate a vulnerability exploitation path matching result.
[0184] For example, if in a series of sub - slices, the vulnerability numbers show a continuity from the initial intrusion vulnerability to the gradual exploitation of other related vulnerabilities for privilege escalation or lateral movement, then the corresponding vulnerability exploitation path can be found in the lateral movement path feature. For example, after the attacker intrudes by exploiting the CVE - 2023 - 5678 vulnerability, the attacker gradually expands the control range in the enterprise network through a series of related vulnerabilities, and the exploitation order of these vulnerabilities in the lateral movement path feature matches the continuity of the vulnerability numbers in the sub - slices.
[0185] Step S455: Based on the conversion frequency of the protocol type within the time overlap region, associate the protocol switching nodes in the lateral movement path feature to generate protocol switching association marks.
[0186] For example, within the time overlap region, if it is found that the protocol type frequently switches from the SMB protocol to other protocols (such as the HTTP protocol for communicating with external servers), then find the corresponding protocol switching node in the lateral movement path feature (such as the node where the attacker switches from an internal network share to interacting with an external server), and establish a protocol switching association mark.
[0187] Step S456: Map the number of repeated occurrences of the malicious payload hash value within the interval region to the payload injection records in the data leakage node feature to generate a payload propagation path.
[0188] For example, if the malicious payload hash value repeatedly appears multiple times within the interval region, this may indicate that the attacker is spreading the malicious payload between different devices. Search for the corresponding payload injection records in the data leakage node feature (such as on which devices the attacker injected the malicious payload to obtain data), so as to construct a payload propagation path.
[0189] Step S457: Superimpose the vulnerability exploitation path matching result, the protocol switching association mark, and the payload propagation path in timestamp order to generate an initial causal relationship chain.
[0190] For example, in chronological order, first the vulnerability exploitation path matching result shows how the attacker gradually exploits vulnerabilities to intrude and escalate privileges, then the protocol switching association mark indicates the protocol conversion of the attacker in different network operations, and finally the payload propagation path shows the propagation of the malicious payload and the data acquisition process.
[0191] Step S458: Detect whether there is a logical break point in the initial causal relationship chain. If there is a logical break point, supplement the missing vulnerability numbers or protocol types according to the attack tool usage records in the full - life - cycle feature to generate a corrected causal relationship chain.
[0192] For example, identify that the jump interval of vulnerability numbers between adjacent sub - slices in the initial causal relationship chain exceeds a preset vulnerability association threshold (assuming the preset threshold is that the maximum interval between adjacent vulnerability numbers in a common attack chain is 2. If it is found that the jump of vulnerability numbers between adjacent sub - slices exceeds this threshold, such as directly jumping from CVE - 2023 - 5678 to a vulnerability that has no relation with the previous one and has a large number gap), the protocol type conversion is not recorded in the lateral movement path feature (such as in the initial causal relationship chain, it shows that the protocol is converted from SMB to a custom protocol that has never appeared in the lateral movement path feature), or the node position where the malicious payload hash value does not appear continuously in the payload propagation path (if there is a missing value in the middle of the malicious payload hash values that should appear continuously in the payload propagation path). If there is a logical breakpoint, supplement the missing vulnerability number or protocol type according to the attack tool usage record in the full - life - cycle feature. For example, if it is found that there is a jump in vulnerability numbers and the attack tool usage record shows that the attacker used a specific vulnerability exploitation tool, supplement the possible vulnerability numbers according to the characteristics of this tool to generate a corrected causal relationship chain.
[0193] Step S459, verify the coherence of the payload propagation path according to the time sequence of the vulnerability exploitation path matching result and the protocol switching association mark in the corrected causal relationship chain. If there is a time - sequence conflict, re - adjust the mapping relationship of the protocol switching association mark, and compare the integrity of the verified corrected causal relationship chain with the attack initial - stage feature, lateral movement path feature, and data - leakage node feature, delete duplicate associated nodes, and merge the operation tracks of the same attack tool to generate a final causal relationship chain.
[0194] Among them, the detection process of the logical breakpoint includes: identifying that the jump interval of vulnerability numbers between adjacent sub - slices in the initial causal relationship chain exceeds a preset vulnerability association threshold, the protocol type conversion is not recorded in the lateral movement path feature, or the node position where the malicious payload hash value does not appear continuously in the payload propagation path.
[0195] If in the corrected causal relationship chain, the vulnerability exploitation path matching result shows that the attacker first exploited a vulnerability on a certain device to obtain a certain privilege, and then the protocol switching association mark indicates that the attacker switched the protocol to communicate with the outside. According to this time sequence, the payload propagation path should start from the device where the privilege was obtained to propagate the malicious payload. If there is a time - sequence conflict (such as the payload propagation path shows that the malicious payload first appears on a device without obtaining the privilege), then re - adjust the mapping relationship of the protocol switching association mark to make the entire corrected causal relationship chain reasonable in terms of time sequence.
[0196] Further, if there are exploit nodes in the corrected causal relationship chain that duplicate the characteristics of the initial stage of the attack, only one is retained; if there are multiple operation traces of the same attack tool (such as using the same hacker script) in the lateral movement path characteristics, they are merged into one continuous operation trace, and finally a complete, concise, and logically coherent final causal relationship chain is obtained. This final causal relationship chain can accurately describe the development process of the entire network attack event and provide an important basis for subsequent security analysis, protection strategy formulation, etc.
[0197] Figure 2 FIG. shows the hardware structure diagram of the remote working protection system 100 provided by the embodiments of the present application. As Figure 2 shown, the remote working protection system 100 may include a processor 110, a machine-readable storage medium 120, a bus 130, and a communication unit 140.
[0198] In a possible design, the remote working protection system 100 may be a single server or a server group. The server group may be centralized or distributed (for example, the remote working protection system 100 may be a distributed system). In some embodiments, the remote working protection system 100 may be local or remote. For example, the remote working protection system 100 may access information and / or data stored in the machine-readable storage medium 120 via a network. Again, for example, the remote working protection system 100 may be directly connected to the machine-readable storage medium 120 to access the stored information and / or data. In some embodiments, the remote working protection system 100 may be implemented on a server. By way of example only, the server may include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an internal cloud, a multi-layer cloud, etc. or any combination thereof.
[0199] The machine-readable storage medium 120 may store data and / or instructions. In some embodiments, the machine-readable storage medium 120 may store data obtained from an external terminal. In some embodiments, the machine-readable storage medium 120 may store the data and / or instructions used by the remote working protection system 100 to execute or use to complete the exemplary methods described in the present application.
[0200] In a specific implementation process, one or more processors 110 execute the computer-executable instructions stored in the machine-readable storage medium 120, so that the processors 110 can execute the remote working network security protection method based on big data in the above method embodiments. The processors 110, the machine-readable storage medium 120, and the communication unit 140 are connected through the bus 130, and the processors 110 may be used to control the transceiver actions of the communication unit 140.
[0201] For the specific implementation process of the processor 110, reference can be made to the various method embodiments executed by the above-mentioned remote work protection system 100. Their implementation principles and technical effects are similar, and will not be elaborated here in this embodiment.
[0202] In addition, an embodiment of the present application also provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, the above-mentioned remote work network security protection method based on big data is implemented.
[0203] It should be noted that, in order to simplify the description of the present application disclosure and thus help the understanding of one or more invention embodiments, in the previous description of the embodiments of the present application, sometimes multiple features are merged into one embodiment, drawing or description thereof.
Claims
1. A remote office network security protection method based on big data, characterized in that, The method includes: Obtaining network activity data generated by multiple terminal devices in a telecommuting scenario, where the network activity data includes traffic data, device status data, and user behavior data; Performing multi-dimensional security feature analysis on the network activity data to generate a dynamic security feature set corresponding to the network activity data, where the dynamic security feature set includes traffic features, device fingerprint features, and behavior pattern features; Based on the dynamic security feature set, using a multi-layer anomaly detection model to identify abnormal behaviors in the network activity data, and outputting an anomaly detection result including anomaly types and anomaly confidence levels; According to the anomaly detection result, matching a target protection policy corresponding to the anomaly type from a preset protection policy library, and dynamically adjusting the weight of the target protection policy based on the anomaly confidence level to generate a real-time protection policy instruction; Executing the real-time protection policy instruction to synchronously regulate the network connection, device permissions, and user access paths in the telecommuting scenario, and feeding back the regulated network status data to the multi-layer anomaly detection model to optimize the anomaly recognition logic; Among them, the multi-layer anomaly detection model includes a primary detection layer, an intermediate aggregation layer, and a high-level decision layer; The step of using a multi-layer anomaly detection model to identify abnormal behaviors in the network activity data based on the dynamic security feature set and outputting an anomaly detection result including anomaly types and anomaly confidence levels includes: Inputting the dynamic security feature set into the primary detection layer, independently performing anomaly scoring on the traffic features, device fingerprint features, and behavior pattern features respectively to obtain a primary anomaly scoring set; Inputting the primary anomaly scoring set into the intermediate aggregation layer, calculating the combined weight of multiple primary anomaly scores according to a preset threat scenario template to generate an intermediate anomaly aggregation index; Inputting the intermediate anomaly aggregation index into the high-level decision layer, combining the pattern matching results in the historical anomaly case library and real-time network environment parameters, calculating the anomaly confidence level and determining the anomaly type; Among them, the threat scenario template defines the following combination logic: when the anomaly score of the traffic feature is higher than the first set score and the anomaly score of the behavior pattern feature is lower than the second set score, it is marked as a camouflage attack scenario; when the anomaly score of the device fingerprint feature continuously increases and is negatively correlated with the anomaly score of the traffic feature, it is marked as a device hijacking scenario.
2. The method for remotely working network security protection based on big data according to claim 1, characterized in that The step of performing multi-dimensional security feature analysis on the network activity data to generate a dynamic security feature set corresponding to the network activity data includes: Extracting protocol type distribution features, packet transmission frequency features, and target address clustering features from the traffic data to form the traffic features; Analyzing device hardware identification features, operating system vulnerability features, and process resource occupancy features from the device status data to form the device fingerprint features; Identifying login time series features, file operation track features, and permission change history features from the user behavior data to form the behavior pattern features; Align the traffic feature, device fingerprint feature, and behavior pattern feature according to a time window, and perform cross-dimensional correlation analysis on the aligned features to generate the dynamic security feature set; Among them, the cross-dimensional correlation analysis includes: identifying abnormal data transmission paths based on the target address clustering feature in the traffic feature and the file operation trace feature in the behavior pattern feature; detecting potential threats of mismatched resource consumption and traffic based on the process resource occupancy feature in the device fingerprint feature and the packet transmission frequency feature in the traffic feature.
3. The method for remotely working network security protection based on big data according to claim 1, wherein, The matching of the target protection policy corresponding to the abnormal type from the preset protection policy library includes: According to the threat level corresponding to the abnormal type, screen a candidate policy set that meets the minimum response level from the protection policy library; Based on the abnormal confidence, sort the policies in the candidate policy set by priority, and select the policy with the highest priority as the basic protection policy; According to the bandwidth load rate, device online status, and user role permissions in the real-time network environment parameters, dynamically adapt the execution parameters of the basic protection policy to generate the target protection policy; Among them, the dynamic adaptation includes: when the bandwidth load rate is greater than the preset critical value, reducing the traffic monitoring frequency and increasing the device fingerprint verification intensity; when a user with a set permission is detected online, restricting external device access while retaining the critical data access channel.
4. The method for remotely working network security protection based on big data according to claim 3, characterized in that, The execution of the real-time protection policy instruction includes: Decompose the real-time protection policy instruction into a network connection control instruction, a device permission reset instruction, and a user access interception instruction; Verify the compatibility of the network connection control instruction with the current network topology. If there is a conflict between the network connection control instruction and the current network topology, reallocate the IP address pool according to the device fingerprint feature; according to the device permission reset instruction, batch modify the file read / write permissions and peripheral interface enable status of the target device, and record the permission difference log before and after the modification; Based on the user access interception instruction, inject a virtual access path into the user behavior data to induce the attacker to trigger the protection mechanism, and at the same time back up an encrypted copy of the real access path; Use the permission difference log and the encrypted copy as the regulated network state data, and input them into the multi-layer anomaly detection model to update the combined weight calculation rule in the threat scenario template.
5. The method for remotely working network security protection based on big data according to claim 2, wherein The extraction of the protocol type distribution feature from the traffic data includes: Capture the nested relationship between the transport layer protocol and the application layer protocol in the traffic data, and identify unconventional protocol combination patterns; Count the proportion of the number of packets of different protocols under the same source address to generate a protocol distribution vector; Compare the similarity of the protocol distribution vector with the known malicious protocol library. If the similarity is greater than the dynamically set score, mark the corresponding protocol as a risk protocol; Among them, the dynamically set score is automatically adjusted according to the proportion of normal service traffic in the current network environment: when the proportion of normal service traffic decreases, reduce the dynamically set score to expand the risk protocol identification range.
6. The method for remote office network security protection based on big data according to claim 3, characterized in that The construction process of the historical anomaly case library includes: Collect the full - life - cycle data of historical network attack events, extract the characteristics of the initial attack stage, the lateral movement path characteristics, and the data leakage node characteristics as the full - life - cycle characteristics; Perform event slicing on the full - life - cycle data to generate multiple independent attack - stage cases; Label each attack - stage case with a corresponding protection - strategy effectiveness label, where the protection - strategy effectiveness label includes the strategy response speed, the resource consumption ratio, and the false - interception rate; Cluster the labeled attack - stage cases according to the attack type to form the retrieval index structure of the historical abnormal case library; Among them, the retrieval index structure adopts multi - dimensional hash mapping, and the dimensions of the multi - dimensional hash mapping include the attack - tool hash, the vulnerability - exploitation hash, and the data - encryption - method hash.
7. The method for remote office network security protection based on big data according to claim 3, characterized in that, The determination steps of the minimum response level include: Define a mapping relationship table between threat levels and response levels, where the device - hijacking scenario corresponds to a first - level response, and the spoofing - attack scenario corresponds to a second - level response; Monitor the resource - occupancy baseline of the available strategies in the current protection - strategy library. If the resource requirements of the available strategies are greater than the upper limit of the system idle resources, then raise the minimum response level to the next level; When multiple abnormal types are detected concurrently, select the response level corresponding to the abnormal type with the highest threat level as the global minimum response level; Among them, the resource - occupancy baseline includes the CPU occupancy rate, the memory consumption, and the network - latency growth coefficient, which are calculated and updated in real time through a sliding - window algorithm.
8. The method for remotely working network security protection based on big data according to claim 4, wherein The verification of the compatibility between the network - connection control instruction and the current network topology structure includes: Parse the IP - address - segment block list and the port - disable list in the network - connection control instruction; Compare the block list with the current active - connection list. If there is an overlapping address segment between the block list and the current active - connection list, then generate an address - conflict report; According to the overlapping address segment in the address - conflict report, query the corresponding device - fingerprint characteristics to determine the department to which the device belongs and the business - importance level; Based on the business - importance level, re - divide the blocked address segment so that the connections of core - business devices are not affected by the network - connection control instruction; Merge the re - divided blocked address segment into the original network - connection control instruction to generate the final control instruction that passes the compatibility verification.
9. The method for remotely working network security protection based on big data according to claim 6, characterized in that The specific steps of the event slicing process include: Take the time point when the attacker first breaks through the network boundary as the starting slicing point; Take the time point when the attacker obtains key data or triggers the protection mechanism as the ending slicing point; Between the starting slicing point and the ending slicing point, perform sub - slicing division at fixed time intervals or key operation events; Extract the following characteristics for each sub - slice: the vulnerability number used, the protocol type adopted for lateral movement, and the hash value of the injected malicious payload; Perform time - series correlation analysis on the sub - slice characteristics and the full - life - cycle characteristics to establish a causal - relationship chain of attack - stage cases; The performing time - series correlation analysis on the sub - slice characteristics and the full - life - cycle characteristics to establish a causal - relationship chain of attack - stage cases includes: Extract the vulnerability number, protocol type, and malicious payload hash value from the sub-slice features, and obtain the timestamp interval of the start slice point and the end slice point corresponding to each sub-slice; Extract the timestamp start point of the attack initial stage feature, the timestamp sequence of the lateral movement path feature, and the timestamp end point of the data leakage node feature from the full life cycle features to generate a full life cycle timeline; Map the timestamp interval of the sub-slice to the full life cycle timeline, identify the time overlap area and the interval area between adjacent sub-slices, and generate an event sequence pattern; Match the vulnerability exploitation path in the lateral movement path feature according to the continuity of the vulnerability number in the event sequence pattern to generate a vulnerability exploitation path matching result; Associate the protocol switching nodes in the lateral movement path feature based on the conversion frequency of the protocol type in the time overlap area to generate a protocol switching association mark; Map the number of repeated occurrences of the malicious payload hash value in the interval area to the payload injection record in the data leakage node feature to generate a payload propagation path; Overlay the vulnerability exploitation path matching result, the protocol switching association mark, and the payload propagation path in timestamp order to generate an initial causal relationship chain; Detect whether there is a logical break point in the initial causal relationship chain. If there is a logical break point, supplement the missing vulnerability number or protocol type according to the attack tool usage record in the full life cycle feature to generate a corrected causal relationship chain; Verify the coherence of the payload propagation path according to the time order of the vulnerability exploitation path matching result and the protocol switching association mark in the corrected causal relationship chain. If there is a time order conflict, readjust the mapping relationship of the protocol switching association mark; Compare the integrity of the verified corrected causal relationship chain with the attack initial stage feature, the lateral movement path feature, and the data leakage node feature, delete duplicate associated nodes, and merge the operation traces of the same attack tool to generate a final causal relationship chain; Among them, the detection process of the logical break point includes: identifying the node positions where the vulnerability number jump interval between adjacent sub-slices in the initial causal relationship chain exceeds the preset vulnerability association threshold, the protocol type conversion is not recorded in the lateral movement path feature, or the malicious payload hash value does not appear continuously in the payload propagation path.
10. A remote working protection system, characterized in that, The remote work protection system includes a processor and a memory. The memory is connected to the processor. The memory is used to store programs, instructions, or codes. The processor is used to execute the programs, instructions, or codes in the memory to implement the big data-based remote work network security protection method according to any one of claims 1-9 above.
Citation Information
Patent Citations
Network security detection method and system
CN118101250A