A network security threat identification method and system based on multivariate event analysis

By generating a scheduling representation array and using network security threat identification networks to identify network security threats, the problem of difficult to fully consider the complex factors of network events in the existing technology is solved, and more accurate identification of network security threats and timely prevention is achieved.

CN119696885BActive Publication Date: 2025-09-02CHONGQING XIXI EVIDENCE SCIENCE RESEARCH INSTITUTE GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411843439.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-14
Publication Date
2025-09-02
Estimated Expiration
2044-12-14

AI Technical Summary

Technical Problem

Existing network security threat identification technologies are difficult to fully consider the various complex factors of network events and their interrelationships, resulting in high false alarm rates and high false alarm rates when processing high-dimensional and large-scale network traffic data, and it is difficult to accurately identify potential network security threats.

Method used

By obtaining multiple multivariate event monitoring data sets in the set monitoring cycle before and after the target multivariate event monitoring data set, combining the set monitoring cycle, the set monitoring time range and the preset array representation channel to generate a to-analytical representation array, and using the network security threat identification network for identification. The network security threat identification network is initialized for debugging and construction based on multiple preset training data sets.

Benefits of technology

It improves the accuracy and reliability of network security threat identification, can reflect network security trends more comprehensively, timely detect potential threats, and protect the security and stability of network systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119696885B_ABST
    Figure CN119696885B_ABST
Patent Text Reader

Abstract

The present application provides a network security threat identification method and system based on multivariate event analysis. Based on multiple second multivariate event monitoring data sets in a set monitoring period before and after multiple preset training data sets and a network security threat priori mark of whether each set training data set contains a network security threat, a neural network is debugged and initialized to obtain a network security threat identification network. Multiple first multivariate event monitoring data sets are obtained to mine and obtain network security related features and information. Characterization arrays are generated for multiple first multivariate event monitoring data sets based on a set monitoring period, a set monitoring time range, and a preset array representation channel to obtain a quasi-analysis characterization array of the network security trend before and after the characterization. The quasi-analysis characterization array is input into the network security threat identification network for network security threat identification to obtain a more reliable network security threat identification result of whether the target multivariate event monitoring data set contains a network security threat.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and more specifically, to a network security threat identification method and system based on multivariate event analysis. Background Art

[0002] With the rapid development of information technology, networks have penetrated every aspect of society, from personal communications to business operations and national security infrastructure. However, this has also brought with it cybersecurity challenges. Cyberattacks are becoming increasingly sophisticated and diverse, including malware distribution, phishing, distributed denial of service (DDoS) attacks, and various vulnerability-based intrusions. These cybersecurity threats not only lead to personal privacy breaches and financial losses for businesses, but can also pose serious threats to national security.

[0003] Accurately identifying network threats is a crucial step in addressing them. Traditional network security detection methods often analyze a single type of network event or a limited set of information sources. For example, some early firewalls primarily performed simple filtering based on pre-defined rules, such as source IP addresses, destination IP addresses, and port numbers in network traffic. This approach can only identify known, characteristic malicious traffic patterns and is often ineffective against new or well-disguised network attacks.

[0004] While intrusion detection systems (IDS) and intrusion prevention systems (IPS) have improved network security detection capabilities to some extent, enabling them to analyze network behavior patterns and detect anomalies, they still have limitations. They often struggle to fully consider the multiple, complex factors and their interrelationships within network activity, and when processing large amounts of network data, they can experience high rates of false positives or omissions. This is because these systems fail to fully exploit the temporal correlations between network events and the inherent connections between different types of network events when identifying threats.

[0005] Furthermore, with the continuous expansion of networks and the increasing complexity of network activities, network traffic data has become highly dimensional and large-scale. Relying solely on simple feature extraction and analysis methods makes it difficult to extract characteristic information truly relevant to network security threats from this massive amount of data. Furthermore, in real-world network environments, network security threats often do not exist in isolation but are closely related to factors such as the network's historical activity and recent changes in network behavior. For example, a seemingly normal network connection may be revealed as a potential network security threat if analyzed in conjunction with factors such as previous and subsequent network connection patterns and changes in relationships between related entities.

[0006] Existing network security threat identification technologies lack an effective method to integrate various network information when dealing with these complex network security situations, and it is difficult to accurately characterize network security trends, resulting in unreliable results of network security threat identification. Summary of the Invention

[0007] The purpose of the present invention is to provide a network security threat identification method and system based on multivariate event analysis. This application is implemented as follows:

[0008] In the first aspect, the present application provides a network security threat identification method based on multivariate event analysis, the method comprising: obtaining multiple first multivariate event monitoring data sets in a set monitoring period before and after a target multivariate event monitoring data set; generating a representation array for the multiple first multivariate event monitoring data sets in combination with the set monitoring period, the set monitoring time range and the preset array representation channel to obtain a quasi-analysis representation array of the target multivariate event monitoring data set, wherein the set monitoring time range is less than the set monitoring period; performing network security threat identification on the quasi-analysis representation array based on a network security threat identification network to obtain a network security threat identification result of whether there is a network security threat in the target multivariate event monitoring data set; the network security threat identification network is obtained by debugging an initialized neural network in combination with multiple second multivariate event monitoring data sets in the set monitoring period before and after multiple preset training data sets and a network security threat priori marker of whether there is a network security threat in each set training data set.

[0009] In a second aspect, the present application provides a computer system comprising: one or more processors; a memory; and one or more computer programs; wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processors, the method described above is implemented.

[0010] The beneficial effects of the present application are as follows: the present application obtains a plurality of first multivariate event monitoring data sets in a set monitoring period before and after a target multivariate event monitoring data set; the plurality of first multivariate event monitoring data sets represent the autocorrelated information in the monitoring period before and after the target multivariate event monitoring data set, so as to mine network security-related features and information, and then, based on the set monitoring period, a set monitoring time range less than the set monitoring period and a preset array representation channel, a characterization array is generated for the plurality of first multivariate event monitoring data sets to obtain a quasi-analysis characterization array of the target multivariate event monitoring data set; the quasi-analysis characterization array represents the network security trend before and after the target multivariate event monitoring data set, thereby enhancing the characterization effect, and the quasi-analysis characterization array is input into a network security threat identification network for network security threat identification, thereby obtaining a network security threat identification result of whether the target multivariate event monitoring data set contains a network security threat. The network security threat identification network is obtained by debugging and initializing a neural network based on a plurality of second multivariate event monitoring data sets in a set monitoring period before and after a plurality of preset training data sets and a network security threat priori marker of whether each set training data set contains a network security threat. The present application generates a representation vector for a target multivariate event monitoring data set, and based on the network security-related features and information before and after the target multivariate event monitoring data set, enables the intended analysis representation array to fuse the traffic sequence information and timing information, so as to make the network security threat identification results of the target multivariate event monitoring data set more reliable under the premise of generating a network security threat identification network of the representation array of fused information. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 This is a flowchart of a network security threat identification method based on multi-event analysis provided in an embodiment of the present application.

[0012] Figure 2 This is a schematic diagram of the composition of a computer system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0013] The execution subject of the network security threat identification method based on multivariate event analysis in the embodiment of the present application is a computer system, including but not limited to a server, a personal computer, a laptop, a tablet computer, a smart phone, etc. The embodiment of the present application provides a network security threat identification method based on multivariate event analysis, which is applied to a computer system, such as Figure 1 As shown, the method includes:

[0014] Step S100: Acquire a plurality of first multivariate event monitoring data sets in a set monitoring period before and after a target multivariate event monitoring data set.

[0015] In step S100 of the embodiment of the present application, the computer system obtains a plurality of first multivariate event monitoring data sets in a set monitoring period before and after the target multivariate event monitoring data set.

[0016] A target multivariate event monitoring dataset is a collection of data generated by a computer system monitoring a series of network-related events in a specific network environment. For example, in an enterprise network environment, a target multivariate event monitoring dataset might include network activity information from various devices on the network (such as servers and employee terminals), such as HTTP requests, file transfers, and logins.

[0017] The set monitoring period is a predefined time range used to determine the time span for data acquisition. The computer system acquires relevant multivariate event monitoring data within this time range. Assuming the set monitoring period is one week, the computer system needs to acquire multivariate event monitoring data from the week before and the week after the target multivariate event monitoring data set. This data constitutes multiple first multivariate event monitoring data sets.

[0018] To obtain this data, computer systems can use network monitoring tools or logging systems. Network monitoring tools capture various event information from network traffic in real time and store it in a specific format. Logging systems record the activity of various components in the system (such as server software and security devices). From these data sources, computer systems can extract data that meets the requirements of the set monitoring period.

[0019] When acquiring data, computer systems must ensure its integrity and accuracy. Data collected by network monitoring tools may require data cleansing to remove invalid data records, such as erroneous data packets caused by network failures. Data in logging systems may require format conversion for unified processing. For example, if different system components use different log formats (e.g., some are text-based, others binary), the computer system needs to convert them into a unified format that is easier to analyze.

[0020] By acquiring multiple first multivariate event monitoring datasets within a set monitoring period before and after the target multivariate event monitoring dataset, the computer system provides a rich data foundation for subsequent analysis. This data contains information about network activity during the time period surrounding the target multivariate event monitoring dataset, reflecting changing trends in network activity and potential cybersecurity-related information. This helps the computer system mine cybersecurity-related features and information in subsequent steps, leading to more accurate cybersecurity threat identification.

[0021] Step S200: generating representation arrays for a plurality of first multivariate event monitoring data sets in combination with a set monitoring period, a set monitoring time range, and a preset array representation channel to obtain a pseudo-analysis representation array of a target multivariate event monitoring data set, wherein the set monitoring time range is less than the set monitoring period.

[0022] In step S200 of the embodiment of the present application, the computer system generates representation arrays for multiple first multivariate event monitoring data sets in combination with the set monitoring period, the set monitoring time range and the preset array representation channel to obtain a pseudo-analysis representation array of the target multivariate event monitoring data set.

[0023] The set monitoring period is a predetermined length of time, such as one day, one week, or one month. During this period, the computer system collects multiple first multivariate event monitoring data sets. The set monitoring time range is a time range that is smaller than the set monitoring period. For example, if the monitoring period is one week, the set monitoring time range can be 12 hours. This means that within this one-week set monitoring period, the computer system will process data in units of 12 hours.

[0024] The preset array represents a channel used to analyze multi-dimensional event monitoring data along specific dimensions, such as the network activity description information channel, the access change data channel, and the entity relationship channel. For example, the network activity description information channel includes information such as the URL, User-Agent string, and request parameters in HTTP requests. The access change data channel includes data such as the same IP address attempting to connect to different ports within a short period of time, using different protocols, or rapidly switching from one geographic location to another. The entity relationship channel focuses on relational data, such as the communication patterns between a specific IP address and multiple other IP addresses.

[0025] When the computer system generates the characterization array, it processes the data according to these parameters. For each set monitoring time range within the set monitoring period, the computer system maps the first multivariate event monitoring data set to different dimensions of the preset array representation channel. Specifically, the computer system can adopt data mapping and encoding technology. For example, for the data in the network activity description information channel, a mapping algorithm can be used to map various types of network activity description information (such as various fields in HTTP requests) to specific positions in a multidimensional array according to predefined rules. Assume that there is a simple mapping algorithm formula: M(x)=(i, j), where x is a certain network activity description information, M is the mapping function, and i and j are indexes in the multidimensional array, indicating the position of the information in the characterization array.

[0026] In this process, the computer system also needs to normalize different types of data to ensure that the data from different channels are on the same order of magnitude. For example, for numerical data such as network traffic quantification data and the frequency of specific behavior, the normalization formula can be used: , where x is the original data, min and max are the minimum and maximum values ​​of the data within the set monitoring time range, and y is the normalized data.

[0027] In this way, the computer system generates a preliminary characterization array for each set monitoring time range. As the processing of each set monitoring time range within the set monitoring cycle is completed, the computer system further integrates these preliminary characterization arrays to obtain a pseudo-analysis characterization array for the target multi-element event monitoring dataset. This pseudo-analysis characterization array integrates information from different set monitoring time ranges and different preset array representation channels within the set monitoring cycle, and can more comprehensively and accurately describe the network security characteristics related to the target multi-element event monitoring dataset, providing a rich and effective data foundation for subsequent network security threat identification.

[0028] Step S300: Based on the network security threat identification network, network security threats are identified on the array to be analyzed, and a network security threat identification result of whether a target multivariate event monitoring data set contains network security threats is obtained; the network security threat identification network is obtained by debugging an initialized neural network in combination with multiple second multivariate event monitoring data sets in a set monitoring period before and after multiple preset training data sets and a network security threat priori mark of whether a network security threat exists in each set training data set.

[0029] In step S300 of the embodiment of the present application, the computer system performs network security threat identification on the target analysis characterization array based on the network security threat identification network, thereby obtaining a network security threat identification result of whether the target multivariate event monitoring data set contains a network security threat.

[0030] The cybersecurity threat identification network is a network model specifically designed to identify cybersecurity threats, derived through a series of training processes. This network model is constructed by a computer system debugging an initialized neural network by combining multiple second multivariate event monitoring datasets from predefined monitoring periods preceding and following multiple predefined training datasets, along with a priori cybersecurity threat markers indicating whether each predefined training dataset contains cybersecurity threats.

[0031] The representation array to be analyzed is generated in the previous step. It integrates relevant information about the target multi-element event monitoring dataset within the set monitoring period, different set monitoring time ranges, and different preset array representation channels. For example, assuming the preset array representation channels include network activity description information channels, access change data channels, and entity relationship channels, the representation array to be analyzed integrates various aspects of information such as URLs in network activities, access changes of IP addresses, and entity relationships between IP addresses. This information, stored in a specific array format, can more comprehensively reflect the network characteristics of the target multi-element event monitoring dataset.

[0032] When a computer system uses a network to identify cybersecurity threats, it uses a series of calculations and analysis methods. The network contains multiple neurons and connection weights and other parameters, which are determined during the training process. When the array of representations to be analyzed is input into the network, each neuron performs a weighted summation operation on the input data. Assume that the input of the neuron is , the corresponding connection weight is , the output y of the neuron can be expressed by the formula Calculated, where f is the activation function (such as Sigmoid function, ReLU function, etc.) and b is the bias term.

[0033] Through such calculations, the cybersecurity threat identification network gradually processes the information in the array of representations to be analyzed, ultimately determining whether a cybersecurity threat exists in the target multivariate event monitoring dataset. For example, if the output of the cybersecurity threat identification network is close to 1, the computer system can determine that a cybersecurity threat exists in the target multivariate event monitoring dataset; if the output is close to 0, it is determined that no cybersecurity threat exists. This determination is based on the cybersecurity threat identification network's learning and analysis of a large amount of training data. It can comprehensively consider the complex relationships between the various network features in the array of representations to be analyzed, thereby making a relatively accurate judgment on the cybersecurity status of the target multivariate event monitoring dataset. This process is crucial for network security management, helping computer systems to promptly detect potential cybersecurity threats, take appropriate preventative measures, and protect the security and stability of network systems.

[0034] As an embodiment, step S200, generating representation arrays for the plurality of first multivariate event monitoring data sets in combination with a set monitoring period, a set monitoring time range, and a preset array representation channel to obtain a pre-analysis representation array of the target multivariate event monitoring data set, may include:

[0035] Step S210: performing data segment interception on a plurality of first multivariate event monitoring data sets according to a set monitoring time range in a set monitoring cycle to obtain a third multivariate event monitoring data set for each set monitoring time range in the set monitoring cycle;

[0036] Step S220: generating a representation array for the third multivariate event monitoring data set of each set monitoring time range in combination with a preset array representation channel to obtain a quasi-analysis representation vector for each set monitoring time range;

[0037] Step S230: fusing the quasi-analysis representation vectors of the multiple set monitoring time ranges in the set monitoring period to obtain a quasi-analysis representation array.

[0038] In step S210, the monitoring period is set to a pre-set, longer time interval, such as 72 hours. This period is set to obtain sufficient network event data to fully reflect network activity. The monitoring time range is set to a shorter time interval within this longer period, such as 12 hours. The computer system performs data segmentation on the multiple first multivariate event monitoring data sets obtained in step S100, using the set monitoring time range as a unit.

[0039] For example, a first multivariate event monitoring dataset contains network activity information from various devices on the network over a 72-hour period, such as server access records and network traffic information. The computer system extracts data from this 72-hour dataset every 12 hours, generating six (72 ÷ 12 = 6) third multivariate event monitoring datasets. Each third multivariate event monitoring dataset represents multivariate event information related to network activity within a 12-hour period.

[0040] When extracting data segments, the computer system can employ timestamp technology. Assuming each network event data record carries a timestamp accurate to the second, the computer system compares the timestamp with the start and end times of the set monitoring time range to select data segments that meet the requirements. Specifically, if a network event data item's timestamp falls within a set 12-hour monitoring time range, the data item is included in the corresponding third-party event monitoring dataset.

[0041] In step S220 , the preset array represents channels for analyzing and characterizing multi-event monitoring data from different dimensions, including a network activity description information channel, an access change data channel, and an entity relationship channel.

[0042] For example, in a 12-hour third-party multivariate event monitoring data set, the network activity description information may include various fields in an HTTP request. The computer system first determines the network activity description string data and network traffic quantitative data corresponding to the network activity description information channel in the data set.

[0043] Network activity description string data, such as the URL in an HTTP request, includes 1000 HTTP requests within a 12-hour period, each with a different URL. These URL strings are part of the network activity description string data. Quantitative network traffic data includes data packet size, transmission rate, and other metrics. For example, at a certain moment, the network monitors an HTTP request with a data packet size of 1024 bytes and a transmission rate of 10 Mbps. These are quantitative network traffic data.

[0044] The computer system performs binning and feature encoding on the quantitative data of network traffic. Binning is the process of dividing continuous numerical data into different intervals (bins). For example, for the packet size, 0-512 bytes are set as one bin, 512-1024 bytes are set as another bin, and so on. Assuming a simple equal-width binning method is used, the formula is: bin number = floor((value - minimum value) / (bin width)), where floor is a rounding function. Feature encoding converts the binning results into an encoding form that the computer can process. For example, using One-Hot Encoding, if there are 5 bins, then the data in the third bin is encoded as [0, 0, 1, 0, 0]. In this way, the first representation vector for each set monitoring time range is obtained.

[0045] For string data describing network activity, the computer system performs data preprocessing and feature encoding. This preprocessing may include removing special characters from the string and converting the string to a uniform encoding format (such as UTF-8). A word embedding model (such as Word2Vec) is then used for feature encoding, converting the string into a vector form to obtain a second representation vector for each set monitoring time range. Finally, the first and second representation vectors are fused, for example, by using a simple concatenation operation to sequentially concatenate the two vectors to obtain a representation vector describing network activity for each set monitoring time range.

[0046] When the preset array indicates that a channel includes an access change data channel, the channel includes a single attribute change channel and multiple attribute change channels. In the 12-hour third multivariate event monitoring data set, for each single attribute change channel, the computer system determines the frequency of occurrence of the specific behavior and the statistical results of the specific attribute value when the corresponding single attribute change occurs.

[0047] For example, the single attribute change channel focuses on attributes related to the source IP address. The frequency of specific behaviors is measured by the number of connection requests from the same source IP address per unit time. For example, if a certain IP address initiates 50 connection requests within a 12-hour period, this represents the frequency of specific behaviors. Specific attribute value statistics can include the maximum, minimum, or average size of packets sent over a period of time. For example, assume that the minimum size of packets sent from this IP address is 128 bytes, the maximum is 1024 bytes, and the average is 512 bytes.

[0048] The computer system bins this data and performs feature encoding. This binning is similar to the binning of network traffic quantification data described above. For example, for the number of connection requests, 0-10 times is binned, 10-20 times is binned, and so on. Feature encoding can also use one-hot encoding to obtain a third representation vector for each set monitoring time range.

[0049] For multiple attribute change channels, the computer system determines the frequency of specific behaviors and the statistical results of specific attribute values ​​when the corresponding multiple attributes are changed. For example, multiple attribute change channels simultaneously focus on changes in the source IP address and the target port. The frequency of specific behaviors may be the number of times a certain IP address switches from one port to another within 12 hours, and the statistical results of specific attribute values ​​may be the average delay time of data packet transmission during the port switching process. These data are binned and feature encoded to obtain the fourth characterization vector for each set monitoring time range. Finally, the third characterization vector and the fourth characterization vector are vector-fused to obtain the access change characterization vector for each set monitoring time range.

[0050] When the preset array indicates that a channel includes an entity relationship channel, the channel includes a set-related entity channel and a set-negative feedback entity channel. In the 12-hour third multivariate event monitoring data set, for each set-related entity channel, the computer system determines a corresponding number of set-related entities and a corresponding ratio of set-related entities.

[0051] For example, if the related entity channel focuses on the relationship between IP addresses, and a certain IP address has communication relationships with five other IP addresses within 12 hours, the number of related entities is set to 5. If there are 10 possible related IP addresses in total, the related entity ratio is set to 5 / 10=0.5.

[0052] The computer system performs binning and feature encoding based on the set number of related entities and the set ratio of related entities. For example, for the number of entities, 0-3 entities are binned, 3-6 entities are binned, and so on. After binning, one-hot encoding is used to obtain the fifth representation vector for each set monitoring time range.

[0053] For a given negative feedback entity channel, the computer system determines the corresponding number of negative feedback entities. For example, if two IP addresses were found on a blacklist within 12 hours (blacklist records are part of the set negative feedback entity channel), the number of negative feedback entities is 2. These entities are then binned and feature-encoded (e.g., 0-1 per bin, 1-3 per bin, etc.) to obtain the sixth representation vector for each set monitoring time range. Finally, the fifth and sixth representation vectors are fused to obtain the entity relationship representation vector for each set monitoring time range.

[0054] In step S230, the computer system generated different types of pseudo-analysis representation vectors for each set monitoring time range in step S220, such as network activity description representation vectors, access change representation vectors, and entity relationship representation vectors. The computer system now merges the pseudo-analysis representation vectors for these multiple set monitoring time ranges within the set monitoring period (e.g., the six 12-hour set monitoring time ranges in the previous example).

[0055] Assume that for each 12-hour monitoring period, a vector of dimension n is generated for each network activity description representation vector. For six monitoring period periods, there are six n-dimensional vectors. Computer systems can employ various fusion methods, such as a simple stacking method, which stacks these six vectors in chronological order into a new vector of dimension 6n. The same stacking method is used for the access change representation vector and the entity relationship representation vector. These three stacked vectors are then concatenated to form a comprehensive quasi-analytic representation array. This quasi-analytic representation array integrates information from different time periods within the monitoring period and from different preset array representation channels. It can more comprehensively reflect the characteristics of the target multivariate event monitoring dataset, providing a richer data foundation for subsequent cybersecurity threat identification.

[0056] As an implementation method, the preset array representation channel includes at least one of a network activity description information channel (for example, the URL, User-Agent string, request parameters, etc. in an HTTP request, and you can give more examples), an access change data channel (for example, the same IP address attempts to connect to different ports, use different protocols, or quickly switch from one geographic location to another within a short period of time, and you can give more examples), and an entity relationship channel (for example, the communication pattern between a certain IP address and multiple other IP addresses, or the frequency with which a specific domain name is accessed by multiple different sources. It may also include tracking of objects such as known malicious IP addresses and blacklisted domain names, and you can give more examples); when the preset array representation channel includes a network activity description information channel, the to-be-analyzed representation vector includes a network activity description representation vector; when the preset array representation channel includes an access change data channel, the to-be-analyzed representation vector includes an access change representation vector; when the preset array representation channel includes an entity relationship channel, the to-be-analyzed representation vector includes an entity relationship representation vector.

[0057] In an embodiment of the present application, the preset array representation channel includes at least one of a network activity description information channel, an access change data channel, and an entity relationship channel. This is a way to characterize network security-related data from multiple key dimensions.

[0058] The network activity description information channel encompasses many elements, such as those in an HTTP request. Taking an HTTP request as an example, in addition to the URL, User-Agent string, and request parameters mentioned above, it also includes information such as the Referer field. For example, in an enterprise network environment, when an employee accesses an internal company website through a browser, the URL in the HTTP request clearly points to the requested web resource, such as "http: / / intranet.example.com / department / report.html." This URL reflects the employee's access target. The User-Agent string identifies the client software and device that initiated the request, such as "Mozilla / 5.0 (Windows NT 10.0; Win64; x64)AppleWebKit / 537.36 (KHTML, like Gecko) Chrome / 90.0.4430.212 Safari / 537.36," indicating that the request was initiated using the Chrome browser on Windows 10. Request parameters may include user authentication information after login, query conditions, and other information. The computer system processes the data from these network activity description information channels to generate a network activity description representation vector. During processing, the computer system may use data extraction technology to accurately extract this relevant information from the network traffic.

[0059] The access change data channel focuses on changes in network access. In addition to situations where the same IP address attempts to connect to different ports, use different protocols, or switch geographical locations within a short period of time, it also includes situations where the login device type of a user account changes frequently within a certain period of time. For example, on an online service platform, an account first logs in from a mobile device within 10 minutes, and then quickly switches to a desktop device to log in. The computer system monitors and analyzes data related to these access changes. In order to quantify this data, the computer system may use counters to count the frequency of specific behaviors, such as counting the number of times a certain IP address connects to different ports within an hour. For statistics on attribute values, for example, calculate the average value of network delays under different protocols in a series of connection attempts. By organizing and analyzing this data, the computer system generates an access change representation vector.

[0060] The entity relationship channel focuses on the relationships between network entities. In addition to the communication patterns between a certain IP address and multiple other IP addresses, the frequency of access to a specific domain name by different sources, and the tracking of malicious IP addresses and blacklisted domain names, it can also include the relationship between a certain network service and the different user groups that call it. For example, in a large network service architecture, the network security protection system finds that a certain network service is frequently called by user groups from several specific departments, which reflects an entity relationship. Computer systems can use association analysis algorithms to determine the relationship between entities, such as the Apriori algorithm, which analyzes a large number of network connection records to find frequently occurring entity association patterns. For the data of the entity relationship channel, the computer system will perform statistics and quantification according to the set rules, such as calculating the proportion of other IP addresses that have a communication relationship with a suspicious IP address to the total number of IP addresses, and then generate an entity relationship representation vector.

[0061] When the preset array representation channel includes a network activity description information channel, the generated network activity description representation vector can reflect the network status from the perspective of the specific description of network activity. When it includes an access change data channel, the access change representation vector provides information from the perspective of the dynamic changes in network access. When it includes an entity relationship channel, the entity relationship representation vector characterizes the network situation from the perspective of the relationship between network entities. These three representation vectors characterize network multi-event monitoring data from different key perspectives. Taken together, they can more comprehensively and accurately describe the overall state of the network, providing a rich and targeted data foundation for network security threat identification, helping computer systems to more effectively identify potential security threats within the network.

[0062] As an embodiment, when the preset array representation channel includes a network activity description information channel, step S220 generates a representation array for the third multivariate event monitoring data set for each set monitoring time range in combination with the preset array representation channel to obtain a quasi-analysis representation vector for each set monitoring time range, which may include:

[0063] Step S221: determining the network activity description character string data and network traffic quantification data corresponding to the network activity description information channel in the third multivariate event monitoring data set within each set monitoring time range;

[0064] Step S222: performing binning and feature encoding on the network traffic quantification data to obtain a first characterization vector for each set monitoring time range;

[0065] Step S223: performing data preprocessing and feature encoding on the network activity description character string data to obtain a second characterization vector for each set monitoring time range;

[0066] Step S224: performing vector fusion on the first characterization vector of each set monitoring time range and the second characterization vector of each set monitoring time range to obtain a network activity description characterization vector of each set monitoring time range.

[0067] In step S221, the monitoring time range is defined as a specific time period within the entire monitoring cycle. The third multivariate event monitoring dataset within this time period contains rich network activity information. The computer system then determines from this dataset two types of data corresponding to the network activity description information channel: network activity description string data and network traffic quantitative data.

[0068] In network activity, there is a lot of data in the form of strings that can describe the characteristics of network behavior. For example, in HTTP requests, the URL (Uniform Resource Locator) is a typical string data describing network activity. When a user enters a URL in a browser, such as "https: / / www.example.com / products / item1.html," this URL contains the location information of the requested target resource. It reflects the server address and specific path of the webpage content the user wants to access. The User-Agent string is also an important string data describing network activity. It identifies the client software and device information that initiated the HTTP request. For example, the string "Mozilla / 5.0 (Windows NT 10.0; Win64; x64)AppleWebKit / 537.36 (KHTML, like Gecko) Chrome / 90.0.4430.212 Safari / 537.36" indicates that the request was initiated by the Chrome browser (Chrome / 90.0.4430.212) running on Windows 10 (Windows NT 10.0) and 64-bit architecture (Win64; x64). It also contains information about the underlying rendering engines (AppleWebKit / 537.36 (KHTML, like Gecko) and Safari / 537.36).

[0069] The sender address and subject line in the email system are also string data describing network activities. For example, the sender address "sender@example.com" clearly indicates the source of the email, while the subject line "Meeting Agenda for NextWeek" roughly describes the content of the email.

[0070] In a DNS (Domain Name System) query, the queried domain name is a string of data describing network activity. For example, when a client queries the IP address of "www.google.com," the domain name "www.google.com" is the string of data describing network activity, reflecting the domain name identifier of the target network resource the client wants to access.

[0071] User IDs and device names in log files also fall into this category. For example, in a corporate network's system log, the login information for user "user123" using device "device-001" is recorded. "user123" and "device-001" are strings describing network activity, helping to track the network activity of a specific user on a specific device.

[0072] Packet size is a common form of quantitative data about network traffic. In network communications, each packet has a specific size. For example, during a file transfer, the packet size might be 1024 bytes. Computer systems can accurately obtain the size of each packet using network monitoring tools.

[0073] Transfer rate is also an important metric for quantifying network traffic. It indicates the amount of data transferred per unit time. For example, on a high-speed network connection, the transfer rate might reach 100 Mbps (megabits per second), reflecting how quickly the network can transmit data.

[0074] Network latency is also part of the quantitative data for network traffic. When a client sends a request to a server, the time between the request and the response is the network latency. For example, on a cross-border network connection, network latency can reach 200 milliseconds due to the distance and intermediary devices.

[0075] Computer systems also need to determine the number of connections or data transfers within a specific time period to quantify network traffic. For example, if a server connects to external clients 500 times in an hour, or if 50MB of data is transferred in that hour, these figures can reflect the frequency and scale of network activity.

[0076] The technical means by which computer systems determine these network activity description string data and network traffic quantitative data primarily include data parsing and extraction techniques. For network traffic data, network monitoring tools (such as network sniffers) can capture network packets and parse relevant information using protocol analysis techniques. For example, for HTTP traffic, tools can parse string data such as the URL and User-Agent, while also obtaining quantitative data such as packet size and transmission rate. For data in log files, computer systems can use log analysis tools to define specific rules to extract string data such as user IDs and device names, as well as relevant quantitative data (such as the number of logins within a specific time period).

[0077] In step S222, binning is the process of dividing continuous network traffic quantization data into different intervals (bins). For example, for packet size data, the computer system can set different bin intervals. Suppose that 0-512 bytes are set as the first bin, 512-1024 bytes are set as the second bin, 1024-2048 bytes are set as the third bin, and so on. The computer system divides each packet size data into the corresponding bin according to the pre-set bin interval. A similar binning operation can also be performed for transmission rate data. For example, 0-10 Mbps is the first bin, 10-50 Mbps is the second bin, 50-100 Mbps is the third bin, and so on. Taking a data with a transmission rate of 30 Mbps as an example, it will be divided into the second bin.

[0078] For network delay time, if 0-50 milliseconds is set as the first box, 50-100 milliseconds is set as the second box, 100-200 milliseconds is set as the third box, etc., then a 120 millisecond network delay time data will be divided into the third box.

[0079] Binning can be done using a simple mathematical formula to determine the bin to which data belongs. For example, for a value x, the bin number n is calculated as: n = floor((x - min) / (width)), where min is the minimum value, width is the bin width (i.e., the size of each bin), and floor is the floor function. For example, for a packet size x = 800 bytes, min = 0, and width = 512, then n = floor((800 - 0) / (512)) = 1, indicating that the data belongs to bin 2.

[0080] After binning, the computer system needs to perform feature encoding on the binned data to convert it into a form suitable for subsequent processing. One feasible feature encoding method is one-hot encoding. Assuming there are three bins for packet size data (such as 0-512 bytes, 512-1024 bytes, and 1024-2048 bytes mentioned above), if a piece of data belongs to the second bin (512-1024 bytes), its one-hot encoding is [0, 1, 0].

[0081] For transmission rate data, if there are three bins, when a piece of data belongs to the third bin (50-100 Mbps), its one-hot encoding is [0, 0, 1]. In this way, the computer system converts the binned network traffic quantitative data into one-hot encoding vectors, and these vectors are combined to obtain the first representation vector for each set monitoring time range.

[0082] In step S223, the network activity description string data first needs to be preprocessed. For URL strings, the computer system may remove some unnecessary parameters. For example, in "https: / / www.example.com / products / item1.html?param1=value1¶m2=value2", the "?param1=value1¶m2=value2" portion may be removed, leaving only "https: / / www.example.com / products / item1.html" to simplify the data and highlight the main resource location information.

[0083] The User-Agent string may be normalized. For example, different browser versions may be standardized to facilitate comparison and analysis. The sender address in the email header may be formatted to ensure it conforms to the standard email address format (e.g., username@domain.com). If it doesn't, it will be corrected or marked as an exception.

[0084] For user IDs and device names in log files, unified encoding conversion may be performed. For example, user IDs in different encoding formats may be converted to a unified UTF-8 encoding to prevent encoding incompatibility issues during subsequent processing.

[0085] After data preprocessing, the computer system needs to encode the features of the string data describing online activity. A common approach is to use word embedding models, such as Word2Vec. Taking a URL string as an example, the preprocessed URL string is used as input. The Word2Vec model maps each word (which can be seen as the part separated by " / " in the URL, such as "products" and "item1.html") into a low-dimensional vector. For example, suppose "products" in "https: / / www.example.com / products / item1.html" is mapped to the vector [0.1, 0.2, -0.3], and "item1.html" is mapped to the vector [-0.2, 0.3, 0.1], and so on. These vectors are then combined to obtain the feature vector for the URL string.

[0086] A similar method can be used to encode features for the User-Agent string. For example, each component of "Mozilla / 5.0 (Windows NT 10.0; Win64; x64) AppleWebKit / 537.36 (KHTML, like Gecko) Chrome / 90.0.4430.212 Safari / 537.36" can be mapped to a vector and then combined to generate a feature vector for the entire User-Agent string. In this way, the computer system processes the network activity description string data to generate a second representation vector for each set monitoring time range.

[0087] Step 4: Perform vector fusion on the first characterization vector of each set monitoring time range and the second characterization vector of each set monitoring time range to obtain a network activity description characterization vector for each set monitoring time range.

[0088] After obtaining the first characterization vector (from processing of network traffic quantification data) and the second characterization vector (from processing of network activity description string data) for each set monitoring time range, the computer system needs to perform a fusion operation on the two vectors.

[0089] For example, suppose the first representation vector is a one-hot encoded vector of length n1, such as [0, 1, 0, 0] (n1=4), and the second representation vector is a vector of length n2 composed of word vectors, such as [0.1, 0.2, -0.3, 0.4] (n2=4). A computer system can perform vector fusion using a simple concatenation method, concatenating the first and second representation vectors in sequence to produce a new vector [0, 1, 0, 0, 0.1, 0.2, -0.3, 0.4]. This new vector represents the network activity description vector for each set monitoring time range. This vector fusion method integrates the features of network traffic quantitative data and network activity description string data, thereby more comprehensively reflecting the information of the network activity description information channel and providing richer feature information for subsequent network security threat identification.

[0090] In one embodiment, when the preset array representation channel includes an access change data channel, and the access change data channel includes a single attribute change channel and multiple attribute change channels, step S220 generates a representation array for the third multivariate event monitoring data set for each set monitoring time range in combination with the preset array representation channel to obtain a quasi-analysis representation vector for each set monitoring time range, including:

[0091] Step S220A: determining the specific behavior occurrence frequency and specific attribute value statistics when a single attribute is changed corresponding to a single attribute change channel in each set monitoring time range of the third multivariate event monitoring data set;

[0092] Step S220B: performing binning and feature coding on the specific behavior occurrence frequency and specific attribute value statistics when a single attribute is changed, to obtain a third characterization vector for each set monitoring time range;

[0093] Step S220C: determining the specific behavior occurrence frequency and specific attribute value statistics when multiple attributes are changed corresponding to multiple attribute change channels in the third multivariate event monitoring data set within each set monitoring time range;

[0094] Step S220D: performing binning and feature coding on the specific behavior occurrence frequency and specific attribute value statistics when multiple attributes are changed, to obtain a fourth characterization vector for each set monitoring time range;

[0095] Step S220E: performing vector fusion on the third characterization vector of each set monitoring time range and the fourth characterization vector of each set monitoring time range to obtain the access change characterization vector of each set monitoring time range.

[0096] In an embodiment of the present application, when the preset array representation channel includes an access change data channel (which includes a single attribute change channel and multiple attribute change channels), an implementation of step S220 includes steps S220A-S220E.

[0097] In step S220A, within a network environment, a single attribute change channel focuses on changes in a specific attribute, and the computer system needs to determine the frequency of specific behaviors associated with that attribute. For example, in network access, using the source IP address as a single attribute, the number of connection requests from the same source IP address per unit time represents the frequency of a specific behavior. Assume that within a set monitoring timeframe (e.g., one hour), a source IP address 192.168.1.100 initiates 50 connection requests to the target server. These 50 requests represent the frequency of this specific behavior, namely, connection requests from that source IP address within that one-hour period.

[0098] Taking access to a specific resource (such as a file or service) as an example, if resource access is used as a single attribute, a computer system can count the number of accesses to this specific resource over a period of time. For example, on an enterprise network, there is a shared file resource " / shared / docs / report.pdf". Within a set monitoring time range of 10 minutes, there are 20 access requests for this file. These 20 times represent the frequency of access to this specific resource.

[0099] For login activity, using the user account as a single attribute, a computer system can count the number of abnormal login attempts detected within the same time period. For example, within a set monitoring period of 30 minutes, three abnormal login attempts were detected for the user account "user123." The three times here represents the frequency of this specific abnormal login attempt behavior.

[0100] For example, when considering network connection attributes, a computer system will measure the maximum, minimum, or average size of packets sent over a period of time. For example, within a 15-minute monitoring period, for connections with source IP address 10.0.0.5, the minimum size of packets sent was 128 bytes, the maximum was 1024 bytes, and the average was 512 bytes. These values ​​represent the statistics for a specific attribute.

[0101] Regarding session-related attributes, using the TCP sequence number as a single attribute, a computer system can calculate the maximum jump in the TCP sequence number during a session. Assuming that the maximum jump in the TCP sequence number during transmission in a network session is 100 (the specific value here is calculated based on the change in sequence numbers in the TCP protocol), this 100 is the statistical result of the specific attribute value (TCP sequence number).

[0102] For a single attribute, the computer system can count the number of unique destination IP addresses that appear within a specific time window. For example, within a 20-minute monitoring window, the source IP address 172.16.0.1 connects to five unique destination IP addresses. These five unique destination IP addresses represent the specific attribute value statistics for the specific attribute, the destination IP address.

[0103] Computer systems can determine the frequency of these specific behaviors and the statistical results of specific attribute values ​​using network monitoring tools and data statistical algorithms. Network monitoring tools can capture various information from network traffic in real time, and then the computer system uses data statistical algorithms to analyze this information. For example, to count the number of connection requests, a counter algorithm can be used, incrementing the counter value each time a new connection request is detected. To count packet sizes, the size of each packet can be recorded during data capture, and then compared to determine the maximum and minimum values, and then calculate the average.

[0104] In step S220B, the computer system bins the frequency data for specific behaviors. For example, using the number of connection requests from the same source IP address per unit time as an example, 0-10 connections are binned in the first bin, 10-20 connections are binned in the second bin, 20-30 connections are binned in the third bin, and so on. For example, if a source IP address receives 15 connection requests within a set monitoring timeframe (e.g., one hour), it will be binned in the second bin.

[0105] Binning is also performed for statistics on specific attribute values, such as packet size. For example, 0-512 bytes are set as the first bin, 512-1024 bytes are set as the second bin, 1024-2048 bytes are set as the third bin, and so on. If the average packet size is 700 bytes within a set monitoring time range, this average value will be placed in the second bin.

[0106] Binning uses a formula to determine the bin to which data belongs. For example, for a value x, the bin number n is calculated as: n = floor((x - min) / (width)), where min is the minimum value, width is the bin width (i.e., the size of each bin), and floor is the floor function. For example, if the number of connection requests x = 15, min = 0, and width = 10, then n = floor((15 - 0) / (10)) = 1, indicating that the data belongs to bin 2.

[0107] After binning, the computer system performs feature encoding on the binned results. A feasible feature encoding method is one-hot encoding. Assuming there are three bins for the number of connection requests (such as 0-10, 10-20, and 20-30), if a piece of data belongs to the second bin (10-20), its one-hot encoding is [0, 1, 0].

[0108] For data packet sizes divided into three bins, when a piece of data falls into the second bin (512-1024 bytes), its one-hot encoding is [0, 1, 0]. In this way, the computer system converts the frequency of specific behaviors and the statistical results of specific attribute values ​​when a single attribute changes after binning into one-hot encoding vectors. These vectors are combined to form a third representation vector for each set monitoring time range.

[0109] In step S220C, multiple attribute change channels simultaneously monitor changes in multiple attributes. For example, when considering both the source IP address and the destination port, the frequency of a specific behavior can be the number of times a source IP address switches its destination port within a certain period of time. Assume that within a set monitoring period of 30 minutes, the source IP address 192.168.1.50 switches its destination port eight times. This eight times represents the frequency of port switching, a specific behavior, when multiple attributes (source IP address and destination port) change.

[0110] Using the source IP address and the protocol used (such as TCP and UDP) as multiple attributes, a computer system can count the number of times a source IP address switches between TCP and UDP protocols over a period of time. For example, within a set monitoring period of 20 minutes, the source IP address 10.0.0.1 switches between TCP and UDP protocols three times. These three times represent the frequency of occurrence of this specific behavior in this case.

[0111] When considering both the source IP address and destination port, a specific attribute value statistic can be the average packet transmission delay during a port switch. For example, if the average packet transmission delay during a series of port switch operations is 150 milliseconds, this 150 milliseconds is the specific attribute value statistic when multiple attributes (source IP address and destination port) are changed.

[0112] For the source IP address and protocol attributes, the specific attribute value statistics can be the average packet size during protocol switching. For example, when the source IP address 172.16.0.2 switches between TCP and UDP, the average packet size is 800 bytes. This 800 bytes is the specific attribute value statistics for this situation.

[0113] The computer system also relies on network monitoring tools and statistical algorithms to determine the frequency of specific behaviors and the statistical results of specific attribute values ​​when these multiple attributes change. Network monitoring tools can capture network event information involving multiple attributes, and the computer system then uses statistical algorithms to organize and calculate this information. For example, the number of port switches can be counted by monitoring the changes in the destination port of each connection from the source IP address. To calculate the average delay of packet transmission, the time difference between sending and receiving the packet during each port switch can be recorded and then averaged.

[0114] In step S220D, the computer system groups the frequency of specific behavior during multiple attribute changes, such as the number of times the source IP address switches to the target port, into bins. For example, 0-5 times is the first bin, 5-10 times is the second bin, 10-15 times is the third bin, and so on. If the source IP address switches to the target port seven times within a set monitoring time range, it will be classified into the second bin.

[0115] Binning is also performed for statistical results of specific attribute values, such as the average delay time of data packet transmission during port switching. For example, 0-100 milliseconds is set as the first bin, 100-200 milliseconds is set as the second bin, and 200-300 milliseconds is set as the third bin. If the average delay time within a set monitoring time range is 180 milliseconds, this average delay time will be classified into the second bin.

[0116] The binning operation can also use the formula n=floor((x-min) / (width)) to determine the bin to which the data belongs, where x is the data to be binned, min is the minimum value, width is the bin width, and floor is the floor function.

[0117] After binning, the computer system performs feature encoding on the binned results. Using one-hot encoding, assume there are three bins for the number of times a source IP address switches to a destination port (e.g., 0-5, 5-10, and 10-15). If a piece of data falls into the second bin (5-10), its one-hot encoding is [0, 1, 0].

[0118] For the average packet transmission delay, if there are three bins, when a piece of data belongs to the second bin (100-200 milliseconds), its one-hot encoding is [0, 1, 0]. In this way, the computer system converts the frequency of specific behaviors and the statistical results of specific attribute values ​​when the multiple binned attributes change into one-hot encoded vectors. These vectors are combined to form the fourth representation vector for each set monitoring time range.

[0119] In step S220E, after obtaining the third characterization vector (from processing data related to a single attribute change channel) and the fourth characterization vector (from processing data related to multiple attribute change channels) for each set monitoring time range, the computer system needs to perform a fusion operation on these two vectors.

[0120] For example, suppose the third representation vector is a one-hot encoded vector of length n1, such as [0, 1, 0] (n1=3), and the fourth representation vector is a one-hot encoded vector of length n2, such as [1, 0, 0] (n2=3). A computer system can use a simple concatenation method to fuse these vectors. This concatenation method sequentially concatenates the third and fourth representation vectors to produce a new vector [0, 1, 0, 1, 0, 0]. This new vector represents the access change representation vector for each set monitoring time range. This vector fusion method integrates data features related to single attribute changes and multiple attribute changes, thereby more comprehensively reflecting the information of the access change data channel and providing richer feature information for subsequent network security threat identification.

[0121] As an embodiment, when the preset array representation channel includes an entity relationship channel, and the entity relationship channel includes a set related entity channel and a set negative feedback entity channel, step S220 generates a representation array for the third multivariate event monitoring data set for each set monitoring time range in combination with the preset array representation channel to obtain a quasi-analysis representation vector for each set monitoring time range, including:

[0122] Step S2201: determining the number and ratio of set related entities corresponding to the set related entity channels in the third multivariate event monitoring data set within each set monitoring time range;

[0123] Step S2202: performing binning and feature coding on the set number of related entities and the set ratio of related entities to obtain a fifth characterization vector for each set monitoring time range;

[0124] Step S2203: determining the number of negative feedback entities corresponding to the set negative feedback entity channel in the third multivariate event monitoring data set within each set monitoring time range;

[0125] Step S2204: performing binning and feature encoding on the number of negative feedback entities to obtain a sixth characterization vector for each set monitoring time range;

[0126] Step S2205: performing vector fusion on the fifth characterization vector of each set monitoring time range and the sixth characterization vector of each set monitoring time range to obtain the entity relationship characterization vector of each set monitoring time range.

[0127] In the embodiment of the present application, when the preset array representation channel includes an entity relationship channel (which includes setting a related entity channel and setting a negative feedback entity channel), an implementation of step S220 includes steps S2201-S2205.

[0128] In step S2201, related entity channels are set within a network environment to monitor the relationships between various entities. Taking IP addresses as an example, within a set monitoring time range (e.g., 1 hour), the computer system will determine the associations between a particular IP address and other IP addresses. For example, suppose the computer system detects that IP address 192.168.1.10 has communication relationships with five other IP addresses (e.g., 192.168.1.11, 192.168.1.12, 192.168.1.13, 192.168.1.14, and 192.168.1.15). The number 5 here represents the set number of related entities.

[0129] For domain name-related entity relationships, if a specific domain name is used as a reference, the computer system will count the number of other domain names that interact with that domain name (such as resolution requests, access to associated services, etc.) within a specific time range. For example, for the specific domain name "example.com", within the set monitoring time range of 30 minutes, there are three other domain names that interact with it. These three are the set number of related entities in this case.

[0130] Regarding user account entity relationships, with a specific user account as the center, the computer system can determine the number of other accounts with which that account has a specific relationship (e.g., shared resource access, joint project participation, etc.). For example, for the user account "user123," within a set monitoring time of 20 minutes, there are four other accounts with which it has a specific relationship. These four accounts constitute the corresponding number of related entities.

[0131] Setting a related entity ratio specifies the proportion of related entities within a specific population. Continuing with the IP address example above, if there are 10 IP addresses in the network environment that could potentially communicate with 192.168.1.10 (these 10 are the total number of potentially related IP addresses determined based on network topology, security policies, etc.), and only 5 actually communicate with it, then the set related entity ratio is 5 / 10 = 0.5.

[0132] For domain names, if there are 8 domain names in a Domain Name System (DNS) zone that may have an interaction relationship with the specific domain name "example.com", and there are 3 domain names that actually have an interaction relationship with it during the monitoring time range, then the relevant entity ratio is set to 3 / 8=0.375.

[0133] In terms of user accounts, if in an enterprise's user account management system, it is determined that there are 10 accounts that may have a specific association with "user123", and there are actually 4 accounts that have a specific association with it, then the relevant entity ratio is set to 4 / 10=0.4.

[0134] The computer system determines the number and proportion of related entities mainly through network monitoring tools and data analysis algorithms. Network monitoring tools can capture the interaction information between various entities in the network, such as the source and destination IP addresses in the network traffic, the relevant domain names in the domain name resolution request, etc. The computer system then organizes and calculates this captured information through data analysis algorithms. For example, for statistics on IP address communication relationships, the source IP addresses and destination IP addresses in the network traffic can be analyzed to determine which IP addresses have communication relationships, and then the number and proportion of related entities can be calculated. For domain name interaction relationships, statistics can be collected by analyzing DNS query and response records. For user account association relationships, they can be determined by analyzing data sources such as internal enterprise resource access logs and records in project collaboration systems.

[0135] In step S2202, the computer system will perform binning for the set number of related entities. For example, 0-3 entities are set as the first bin, 3-6 entities are set as the second bin, 6-9 entities are set as the third bin, and so on. Assuming that within a set monitoring time range (e.g., 1 hour), the number of related entities determined to be set is 4, then it will be classified into the second bin.

[0136] The same binning operation is performed for the set related entity ratio. For example, 0-0.3 is set as the first bin, 0.3-0.6 is set as the second bin, 0.6-0.9 is set as the third bin, etc. If the calculated set related entity ratio is 0.4 within a set monitoring time range, then this ratio will be divided into the second bin.

[0137] Binning uses a formula to determine the bin to which data belongs. For a value x, the bin number n is calculated as: n = floor((x - min) / (width)), where min is the minimum value, width is the bin width (i.e., the size of each bin), and floor is the floor function. For example, if the number of related entities x = 4, min = 0, and width = 3, then n = floor((4 - 0) / (3)) = 1, indicating that the data belongs to bin 2.

[0138] After binning, the computer system performs feature encoding on the binned results. A common feature encoding method is One-Hot Encoding (IHE). Assuming there are three bins for the number of related entities (e.g., 0-3, 3-6, and 6-9 above), if a data item belongs to the second bin (3-6), its IHE encoding is [0, 1, 0].

[0139] For the set relevant entity ratio, if there are three bins, when a data point belongs to the second bin (0.3-0.6), its one-hot encoding is [0, 1, 0]. In this way, the computer system converts the number of set relevant entities and the set relevant entity ratio after binning into one-hot encoding vectors. These vectors are combined to obtain the fifth representation vector for each set monitoring time range.

[0140] In step S2203, in the field of network security, setting a negative feedback entity channel encompasses various scenarios. For example, within a set monitoring timeframe (e.g., 30 minutes), the computer system will count the number of IP addresses or domain names on the blacklist. Assuming that within this timeframe, three IP addresses on the network security system's blacklist are identified as malicious, these three represent the number of negative feedback entities in this scenario.

[0141] In the case of user reports, if a user reports suspicious activity, such as a spam source, the computer system will count the number of related entities reported within a specific monitoring timeframe. For example, if two IP addresses are reported as spam sources within a 15-minute monitoring timeframe, these two IP addresses would be considered the number of negative feedback entities associated with the user report.

[0142] Security event logs, such as those generated by firewalls or intrusion detection systems (IDS), contain information about entities flagged as potential threats. If, within a 20-minute monitoring period, the security event logs show four IP addresses flagged for anomalous behavior, these four entities represent the number of negative feedback entities associated with the security event logs.

[0143] The computer system determines the number of entities receiving negative feedback primarily by querying and counting various security-related data sources. For blacklist records, the computer system can directly query the blacklist database and count the number of entities that were newly added or remain on the blacklist within the set monitoring timeframe. For user report information, the computer system parses the user report records and counts the number of related entities. For security event logs, the computer system analyzes the log files, extracts information about entities marked as anomalies or potential threats, and then counts the number of entities.

[0144] In step S2204, the computer system divides the number of negative feedback entities into bins. For example, 0-1 is set as the first bin, 1-3 is set as the second bin, 3-5 is set as the third bin, and so on. If the number of negative feedback entities determined within a set monitoring time range (such as 30 minutes) is 2, then it will be classified into the second bin.

[0145] The binning operation can also use the formula n=floor((x-min) / (width)) to determine the bin to which the data belongs, where x is the number of negative feedback entities, min is the minimum value, and width is the bin width. For example, if the number of negative feedback entities x=2, min=0, and width=1, then n=floor((2-0) / (1))=2, which means that the data belongs to the second bin.

[0146] After binning, the computer system performs feature encoding on the binned results. Using one-hot encoding, assuming there are three bins for the number of negative feedback entities (e.g., 0-1, 1-3, and 3-5), if a piece of data falls into the second bin (1-3), its one-hot encoding is [0, 1, 0]. In this way, the computer system converts the binned number of negative feedback entities into a one-hot encoded vector, which becomes the sixth representation vector for each set monitoring time range.

[0147] In step S2205, after obtaining the fifth characterization vector (from processing the set relevant entity channel related data) and the sixth characterization vector (from processing the set negative feedback entity channel related data) for each set monitoring time range, the computer system needs to perform a fusion operation on these two vectors.

[0148] For example, suppose the fifth representation vector is a one-hot encoded vector of length n1, such as [0, 1, 0] (n1=3), and the sixth representation vector is a one-hot encoded vector of length n2, such as [1, 0, 0] (n2=3). A computer system can perform vector fusion using a simple concatenation method, concatenating the fifth and sixth representation vectors in sequence to produce a new vector [0, 1, 0, 1, 0, 0]. This new vector is the entity relationship representation vector for each set monitoring time range. This vector fusion method integrates the information features of the set relevant entity channel and the set negative feedback entity channel, thereby more comprehensively reflecting the information of the entity relationship channel and providing richer feature information for subsequent network security threat identification.

[0149] As an implementation method, the network security threat identification network is trained through the following steps:

[0150] Step S10: Acquire multiple second multivariate event monitoring data sets in set monitoring periods before and after the multiple preset training data sets and a network security threat priori marker indicating whether each set training data set contains a network security threat;

[0151] Step S20: For each set training data set, generating a representation array for a plurality of second multivariate event monitoring data sets in combination with the set monitoring period, the set monitoring time range, and the preset array representation channel to obtain a training data set representation array of the preset training data set;

[0152] Step S30: performing network security threat identification on the training data set representation array based on the initialized neural network to obtain a network security threat identification result of the preset training data set;

[0153] Step S40: If the network security threat identification result is inconsistent with the network security threat priori label, then repeatedly debugging the network parameters of the initialized neural network in combination with the training error value of the initialized neural network until a preset debugging stop condition is met;

[0154] Step S50: Determine the initialized neural network after debugging as a network security threat identification network.

[0155] In the embodiment of the present application, the training of the network security threat identification network is completed through steps S10-S50.

[0156] In step S10, the preset training dataset is a collection of data used to train the network security threat identification network. These datasets are carefully selected and prepared, and contain information on multiple events related to network activity. For example, in an enterprise network environment, the preset training dataset may include a combination of multiple data sources such as server access logs, network traffic data, and user login information. This data reflects various network behavior patterns under both normal and abnormal (security threat-involved) conditions.

[0157] The preset training datasets can be divided into different types, such as active training datasets and passive training datasets. Active training datasets include past multivariate event monitoring datasets of network security, while passive training datasets include past multivariate event monitoring datasets of pre-set network vulnerabilities.

[0158] The set monitoring period is a predefined time range. For example, the set monitoring period can be one week. The computer system needs to obtain multiple second multivariate event monitoring data sets within the week before and after each set training data set. This means that for each set training data set, the computer system needs to collect relevant network event data from the week before and the week after.

[0159] The purpose of this is to obtain network activity information for the time periods before and after the preset training dataset, so as to better learn the characteristics and patterns of network security threats. For example, before a network security vulnerability is discovered, there may be some early signs of abnormal network activity. After the vulnerability is fixed, network activity will return to normal or new normal pattern characteristics will emerge.

[0160] The second multivariate event monitoring dataset is network activity data collected around a pre-set training dataset during a set monitoring period. This data contains various network event information, such as network connection status, service access status, and data transmission status. Taking network connection status as an example, the second multivariate event monitoring dataset may include information such as source IP address, destination IP address, port number used, connection establishment time, and connection duration.

[0161] Service access information may include the name of the service being accessed, the type of service (e.g., web service, database service), access time, access frequency, etc. The diversity of this data helps to comprehensively describe the state of network activity, thus providing rich learning materials for network security threat identification.

[0162] The cybersecurity threat prior label is a pre-determined indicator of whether each training dataset contains a cybersecurity threat. For an active training dataset, its cybersecurity threat prior label indicates that it contains no cybersecurity threats. For example, an active training dataset consisting of network activity data collected from an enterprise network during normal operation with no security incidents would have a corresponding cybersecurity threat prior label of "no threat."

[0163] For a passive training dataset, its cybersecurity threat prior label indicates the presence of a cybersecurity threat. For example, a passive training dataset is network activity data collected from an enterprise network during a specific cyberattack (such as a SQL injection attack), and its corresponding cybersecurity threat prior label is "threat present." These prior labels provide supervisory information for the training of the cybersecurity threat identification network, enabling the network to learn how to distinguish between network activity patterns that indicate a security threat and those that do not.

[0164] The technical means by which computer systems acquire these data and markers include extracting data from data sources such as network monitoring tools and log storage systems. For example, network monitoring tools (such as network sniffers) can capture network traffic data and store it as part of the second multivariate event monitoring dataset. Log storage systems (such as server logs and firewall logs) can provide data such as user login information and security event records. Computer systems can extract the required information from this data and determine a priori markers of network security threats based on predefined rules.

[0165] In step S20, the monitoring time range is set to a time interval that is less than the set monitoring period. For example, if the monitoring period is set to one week (7 days), the set monitoring time range may be one day. The computer system will process the plurality of second multivariate event monitoring data sets according to the one-day time range within each set monitoring period (one week).

[0166] The benefit of this approach is that it allows analysis of network activity data at different time granularities, capturing both short-term and long-term network activity characteristics. For example, within a day, there may be specific peaks and valleys in network activity. These short-term characteristics may be related to network security threats. At the same time, there may also be overall network activity trends within a week. Combining these two helps to more comprehensively describe the status of network activity.

[0167] The preset array representation channels include network activity description information channel, access change data channel, entity relationship channel, etc. These channels analyze and represent network activity data from different dimensions.

[0168] Taking the network activity description information channel as an example, the computer system will extract information such as the URL, User-Agent string, and request parameters in HTTP requests, as well as quantitative network traffic data (such as packet size and transmission rate) from the second multivariate event monitoring data set. For the access change data channel, the computer system will focus on situations such as the same IP address attempting to connect to different ports within a short period of time, using different protocols, or quickly switching from one geographic location to another. For the entity relationship channel, the computer system will analyze relational data such as the communication patterns between a specific IP address and multiple other IP addresses, the frequency with which a specific domain name is accessed by multiple different sources, and so on.

[0169] The computer system processes each second multivariate event monitoring data set within a set monitoring time range according to a preset array representation channel. For example, for the network activity description information channel, the computer system determines network activity description string data and network traffic quantitative data. The network traffic quantitative data is then binned and feature encoded, while the network activity description string data is preprocessed and feature encoded. Finally, the two are combined to generate a network activity description representation vector.

[0170] Similar operations are performed on the access change data channel and the entity relationship channel, respectively obtaining access change representation vectors and entity relationship representation vectors. The computer system then fuses these representation vectors within each set monitoring time range to obtain a comprehensive representation vector for each set monitoring time range. Finally, the comprehensive representation vectors for multiple set monitoring time ranges within the set monitoring cycle are fused to obtain a training dataset representation array for the preset training dataset.

[0171] For example, suppose there are seven set monitoring time ranges (one per day) within a set monitoring cycle (one week). For each set monitoring time range, the computer system generates a comprehensive representation vector through the above processing. These vectors are combined through splicing and other methods to form the training dataset representation array. This array integrates information from different time ranges and different channels, and can comprehensively reflect the network activity characteristics related to the preset training dataset.

[0172] In step S30, the initialized neural network is a pre-built network model with a certain structure, which includes multiple neurons and connection weights and other parameters. For example, the initialized neural network can be a multi-layer perceptron (MLP) structure, including an input layer, several hidden layers, and an output layer.

[0173] The number of neurons in the input layer depends on the dimensionality of the array representing the training dataset. For example, if the dimension of the array representing the training dataset is n, then the input layer will have n neurons. The number of neurons and the number of layers in the hidden layer can be determined empirically or through preliminary experiments. The output layer typically has one or more neurons. In the case of cybersecurity threat identification, the output layer may have a single neuron, whose output value indicates whether a cybersecurity threat exists (e.g., an output close to 1 indicates the presence of a threat, and close to 0 indicates the absence of a threat).

[0174] When the computer system inputs the training data set representation array into the initialized neural network, the connection weights between neurons perform a weighted summation operation on the input data. For example, for the input value of the i-th neuron in the input layer , the connection weight connected to the jth neuron in the hidden layer is , the input of the jth neuron in the hidden layer is ,in is the bias term of the jth neuron in the hidden layer.

[0175] Then, the neurons in the hidden layer process the input through the activation function (such as Sigmoid function, ReLU function, etc.) to obtain the output. For example, for the Sigmoid function , the output of the hidden layer neurons is This process is carried out in each layer of the neural network, and finally a prediction result is obtained at the output layer about whether there is a network security threat in the preset training data set, that is, the network security threat identification result.

[0176] In step S40, the training error value is an indicator used to measure the difference between the network security threat identification result and the network security threat prior label. A feasible method for calculating the training error value includes mean square error (MSE). For example, for a preset training data set, the network security threat prior label is y (e.g., y = 0 means no threat, y = 1 means threat), and the network security threat identification result is , then the mean square error , where N is the number of samples in the training dataset.

[0177] The computer system calculates the training error value to determine the degree of deviation between the network's prediction results and the actual label. If the training error value is large, it means that the network's prediction effect is poor and the network parameters need to be debugged.

[0178] Network parameters primarily refer to parameters such as connection weights and biases in a neural network. When cybersecurity threat identification results differ from prior cybersecurity threat labels, the computer system adjusts these parameters based on the training error. For example, in the backpropagation algorithm, the computer system calculates the gradient of the error with respect to each parameter and then updates the parameter at a specific learning rate.

[0179] Assume that the update formula of connection weight w is ,in is the learning rate, and E is the training error. By continuously adjusting the network parameters based on the training error, the network’s prediction results gradually approach the prior markers of network security threats.

[0180] The preset debugging stop conditions are pre-set criteria for determining whether neural network training is complete. These conditions may include reaching a certain number of training rounds (such as 100 rounds), the training error value being less than a certain threshold (such as 0.01), or the accuracy on the validation set reaching a certain level.

[0181] For example, if the training error value is set to be less than 0.01 as the debugging stop condition, when the computer system finds that the calculated MSE is less than 0.01 during the debugging process, it is considered that the neural network has been trained to a level that meets the requirements and the debugging process can be stopped.

[0182] In step S50, when the computer system completes debugging of the initialized neural network, i.e., when the preset debugging stop condition is met, the debugged neural network is determined to be a network security threat identification network. This network has acquired the ability to identify network security threats by learning from multiple preset training datasets, their associated second multivariate event monitoring datasets, and prior signatures of network security threats. In subsequent actual network security threat identification processes, this network will be used to process the target multivariate event monitoring dataset's array of pseudo-analysis representations to determine whether the target multivariate event monitoring dataset contains network security threats.

[0183] As an implementation method, multiple preset training data sets include multiple positive training data sets and multiple negative training data sets, the multiple positive training data sets include multiple past multivariate event monitoring data sets of network security, and the network security threat prior mark of the positive training data set indicates that there is no network security threat in the positive training data set; the multiple negative training data sets include multiple past multivariate event monitoring data sets with set network vulnerabilities, and the network security threat prior mark of the negative training data set indicates that there is a network security threat in the negative training data set.

[0184] In the embodiment of the present application, the preset training data set is divided into multiple positive training data sets and multiple negative training data sets, which is an effective way to organize training data.

[0185] The active training dataset consists of multiple historical multivariate network security event monitoring datasets. For example, within an enterprise network environment, computer systems collect network operation data during normal working hours. This data covers multiple aspects, such as normal network traffic information, including legitimate HTTP requests, normal file transfers, and other multivariate events. Regarding network traffic, information such as packet size, transmission rate, and source and destination IP addresses are all within normal ranges and patterns. For HTTP requests, for example, the URL is a normal address for accessing internal enterprise resources, and the User-Agent string indicates that the request was initiated by a legitimate browser or client software. Login activities are all legitimate users logging in at normal working hours and locations using the correct account and password. These historical multivariate network security event monitoring datasets are combined into the active training dataset and labeled as free of network threats. This labeling serves as a priori marker of network threats, providing a reference for the characteristics of network activity in a threat-free state for training the network for network threat identification.

[0186] The passive training dataset contains multiple past multivariate event monitoring datasets with set network vulnerabilities. For example, when an enterprise network suffers a SQL injection attack, the computer system will collect network data during the attack as part of the passive training dataset. In this case, abnormal HTTP requests may appear in the network traffic, such as URL requests containing malicious SQL statements. The purpose of these requests is to obtain illegal data or perform malicious operations by exploiting vulnerabilities in database query statements. In addition, abnormal login attempts may occur, such as frequent attempts to log in to an account using brute force password cracking, or access requests from abnormal IP addresses (such as hacker-controlled IP addresses) to specific internal resources of the enterprise (such as servers containing sensitive information). These multivariate event monitoring datasets during the period of the set network vulnerabilities constitute the passive training dataset and are marked as containing network security threats. This prior labeling of network security threats provides a reference for the characteristics of network activities under threat for the training of the network security threat identification network.

[0187] When constructing a positive training dataset, a computer system can obtain data by screening network logs during normal operation. For network traffic data, network monitoring tools can be used to collect data during normal time periods. When constructing a negative training dataset, the computer system can search security event logs for data from periods when network vulnerabilities existed, while also incorporating logs from network intrusion detection systems (IDS) or intrusion prevention systems (IPS) during attacks. In this way, the computer system can utilize both positive and negative training datasets to provide non-threatening and threatening network activity patterns, respectively, effectively training the network security threat identification network to accurately distinguish between situations where network threats are present and those where they are absent. This organization of training data helps improve the accuracy and reliability of the network security threat identification network, enabling it to make more accurate judgments when faced with real-world network security threat identification tasks.

[0188] As an implementation method, multiple active training data sets are collected by the following steps:

[0189] Step S11: obtaining a first set of monitoring data sets corresponding to a plurality of past multivariate event monitoring data sets of network security;

[0190] Step S12: selecting from the first monitoring data set a past multivariate event monitoring data set that has multiple multivariate event monitoring data sets in a set monitoring period before and after the first monitoring data set to obtain a second monitoring data set set;

[0191] Step S13: arbitrarily extracting the second monitoring data set to obtain multiple active training data sets;

[0192] Multiple negative training datasets are collected through the following steps, including:

[0193] Step S14: obtaining a third monitoring data set set including a plurality of past multivariate event monitoring data sets in which a set network vulnerability exists;

[0194] Step S15: classifying the past multivariate event monitoring datasets in the third monitoring dataset set in combination with the entity types in the past multivariate event monitoring datasets to obtain a plurality of fourth monitoring dataset sets;

[0195] Step S16: arbitrarily extract each fourth monitoring data set to obtain multiple negative training data sets.

[0196] In step S11, the past multi-event monitoring dataset is a collection of data obtained from monitoring network activities. It contains information on multiple aspects of the network. For example, in an enterprise network environment, these datasets may include network traffic data, server logs, user login information, and other types of data.

[0197] Network traffic data includes information such as source IP address, destination IP address, port number, protocol type, packet size, and transmission time. For example, in a simple HTTP request, the source IP address might be the IP address of an internal employee's office computer, the destination IP address might be the IP address of an internal web server, the port number might be 80 (the default HTTP port), and the protocol type might be TCP. The packet size varies depending on the request content, and the transmission time reflects network latency.

[0198] Server logs record various server activities, such as file access requests, service start and stop times, error messages, etc. For example, when an employee downloads a file from the server, the server log will record the employee's account information, the downloaded file name, the download time, etc.

[0199] User login information includes the login account, login time, login IP address, whether multi-factor authentication is used, etc. For example, if an employee uses a username and password to log into the enterprise resource management system from a specific IP address on the company's internal network at 9 a.m., all this information will be recorded.

[0200] The first monitoring dataset is composed of multiple historical multivariate event monitoring datasets corresponding to network security. Computer systems can obtain this data in a variety of ways. Network traffic data can be obtained using network monitoring tools, such as network sniffers. Network sniffers capture data packets on the network and extract relevant information, storing it as part of the network traffic data.

[0201] For server logs, the computer system can directly read them from the server's log storage system. For example, in Linux systems, server logs are typically stored in a specific directory (such as / var / log), and the computer system can use file read operations to obtain these log data.

[0202] User login information may be stored in an enterprise's authentication server or a dedicated user management system. Computer systems can interact with these systems and obtain user login information using database queries and other technical means. Data from these various sources is integrated to form the first monitoring dataset, providing the foundation for the subsequent construction of the active training dataset.

[0203] In step S12, the monitoring period is set to a predetermined time range. For example, the monitoring period may be set to one week (7 days). This time period is set to obtain sufficient network activity data to reflect changes in network status.

[0204] In a network environment, different network activities exhibit distinct characteristics during different time periods. For example, in an enterprise network, network traffic patterns may differ between weekdays and non-weekdays. During weekdays, there may be more business-related network activity, such as employees accessing internal business systems and sharing files. Non-weekdays may also feature system maintenance-related network activity or relatively low network usage.

[0205] The computer system selects from the first set of monitoring data sets. For each past multivariate event monitoring data set, the computer system checks whether there are sufficient multivariate event monitoring data sets in the set monitoring periods before and after the data set. For example, if the set monitoring period is one week, the computer system checks whether there is sufficient monitoring data for the network activity corresponding to a past multivariate event monitoring data set in the week before and the week after the data set.

[0206] Suppose the first monitoring dataset contains a historical multivariate event monitoring dataset covering network activity on an enterprise network on a specific date. The computer system checks to see if the network traffic data, server logs, and user login information from the week before and the week after this date are complete and sufficient for analysis. If these conditions are met, the historical multivariate event monitoring dataset is selected for inclusion in the second monitoring dataset. This approach aims to obtain contextually relevant network activity data for better analysis of network security.

[0207] In step S13, the computer system randomly extracts data from the second monitoring dataset set. This means that, under certain conditions, a data subset is randomly selected from the second monitoring dataset set as an active training dataset. For example, if the second monitoring dataset set contains 100 historical multivariate event monitoring datasets, the computer system may randomly extract a preset number (e.g., 20).

[0208] This random sampling approach helps ensure diversity in the active training dataset. This is because random sampling avoids the presence of specific patterns or biases in the dataset. For example, if the dataset is not randomly selected but selected in a certain order (such as chronological order), the training dataset may be overly concentrated on the network activity characteristics of a specific time period, while ignoring different network activity patterns that may occur in other time periods.

[0209] These extracted datasets constitute multiple active training datasets. These active training datasets contain historical multivariate cybersecurity event monitoring datasets, and by definition, their cybersecurity threat prior labels indicate the absence of cybersecurity threats. This is because these datasets are extracted from a normally operating network environment, unaffected by cybersecurity threats. For example, within an enterprise network, the network activities in these datasets are all legitimate and normal network operations that comply with enterprise security policies, such as normal employee business operations and system maintenance activities.

[0210] In step S14, these data sets are records of network activity data during the period when a specific vulnerability exists in the network. For example, when a SQL injection vulnerability exists in an enterprise network, network traffic data, server logs, user login information, etc. during the period when the vulnerability exists will be affected.

[0211] In terms of network traffic, abnormal HTTP requests may appear, which attempt to exploit SQL injection vulnerabilities to obtain sensitive information from the database. For example, a malicious URL may contain some special SQL statements such as "' or 1=1--" to bypass the database authentication mechanism.

[0212] Server logs may record unusual database query errors, such as syntax errors caused by malicious SQL injection attempts. User login information may also show unusual login attempts, such as frequent login attempts from suspicious external IP addresses. This could be an attempt by hackers to exploit SQL injection vulnerabilities to gain login privileges and further infiltrate the system.

[0213] The computer system obtains these past multi-event monitoring data sets involving the specified network vulnerabilities and forms them into a third set of monitoring data sets. The computer system can search for these data sets from the enterprise's security event records. For example, when an enterprise's intrusion detection system (IDS) or intrusion prevention system (IPS) detects a security event related to a network vulnerability, it will record the network activity data at that time. The computer system can extract relevant data from these records to form the third set of monitoring data sets.

[0214] In addition, when analyzing security incidents, the company's network security team may manually collect some relevant data, such as detailed logs of the server during the vulnerability period, detailed analysis reports of network traffic, etc. These data will also be included in the third monitoring data set.

[0215] In step S15, in the field of network security, entity types may include IP addresses, domain names, user accounts, service types, etc. For example, using IP addresses as entity types, the computer system may classify data based on the source IP addresses and destination IP addresses involved in network activities.

[0216] For domain name entity types, computer systems can categorize them based on the domain name involved in network requests (e.g., internal enterprise domain name, external partner domain name, etc.). User account entity types are categorized based on the logged-in user account. Different user accounts may have different permissions and network activity patterns. Service type entity types can be categorized based on the service involved in the network activity, such as web services, database services, and email services.

[0217] The computer system categorizes the past multivariate event monitoring datasets in the third set of monitoring datasets. For example, if network traffic data from a period during which a network vulnerability exists involves a specific IP address range (e.g., IP addresses of a subnet within an enterprise), the computer system will categorize the data into a fourth set of monitoring datasets identified by the IP address range.

[0218] Taking domain names as an example, if there was unusual activity targeting a specific domain name (such as the domain name of a company's financial system) during the vulnerability period, the relevant data would be classified into the fourth monitoring data set identified by that domain name. Regarding user accounts, if a user account was found to have unusual login or operation behavior during the vulnerability period, the network activity data associated with that account would be classified into the fourth monitoring data set identified by that user account. Regarding service types, if the vulnerability primarily affected database services, the network activity data related to the database service (such as database query requests, database connection attempts, etc.) would be classified into the fourth monitoring data set identified by the database service.

[0219] In step S16, similar to constructing the active training dataset, the computer system randomly extracts data from each fourth monitoring dataset set. For example, assuming that a fourth monitoring dataset set contains 50 past multivariate event monitoring datasets, the computer system randomly extracts data based on a preset number of samples (e.g., 10).

[0220] This random extraction method also helps ensure the diversity of the negative training dataset. Since each fourth monitoring dataset is classified according to entity type, there may be some internal similarities. Random extraction can avoid over-reliance on a specific pattern or situation.

[0221] The extracted datasets form multiple passive training datasets. These passive training datasets contain historical multivariate event monitoring datasets involving the presence of pre-defined network vulnerabilities, and their cybersecurity threat prior markers indicate the presence of cybersecurity threats. This is because these datasets were collected from situations where networks were vulnerable to vulnerabilities and security threats, and contain various network activity characteristics related to cybersecurity threats, such as abnormal network requests and malicious operations. These characteristics can be learned by the cybersecurity threat recognition network, enabling it to accurately identify situations where cybersecurity threats exist.

[0222] In one embodiment, when the initialized neural network includes a first detection component and a second detection component, the first detection component includes a plurality of first detection modules, and the second detection component includes a second detection module, step S300 performs network security threat identification on the to-be-analyzed representation array based on the network security threat identification network to obtain a network security threat identification result indicating whether a target multivariate event monitoring dataset contains a network security threat, which may include:

[0223] Step S310: performing network security threat identification on the target analysis representation array based on the multiple first detection modules to obtain multiple network security threat identification results;

[0224] Step S320: Integrate the multiple network security threat identification results to obtain an integrated result;

[0225] Step S330: performing network security threat identification on the integration result based on the second detection module to obtain a network security threat identification result.

[0226] In step S310, the first detection module in the initialized neural network is a component specifically designed to identify network security threats from different angles for the array of representations to be analyzed. Each first detection module has its own specific structure and function, and may focus on analyzing different parts or features in the array of representations to be analyzed. For example, in a network security threat identification network, the first detection module can be divided into a detection module for characteristics related to the network activity description information channel, a detection module for characteristics related to the access change data channel, and a detection module for characteristics related to the entity relationship channel, etc.

[0227] For example, the first detection module, which targets features related to the network activity description information channel, might have an internal structure consisting of a sub-neural network containing multiple neurons. These neurons are designed to recognize specific patterns within the network activity description representation vector. Assuming the network activity description representation vector includes encoded features of information such as the URL and User-Agent string in an HTTP request, the neurons in this first detection module might detect specific keywords in the URL (e.g., certain character combinations that may be associated with malicious activity) or unusual identifiers in the User-Agent string (e.g., a forged browser identifier).

[0228] The first detection module, which targets features related to access change data channels, may focus on analyzing the information in the access change representation vector. For example, if the access change representation vector contains feature codes related to a single attribute change (such as a change in the connection frequency of the source IP address) and multiple attribute changes (such as a simultaneous change in the source IP address and the target port), the first detection module will determine whether these changes conform to the pattern of network security threats based on predefined rules and algorithms. For example, an abnormally high frequency of connections from the same source IP address to different target ports within a short period of time may be considered a potential network security threat pattern, and the first detection module will identify this pattern.

[0229] The first detection module, which targets entity relationship channel-related features, primarily processes entity relationship representation vectors. For example, if the entity relationship representation vector contains information about the association between a certain IP address and an IP address on a blacklist (obtained by encoding the negative feedback entity channel-related features), the first detection module will identify whether this association indicates a network security threat. If a normal internal enterprise IP address suddenly establishes communication relationships with multiple blacklisted IP addresses, this could be a sign of malicious control or a security risk, and this module can detect this situation.

[0230] When a computer system inputs a pseudo-analysis representation array into a network security threat identification network comprising multiple first detection modules, each first detection module independently processes the pseudo-analysis representation array. For example, a first detection module associated with a network activity description information channel receives the portion of the pseudo-analysis representation array corresponding to the network activity description representation vector. Assuming the pseudo-analysis representation array is a multidimensional array, in which the portion corresponding to the network activity description representation vector is a subarray, the first detection module performs calculations on this subarray based on its internal neuron connections and weight settings.

[0231] Take a simple neuron calculation as an example, assuming that the input of the neuron is , the corresponding connection weight is , the output y of the neuron can be expressed by the formula Calculated, where f is the activation function (such as Sigmoid function Or the ReLU function , b is the bias term. Multiple neurons in this first detection module perform such calculations, ultimately obtaining a network security threat identification result related to the network activity description information channel.

[0232] Similarly, the first detection modules associated with the access change data channel and the entity relationship channel will also perform similar calculations, respectively obtaining cybersecurity threat identification results for access changes and entity relationships. Thus, through the parallel processing of multiple first detection modules, the computer system will obtain multiple cybersecurity threat identification results, each of which reflects from a different perspective whether the target multivariate event monitoring dataset represented by the target representation array to be analyzed contains cybersecurity threats.

[0233] In step S320, since the multiple network security threat identification results obtained in step S310 are obtained from different first detection modules, each result focuses on different aspects or features of the characterization array to be analyzed, so these results may have certain differences or incompleteness. For example, a first detection module based on a network activity description information channel may obtain a less accurate preliminary network security threat identification result due to the lack of some information or noise interference in the network activity description characterization vector. Similarly, the first detection module based on the access change data channel and the entity relationship channel may also face similar situations. In order to obtain a more comprehensive and accurate judgment on whether the target multi-event monitoring data set contains a network security threat, the computer system needs to integrate these multiple network security threat identification results.

[0234] A feasible integration operation method is the weighted average method. The computer system assigns a weight to each network security threat identification result. This weight reflects the importance or reliability of the corresponding first detection module in the overall network security threat identification. For example, suppose that after preliminary experiments or empirical analysis, the weight of the first detection module based on the network activity description information channel is determined to be =0.3, the weight of the first detection module based on the access change data channel is =0.3, the weight of the first detection module based on the entity relationship channel is =0.4.

[0235] If the network security threat identification results obtained by the three first detection modules are (here It can be a value between 0 and 1, indicating the probability of no network security threat to the existence of network security threat). Then the integration result R can be obtained by the formula For example, if =0.2 (indicates that from the perspective of the network activity description information channel, there is a small possibility of network security threat), =0.3 (indicates that from the perspective of accessing and changing data channels, there is a certain possibility of network security threats), =0.5 (indicates that from the perspective of entity relationship channel, there is a greater possibility of network security threat), then the integration result .

[0236] In addition to the weighted average method, other integration methods can be used, such as voting. In this voting method, if more than half of the multiple first detection modules determine that a network security threat exists, the integrated result is considered to indicate that a network security threat exists. For example, if there are five first detection modules, and three of them identify the network security threat as existing (denoted as 1), and two modules identify the threat as not existing (denoted as 0), then according to the voting method, the integrated result is considered to indicate that a network security threat exists (i.e., 1).

[0237] In step S330, the second detection module is another important component in the network security threat identification network. Its function is to perform final network security threat identification based on the integrated results. The second detection module typically has a more complex structure or higher-level feature analysis capabilities. For example, it may be a multi-layer perceptron (MLP) structure with multiple hidden layers.

[0238] Compared to the first detection module, the second detection module receives an integrated result as input. This result combines cybersecurity threat identification information from the different channel characteristics in the array to be analyzed. The second detection module's task is to extract deeper cybersecurity threat characteristics or patterns from this integrated result. For example, it may analyze the meaning of the integrated result within different value ranges and the relationship between the integrated result and typical threat patterns in historical cybersecurity data.

[0239] When the computer system inputs the integrated results into the second detection module, it performs calculations based on its internal neuron connections and weight settings. For example, suppose the second detection module has a single neuron in its input layer that receives the integrated result R. This information is then passed through multiple neurons in the hidden layer for feature extraction. The hidden layer neurons perform calculations similar to the previously mentioned neurons, using a weighted summation followed by an activation function to produce an output.

[0240] Assume that the hidden layer has m neurons and the connection weight of the input to the jth neuron in the hidden layer is , the bias term of the jth neuron in the hidden layer is , then the input of the jth neuron in the hidden layer is (Here i=1 because the input layer has only one neuron), the output is , where f is the activation function. After processing by the hidden layer, the output layer will obtain the final cybersecurity threat identification result. For example, if the output value of the neuron in the output layer is close to 1, the computer system determines that a cybersecurity threat exists in the target multivariate event monitoring dataset; if the output value is close to 0, it determines that no cybersecurity threat exists. In this way, through processing by the second detection module, the computer system obtains the final cybersecurity threat identification result regarding whether a cybersecurity threat exists in the target multivariate event monitoring dataset. This result comprehensively considers the multiple features in the target analysis representation array and the analysis results of different detection modules, and has high accuracy and reliability.

[0241] The embodiment of the present application provides a computer system, such as Figure 2 As shown, computer system 100 includes: a processor 101 and a memory 103. Processor 101 and memory 103 are connected, for example, via bus 102. Optionally, computer system 100 may further include a transceiver 104. It should be noted that in actual applications, the number of transceivers 104 is not limited to one, and the structure of computer system 100 does not constitute a limitation on the embodiments of this application.

[0242] An embodiment of the present application provides a computer system. The computer system in the embodiment of the present application includes: one or more processors; a memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors. When the one or more programs are executed by the processor, the above method is implemented.

Claims

1. A network security threat identification method based on multivariate event analysis, characterized in that: The method comprises: Acquire a plurality of first multivariate event monitoring data sets in a set monitoring period before and after the target multivariate event monitoring data set; Performing data segment interception on the plurality of first multivariate event monitoring data sets according to the set monitoring time range in the set monitoring period to obtain a third multivariate event monitoring data set for each set monitoring time range in the set monitoring period; When the preset array representation channel includes a network activity description information channel, determining the network activity description character string data and network traffic quantification data corresponding to the network activity description information channel in the third multivariate event monitoring data set within each set monitoring time range; Performing binning and feature encoding on the network traffic quantification data to obtain a first characterization vector for each set monitoring time range; Performing data preprocessing and feature encoding on the network activity description character string data to obtain a second characterization vector for each set monitoring time range; Performing vector fusion on the first characterization vector of each set monitoring time range and the second characterization vector of each set monitoring time range to obtain a network activity description characterization vector of each set monitoring time range; fusing quasi-analysis characterization vectors of a plurality of set monitoring time ranges in the set monitoring period to obtain the quasi-analysis characterization array, wherein the set monitoring time range is smaller than the set monitoring period; Based on the network security threat identification network, network security threats are identified on the quasi-analysis characterization array to obtain a network security threat identification result of whether the target multivariate event monitoring data set contains network security threats; the network security threat identification network is obtained by debugging an initialized neural network in combination with multiple second multivariate event monitoring data sets in a set monitoring period before and after multiple preset training data sets and a network security threat priori mark of whether each set training data set contains network security threats.

2. The method according to claim 1, characterized in that The preset array representation channel includes at least one of a network activity description information channel, an access change data channel, and an entity relationship channel; when the preset array representation channel includes the network activity description information channel, the pseudo-analysis representation vector includes a network activity description representation vector; when the preset array representation channel includes the access change data channel, the pseudo-analysis representation vector includes an access change representation vector; when the preset array representation channel includes the entity relationship channel, the pseudo-analysis representation vector includes an entity relationship representation vector.

3. The method according to claim 2, characterized in that When the preset array representation channel includes the access change data channel, and the access change data channel includes a single attribute change channel and multiple attribute change channels, generating a representation array for the third multivariate event monitoring data set for each set monitoring time range in combination with the preset array representation channel to obtain a quasi-analysis representation vector for each set monitoring time range includes: Determine the specific behavior occurrence frequency and specific attribute value statistics when the single attribute corresponding to the single attribute change channel in the third multivariate event monitoring data set within each set monitoring time range is changed; performing binning and feature coding on the specific behavior occurrence frequency and specific attribute value statistics when the single attribute is changed, to obtain a third characterization vector for each set monitoring time range; Determining the specific behavior occurrence frequency and specific attribute value statistics when the multiple attributes corresponding to the multiple attribute change channels in the third multivariate event monitoring data set within each set monitoring time range are changed; performing binning and feature coding on the specific behavior occurrence frequencies and specific attribute value statistics when the multiple attributes are changed, to obtain a fourth characterization vector for each set monitoring time range; Vector fusion is performed on the third characterization vector of each set monitoring time range and the fourth characterization vector of each set monitoring time range to obtain the access change characterization vector of each set monitoring time range.

4. The method according to claim 2, characterized in that When the preset array representation channel includes the entity relationship channel, and the entity relationship channel includes a setting related entity channel and a setting negative feedback entity channel, the representation array generation is performed on the third multivariate event monitoring data set of each set monitoring time range in combination with the preset array representation channel to obtain a quasi-analysis representation vector for each set monitoring time range, including: Determining the number and ratio of setting-related entities corresponding to the setting-related entity channels in the third multivariate event monitoring data set within each setting monitoring time range; performing binning and feature coding on the set number of related entities and the set ratio of related entities to obtain a fifth characterization vector for each set monitoring time range; Determining the number of negative feedback entities corresponding to the set negative feedback entity channel in the third multivariate event monitoring data set within each set monitoring time range; performing binning and feature encoding on the number of negative feedback entities to obtain a sixth characterization vector for each set monitoring time range; Vector fusion is performed on the fifth characterization vector of each set monitoring time range and the sixth characterization vector of each set monitoring time range to obtain the entity relationship characterization vector of each set monitoring time range.

5. The method according to claim 1, wherein The network security threat identification network is trained by the following steps: Acquire multiple second multivariate event monitoring data sets in set monitoring periods before and after the multiple preset training data sets and a network security threat priori marker indicating whether each of the set training data sets contains a network security threat; For each of the set training data sets, generating a representation array for the plurality of second multivariate event monitoring data sets in combination with the set monitoring period, the set monitoring time range, and the preset array representation channel to obtain a training data set representation array of the preset training data set; Performing network security threat identification on the training data set representation array based on the initialized neural network to obtain a network security threat identification result of the preset training data set; If the network security threat identification result is inconsistent with the network security threat priori label, repeatedly debugging the network parameter variables of the initialized neural network in combination with the training error value of the initialized neural network until a preset debugging stop condition is met; Determining the initialized neural network after debugging as the network security threat identification network; The plurality of preset training data sets include a plurality of active training data sets and a plurality of passive training data sets, the plurality of active training data sets include a plurality of past multivariate network security event monitoring data sets, and the network security threat priori markers of the active training data sets indicate that the active training data sets do not contain network security threats; The multiple negative training data sets include multiple past multivariate event monitoring data sets containing set network vulnerabilities, and the network security threat priori labels of the negative training data sets indicate that the negative training data sets contain network security threats.

6. The method according to claim 1, characterized in that The multiple active training data sets are collected by the following steps: Obtaining a first set of monitoring data sets corresponding to a plurality of past multivariate event monitoring data sets of network security; Selecting, from the first monitoring data set, past multivariate event monitoring data sets that have multiple multivariate event monitoring data sets in a set monitoring period before and after the first monitoring data set to obtain a second monitoring data set set; arbitrarily extracting the second monitoring data set to obtain the plurality of active training data sets; The plurality of negative training data sets are collected by the following steps, including: Obtaining a third monitoring dataset set including a plurality of past multivariate event monitoring datasets having a set network vulnerability; Classifying the past multivariate event monitoring datasets in the third monitoring dataset set in combination with entity types in the past multivariate event monitoring dataset to obtain a plurality of fourth monitoring dataset sets; Each fourth monitoring data set is randomly sampled to obtain the multiple negative training data sets.

7. The method according to any one of claims 1 to 6, characterized in that When the initialized neural network includes a first detection component and a second detection component, the first detection component includes a plurality of first detection modules, and the second detection component includes a second detection module, the network security threat identification network is used to perform network security threat identification on the target analysis representation array to obtain a network security threat identification result of whether the target multivariate event monitoring data set contains a network security threat, including: Performing network security threat identification on the to-be-analyzed characterization array based on the multiple first detection modules to obtain multiple network security threat identification results; Performing an integration operation on the multiple network security threat identification results to obtain an integration result; The network security threat identification is performed on the integration result based on the second detection module to obtain the network security threat identification result.

8. A computer system, characterized in that: include: one or more processors; Memory; one or more computer programs; The one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processors, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Data information security processing method and system

    CN117473571A

  • Network security threat identification method and system based on multivariate event analysis

    CN117792801A