Internet attack detection method and device, non-volatile storage medium, and processor
Through the unsupervised learning model, a learning model is generated for detecting distributed denial of service attacks, which solves the problems of high labor costs and long learning time in the existing technology, and realizes accurate judgment of real-time update attacks and improves industrial Internet network security.
Patent Information
- Application Number
- CN202211152962.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-09-21
AI Technical Summary
When detecting distributed denial of service attacks, the prior art needs to extract feature data through supervised learning, resulting in high labor costs and long learning time, and the inability to accurately judge continuous update attacks.
By determining the two-dimensional vector of the Internet protocol address to be detected within the first time range, an unsupervised learning model is used to cluster the two-dimensional sample vectors of multiple Internet protocol addresses in the continuous time range, and a first learning model is generated to analyze the two-dimensional vector, and then determining whether there is an Internet attack.
It realizes accurate judgment of real-time updates and diverse distributed denial of service attacks, reduces labor costs and learning time, and improves the security of industrial Internet networks.
Smart Images

Figure CN115499226B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of network security, and specifically, to a method and device for detecting Internet attacks, a non-volatile storage medium, and a processor. Background Art
[0002] In the era of Industry 4.0, resources, information, objects and people are bound together through the Internet, and information security faces new challenges. Among them, Distributed Denial of Service (DDoS) attacks are one of the current mainstream attack methods.
[0003] In the related technologies, products that defend against distributed denial of service attacks, such as routers and industrial firewalls, detect attacks by performing traffic statistics per unit time and network 5-tuple statistics from the dimensions of network layer 3, 4, and 7 protocols; however, they do not defend against distributed denial of service attacks from the dimensions of public industrial protocols. Industrial control systems communicate through public industrial protocols, such as the OPC (OLE for Process Control) protocol. Attackers can use public industrial protocols to interact with industrial control systems, causing them to perform a large number of invalid operations, thereby exhausting host resources and affecting normal service capabilities.
[0004] The related technology for detecting distributed denial of service attacks is to extract the characteristic data of distributed denial of service, perform deep learning on the characteristic data of distributed denial of service to obtain a model, and then perform matching operations and detect attacks based on real-time data and the model. The related technology for detecting distributed denial of service attacks has the following defects: 1. The industrial distributed denial of service characteristic data used for deep learning is manually collected, which belongs to supervised learning, with high labor costs and long learning time; 2. The model after learning is only a priori model and cannot make accurate judgments on continuously updated attacks.
[0005] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention
[0006] The embodiments of the present application provide a method and device for detecting Internet attacks, a non-volatile storage medium, and a processor, so as to at least solve the technical problems that the distributed denial of service attack intrusion detection method needs to use supervised learning to extract feature data, resulting in high labor costs, long learning time, and inability to make accurate judgments on continuously updated distributed denial of service attacks.
[0007] According to one aspect of an embodiment of the present application, a method for detecting an Internet attack is provided, comprising: determining a two-dimensional vector of an Internet Protocol address to be detected within a first time range, wherein the two-dimensional vector is a feature vector determined based on the information entropy of an error code of reply data and the information entropy of a request code of the Internet Protocol address to be detected within the first time range; analyzing the two-dimensional vector using a first learning model to obtain a detection result, wherein the first learning model is obtained by training clustering results of two-dimensional sample vectors of multiple Internet Protocol addresses within a continuous time range, and the two-dimensional sample vector includes: a vector determined by the information entropy of an error code of reply data and the information entropy of a request code within each time range within the continuous time range; wherein, when the distance between the two-dimensional vector and the center of a vector cluster corresponding to the first learning model is greater than a target distance, the detection result is determined to be a first detection result, and when the distance between the two-dimensional vector and the center of the vector cluster is not greater than the target distance, the detection result is determined to be a second detection result, wherein the first detection result is used to indicate that an Internet attack exists within the first time range, and the second detection result is used to indicate that an Internet attack does not exist within the first time range.
[0008] Optionally, before the first learning model is used to analyze the two-dimensional vector and obtain the detection result, the method also includes: traversing the second learning model located in the time range before the first time range in sequence from back to front until the absence of Internet attack is detected in the time range before the first time range, and using the second learning model corresponding to the time range in which the absence of Internet attack is detected as the first learning model.
[0009] Optionally, the two-dimensional sample vector is converted into a first sample vector and a second sample vector, wherein the first sample vector is a vector corresponding to the information entropy of the error code, and the second sample vector is a vector corresponding to the information entropy of the request code; when the number of first sample vectors is greater than the number of second sample vectors, the vector cluster formed by the first sample vectors is taken as the clustering result; when the number of first sample vectors is less than the number of second sample vectors, the vector cluster formed by the second sample vectors is taken as the clustering result.
[0010] Optionally, before determining the two-dimensional vector of the Internet Protocol address to be detected within the first time range, the method also includes: obtaining an error code of a reply data packet during a server reply process, obtaining a request code of a request data packet during a client request process, and assigning values to the error code and the request code; obtaining information entropy of the error code based on the assignment result corresponding to the error code and the number of occurrences of the error code within the first time range; obtaining information entropy of the request code based on the assignment result corresponding to the request code and the number of occurrences of the request code within the first time range.
[0011] Optionally, the server and the client exchange data based on a public industrial control protocol.
[0012] Optionally, the reply data packet and the request data packet are network data packets mirrored by a router, wherein the router is a router protected by a firewall, and the router and the client perform data exchange based on a public industrial control protocol.
[0013] Optionally, the error code includes at least one of the following: node does not exist, subscription failed, no read permission, no write permission, internal error, insufficient memory, encoding failure, decoding failure, and the request code includes at least one of the following: traversing nodes, reading node values, reading node attributes, subscribing to nodes, subscribing to events, subscribing to alarms, writing node values, canceling subscription, executing functions, closing sessions.
[0014] According to another aspect of the embodiments of the present application, a non-volatile storage medium is provided, the storage medium includes a stored program, wherein when the program is run, the device where the storage medium is located is controlled to execute the above-mentioned Internet attack detection method.
[0015] According to yet another aspect of the embodiments of the present application, a processor is provided, which is used to run a program stored in a memory, wherein the above Internet attack detection method is executed when the program is running.
[0016] In an embodiment of the present application, a two-dimensional vector of an Internet Protocol address to be detected within a first time range is determined, wherein the two-dimensional vector is a feature vector determined based on the information entropy of an error code of reply data of the Internet Protocol address to be detected within the first time range and the information entropy of a request code; a first learning model is used to analyze the two-dimensional vector to obtain a detection result, wherein the first learning model is trained on clustering results of two-dimensional sample vectors of multiple Internet Protocol addresses within a continuous time range, and the two-dimensional sample vector includes: a vector determined by the information entropy of an error code of reply data and the information entropy of a request code within each time range within the continuous time range; wherein, when the distance between the two-dimensional vector and the center of the vector cluster corresponding to the first learning model is greater than the target distance, the detection result is determined to be the first detection result. When the distance between the two-dimensional vector and the center of the vector cluster is not greater than the target distance, the detection result is determined to be the second detection result, wherein the first detection result is used to indicate that there is an Internet attack within the first time range, and the second detection result is used to indicate that there is no Internet attack within the first time range. The distributed denial of service attacks carried out using public industrial protocols are detected by an unsupervised learning model, thereby achieving the purpose of making accurate judgments on real-time updated and diverse distributed denial of service attacks, thereby achieving the technical effect of ensuring the security of the industrial Internet network, and further solving the technical problems of high labor costs and long learning time caused by the need to use supervised learning to extract feature data in the distributed denial of service attack intrusion detection method, and being unable to make accurate judgments on the continuously updated distributed denial of service attacks. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0018] Figure 1 is a flow chart of a method for detecting Internet attacks according to an embodiment of the present application;
[0019] Figure 2 is a flow chart of another method for detecting Internet attacks according to an embodiment of the present application;
[0020] Figure 3 is a network environment diagram of a distributed denial of service attack detection service according to an embodiment of the present application;
[0021] Figure 4 is a clustering result diagram according to an embodiment of the present application;
[0022] Figure 5 is another clustering result diagram according to an embodiment of the present application;
[0023] Figure 6 is another clustering result diagram according to an embodiment of the present application;
[0024] Figure 7 is a structural diagram of an Internet attack detection device according to an embodiment of the present application;
[0025] Figure 8 It is a hardware structure block diagram of a computer terminal (or electronic device) according to an Internet attack detection method provided in an embodiment of the present application. DETAILED DESCRIPTION
[0026] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0028] According to an embodiment of the present application, a method embodiment of a method for detecting an Internet attack is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0029] Figure 1 is a flow chart of a method for detecting Internet attacks according to an embodiment of the present application, such as Figure 1 As shown, the method comprises the following steps:
[0030] Step S102, determining a two-dimensional vector of the Internet Protocol address to be detected within the first time range, wherein the two-dimensional vector is a feature vector determined based on the information entropy of the error code of the reply data of the Internet Protocol address to be detected within the first time range and the information entropy of the request code.
[0031] According to an optional embodiment of the present application, the error code information entropy and request code information entropy of the Internet Protocol address to be detected are calculated in real time every 5 seconds, that is, a two-dimensional vector [CodeS ti ,ReqS ti ], where ti is used to indicate the time, such as t1 for the first 5 seconds; CodeS ti and ReqS ti They represent the information entropy of the error code and the information entropy of the request code within the i-th 5 seconds. It should be noted that the first time range is a unit statistical time, and in the embodiment of the present application, the unit statistical time is 5 seconds.
[0032] Step S104, using the first learning model to analyze the two-dimensional vector to obtain a detection result, wherein the first learning model is trained by clustering results of two-dimensional sample vectors of multiple Internet Protocol addresses within a continuous time range, and the two-dimensional sample vector includes: a vector determined by the information entropy of the error code of the reply data and the information entropy of the request code within each time range within the continuous time range; wherein, when the distance between the two-dimensional vector and the center of the vector cluster corresponding to the first learning model is greater than the target distance, the detection result is determined to be a first detection result, and when the distance between the two-dimensional vector and the center of the vector cluster is not greater than the target distance, the detection result is determined to be a second detection result, wherein the first detection result is used to indicate that an Internet attack exists within the first time range, and the second detection result is used to indicate that an Internet attack does not exist within the first time range.
[0033] According to another optional embodiment of the present application, for each Internet Protocol address, the information entropy of the error code and the information entropy of the request code can be calculated within each unit time of 5 seconds, that is, a set of two-dimensional vectors of continuous time series can be obtained:
[0034] V ti =[CodeS ti ,ReqS ti ]
[0035] Among them, ti is used to indicate the time, such as t1 represents the first 5 seconds; CodeS ti and ReqS ti Respectively represent the information entropy of the error code and the information entropy of the request code within the i-th 5 seconds, such as: V t1 =[0.4232,0.5558], V t2 =[0.4124,0.5018], V t3 =[0.4099,0.5191], all V ti It is a learning sample of the first learning model of unsupervised learning.
[0036] According to an optional embodiment of the present application, the first learning model is obtained by training the clustering results of two-dimensional sample vectors of multiple two-dimensional Internet Protocol addresses within a continuous time range, and the continuous time range is a continuous time range within a statistical period. For example, data within 20 minutes of the Internet Protocol address is collected to obtain 240 two-dimensional vectors (each vector generation takes 5 seconds), where 20 minutes is a specified statistical period and the continuous time range is 5 consecutive seconds. K-means clustering operation is performed on the 240 samples, and the vector with more circular vectors and star vectors in each sample is retained.
[0037] Figure 4is a clustering result diagram according to an embodiment of the present application, the clustering result of the two-dimensional sample vector of the Internet Protocol address within a continuous time range is as follows Figure 4 As shown, the samples are clustered into two categories, one is circular vectors and the other is star vectors. If there are more circular vectors and fewer star points, all the circular vectors are retained as a model of the normal situation to detect attacks.
[0038] According to an optional embodiment of the present application, after analyzing the two-dimensional vector using the first learning model, when the distance between the two-dimensional vector and the center of the vector cluster corresponding to the first learning model is greater than the target distance, that is, it is isolated from the vector cluster, then a first detection result is generated, that is, an Internet attack occurred within the 5 seconds; when the distance between the two-dimensional vector and the center of the vector cluster is not greater than the target distance, that is, it is clustered with the vector cluster, then a second detection result is generated, that is, no Internet attack occurred within the 5 seconds. Because the information entropy of the error code characterizes the discrete degree of the server's processing events within 5 seconds. When the server is attacked by DDoS using the OPC protocol, such as "operating on non-existent nodes", the events it processes will be more concentrated, so the discrete degree of the error codes replied within the specified time range will also be concentrated, and its information entropy will be higher, which is different from the normal value.
[0039] Figure 5 is another clustering result diagram according to an embodiment of the present application, such as Figure 5 As shown, the star vector is a two-dimensional vector to be detected. After the star vector is analyzed by the first learning model, the star vector is isolated and clustered into one category. At this time, it is considered that an Internet attack has occurred in the 5 seconds.
[0040] Figure 6 is another clustering result diagram according to an embodiment of the present application, such as Figure 6 As shown, the "X"-shaped vector is the detected two-dimensional vector. After the "X"-shaped vector is analyzed using the first learning model, the "X"-shaped vector and the two-dimensional vector of the first learning model are clustered into two categories. The "X"-shaped vector is in the star-shaped vector category, and it is considered that no Internet attack occurred in the 5 seconds.
[0041] In some optional embodiments of the present application, the model of each Internet Protocol address is calculated every 20 minutes. The current 20-minute detection uses the model calculated in the previous 20 minutes, and the current 20-minute data is used to calculate a new model for the next 20-minute detection. If an attack occurs in the current 20 minutes, the model calculation is not performed, and the next 20-minute monitoring model uses the most recently calculated model.
[0042] According to the above steps, the distributed denial of service attacks carried out using public industrial protocols are detected through unsupervised learning models, so as to achieve the purpose of making accurate judgments on the real-time updated and diverse distributed denial of service attacks, thereby achieving the technical effect of ensuring the security of the industrial Internet network, and further solving the technical problems of high labor costs and long learning time caused by the need to use supervised learning to extract feature data for distributed denial of service attack intrusion detection methods, and being unable to make accurate judgments on continuously updated distributed denial of service attacks.
[0043] According to an optional embodiment of the present application, before the two-dimensional vector is analyzed by the first learning model and the detection result is obtained, the method also includes: traversing the second learning model located in the time range before the first time range in sequence from back to front until the absence of Internet attack is detected in the time range before the first time range, and using the second learning model corresponding to the time range in which the absence of Internet attack is detected as the first learning model.
[0044] According to another optional embodiment of the present application, when the detection result of the learning model corresponding to the n-1 time range is that there is no Internet attack, the learning model corresponding to the n-1 time range is used to analyze the two-dimensional vector of the n time range; when the detection result of the learning model corresponding to the n-1 time range is that there is an Internet attack, the two-dimensional vector of the n time range is analyzed using the learning model that is closest to the previous n time range and whose detection result is that there is no Internet attack, where n is a positive number.
[0045] The current 20-minute detection uses the model calculated in the previous 20 minutes, and uses the current 20-minute data to calculate a new model for the next 20-minute detection. If an attack occurs in the current 20 minutes, the model calculation is not performed, and the monitoring model for the next 20 minutes uses the most recently calculated model without an attack. In other words, the model used by the Internet Protocol address to be detected is the latest model without an attack.
[0046] Through the above steps, the present application uses a clustering algorithm to implement an unsupervised learning generation model, autonomous learning, and no manual labeling costs. Moreover, the learning samples of the learning model of the present application come from real-time data, and the model is updated in real time to cope with changes in scenarios.
[0047] In some optional embodiments of the present application, the two-dimensional sample vector is converted into a first sample vector and a second sample vector, wherein the first sample vector is a vector corresponding to the information entropy of the error code, and the second sample vector is a vector corresponding to the information entropy of the request code; when the number of first sample vectors is greater than the number of second sample vectors, the vector cluster formed by the first sample vectors is taken as the clustering result; when the number of first sample vectors is less than the number of second sample vectors, the vector cluster formed by the second sample vectors is taken as the clustering result.
[0048] According to another optional embodiment of the present application, data of Internet Protocol addresses within 20 minutes are collected to obtain 240 two-dimensional sample vectors. A clustering operation is performed on the 240 two-dimensional sample vectors to obtain a first sample vector and a second sample vector.
[0049] Figure 4 is a clustering result diagram according to an embodiment of the present application, the clustering result of the two-dimensional sample vector of the Internet Protocol address within a continuous time range is as follows Figure 4 As shown, the samples are clustered into two categories, one is the circular vector, that is, the vector corresponding to the information entropy of the error code, and the other is the star vector, that is, the vector corresponding to the information entropy of the request code. If there are many circular vectors and few star points, all the circular vectors will be retained as a model of the normal situation to detect attacks.
[0050] In some optional embodiments of the present application, before determining the two-dimensional vector of the Internet Protocol address to be detected within the first time range, the method also includes: obtaining an error code of a reply data packet during a server reply process, obtaining a request code of a request data packet during a client request process, and assigning values to the error code and the request code; obtaining information entropy of the error code based on the assignment result corresponding to the error code and the number of occurrences of the error code within the first time range; obtaining information entropy of the request code based on the assignment result corresponding to the request code and the number of occurrences of the request code within the first time range.
[0051] As another optional embodiment of the present application, the public industrial protocol server responds to the public industrial protocol client request and generates a reply, and the reply data packet contains error codes representing various meanings. The information entropy of the error code of the server replying to a client Internet Protocol address within 5 seconds is counted, and this value is a characteristic value. The calculation formula of information entropy is as follows:
[0052]
[0053] Among them, P i The weight of the number of occurrences of the i-th error code or request code within the specified time range, where i is a positive integer.
[0054] For example, within 5 seconds, the public industrial protocol server has the following public industrial protocol reply:
[0055]
[0056]
[0057] For calculation convenience, a value is defined for each common industrial protocol reply:
[0058] Error code Numeric success 100 Node does not exist 101 Inconsistent data types 102 Subscription failed 103 No read permission 104 No write permission 105 Encryption method not supported 106 The number of sessions has exceeded the upper limit 107 Invalid username or password 108 Certificate not trusted 109 The node already exists 110 Internal error 111 Out of Memory 112 Encoding failed 113 Decoding failed 114 The message length exceeds the upper limit 115 Operation timed out 116 The request does not carry an operation object 117 Data length exceeds the limit 118 Data value out of range 119 The operation is not completed 120 Subscription deleted 121 The safe passage has been closed 122 The token has expired 123 Illegal message header 124 Unknown error 125
[0059] Then, within the 5 seconds, the information entropy of the error code of "201.192.1.10" is calculated as follows:
[0060] H(U)=-(8 / 24)×log(8 / 24)-(12 / 24)×log(12 / 24)-(4 / 24)×log(4 / 24)=0.4232
[0061] The information entropy of the reply error code represents the discrete degree of the server's processing events within 5 seconds. When the server is attacked by DDoS using public industrial protocols, such as "operating on non-existent nodes", the events it processes will be more concentrated, so the discrete degree of the error codes replied within a specified time range will also be concentrated, and its information entropy will be higher and different from the normal value.
[0062] The information entropy from public industrial protocol client requests within 5 seconds is counted. This value is a characteristic value.
[0063] For example, within 5 seconds, the public industrial protocol server receives the following public industrial protocol request:
[0064] IP Request Code Occurrence 201.192.1.10 Subscribe to Node(4) 4 201.192.1.10 Read node attributes (3) 8 201.192.1.10 Read node value (2) 6 201.192.1.10 Traversing nodes(1) 2
[0065] For calculation convenience, a value is defined for each public industrial protocol request:
[0066] Request Code Numeric Traversing nodes 1 Read node value 2 Read node attributes 3 Subscribe to Node 4 Subscribing to Events 5 Subscribe to alerts 6 Write node value 7 Establish a session using username and password 8 Using certificates to establish a secure channel 9 Unsubscribe 10 Execute function 11 Close Session 12 Read historical data 13 unknown 14
[0067] Then, within the 5 seconds, the information entropy of the request code "201.192.1.10" is calculated as follows:
[0068] H(U)=-(4 / 20)×log(4 / 20)-(8 / 20)×log(8 / 20)-(6 / 20)×log(6 / 20)-(2 / 20)×log(2 / 20)=0.5558
[0069] According to an optional embodiment of the present application, the server and the client perform data exchange based on a public industrial control protocol.
[0070] The related technologies can only cover the scenarios of conventional attacks and cannot detect distributed denial of service attacks using public industrial protocols. The server and client of this application interact with data based on the public industrial control protocol and can detect DDoS attacks using public industrial protocols.
[0071] In some optional embodiments of the present application, the reply data packet and the request data packet are network data packets mirrored by a router, wherein the router is a router protected by a firewall, and the router and the client perform data exchange based on a public industrial control protocol.
[0072] Figure 3 is a network environment diagram of a distributed denial of service attack detection service according to an embodiment of the present application, such as Figure 3 As shown, the method provided by the present application is implemented in the "DDoS attack detection service". It is deployed behind the firewall and can collect the required feature content from the network data packets mirrored by the router for learning modeling and attack detection.
[0073] In some optional embodiments of the present application, the error code includes at least one of the following: node does not exist, subscription failed, no read permission, no write permission, internal error, insufficient memory, encoding failure, decoding failure, and the request code includes at least one of the following: traversing nodes, reading node values, reading node attributes, subscribing to nodes, subscribing to events, subscribing to alarms, writing node values, canceling subscription, executing functions, closing sessions.
[0074] Figure 2 is a flow chart of another method for detecting Internet attacks according to an embodiment of the present application, such as Figure 2 As shown, the model of each IP is calculated once in each statistical period. The detection in the current statistical period uses the model calculated in the previous statistical period, and uses the data of the current statistical period to calculate a new model for the detection in the next statistical period. If an attack occurs in the current statistical period, the model calculation is not performed, and the monitoring model of the next statistical period continues to use the most recently calculated model. For example, calculation is performed every 20 minutes. The detection in the current 20 minutes uses the model calculated in the previous 20 minutes, and uses the data of the current 20 minutes to calculate a new model for the detection in the next 20 minutes. If an attack occurs in the current 20 minutes, the model calculation is not performed, and the monitoring model of the next 20 minutes continues to use the most recently calculated model.
[0075] Through the above steps, the learning samples of the model of the present application come from real-time data, and the model is updated in real time to cope with scene changes.
[0076] Figure 7 is a structural diagram of a positioning device for a probe according to an embodiment of the present application, such as Figure 7 As shown, the device comprises:
[0077] A determination module 70, configured to determine a two-dimensional vector of the Internet Protocol address to be detected within a first time range, wherein the two-dimensional vector is a feature vector determined based on information entropy of an error code of reply data of the Internet Protocol address to be detected within the first time range and information entropy of a request code;
[0078] The detection module 72 is used to analyze the two-dimensional vector using a first learning model to obtain a detection result, wherein the first learning model is obtained by training the clustering results of the two-dimensional sample vectors of multiple two-dimensional Internet Protocol addresses within a continuous time range, and the two-dimensional sample vector includes: a vector determined by the information entropy of the error code and the information entropy of the request code of the reply data within each time range within the continuous time range; wherein, when the distance between the two-dimensional vector and the center of the vector cluster corresponding to the first learning model is greater than the target distance, the detection result is determined to be a first detection result, and when the distance between the two-dimensional vector and the center of the vector cluster is not greater than the target distance, the detection result is determined to be a second detection result, wherein the first detection result is used to indicate that an Internet attack exists within the first time range, and the second detection result is used to indicate that an Internet attack does not exist within the first time range.
[0079] It should be noted that Figure 7 The preferred implementation of the illustrated embodiment can be found in Figure 1 The relevant description of the illustrated embodiment will not be repeated here.
[0080] Figure 8 According to an embodiment of the present application, a hardware structure block diagram of a computer terminal (or electronic device) for a method for detecting an Internet attack is provided. Figure 8 As shown, the computer terminal 80 (or electronic device 80) may include one or more (802a, 802b, ..., 802n are used to illustrate) processors (the processor may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 804 for storing data, and a transmission module 806 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 8 The structure shown is for illustration only and does not limit the structure of the above electronic device. Figure 8 More or fewer components as shown, or with Figure 8 Different configurations are shown.
[0081] It should be noted that the one or more processors and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 80 (or electronic device). As described in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0082] The memory 804 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the detection method of Internet attacks in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 804, that is, realizing the above-mentioned detection method of Internet attacks. The memory 804 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 804 may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the computer terminal 80 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0083] The transmission module 806 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 80. In one example, the transmission device 806 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 806 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0084] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 80 (or electronic device).
[0085] It should be noted that, in some optional embodiments, the above Figure 8 The computer device (or electronic device) shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. It should be noted that Figure 8This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the above-described computer device (or electronic device).
[0086] It should be noted that Figure 8 The electronic device shown is used to perform Figure 1 The detection method of Internet attacks shown in the figure, therefore the relevant explanations in the execution method of the above command are also applicable to the electronic device and will not be repeated here.
[0087] An embodiment of the present application further provides a non-volatile storage medium, which includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute the above-mentioned Internet attack detection method.
[0088] A program for a non-volatile storage medium to perform the following functions: determining a two-dimensional vector of an Internet Protocol address to be detected within a first time range, wherein the two-dimensional vector is a feature vector determined based on the information entropy of an error code of reply data and the information entropy of a request code of the Internet Protocol address to be detected within the first time range; analyzing the two-dimensional vector using a first learning model to obtain a detection result, wherein the first learning model is obtained by training the clustering results of two-dimensional sample vectors of multiple Internet Protocol addresses within a continuous time range, and the two-dimensional sample vector includes: a vector determined by the information entropy of the error code of reply data and the information entropy of the request code within each time range within the continuous time range; wherein when the distance between the two-dimensional vector and the center of the vector cluster corresponding to the first learning model is greater than a target distance, determining the detection result to be a first detection result, and when the distance between the two-dimensional vector and the center of the vector cluster is not greater than the target distance, determining the detection result to be a second detection result, wherein the first detection result is used to indicate that an Internet attack exists within the first time range, and the second detection result is used to indicate that an Internet attack does not exist within the first time range.
[0089] The embodiment of the present application further provides a processor, which is used to run a program stored in a memory, wherein the above-mentioned Internet attack detection method is executed when the program is running.
[0090] The processor is used to run a program that performs the following functions: determining a two-dimensional vector of an Internet Protocol address to be detected within a first time range, wherein the two-dimensional vector is a feature vector determined based on the information entropy of an error code of reply data and the information entropy of a request code of the Internet Protocol address to be detected within the first time range; using a first learning model to analyze the two-dimensional vector to obtain a detection result, wherein the first learning model is obtained by training the clustering results of two-dimensional sample vectors of multiple Internet Protocol addresses within a continuous time range, and the two-dimensional sample vector includes: a vector determined by the information entropy of an error code of reply data and the information entropy of a request code within each time range within the continuous time range; wherein when the distance between the two-dimensional vector and the center of a vector cluster corresponding to the first learning model is greater than a target distance, determining the detection result to be a first detection result, and when the distance between the two-dimensional vector and the center of the vector cluster is not greater than the target distance, determining the detection result to be a second detection result, wherein the first detection result is used to indicate that an Internet attack exists within the first time range, and the second detection result is used to indicate that an Internet attack does not exist within the first time range.
[0091] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0092] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0093] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0094] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0095] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0096] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the relevant technology or all or part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk, etc. Various media that can store program codes.
[0097] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for detecting Internet attacks, characterized in that: include: Determine a two-dimensional vector of the Internet Protocol address to be detected within a first time range, wherein the two-dimensional vector is a feature vector determined according to information entropy of an error code of reply data of the Internet Protocol address to be detected within the first time range and information entropy of a request code; The two-dimensional vector is analyzed by using a first learning model to obtain a detection result, wherein the first learning model is obtained by training a clustering result of two-dimensional sample vectors of a plurality of Internet Protocol addresses within a continuous time range, and the two-dimensional sample vector includes: a vector determined by the information entropy of an error code of reply data and the information entropy of a request code within each time range within the continuous time range; Wherein, when the distance between the two-dimensional vector and the center of the vector cluster corresponding to the first learning model is greater than the target distance, the detection result is determined to be a first detection result, and when the distance between the two-dimensional vector and the center of the vector cluster is not greater than the target distance, the detection result is determined to be a second detection result, wherein the first detection result is used to indicate that an Internet attack exists within the first time range, and the second detection result is used to indicate that an Internet attack does not exist within the first time range; The two-dimensional sample vector is converted into a first sample vector and a second sample vector, wherein the first sample vector is a vector corresponding to the information entropy of the error code, and the second sample vector is a vector corresponding to the information entropy of the request code; when the number of the first sample vectors is greater than the number of the second sample vectors, the vector cluster formed by the first sample vectors is used as the clustering result; when the number of the first sample vectors is less than the number of the second sample vectors, the vector cluster formed by the second sample vectors is used as the clustering result.
2. The method according to claim 1, characterized in that Before analyzing the two-dimensional vector using the first learning model to obtain the detection result, the method further includes: The second learning model in the time range before the first time range is traversed sequentially from back to front until no Internet attack is detected in the time range before the first time range, and the second learning model corresponding to the time range in which no Internet attack is detected is used as the first learning model.
3. The method according to claim 1, characterized in that Before determining the two-dimensional vector of the Internet Protocol address to be detected within the first time range, the method further includes: Obtain an error code of a reply data packet during a server reply process, obtain a request code of a request data packet during a client request process, and assign values to the error code and the request code; Obtaining information entropy of the error code according to the assignment result corresponding to the error code and the number of occurrences of the error code within the first time range; The information entropy of the request code is obtained according to the assignment result corresponding to the request code and the number of occurrences of the request code within the first time range.
4. The method according to claim 3, characterized in that The server and the client perform data exchange based on a public industrial control protocol.
5. The method according to claim 3, characterized in that: The reply data packet and the request data packet are network data packets mirrored by a router, wherein the router is a router protected by a firewall, and the router and the client perform data exchange based on a public industrial control protocol.
6. The method according to claim 5, characterized in that The error code includes at least one of the following: node does not exist, subscription failed, no read permission, no write permission, internal error, insufficient memory, encoding failure, decoding failure; the request code includes at least one of the following: traversing nodes, reading node values, reading node attributes, subscribing to nodes, subscribing to events, subscribing to alarms, writing node values, canceling subscriptions, executing functions, and closing sessions.
7. A device for detecting Internet attacks, characterized in that: include: A determination module, used to determine a two-dimensional vector of an Internet Protocol address to be detected within a first time range, wherein the two-dimensional vector is a feature vector determined based on information entropy of an error code of reply data of the Internet Protocol address to be detected within the first time range and information entropy of a request code; A detection module, configured to analyze the two-dimensional vector using a first learning model to obtain a detection result, wherein the first learning model is obtained by training the clustering results of two-dimensional sample vectors of multiple two-dimensional Internet Protocol addresses within a continuous time range, and the two-dimensional sample vector includes: a vector determined by the information entropy of an error code and an information entropy of a request code for reply data within each time range within the continuous time range; wherein, when the distance between the two-dimensional vector and the center of the vector cluster corresponding to the first learning model is greater than a target distance, the detection result is determined to be a first detection result, and when the distance between the two-dimensional vector and the center of the vector cluster is not greater than the target distance, the detection result is determined to be a second detection result, wherein the first detection result is used to indicate that an Internet attack exists within the first time range, and the second detection result is used to indicate that an Internet attack does not exist within the first time range; The detection device for Internet attacks is also used to perform the following steps: converting the two-dimensional sample vector into a first sample vector and a second sample vector, wherein the first sample vector is a vector corresponding to the information entropy of the error code, and the second sample vector is a vector corresponding to the information entropy of the request code; when the number of the first sample vectors is greater than the number of the second sample vectors, taking the vector cluster formed by the first sample vectors as the clustering result; when the number of the first sample vectors is less than the number of the second sample vectors, taking the vector cluster formed by the second sample vectors as the clustering result.
8. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the Internet attack detection method according to any one of claims 1 to 6.
9. An electronic device, characterized in that: The invention comprises a processor and a memory, wherein the processor is used to run a program stored in the memory, wherein the program, when running, executes the method for detecting Internet attacks as described in any one of claims 1 to 6.
Citation Information
Patent Citations
DDOS attack detection method and device
CN111224916A
Distributed denial of service attack monitoring method, system and device and storage medium
CN112804230A