A message interception method, device, electronic device and storage medium

By obtaining the domain name or last character of segmented packets in advance in the domain name interception system, the problems of interception failure and memory consumption in the existing technology caused by large HTTP packet splitting are solved, and more efficient packet interception and performance improvement are achieved.

CN118555148BActive Publication Date: 2025-05-16HANGZHOU YOUYUN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411026579.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2025-05-16
Estimated Expiration
2044-07-29

AI Technical Summary

Technical Problem

When the existing domain name intercepting system processes HTTP packets larger than the maximum segment length, the Host field is split through TCP segmentation processing, resulting in the interception failure, which is inefficient and has a high memory consumption.

Method used

When receiving the first segmented message, if there is a Host field and the domain name is an incomplete domain name, the first information is obtained, including the segmented domain name; if there is no Host field, the second information is obtained, including the last set number of characters of the data part. The second segmented message is received and the target domain name is determined based on its data portion and the first information or the second information. If the target domain name is not in the preset domain name whitelist, send a connection reset message to the sender.

Benefits of technology

By obtaining the last characters of the segmented domain name or data part in advance, the splicing and identification process of segmented message domain names can be accelerated, the efficiency and performance of packet interception can be improved, and memory consumption can be reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118555148B_ABST
    Figure CN118555148B_ABST
Patent Text Reader

Abstract

The present application provides a message interception method, device, electronic device and storage medium. When the data part of a first segmented message conforms to the HTTP request format, if the data part has a Host field and an incomplete domain name, first information, i.e., the segmented domain name of the data part, is obtained; if the data part does not have a Host field, second information, i.e., the last set number of characters of the data part, is obtained; a second segmented message is received, and a target domain name is determined according to the data part of the second segmented message and the first or second information; if the target domain name does not exist in the domain name whitelist, a connection reset message is sent to the sender of the first segmented message, and the target domain name is determined according to the data part of the second segmented message and the segmented domain name or the last set number of characters in the data part of the first segmented message. This can not only accelerate the recognition and splicing process of the segmented message domain name, improve the message interception efficiency and interception performance, but also reduce memory consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network security technology, and in particular to a message interception method, device, electronic device and storage medium. Background Art

[0002] In the field of network security, domain name interception systems play a vital role. Current domain name interception systems usually use bypass mirroring mode to monitor HTTP messages, thereby intercepting malicious or unauthorized domain names. During data transmission, when the length of the HTTP message generated by the application layer is greater than the MSS (Maximum Segment Size), the HTTP message needs to be processed by TCP segmentation, that is, the HTTP message is split into multiple TCP segments. The data part of each TCP segment contains a part of the HTTP message, and each TCP segment corresponds to a segmented message.

[0003] When an HTTP message is split into multiple segmented messages for transmission, for example, when it is split into two segmented messages for transmission, the Host field may be split into the data part of the two segmented messages, or all of it may be located in the data part of the second segmented message. Currently, the domain name interception system reconstructs the complete HTTP message by splicing the data parts of all segmented messages, and then identifies and intercepts the complete HTTP message. However, after the complete HTTP message is reconstructed, the client may have received the response, resulting in interception failure. Therefore, this method not only has low message interception efficiency and cannot effectively intercept, but also causes unnecessary memory consumption to reconstruct the complete HTTP message. Summary of the invention

[0004] In view of this, the present application provides a message interception method, device, electronic device and storage medium to solve the deficiencies in the related art.

[0005] In a first aspect of the present application, a message interception method is provided, comprising:

[0006] In a case where the data portion of the received first segmented message conforms to the HTTP request format, if the data portion of the first segmented message has a Host field and the domain name contained therein is an incomplete domain name, obtaining first information, the first information at least including the segmented domain name in the data portion of the first segmented message;

[0007] If the Host field does not exist in the data part of the first segmented message, obtain second information, where the second information includes at least the last set number of characters of the data part of the first segmented message;

[0008] receiving a second segmented message, where the second segmented message and the first segmented message are continuous messages in the same direction in the same TCP stream;

[0009] Determine a target domain name according to the data portion of the second segment message and the first information or the second information;

[0010] When it is determined that the target domain name does not exist in the preset domain name whitelist, a connection reset message is sent to the sender of the first segment message.

[0011] According to one embodiment of the present application, the first information further includes a uniform resource identifier URI in the data portion of the first segmented message, and the method further includes:

[0012] Using the triplet information of the first segmented message and the first information as keywords, querying the first target node in a preset domain name prediction table, the triplet information including the source and destination IP addresses and the destination port;

[0013] If the first target node exists, and the interception flag in the first target node indicates that an interception operation is performed, a connection reset message is sent to the sender of the first segmented message.

[0014] According to one embodiment of the present application, the first information further includes a URI in a data portion of the first segmented message, and the method further includes:

[0015] Determine an interception mark according to the target domain name and the domain name whitelist, wherein the interception mark is used to indicate whether to perform an interception operation;

[0016] The interception mark is stored in a node of a preset domain name prediction table using the triplet information of the second segmented message and the first information as keywords, wherein the triplet information includes a source destination IP address and a destination port.

[0017] According to an embodiment of the present application, a first timestamp is further stored in the node of the domain name prediction table, wherein the first timestamp is a timestamp when the interception mark is stored in the node of the domain name prediction table using the triple information of the second segmented message and the first information as keywords; the method further includes:

[0018] Determine whether the difference between the first timestamp and the current timestamp in each node of the domain name prediction table is greater than a preset value;

[0019] If the difference between the first timestamp and the current timestamp in a node is greater than a preset value, the node is deleted from the domain name prediction table.

[0020] According to one embodiment of the present application, the method further includes:

[0021] The first information or the second information is stored in a node of a preset segmented domain name table using the four-tuple information of the first segmented message as a keyword, wherein the four-tuple information includes a source-destination IP address and a source-destination port.

[0022] According to an embodiment of the present application, determining the target domain name according to the data portion of the second segmented message and the first information or the second information includes:

[0023] Using the quadruple information of the second segment message as a keyword, querying the second target node in the segment domain name table;

[0024] The target domain name is determined according to the data portion of the second segmented message and the first information or the second information in the second target node.

[0025] According to one embodiment of the present application, the method further includes:

[0026] If the data part of the first segmented message has a Host field, determine whether there is a line terminator after the Host field, where the line terminator is used to separate the fields in the data part of the first segmented message;

[0027] If the line terminator does not exist, it is determined that the domain name included in the data portion of the first segmented message is an incomplete domain name.

[0028] According to one embodiment of the present application, the first information and the second information further include a target TCP sequence number, and the target TCP sequence number is determined based on the TCP sequence number and the length of the data part of the first segmented message.

[0029] According to one embodiment of the present application, the method further includes:

[0030] Determine whether the TCP sequence number of the second segment message is the same as the target TCP sequence number in the first information or the second information;

[0031] If they are the same, it is determined that the second segmented message and the first segmented message are continuous messages in the same direction in the same TCP stream.

[0032] In a second aspect of the present application, a message interception device is provided, comprising:

[0033] A first acquisition unit is configured to acquire first information when the data portion of the received first segmented message conforms to the HTTP request format and if the data portion of the first segmented message has a Host field and the domain name contained therein is an incomplete domain name, wherein the first information at least includes a segmented domain name in the data portion of the first segmented message;

[0034] A second acquiring unit, configured to acquire second information if the data part of the first segmented message does not have a Host field, the second information comprising at least a set number of characters at the end of the data part of the first segmented message;

[0035] A receiving unit, configured to receive a second segmented message, wherein the second segmented message and the first segmented message are continuous messages in the same direction in the same TCP stream;

[0036] a determining unit, configured to determine a target domain name according to a data portion of the second segmented message and the first information or the second information;

[0037] The interception unit is used to send a connection reset message to the sender of the first segment message when it is determined that the target domain name does not exist in the preset domain name whitelist.

[0038] According to one embodiment of the present application, the first information further includes a uniform resource identifier URI in the data portion of the first segmented message, and the device further includes:

[0039] A prediction unit, configured to query a first target node in a preset domain name prediction table using the triplet information of the first segmented message and the first information as keywords, wherein the triplet information includes a source IP address and a destination port;

[0040] If the first target node exists, and the interception flag in the first target node indicates that an interception operation is performed, a connection reset message is sent to the sender of the first segmented message.

[0041] According to one embodiment of the present application, the first information further includes a URI in a data portion of the first segmented message, and the apparatus further includes:

[0042] A storage unit, used to determine an interception mark according to the target domain name and the domain name whitelist, wherein the interception mark is used to indicate whether to perform an interception operation;

[0043] The interception mark is stored in a node of a preset domain name prediction table using the triplet information of the second segmented message and the first information as keywords, wherein the triplet information includes a source destination IP address and a destination port.

[0044] According to an embodiment of the present application, a first timestamp is further stored in the node of the domain name prediction table, and the first timestamp is a timestamp when the interception mark is stored in the node of the domain name prediction table using the triple information of the second segmented message and the first information as keywords; the device also includes:

[0045] An updating unit, used to determine whether the difference between the first timestamp and the current timestamp in each node of the domain name prediction table is greater than a preset value;

[0046] If the difference between the first timestamp and the current timestamp in a node is greater than a preset value, the node is deleted from the domain name prediction table.

[0047] In a third aspect of the present application, an electronic device is provided, including a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor is used to execute the machine executable instructions to implement the steps of the method proposed in the above embodiment.

[0048] In a fourth aspect of the present application, a machine-readable storage medium is provided, wherein the machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by a processor, the steps of the method proposed in the above embodiment are implemented.

[0049] In a fifth aspect of the present application, a computer program product is provided, comprising a computer program / instruction, which implements the steps of the method proposed in the above embodiment when executed by a processor.

[0050] It can be seen from the above technical solution that, when the data part of the received first segmented message conforms to the HTTP request format, if the data part of the first segmented message has a Host field and the domain name contained is an incomplete domain name, then the first information is obtained, and the first information at least includes the segmented domain name in the data part of the first segmented message; if the data part of the first segmented message does not have a Host field, then the second information is obtained, and the second information at least includes the last set number of characters of the data part of the first segmented message; the second segmented message is received, and the second segmented message and the first segmented message are continuous messages in the same direction in the same TCP stream, and the target domain name is determined according to the data part of the second segmented message and the first information or the second information. When it is determined that the target domain name does not exist in the preset domain name whitelist, a connection reset message is sent to the sender of the first segmented message, and the target domain name is determined by the data part of the second segmented message and the segmented domain name or the last set number of characters in the data part of the first segmented message. This can not only accelerate the splicing and identification process of the segmented message domain name, improve the message interception efficiency and interception performance, but also reduce memory consumption.

[0051] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1is a schematic diagram of an exemplary bypass mirror mode provided in an embodiment of the present application;

[0053] Figure 2 It is a flow chart of a message interception method provided in an embodiment of the present application;

[0054] Figure 3 It is a schematic diagram of a process of updating a domain name prediction table provided in an embodiment of the present application;

[0055] Figure 4 It is a flowchart of a message interception method provided by another embodiment of the present application;

[0056] Figure 5 It is a structural schematic diagram of a message interception device provided in an embodiment of the present application;

[0057] Figure 6 It is a schematic diagram of the hardware structure of an electronic device shown in an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0058] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0059] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "said" and "the" used in this application and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.

[0060] In order to enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application and to make the above-mentioned purposes, features and advantages of the embodiments of the present application more obvious and understandable, the technical solutions in the embodiments of the present application are further described in detail below in conjunction with the accompanying drawings.

[0061] In the field of network security, domain name interception systems play a vital role. Current domain name interception systems usually use bypass mirroring mode to monitor HTTP messages, thereby intercepting malicious or unauthorized domain names. Figure 1 As shown, Figure 1It is a schematic diagram of an exemplary bypass mirroring mode provided in an embodiment of the present application, wherein the switching device 102 can receive from a first port an HTTP request message sent by the client 101 when requesting a target service to the server 103, and forward the HTTP request message to the server 103 from a second port. The switching device 102 is configured in a bypass mirroring mode. In the bypass mirroring mode, the switching device 102 can also mirror the HTTP request message received from the first port to generate a completely identical mirrored message, and forward the mirrored message from a third port to the electronic device 104.

[0062] During data transmission, when the length of the HTTP message generated by the application layer is greater than the MSS (Maximum Segment Size), the HTTP message needs to be segmented by TCP, that is, the HTTP message is split into multiple TCP segments. The data part of each TCP segment contains a part of the HTTP message, and each TCP segment corresponds to a segmented message. For example, if the MTU (Maximum Transmission Unit) is 1500 bytes, then the MSS is usually 1500 minus the size of the IP and TCP headers (20 bytes each), that is, 1460 bytes. If the length of the HTTP message generated by the application layer exceeds 1460 bytes, it needs to be split into multiple TCP segments.

[0063] When an HTTP message is split into multiple segmented messages for transmission, for example, when it is split into two segmented messages for transmission, the Host field may be split into the data part of the two segmented messages, or all of it may be located in the data part of the second segmented message. Currently, the domain name interception system reconstructs the complete HTTP message by splicing the data parts of all segmented messages, and then identifies and intercepts the complete HTTP message. However, after the complete HTTP message is reconstructed, the client may have received the response, resulting in interception failure. Therefore, this method not only has low message interception efficiency and cannot effectively intercept, but also causes unnecessary memory consumption to reconstruct the complete HTTP message.

[0064] In view of this, an embodiment of the present application discloses a message interception method to solve the deficiencies in the related technology.

[0065] like Figure 2 As shown, Figure 2 The embodiment of the present application does not limit the execution subject of the message interception method process. Optionally, the execution subject can be any electronic device, such as a server for message interception, or a terminal that needs to perform message interception.

[0066] In some embodiments, the message interception method can be applied to Figure 1 In the electronic device 104 shown, the electronic device 104 can receive the mirror message sent by the client 101 when requesting the target service from the server 103.

[0067] Exemplarily, the electronic device may store a preset domain name whitelist, where the domain name whitelist refers to a list of domain names registered through domain name registration, and is used to allow specific domain names to be released when accessed.

[0068] The message interception method may include the following steps:

[0069] Step 201: When the data part of the received first segmented message conforms to the HTTP request format, if the data part of the first segmented message has a Host field and the domain name contained therein is an incomplete domain name, then obtain first information, which includes at least the segmented domain name in the data part of the first segmented message.

[0070] In some embodiments, a first segmented message is received, and the first segmented message can be determined by determining whether a data portion of the received initial segmented message conforms to an HTTP request format.

[0071] After receiving the initial segmented message, it can be determined whether the data portion of the initial segmented message conforms to the HTTP request format. If the data portion of the initial segmented message conforms to the HTTP request format, the initial segmented message is determined to be the first segmented message, thereby determining that the data portion of the received first segmented message conforms to the HTTP request format.

[0072] In some embodiments, after receiving the initial segment message, it is possible to determine whether the data portion of the initial segment message conforms to the HTTP request format by determining whether the data portion of the initial segment message includes a target field. For example, the target field may be an HTTP request method, a URI (Uniform Resource Identifier) ​​and / or an HTTP version.

[0073] If the data portion of the initial segmented message contains a target field, it can be determined that the data portion of the initial segmented message conforms to the HTTP request format, thereby determining that the initial segmented message is the first segmented message, thereby determining that the data portion of the received first segmented message conforms to the HTTP request format.

[0074] For example, the data portion of the initial segment message is:

[0075] GET / index.html HTTP / 1.1\r\nUser-Agent: curl / 7.29.0\r\nHost: www.te

[0076] It can be determined that the data portion of the initial segmented message includes the HTTP request method (GET), URI ( / index.html) and HTTP version (HTTP / 1.1), so it can be determined that the initial segmented message is the first segmented message, and thus it can be determined that the data portion of the first segmented message conforms to the HTTP request format.

[0077] Of course, the embodiments of the present application do not limit the specific method of determining whether the data portion of the first segment message conforms to the HTTP request format. For example, it can also be determined by syntax verification, regular expression matching, etc. whether the data portion of the first segment message conforms to the HTTP request format.

[0078] In some embodiments, the received first segmented message may be parsed to obtain target information, which may include source IP, destination IP, source port, destination port, TCP sequence number, data portion, and data portion length.

[0079] After receiving the first segmented message from the network card, the link layer is parsed first, and the header and tail of the Ethernet frame are removed to expose the data of the network layer; then the network layer (IP layer) is parsed to obtain the source IP and destination IP contained in the IP header, and then the transport layer (TCP layer) is parsed to obtain the source port, destination port, TCP sequence number, data part and data part length. Among them, the TCP sequence number is used to ensure the order and integrity of data during network transmission. Each TCP segment can carry a unique TCP sequence number, which indicates the position of the first byte of the data part in the TCP segment in the sender's data stream.

[0080] In some embodiments, the first information may further include a target TCP sequence number, where the target TCP sequence number is determined based on the TCP sequence number and the length of the data portion of the first segmented message.

[0081] Exemplarily, the target TCP sequence number can be obtained by adding the TCP sequence number of the first segment message to the length of the data portion. As for the specific role of the target TCP sequence number in the first information, see step 203, which will not be described in detail here.

[0082] In some embodiments, whether the data part of the first segmented message has a Host field can be determined by judging whether the data part of the first segmented message has a first target string. For example, the first target string can be "Host: " (with a space after the colon).

[0083] If the data portion of the first segmented message contains the first target character string, it can be determined that the data portion of the first segmented message contains the Host field.

[0084] For example, the data part of the first segment message is:

[0085] GET / index.html HTTP / 1.1\r\nUser-Agent: curl / 7.29.0\r\nHost: www.te

[0086] It can be determined that the data part of the first segmented message contains the first target string "Host: " (with a space after the colon), and therefore it can be determined that the data part of the first segmented message contains the Host field.

[0087] Of course, the embodiments of the present application do not limit the specific method of determining whether the Host field exists in the data part of the first segment message.

[0088] In some embodiments, whether the domain name contained in the data portion of the first segmented message is an incomplete domain name can be determined by judging whether there is a second target character string after the Host field of the data portion of the first segmented message. The incomplete domain name in the data portion is the segmented domain name in the data portion. The second target character string can be any character string used to identify field boundaries in a specific context, such as a line terminator "\r\n", "\r" or other specific end marks.

[0089] In some embodiments, if the data portion of the first segmented message has a Host field, determining whether there is a line terminator after the Host field, the line terminator being used to separate the fields in the data portion of the first segmented message;

[0090] If the line terminator does not exist, it is determined that the domain name included in the data part of the first segmented message is an incomplete domain name; if the line terminator exists, it is determined that the domain name included in the data part of the first segmented message is a complete domain name.

[0091] For example, the data part of the first segment message is:

[0092] GET / index.html HTTP / 1.1\r\nUser-Agent: curl / 7.29.0\r\nHost: www.te

[0093] It can be determined that there is no line terminator "\r\n" after the Host field in the data part of the first segmented message. Therefore, it can be determined that the domain name contained in the data part of the first segmented message is an incomplete domain name, and the incomplete domain name is "www.te", that is, the segmented domain name in the data part is "www.te".

[0094] For example, the data part of the first segment message is:

[0095] GET / HTTP / 1.1\r\nUser-Agent: curl / 7.29.0\r\nHost: www.test.com\r\nAc

[0096] It can be determined that there is a line terminator "\r\n" after the Host field in the data part of the first segment message, so it can be determined that the domain name included in the data part of the first segment message is a complete domain name, and the complete domain name is "www.test.com".

[0097] Of course, the embodiments of the present application do not limit the specific method of determining that the domain name included in the data portion of the first segment message is an incomplete domain name. For example, the domain name included in the data portion of the first segment message can also be determined as an incomplete domain name by regular expressions, DNS resolution, etc.

[0098] Step 202: If the Host field does not exist in the data portion of the first segmented message, obtain second information, where the second information includes at least the last set number of characters of the data portion of the first segmented message.

[0099] In an embodiment of the present application, if the data part of the first segmented message does not have a Host field, it indicates that the Host field is truncated or there is no Host field itself, and the last set number of characters of the data part of the first segmented message can be obtained.

[0100] In some embodiments, whether the data part of the first segmented message has a Host field can be determined by judging whether the data part of the first segmented message has a first target string. For example, the first target string can be "Host: " (with a space after the colon).

[0101] If the data portion of the first segmented message does not contain the first target character string, it can be determined that the data portion of the first segmented message does not contain the Host field.

[0102] For example, the data part of the first segment message is:

[0103] GET / index.html HTTP / 1.1\r\nUser-Agent: curl / 7.29.0\r\nHo

[0104] It can be determined that the data portion of the first segmented message does not contain the first target string "Host: " (there is a space after the colon), so it can be determined that the data portion of the first segmented message does not contain the Host field.

[0105] The above-mentioned set number can be preset according to actual needs, and the embodiments of the present application do not specifically limit the set number.

[0106] Exemplarily, the set number may be preset to 6.

[0107] In some embodiments, the set number may be within a target numerical range, and the target numerical range may be determined according to the number of characters in the first target character string and a preset threshold.

[0108] Exemplarily, the target value range may be greater than or equal to the above number of characters, and less than or equal to the sum of the above number of characters and a preset threshold.

[0109] For example, when the first target string is "Host: " (with a space after the colon) and the preset threshold is 3, the number of characters in the first target string is 6, and the set number can be within the target numerical range [6, 9], that is, the set number can be 6, 7, 8 or 9.

[0110] For example, when the first target character string is “Host: ” (with a space after the colon) and the preset threshold is 0, the set number may be 6.

[0111] In some embodiments, the second information may further include a target TCP sequence number, where the target TCP sequence number is determined based on the TCP sequence number and data portion length of the first segmented message.

[0112] Exemplarily, the target TCP sequence number can be obtained by adding the TCP sequence number of the first segment message to the length of the data portion. As for the specific role of the target TCP sequence number in the second information, see step 203, which will not be described in detail here.

[0113] In an embodiment of the present application, if the Host field does not exist in the data portion of the first segmented message, the second information is obtained, and the second information includes at least the last set number of characters of the data portion of the first segmented message. By obtaining the last set number of characters of the data portion of the first segmented message, the memory usage of the segmented domain name concatenation and identification process can be reduced, thereby reducing memory consumption.

[0114] Step 203: Receive a second segmented message, where the second segmented message and the first segmented message are continuous messages in the same direction in the same TCP stream.

[0115] In some embodiments, a second segmented message is received, and the second segmented message can be determined by determining whether a data portion of the received initial segmented message conforms to an HTTP request format.

[0116] After receiving the initial segment message, it can be determined whether the data portion of the initial segment message conforms to the HTTP request format. If the data portion of the initial segment message does not conform to the HTTP request format, the initial segment message is determined to be the second segment message.

[0117] In some embodiments, whether the data portion of the initial segment message conforms to the HTTP request format can be determined by determining whether the data portion of the initial segment message includes a target field. For example, the target field can be an HTTP request method, a URI (Uniform Resource Identifier) ​​and / or an HTTP version.

[0118] If the data portion of the initial segmented message does not include a target field, it can be determined that the data portion of the initial segmented message does not conform to the HTTP request format, and thus it can be determined that the initial segmented message is the second segmented message, that is, it can be determined that the second segmented message is received.

[0119] For example, the data portion of the initial segment message is:

[0120] st: www.test.com\r\nAccept: * / *\r\n\r\n

[0121] It can be determined that the data portion of the initial segmented message does not include the HTTP request method, URI, and HTTP version, so it can be determined that the initial segmented message is the second segmented message, that is, it can be determined that the second segmented message is received.

[0122] In some embodiments, the first information obtained in step 201 and the second information obtained in step 202 also include a target TCP sequence number, and the target TCP sequence number is determined based on the TCP sequence number and the length of the data portion of the first segmented message.

[0123] The second segment message and the first segment message are continuous messages in the same direction in the same TCP stream. Among them, "same direction" indicates that the second segment message and the first segment message are in the same communication direction, such as the sending direction or the receiving direction. "TCP stream" represents a logically continuous data sequence. The HTTP message generated by the application layer may be divided into multiple TCP segments during the data transmission process. These TCP segments are all in the same TCP stream. Each TCP segment has a unique TCP sequence number to ensure the order and integrity of the data when the receiving end reassembles the data. If the TCP sequence number of the second segment message is immediately followed by the TCP sequence number of the first segment message, this indicates that the second segment message is a direct successor to the first segment message, thereby confirming their continuity.

[0124] It can be determined in the following manner that the second segmented message and the first segmented message are continuous messages in the same direction in the same TCP stream.

[0125] Determine whether the TCP sequence number of the second segment message is the same as the target TCP sequence number in the first information or the second information.

[0126] If they are the same, it can be determined that the second segmented message and the first segmented message are continuous messages in the same direction in the same TCP stream.

[0127] Of course, the embodiments of the present application do not limit the specific method of determining that the second segmented message and the first segmented message are continuous messages in the same direction in the same TCP stream. For example, it is also possible to determine that the second segmented message and the first segmented message are continuous messages in the same direction in the same TCP stream through four-tuple information, segment offset, etc.

[0128] Step 204: Determine the target domain name according to the data portion of the second segment message and the first information or the second information.

[0129] In an embodiment of the present application, the domain name can be spliced ​​and identified based on the data portion of the second segmented message and the first information or the second information, so as to determine the target domain name, that is, determine the complete domain name.

[0130] In some embodiments, the data portion of the second segmented message may be concatenated and identified with the segmented domain name in the data portion of the first segmented message in the first information, thereby determining the target domain name.

[0131] In some embodiments, the segmented domain name in the data portion of the second segmented message can be identified and intercepted based on the first second target character string in the data portion of the second segmented message, and the segmented domain name in the data portion of the second segmented message can be concatenated with the segmented domain name in the data portion of the first segmented message in the first information to determine the target domain name.

[0132] The second target character string may be any character string used to identify field boundaries in a specific context, such as a line terminator “\r\n”, “\r” or other specific end marks.

[0133] In some embodiments, the data portion of the second segmented message may be concatenated and identified with the last set number of characters of the data portion of the first segmented message in the second information, thereby determining the target domain name.

[0134] In some embodiments, after the data portion of the second segmented message is concatenated with the last set number of characters of the data portion of the first segmented message in the second information, the target domain name can be determined by identifying the host field and the second target character string in the concatenated data.

[0135] Step 205: When it is determined that the target domain name does not exist in the preset domain name whitelist, a connection reset message is sent to the sender of the first segment message.

[0136] In an embodiment of the present application, the target domain name may be compared with a preset domain name whitelist, and when it is determined that the target domain name does not exist in the domain name whitelist, a connection reset message is sent to the sender of the first segmented message.

[0137] In an embodiment of the present application, when the data part of the received first segmented message conforms to the HTTP request format, if the data part of the first segmented message has a Host field and the domain name contained is an incomplete domain name, then first information is obtained, and the first information includes at least the segmented domain name in the data part of the first segmented message; if the data part of the first segmented message does not have a Host field, then second information is obtained, and the second information includes at least the last set number of characters of the data part of the first segmented message; a second segmented message is received, and the second segmented message and the first segmented message are continuous messages in the same direction in the same TCP stream, and a target domain name is determined according to the data part of the second segmented message and the first information or the second information, and when it is determined that the target domain name does not exist in the preset domain name whitelist, a connection reset message is sent to the sender of the first segmented message, and the target domain name is determined by the data part of the second segmented message and the segmented domain name or the last set number of characters in the data part of the first segmented message, which can not only accelerate the splicing and identification process of the segmented message domain name, improve the message interception efficiency and interception performance, but also reduce memory consumption.

[0138] In some embodiments, when the data portion of the received message conforms to the HTTP request format, if the data portion of the message contains a Host field and the domain name contained therein is a complete domain name, the complete domain name is compared with the domain name whitelist, and when it is determined that the complete domain name does not exist in the domain name whitelist, a connection reset message is sent to the sender of the message.

[0139] In some embodiments, whether the data portion of the message conforms to the HTTP request format can be determined by determining whether the data portion of the message includes a target field. For example, the target field can be an HTTP request method, a URI (Uniform Resource Identifier) ​​and / or an HTTP version.

[0140] In some embodiments, whether the data part of the message has a Host field can be determined by judging whether the data part of the message has a first target string. For example, the first target string can be "Host: " (with a space after the colon).

[0141] In some embodiments, whether the domain name included in the data portion of the message is an incomplete domain name can be determined by determining whether a second target string exists after the Host field of the data portion of the message. The second target string can be any string used to identify field boundaries in a specific context, such as a line terminator "\r\n", "\r" or other specific end markers.

[0142] In the embodiments of the present application, domain name identification and interception processing can also be performed on ordinary messages, that is, non-segmented messages, thereby achieving the universality of message interception.

[0143] In some embodiments, the electronic device may also store a preset segmented domain name table, which may query a node using the four-tuple information as a keyword, and the node may store the first information or the second information, wherein the four-tuple information includes the source and destination IP addresses and the source and destination ports.

[0144] In some embodiments, the node of the segmented domain name table includes at least a domain name field and a message field, wherein the domain name field is used to store the segmented domain name in the data part of the first segmented message in the first information, and the message field is used to store the last set number of characters in the data part of the first segmented message.

[0145] In some embodiments, the first information or the second information may be stored in a node of a preset segmented domain name table using the four-tuple information of the first segmented message as a keyword, wherein the four-tuple information includes a source-destination IP address and a source-destination port.

[0146] In some embodiments, the second target node may be queried in the segment domain name table using the quadruple information of the second segment message as a keyword;

[0147] The target domain name is determined according to the data portion of the second segmented message and the first information or the second information in the second target node.

[0148] In some embodiments, the first information and the second information may further include a target TCP sequence number, so the node may further include a next seq (Sequence number, sequence number) field, and the next seq field is used to store the target TCP sequence number.

[0149] In some embodiments, the first information may further include a URI in the data portion of the first segmented message, so the node may further include a request URI field, which is used to store the URI in the first information. As for the specific role of the URI in the data portion of the first segmented message, see below, which will not be described in detail here.

[0150] In an embodiment of the present application, the first information or the second information is stored in a node of a preset segmented domain name table by using the four-tuple information of the first segmented message as a keyword, and the second target node is queried in the segmented domain name table by using the four-tuple information of the second segmented message as a keyword. The target domain name is determined according to the data part of the second segmented message and the first information or the second information in the second target node. Since the first information contains the segmented domain name in the data part of the first segmented message, and the second information contains the last set number of characters of the data part of the first segmented message, the memory usage of the segmented domain name splicing and identification process can be reduced, thereby reducing memory consumption.

[0151] In some embodiments, the first information further includes a URI in a data portion of the first segmented message.

[0152] In some embodiments, the electronic device may also store a preset domain name prediction table, which may use the triple information and the first information (including the segment domain name and URI in the data portion of the first segment message) as keywords to query the node, and the node may contain at least an interception mark. The triple information includes the source and destination IP addresses and the destination port, and the interception mark is used to indicate whether to perform an interception operation.

[0153] In some embodiments, the first target node may be queried in a preset domain name prediction table using the triplet information of the first segmented message and the first information (including the segmented domain name and URI in the data portion of the first segmented message) as keywords, wherein the triplet information includes the source destination IP address and the destination port;

[0154] If the first target node exists, and the interception flag in the first target node indicates that an interception operation is performed, a connection reset message is sent to the sender of the first segmented message.

[0155] In an embodiment of the present application, the domain names of messages initiated with the same triplet information (source IP / destination IP / destination port), access domain name (segmented domain name in the data part of the first segmented message) and access URI (URI in the data part of the first segmented message) can be considered to be consistent, so that the domain name segmentation of this request can be predicted to be consistent with the previously recorded one through the source IP, destination IP, destination port, access domain name and access URI, and then the message interception and release can be accelerated in a predictive manner, thereby improving the message interception / release efficiency.

[0156] In some embodiments, the first information further includes a URI in a data portion of the first segmented message.

[0157] In some embodiments, the electronic device may also store a preset domain name prediction table, which may use the triple information and the first information (including the segment domain name and URI in the data portion of the first segment message) as keywords to query the node, and the node may contain at least an interception mark. The triple information includes the source and destination IP addresses and the destination port, and the interception mark is used to indicate whether to perform an interception operation.

[0158] In some embodiments, an interception mark may be determined according to the target domain name and the domain name whitelist, and the interception mark is used to indicate whether to perform an interception operation;

[0159] Using the triplet information of the second segmented message and the first information (including the segmented domain name and URI in the data part of the first segmented message) as keywords, the interception mark is stored in the node of the preset domain name prediction table, and the triplet information includes the source and destination IP addresses and destination ports.

[0160] Exemplarily, the interception mark can be determined by judging whether the target domain name exists in the domain name whitelist. If the target domain name exists in the domain name whitelist, the interception mark of the target domain name is a mark indicating that the interception operation is not performed, for example, the interception mark of the target domain name can be "0", "true", etc.; if the target domain name does not exist in the domain name whitelist, the interception mark of the target domain name is a mark indicating that the interception operation is performed, for example, the interception mark of the target domain name can be "1", "false", etc.

[0161] In some embodiments, after the interception mark is stored in the node of the preset domain name prediction table using the triplet information of the second segmented message and the first information (including the segmented domain name and URI in the data part of the first segmented message) as keywords, the node of the domain name prediction table stores information that can be used for interception / release prediction. For the next first segmented message, when the data part of the received first segmented message conforms to the HTTP request format, if the data part of the first segmented message has a Host field, and the domain name contained is an incomplete domain name, the triplet information of the first segmented message and the first information (including the segmented domain name and URI in the data part of the first segmented message) can be used as keywords to query the first target node in the domain name prediction table. If the first target node exists and the interception mark in the first target node indicates that an interception operation is to be performed, a connection reset message is sent to the sender of the first segmented message.

[0162] In an embodiment of the present application, the domain names of messages initiated with the same triplet information (source IP / destination IP / destination port), access domain name (segmented domain name in the data part of the first segmented message) and access URI (URI in the data part of the first segmented message) can be considered to be consistent, so that the domain name segmentation of this request can be predicted to be consistent with the previously recorded one through the source IP, destination IP, destination port, access domain name and access URI, and then the message interception and release can be accelerated in a predictive manner, thereby improving the message interception / release efficiency.

[0163] In some embodiments, a first timestamp may also be stored in the node of the domain name prediction table, wherein the first timestamp is a timestamp when the interception mark is stored in the node of the domain name prediction table using the triple information of the second segmented message and the first information as keywords;

[0164] In the case where the interception mark and the first timestamp are stored in the nodes of the domain name prediction table, it can be determined whether the difference between the first timestamp in each node of the domain name prediction table and the current timestamp is greater than a preset value;

[0165] If the difference between the first timestamp and the current timestamp in a node is greater than a preset value, the node is deleted from the domain name prediction table.

[0166] Exemplarily, the above-mentioned preset value can be set according to actual needs. The embodiments of the present application do not specifically limit the preset value. For example, the preset value can be 1 second, 2 seconds, etc.

[0167] like Figure 3 As shown, Figure 3This is a schematic diagram of a domain name prediction table update process provided by an embodiment of the present application, which starts a periodic task, executes once per second, traverses each node of the domain name prediction table every second, and determines whether the difference between the first timestamp in the node and the current timestamp is greater than a preset value, if not, ends, and if so, deletes the node from the domain name prediction table. Among them, the first timestamp is the timestamp when the interception mark is stored in the node of the domain name prediction table.

[0168] In an embodiment of the present application, a periodic task can be started to periodically detect and clean up the nodes in the domain name prediction table, because it is possible that a domain name was not in the domain name whitelist at the previous moment, but the domain name is in the domain name whitelist at the next moment. By periodically detecting and cleaning up the nodes in the domain name prediction table, the accuracy of message interception / release in a predictive manner can be improved.

[0169] See also Figure 4 , Figure 4 It is a flow chart of a message interception method provided by another embodiment of the present application. The message interception method can be applied to electronic devices, such as domain name interception systems, firewalls, etc. The electronic devices pre-store a domain name whitelist, a segmented domain name table, and a domain name prediction table.

[0170] Among them, the domain name whitelist refers to a list of domain names that have passed the domain name registration, which is used to allow specific domain names to be released when accessed. The content in the domain name whitelist is added by the user, will be solidified locally, and will be reloaded when the process is restarted.

[0171] The segmented domain name table can query nodes using four-tuple information as keywords. The node can contain a domain name field, a message field, a next seq field, and a request URI field. The above four-tuple information includes the source IP address and the source port.

[0172] The domain name prediction table can query nodes with triple information, access domain name (segmented domain name in the data part of the message) and access URI as keywords. The node can contain an interception mark and a first timestamp. The above triple information includes the source IP address and the destination port. The above first timestamp is the timestamp when the interception mark is stored in the node of the domain name prediction table. This table creates a blank table when the process is started.

[0173] The message interception method may include the following steps:

[0174] 1 Get the message from the network card and start parsing;

[0175] 2. Parse the network layer (IP layer) to obtain the source IP address and destination IP address;

[0176] 3. Parse the transport layer (TCP layer) to obtain the source port, destination port, TCP sequence number, data part, and data part length;

[0177] 4. Determine whether the data part of the message conforms to the HTTP request format;

[0178] 5. If it conforms to the HTTP request format, go to 6. If it does not conform to the HTTP request format, go to 16.

[0179] 6Query the Host field;

[0180] 7 If the corresponding Host field is found, go to 8; if the corresponding Host field is not found, it means that the message may be truncated or there is no Host in this message, go to 14;

[0181] 8 Check whether there is "\r" after the Host field. If it exists, go to 9; otherwise, go to 10.

[0182] 9. Obtain the complete domain name and compare it with the domain name whitelist. If the complete domain name is in the domain name whitelist, the process is released and ends. If the complete domain name is not in the domain name whitelist, the process is intercepted and ends.

[0183] 10. If there is no "\r" after the Host field, it means that the domain name is truncated. Get the domain name of the truncated part and the URI in the data part of the message. The domain name of the truncated part is the segmented domain name in the data part of the message.

[0184] 11. Use the triple information of the message (source IP / destination IP / destination port), the segmented domain name in the data part of the message, and the URI in the data part of the message as the key to query the domain name prediction table. If the node exists, determine the interception mark in the node (0 release / 1 interception). If the interception mark is 0, release and end the process. If the interception mark is 1, intercept and end the process. If the node does not exist, execute 12.

[0185] 12. Use the four-tuple information (source IP / destination IP / source port / destination port) of the message as the keyword key to query the segmented domain name table. If the node exists, clear the information in the node and execute 13 to end the process. If the node does not exist, create a node in the segmented domain name table using the four-tuple information (source IP / destination IP / source port / destination port) of the message as the keyword key and execute 13 to end the process.

[0186] 13. Store the segment domain name in the data part of the message in the domain name field of the node, store the URI in the data part of the message in the request URI field, store the target TCP sequence number in the next seq field, and leave the message field blank, wherein the target TCP sequence number is calculated as follows: add the TCP sequence number of the current message to the length of the data part of the current message;

[0187] 14. Use the four-tuple information of the message as the keyword key to query the segment domain name table. If the node exists, clear the information in the node, and execute 15 to end the process. If the node does not exist, create a node in the segment domain name table using the four-tuple information of the message as the keyword key, and execute 15 to end the process.

[0188] 15 Leave the domain name field and request URI field of the node blank, store the last 6 characters of the message in the message field, and store the target TCP sequence number in the next seq field. The target TCP sequence number is calculated as follows: add the TCP sequence number of the current message to the length of the data part of the current message;

[0189] 16. Use the four-tuple information of the message as the keyword key to query the segmented domain name table. If the node does not exist, the process ends. If the node exists, execute 17.

[0190] 17. Determine whether the TCP sequence number of the message is consistent with the target TCP sequence number in the next seq field of the node. If they are inconsistent, it indicates that the message is not the next segment message, and the process ends. If they are consistent, it indicates that the message is the next segment message, and execute 18.

[0191] 18: Determine whether the domain name field in the node is empty. If it is empty, execute 19. If it is not empty, execute 22.

[0192] 19. Concatenate the data in the message field of the node with the data part of the current message;

[0193] 20. Query the Host field for the concatenated data. If the corresponding Host field cannot be found, the message ends the process abnormally. If the corresponding Host field is found, execute 21.

[0194] 21 Check whether there is "\r" after the Host field. If not, the message ends abnormally. If so, execute 9.

[0195] 22 Check whether "\r" exists in the data part of the message. If not, the message ends abnormally. If so, execute 23.

[0196] 23. intercept the field before the first "\r" in the data part of the message, and concatenate the segmented domain name in the domain name field of the node with the intercepted field to obtain the complete domain name;

[0197] 24 Compare the complete domain name with the domain name whitelist. If the complete domain name is in the domain name whitelist, set the corresponding interception flag, which is 0 at this time. If the complete domain name is not in the domain name whitelist, set the corresponding interception flag, which is 1 at this time.

[0198] 25 Using the triple information of the message, the segmented domain name in the domain name field of the node, and the URI in the request URI field of the node as keywords, create a node in the domain name prediction table, and store the interception mark and the first timestamp in the node, where the first timestamp refers to the timestamp when the interception mark is stored in the node of the domain name prediction table;

[0199] 26 Return the interception mark, intercept or release according to the interception mark, and end the process.

[0200] For example, the web request message is divided into two parts and sent to the electronic device and the Host is cut off.

[0201] The first segment is "GET / HTTP / 1.1\r\nUser-Agent: curl / 7.29.0\r\nHo"

[0202] The second segment is "st: www.test.com\r\nAccept: * / *\r\n\r\n"

[0203] After obtaining the first segment of the message, the four-tuple information is parsed. According to "GET / HTTP / 1.1", it can be determined that it conforms to the HTTP request format. When parsing the HTTP protocol, the Host field is not recognized. ".0\r\nHo" is stored in the message field of the node with the four-tuple information as the keyword key, the domain name field of the node is set to empty, and the target TCP sequence number is calculated and stored in the next seq field of the node.

[0204] After obtaining the second segment of the message, parse the four-tuple information, and determine that it does not conform to the HTTP request format. Then use the four-tuple information as the keyword key to find the corresponding node, that is, the node in the previous step. After successfully comparing the TCP sequence number, concatenate ".0\r\nHo" with the current data part to get ".0\r\nHost: www.test.com\r\nAccept: * / *\r\n\r\n", identify the domain name and intercept or release it according to the domain name whitelist.

[0205] For example, the web request message is divided into two segments and sent to the electronic device, and the domain name is truncated.

[0206] The first segment is "GET / index.html HTTP / 1.1\r\nUser-Agent: curl / 7.29.0\r\nHost:www.te"

[0207] The second segment is "st.com\r\nAccept: * / *\r\n\r\n"

[0208] When this web request is requested for the first time:

[0209] After obtaining the first segment of the message, the four-tuple information is parsed. According to "GET / index.html HTTP / 1.1", it can be determined that it conforms to the HTTP request format. When parsing the HTTP protocol, the Host field is identified but the domain name is truncated. The request URI " / index.html" and the segmented domain name "www.te" of the data part are obtained, and "www.te" is stored in the domain name field of the node with the four-tuple information as the keyword key. " / index.html" is stored in the request URI field of the node. The target TCP sequence number is calculated and stored in the next seq field of the node, and the message field of the node is set to empty.

[0210] After obtaining the second segment of the message, the four-tuple information is parsed. After determining that it does not conform to the HTTP request format, the corresponding node is found using the four-tuple information as the keyword key, that is, the node in the previous step. After successfully comparing the TCP sequence number, the domain name "st.com" after the data part is obtained, and it is spliced ​​with the segmented domain name "www.te" in the domain name field of the node, and a domain name whitelist comparison is performed. The interception mark is set to intercept or release, and at the same time, a node is created with the triplet information, "www.te", and " / index.html" as the keyword key, and the interception mark and the first timestamp are stored in the node.

[0211] When this similar web request is requested for the second time:

[0212] After obtaining the first segment of the message, the four-tuple information is parsed. According to "GET / index.html HTTP / 1.1", it can be determined that it conforms to the HTTP request format. When parsing the HTTP protocol, the Host field is identified but the domain name is truncated. The request URI " / index.html" and the segmented domain name "www.te" of the data part are obtained. The domain name prediction table is queried with the triple information, "www.te" and " / index.html" as keywords to obtain the node, and intercept or release it according to the interception mark in the node.

[0213] In an embodiment of the present application, the target domain name is determined according to the segmented domain name or the last set number of characters in the second segmented message data part and the first segmented message data part, which can not only accelerate the recognition and splicing process of the segmented message domain name and improve the message interception efficiency and interception performance, but also reduce memory consumption, and the messages initiated with the same triplet information (source IP / destination IP / destination port), access domain name (segmented domain name in the data part of the message) and access URI (URI in the data part of the message) can be considered to have the same domain name, so that the source IP, destination IP, destination port, access domain name and access URI can predict that the domain name segmentation of this request is consistent with the previously recorded one, and then the message interception and release can be accelerated in a predictive manner, thereby improving the message interception / release efficiency.

[0214] The method provided by the present application is described above. The device provided by the present application is described below:

[0215] See also Figure 5 , is a structural diagram of a message interception device provided in an embodiment of the present application, such as Figure 5 As shown, the message interception device may include:

[0216] The first acquisition unit 510 is configured to, when the data part of the received first segmented message conforms to the HTTP request format, acquire first information if the data part of the first segmented message has a Host field and the domain name contained is an incomplete domain name, wherein the first information at least includes the segmented domain name in the data part of the first segmented message;

[0217] A second acquisition unit 520 is configured to acquire second information if the data part of the first segmented message does not have a Host field, where the second information includes at least a set number of characters at the end of the data part of the first segmented message;

[0218] The receiving unit 530 is used to receive a second segmented message, where the second segmented message and the first segmented message are continuous messages in the same direction in the same TCP stream;

[0219] A determining unit 540, configured to determine a target domain name according to a data portion of the second segment message and the first information or the second information;

[0220] The interception unit 550 is used to send a connection reset message to the sender of the first segment message when it is determined that the target domain name does not exist in the preset domain name whitelist.

[0221] In some embodiments, the first information further includes a uniform resource identifier (URI) in a data portion of the first segmented message, and the apparatus further includes:

[0222] A prediction unit, configured to query a target node in a preset domain name prediction table using the triplet information of the first segmented message and the first information as keywords, wherein the triplet information includes a source IP address and a destination port of the first segmented message;

[0223] If the target node exists, and the interception flag in the target node indicates that an interception operation is to be performed, a connection reset message is sent to the sender of the first segmented message.

[0224] In some embodiments, the first information further includes a URI in a data portion of the first segmented message, and the apparatus further includes:

[0225] A storage unit, used to determine an interception mark according to the target domain name and the domain name whitelist, wherein the interception mark is used to indicate whether to perform an interception operation;

[0226] The interception mark is stored in a node of a preset domain name prediction table using the triplet information of the second segmented message and the first information as keywords, wherein the triplet information includes the source and destination IP addresses and destination ports of the second segmented message.

[0227] In some embodiments, a first timestamp is further stored in the node of the domain name prediction table, and the first timestamp is a timestamp when the interception mark is stored in the node of the domain name prediction table using the triple information of the second segmented message and the first information as keywords; the device further includes:

[0228] An updating unit, used to determine whether the difference between the first timestamp and the current timestamp in each node of the domain name prediction table is greater than a preset value;

[0229] If the difference between the first timestamp and the current timestamp in a node is greater than a preset value, the node is deleted from the domain name prediction table.

[0230] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.

[0231] The present application embodiment also provides a hardware structure. Figure 6 , Figure 6 This is a structural diagram of an electronic device provided in an embodiment of the present application. Figure 6 As shown, the hardware structure may include: a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the method disclosed in the above example of this application.

[0232] Based on the same application concept as the above method, an embodiment of the present application also provides a machine-readable storage medium, on which a number of computer instructions are stored. When the computer instructions are executed by a processor, the method disclosed in the above example of the present application can be implemented.

[0233] Exemplarily, the above-mentioned machine-readable storage medium may be any electronic, magnetic, optical or other physical storage device, which may contain or store information, such as executable instructions, data, etc. For example, the machine-readable storage medium may be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drive (such as hard disk drive), solid state drive, any type of storage disk (such as optical disk, DVD, etc.), or similar storage medium, or a combination thereof.

[0234] It should be noted that, in this article, relational terms such as target and target are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0235] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A message interception method, characterized in that: The method includes: In a case where the data portion of the received first segmented message conforms to the HTTP request format, if the data portion of the first segmented message has a Host field and the domain name contained therein is an incomplete domain name, obtaining first information, the first information including at least a segmented domain name and a uniform resource identifier URI in the data portion of the first segmented message; Using the triplet information of the first segmented message, the segmented domain name and the URI as keywords, querying a preset domain name prediction table; if a node exists and the interception mark in the node indicates that an interception operation is performed, sending a connection reset message to the sender of the first segmented message, wherein the triplet information includes a source and destination IP address and a destination port; If the Host field does not exist in the data part of the first segmented message, obtain second information, where the second information includes at least the last set number of characters of the data part of the first segmented message, and the absence of the Host field in the data part of the first segmented message indicates that the Host field is truncated; receiving a second segmented message, where the second segmented message and the first segmented message are continuous messages in the same direction in the same TCP stream; Determine the target domain name based on the data part of the second segmented message and the first information, and when it is determined that the target domain name does not exist in the preset domain name whitelist, set an interception mark representing the execution of an interception operation, and create a node in the domain name prediction table with the triple information of the first segmented message, the segmented domain name and the URI as keywords, and after storing the interception mark in the created node, send a connection reset message to the sender of the first segmented message; or, determine the target domain name based on the data part of the second segmented message and the second information, and when it is determined that the target domain name does not exist in the domain name whitelist, send a connection reset message to the sender of the first segmented message.

2. The method according to claim 1, characterized in that The node of the domain name prediction table further stores a first timestamp, and the first timestamp is the timestamp when the interception mark is stored in the node of the domain name prediction table; the method further includes: Determine whether the difference between the first timestamp and the current timestamp in each node of the domain name prediction table is greater than a preset value; If the difference between the first timestamp and the current timestamp in a node is greater than a preset value, the node is deleted from the domain name prediction table.

3. The method according to claim 1, characterized in that The method further comprises: Using the four-tuple information of the first segmented message as a keyword, the first information or the second information is stored in a node of a preset segmented domain name table, wherein the four-tuple information includes a source-destination IP address and a source-destination port.

4. The method according to claim 3, characterized in that The determining the target domain name according to the data portion of the second segment message and the first information or the second information includes: Using the quadruple information of the second segment message as a keyword, querying the second target node in the segment domain name table; The target domain name is determined according to the data portion of the second segmented message and the first information or the second information in the second target node.

5. The method according to claim 1, characterized in that: The method further comprises: If the data part of the first segmented message has a Host field, determine whether there is a line terminator after the Host field, where the line terminator is used to separate the fields in the data part of the first segmented message; If the line terminator does not exist, it is determined that the domain name included in the data portion of the first segmented message is an incomplete domain name.

6. The method according to claim 1, characterized in that The first information and the second information also include a target TCP sequence number, and the target TCP sequence number is determined according to the TCP sequence number and the length of the data part of the first segmented message.

7. The method according to claim 6, characterized in that The method further comprises: Determine whether the TCP sequence number of the second segment message is the same as the target TCP sequence number in the first information or the second information; If they are the same, it is determined that the second segmented message and the first segmented message are continuous messages in the same direction in the same TCP stream.

8. A message interception device, characterized in that: The device includes: A first acquisition unit is configured to, when the data portion of the received first segmented message conforms to the HTTP request format, acquire first information if a Host field exists in the data portion of the first segmented message and the domain name contained therein is an incomplete domain name, wherein the first information at least includes a segmented domain name and a uniform resource identifier URI in the data portion of the first segmented message; Using the triplet information of the first segmented message, the segmented domain name and the URI as keywords, querying a preset domain name prediction table; if a node exists and the interception mark in the node indicates that an interception operation is performed, sending a connection reset message to the sender of the first segmented message, wherein the triplet information includes a source and destination IP address and a destination port; a second acquiring unit, configured to acquire second information if the data portion of the first segmented message does not contain the Host field, the second information comprising at least a set number of characters at the end of the data portion of the first segmented message, the absence of the Host field in the data portion of the first segmented message indicating that the Host field is truncated; A receiving unit, configured to receive a second segmented message, wherein the second segmented message and the first segmented message are continuous messages in the same direction in the same TCP stream; An interception unit is used to determine a target domain name based on the data portion of the second segmented message and the first information, and to set an interception mark representing the execution of an interception operation when it is determined that the target domain name does not exist in a preset domain name whitelist, and to create a node in the domain name prediction table with the triple information of the first segmented message, the segmented domain name and the URI as keywords, and after storing the interception mark in the created node, send a connection reset message to the sender of the first segmented message; or, to determine a target domain name based on the data portion of the second segmented message and the second information, and to send a connection reset message to the sender of the first segmented message when it is determined that the target domain name does not exist in the domain name whitelist.

9. The device according to claim 8, characterized in that The node of the domain name prediction table also stores a first timestamp, and the first timestamp is the timestamp when the interception mark is stored in the node of the domain name prediction table; the device also includes: An updating unit, used to determine whether the difference between the first timestamp and the current timestamp in each node of the domain name prediction table is greater than a preset value; If the difference between the first timestamp and the current timestamp in a node is greater than a preset value, the node is deleted from the domain name prediction table.

10. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor is used to execute the machine executable instructions to implement the method according to any one of claims 1 to 7.

11. A machine-readable storage medium, characterized in that: The machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

12. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Message filtering method and device

    CN103401850A