Compromised Terminal Identification Method, System and Medium for Metropolitan Area Network
By spectroscopic processing of Internet traffic and matching threat intelligence databases in the metropolitan area network, identifying and positioning the lost terminals, the problem that traditional security protection systems cannot identify in time is solved, and fast and accurate positioning and protection of the lost terminals is achieved.
Patent Information
- Application Number
- CN202510647531.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The prior art cannot identify and locate lost terminals controlled by attackers in the metropolitan area network in a timely and accurately manner, especially when facing complex attacks such as advanced persistent threats and zombie worms, traditional security protection systems are difficult to respond quickly.
By spectroscopic processing of the uplink Internet access traffic of the operator's core network equipment at the metropolitan area network interface, filtering and formatting the domain name request to resolve the message flow, using the threat intelligence library for double-order iterative matching, combining the unit IP library and secure DNS equipment, identifying and positioning the lost terminal.
It realizes fast and accurate identification and positioning of lost terminals, improves the protection efficiency and accuracy of complex attack modes, and ensures real-time monitoring and response capabilities of network security.
Smart Images

Figure CN120165990B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of digital communication technologies, and particularly to a method, system, and medium for identifying compromised terminals for a metropolitan area network (MAN). Background Art
[0002] With the continuous upgrading of network attack means, traditional security protection methods have been difficult to effectively cope with complex attacks such as advanced persistent threats (APTs) and botnets. These attacks usually spread horizontally by controlling a large number of terminal devices (compromised terminals), greatly increasing the difficulty of detection and prevention. Existing security protection systems mainly rely on static rules and external threat intelligence for protection. However, due to the lack of comprehensive monitoring of large-scale network traffic and real-time detection of internal terminals, it is difficult to quickly and accurately identify the devices that have been attacked and controlled. Therefore, there is an urgent need for a security protection technology that can monitor large-scale network traffic in real time and accurately identify compromised terminals. Summary of the Invention
[0003] This application provides a method, system, and medium for identifying compromised terminals for a metropolitan area network, which are used to solve the technical problem in the prior art that traditional network security protection cannot timely and accurately identify and locate compromised terminals controlled by attackers in a metropolitan area network.
[0004] In the first aspect of this application, a method for identifying compromised terminals for a metropolitan area network is provided. The method includes: at the interface of the metropolitan area network, performing optical splitting on the upstream Internet access traffic of the operator's core network devices, importing the optical signal into a splitter, filtering and forwarding the Internet access traffic in the splitter to obtain a set of domain name request resolution message streams; sending the set of domain name request resolution message streams to a collector for formatting processing to obtain a set of domain name request resolution records; performing two-stage iterative matching on the domain name fields in the set of domain name request resolution records with malicious domain names in a threat intelligence database, marking the threat level of the domain name request resolution records according to the matching results, and summarizing the domain name request resolution records marked as serious threats into a set of risk domain name request resolution records; retrieving based on the set of risk domain name request resolution records and a unit IP library respectively to determine a set of compromised units, and deploying a secure DNS device at the Internet exit of each compromised unit to collect internal domain name request records to obtain a cluster of internal domain name request records; matching the set of risk domain name request resolution records with the corresponding set of internal domain name request records in the cluster of internal domain name request records to obtain a set of compromised terminals.
[0005] The second aspect of the present application provides a compromised terminal identification system for a metropolitan area network. The system includes: a domain name request resolution module, which is used to perform optical splitting on the upstream Internet access traffic of the core network equipment of the operator at the interface of the metropolitan area network, import the optical signal into a splitter, filter and forward the Internet access traffic in the splitter to obtain a set of domain name request resolution message streams; a formatting processing module, which is used to send the set of domain name request resolution message streams to a collector for formatting processing to obtain a set of domain name request resolution records; a risk domain name identification module, which is used to perform two-stage iterative matching on the domain name fields in the set of domain name request resolution records with the malicious domain names in the threat intelligence library, identify the threat level of the domain name request resolution records according to the matching results, and summarize the domain name request resolution records marked as serious threats into a set of risk domain name request resolution records; an internal domain name request record collection module, which is used to retrieve based on the set of risk domain name request resolution records and the unit IP library of each unit to determine a set of compromised units, and deploy a secure DNS device at the Internet exit of each compromised unit to collect internal domain name request records to obtain a cluster of internal domain name request records; a compromised terminal matching module, which is used to match the set of risk domain name request resolution records with the corresponding set of internal domain name request records in the cluster of internal domain name request records to obtain a set of compromised terminals.
[0006] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method of the first aspect is implemented.
[0007] One or more technical solutions provided in the present application have at least the following technical effects or advantages:
[0008] The method, system and medium for identifying compromised terminals for a metropolitan area network provided in the present application relate to the field of digital communication technologies. By performing optical splitting on the metropolitan area network traffic, filtering and formatting the domain name request resolution message streams, using two-stage matching and the threat intelligence library to identify risk domain names, matching with the unit IP library to determine compromised units, and deploying a secure DNS device at their Internet exits to collect internal domain name requests, and finally identifying compromised terminal devices by matching internal records, it solves the technical problem in the prior art that traditional network security protection cannot timely and accurately identify and locate compromised terminals controlled by attackers in the metropolitan area network, and realizes the technical effect of quickly and accurately identifying and locating compromised terminals through multi-level data matching and internal and external network traffic analysis. Description of the Drawings
[0009] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0010] Figure 1 Schematic flow diagram of the compromised terminal identification method for the metropolitan area network provided by the embodiments of the present application;
[0011] Figure 2 Schematic structural diagram of the compromised terminal identification system for the metropolitan area network provided by the embodiments of the present application.
[0012] Explanation of reference numerals: Domain name request resolution module 11, formatting processing module 12, risk domain name identification module 13, internal domain name request record collection module 14, compromised terminal matching module 15. Detailed implementation manners
[0013] The present application provides a compromised terminal identification method, system and medium for the metropolitan area network, which is used to solve the technical problem that traditional network security protection in the prior art cannot timely and accurately identify and locate compromised terminals controlled by attackers in the metropolitan area network.
[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0015] It should be noted that the terms "first", "second", etc. in the specification of the present application and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server including a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or modules not clearly listed or inherent to these processes, methods, products or devices.
[0016] Embodiment 1, as Figure 1 shown, the present application provides a compromised terminal identification method for the metropolitan area network, and the method includes:
[0017] P10: At the interface of the metropolitan area network, perform optical splitting on the upstream Internet access traffic of the operator's core network devices, import the optical signal into a splitter, filter and forward the Internet access traffic in the splitter, and obtain a set of domain name request resolution message flows.
[0018] Among them, in the splitter, use the five-tuple to filter the traffic in the Internet access traffic that does not meet the preset requirements to obtain the set of domain name request resolution message flows. Among them, the preset requirement is that the domain name request traffic must conform to the UDP or TCP protocol, and the destination port is the corresponding UDP port or TCP port 53.
[0019] It should be understood that at the interface of the metropolitan area network, first perform optical splitting on the upstream Internet access traffic of the operator's core network devices. The purpose is to separate the full network traffic from different operators. Specifically, the optical signal is extracted from the operator's core network devices and imported into the splitter. Through this optical splitting process, all Internet access traffic from different operators is concentrated in the splitter for unified traffic screening and processing.
[0020] In the splitter, screen the received Internet traffic through the five-tuple filtering technology. The five-tuple filtering method classifies and identifies data streams based on the five-tuple (source IP address, source port, destination IP address, destination port, and protocol type). In this step, the preset five-tuple filtering requirement is that the traffic must conform to the UDP protocol or TCP protocol, and the destination port must be UDP port 53 or TCP port 53. These two ports are specifically used for DNS (Domain Name System) request resolution.
[0021] During the filtering process, by checking the five-tuple, the splitter filters out all traffic that does not meet the above conditions. For example, if the destination port of a certain data stream is not 53, or its protocol is not UDP or TCP, then this traffic will be excluded, thus avoiding interference from irrelevant traffic to subsequent processing. After this screening, what is finally left is the traffic that conforms to the domain name request resolution message flow, and these message flows contain key data related to DNS requests.
[0022] Through the above screening, the splitter finally obtains a set of domain name request resolution message flows, which includes all DNS requests and resolution messages that meet the conditions. The key data included in these message flows includes fields such as domain name request time, source IP address, destination IP address, resolved domain name, and resolved IP address. These domain name request resolution records will be transmitted to the downstream collector for further processing and storage for risk analysis, entity identification, and terminal location in subsequent steps.
[0023] Through precise screening by the splitter, we ensure that only data flows related to DNS domain name request resolution are collected, and avoid wasting non-target traffic. This method not only improves the accuracy of data collection, but also enables subsequent security analysis to focus on high-risk domain name requests, thereby improving the efficiency and accuracy of the system.
[0024] P20: Send the domain name request resolution message flow set to the collector for formatting, and obtain a domain name request resolution record set.
[0025] The domain name request resolution record set includes at least the domain name request time, domain name request source IP address, source port, domain name server IP address, destination port, requested resolution domain name, domain name resolution response time and resolution IP address.
[0026] Optionally, in this step, the domain name request resolution message flow set is sent to the collector for formatting. After receiving these message flows, the collector will parse them, extract key information and convert them into a standardized domain name request resolution record set. These records contain multiple fields to ensure that the detailed information of each domain name request can be fully reflected, which is convenient for subsequent analysis and tracing.
[0027] First, the collector extracts the domain name request time and records the specific time when each request occurs. This is crucial for analyzing the time distribution of requests and whether there is abnormal traffic or attack mode. Next, the domain name request source IP address is extracted. This field can help trace back to the device or terminal that initiated the request and further identify potential attack sources. The source port is also necessary information. It identifies the port used when the request was initiated, which is crucial for locating network sessions and analyzing the source of requests.
[0028] At the same time, the domain name server IP address will be recorded, which points to the DNS server that performs domain name resolution, helping to determine which server the request is resolved through. This field helps analyze whether there is an abnormality in the DNS server itself or the possibility of being attacked. The destination port field is usually port 53, which is the standard port of the DNS protocol, ensuring that the request meets the basic requirements of DNS requests. Next, the collector will extract the domain name requested for resolution, that is, the target domain name that initiated the request. This field is the core information of the request and can directly reflect the attacker's goal or attack intention.
[0029] The collector also records the domain name resolution response time, which reflects the response speed of the DNS service and can help identify abnormal performance of the DNS service. Especially in attack scenarios, abnormal response times may indicate problems such as denial-of-service (DoS) attacks. Finally, the collector extracts the resolved IP address, which is the result returned by the DNS server and identifies the target server or device to which the domain name is ultimately resolved.
[0030] Through formatting, the collector converts these raw message streams into structured records and stores them in the server database. These records will include all key information about the domain name requests, ensuring that subsequent analysis can perform effective risk assessment, threat identification, and terminal location based on this data. At the same time, the records stored by the collector need to have a sufficient traceability duration to enable retrospective analysis in the event of subsequent security incidents. Table 1 shows the specific descriptions of each field in the domain name request resolution record set:
[0031] Table 1: Domain Name Resolution Record Table (Resolution)
[0032] Field Name Field Name Meaning Request_Time Domain Name Request Time Source_IP Domain Name Request Source IP Address Source_Port Source Port DNS_IP Domain Name Server IP Destination_Port Destination Port DomainName Domain Name to be Resolved Response_Time Domain Name Resolution Response Time Resolved_IP Resolved IP Address Threat_Type Threat Type Threat_Level Threat Level
[0033] The formatted domain name request resolution record set not only improves the readability and operability of the data but also provides accurate data support for subsequent threat detection, entity identification, and terminal location in the system.
[0034] P30: Perform a two-stage iterative match between the domain name field in the domain name request resolution record set and the malicious domain names in the threat intelligence database. Based on the matching results, mark the threat level of the domain name request resolution records, and summarize the domain name request resolution records marked as serious threats into a risk domain name request resolution record set.
[0035] Furthermore, step P30 of the embodiment of the present application further includes:
[0036] P31: Perform two - stage keyword extraction on the malicious domain names in the threat intelligence library to obtain a set of first - order keywords and a set of second - order mapped keyword groups; P32: Use an encoder to perform first - order matching on the domain name fields in the domain name request parsing record set according to the set of first - order keywords to obtain a first - order matching result; P33: Use the set of second - order mapped keyword groups to perform implicit iterative matching recognition on the domain name request parsing records corresponding to the first - order matching result to obtain a set of second - order matching information, where each piece of second - order matching information corresponds to a domain name request parsing record in the first - order matching result; P34: Invoke a threat level identifier to perform threat level recognition on the set of second - order matching information, and perform mapping identification on the domain name request parsing records in the first - order matching result according to the recognition result to obtain a completed - identification first - order matching result; P35: Aggregate the domain name request parsing records in the completed - identification first - order matching result that are marked as serious threats to obtain a set of risk domain name request parsing records.
[0037] Specifically, perform two - stage iterative matching on the domain name fields in the domain name request parsing record set and the malicious domain names in the threat intelligence library. The core of this process lies in accurately identifying those domain name requests that may be related to malicious activities through a multi - level keyword extraction and matching mechanism, assigning corresponding threat levels to these domain name requests, and finally screening out the domain name requests with serious threats and aggregating them into a set of risk domain name request parsing records.
[0038] Specifically, first perform two - stage keyword extraction on the malicious domain names in the threat intelligence library. The purpose of this process is to extract a set of first - order keywords and a set of second - order mapped keyword groups from the malicious domain names. The set of first - order keywords includes the common keyword parts in the malicious domain names, usually the explicit parts in the domain names (for example, "malware", "phishing", etc.), while the set of second - order mapped keyword groups is the implicit features derived from the first - order keywords, which may be domain name suffixes, IP address ranges, etc. related to malicious activities. Exemplarily, collect the intelligence of multiple malicious domain names, and the threat intelligence of each manufacturer must be aggregated into a threat intelligence record table, and the fields included are at least as listed in Table 2:
[0039] Table 2: Threat Intelligence Record Table (Threat)
[0040] Field Name Field Name Meaning DomainName Domain Name CNAME Alias DomainName_Owner Domain Name Owner Registration_Time Domain Name Registration Time Expiration_Time Domain Name Expiration Time Service_Provider Domain Name Service Provider Resolved_IP Resolved IP IP_Location Geographical Location of Resolved IP Threat_Information Threat Intelligence Information Threat_Type Threat Type Threat_Level Threat Level Information_Update_Time Intelligence Update Time Information_Source Intelligence Source
[0041] Next, use an encoder to process the set of first - order keywords and perform first - order matching on the domain name fields in the domain name request parsing record set. The purpose of this step is to preliminarily screen out those records that contain first - order keywords in the domain name. The encoder generates a preliminary matching result through the matching of the domain name fields and identifies the domain name request records that may pose threats.
[0042] Furthermore, the second-order mapping keyword group set is used to perform implicit iterative matching recognition on the domain name request records in the first-order matching results. This process conducts a deeper analysis and recognition of the first-order matching results, taking into account implicit features and more complex attack patterns. Through this implicit matching, more potential threat information can be identified, forming a second-order matching information set, where each element in the set corresponds to a domain name request resolution record in the first-order matching results.
[0043] Next, the threat level identifier is called to identify the threat level of the second-order matching information set. The identifier determines the threat level of the domain name request based on the data characteristics in the second-order matching information and maps this level information back to the domain name request records in the first-order matching results. In this way, each domain name request record can be tagged with a threat level, providing a basis for subsequent risk assessment and defense.
[0044] Finally, the domain name request resolution records marked as serious threats are extracted from the first-order matching results and summarized to form a set of risk domain name request resolution records. These records represent the domain name requests determined to be of high risk during the entire matching process, usually domain names related to network attacks or malicious behavior.
[0045] Exemplarily, for the risk analysis of requested domain names: Each domain name request resolution record is matched one by one with the domain name and threat intelligence record table, and fields such as the threat type and threat level of the domain name request resolution record are updated.
[0046] The reference SQL statements for database operations are as follows (Domain name resolution record table: Table A, Threat intelligence record table: Table B, Domain name field: Field D, Threat type field: Field X, Threat level field: Field Y): UPDATE Table A; SET Field X = (SELECT b.Field X FROM Table B b WHERE b.Field D = Table A.Field D); Field Y = (SELECT b.Field Y FROM Table B b WHERE b.Field A = Table A.Field A) WHERE EXISTS (SELECT 1 FROM Table B b WHERE b.Field D = Table A.Field D); The reference SQL statement for finding the source IP address of the domain name request with a threat level of "serious threat" is as follows (Domain name resolution record table: Table A, Source IP address field of domain name request: Field S, Threat level field: Field Y): SELECT Field S FROM Table A; WHERE Field Y = "serious threat"; In this way, the IP addresses infected with malicious programs at the "serious threat" level are output.
[0047] Through the above steps, this method can not only achieve accurate threat detection through two-level keyword matching, but also ensure that the system can timely and accurately identify potential network security threats through multi-level matching and threat level identification. This method can provide strong data support for network security protection, especially in the face of complex and changeable malicious attacks, providing an efficient detection method.
[0048] Furthermore, step P31 of the embodiment of the present application further includes:
[0049] P31-1: extract malicious domain names from the threat intelligence database to obtain a first-order keyword set; P31-2: retrieve the records of the first-order keyword set in the threat intelligence according to the preset implicit malicious domain name keyword type set to determine the second-order mapping keyword group set. The preset implicit malicious domain name keyword type set includes at least DNS resolution, historical deformation, alias, and jump pointing.
[0050] It should be understood that the processing of malicious domain names in the threat intelligence library can be further refined to extract a more targeted and accurate set of keywords.
[0051] First, extract all malicious domain name information from the threat intelligence database and analyze these malicious domain names to generate a first-order keyword set. This set consists of characteristic words that appear explicitly in malicious domain names. These characteristics are usually iconic words in domain names or words related to attack types, such as "malware" and "phishing". These words reflect the characteristics of specific malicious activities or attack families. By extracting these keywords, we can lay the foundation for subsequent malicious domain name matching and identification.
[0052] Next, the system further retrieves and analyzes the records in the first-order keyword set according to the preset set of hidden malicious domain name keyword types. The set of hidden malicious domain name keyword types includes multiple features, which may not appear explicitly in the domain name, but can effectively help identify malicious behavior. First, the system analyzes the resolution behavior of domain name requests based on the feature of DNS resolution. Malicious domain names often use specific resolution methods, such as using non-standard DNS requests or adopting special DNS query modes. These behaviors can help the system identify potential attack domain names. Secondly, historical deformation is another important analysis point. Malicious domain names often perform character replacement or spelling deformation in order to evade static matching rules or increase the concealment of attacks. By looking back at historical data, the system can identify these deformed domain names and further increase the accuracy of matching.
[0053] In addition, malicious domains usually have aliases, that is, multiple domains point to the same malicious control point. Attackers often use multiple domains or IP addresses to control victim terminals, and these aliases usually have a high correlation with the main domain name. By extracting this alias information, the system can identify more potential malicious activities. Finally, the characteristics of the jump destination also provide important clues for the identification of malicious domains. Attackers may use DNS redirection to resolve one domain name to another domain name or IP address. This redirection is usually used to avoid detection and increase the concealment of the attack. By analyzing the jump destination of the domain name, the system can trace these jump paths and further reveal the potential threats of malicious domains.
[0054] By retrieving and analyzing these types of implicit keywords, the system can mine more implicit malicious features from the first-order keyword set, and finally form a second-order mapped keyword group set. These mapped keywords represent the malicious features of the domain name at a deeper level and can help the system identify more complex and concealed malicious domains.
[0055] After this process is completed, the system will perform implicit iterative matching on the domain name request resolution records in the first-order matching results based on the second-order mapped keyword group set, so as to obtain a more accurate malicious domain name identification result. This iterative matching process can not only further improve the matching accuracy of malicious domain names, but also discover some malicious domain names that evade traditional matching rules through means such as deformation, aliasing, or jump destination, providing a more detailed basis for subsequent threat level identification.
[0056] In summary, through the extraction of first-order keywords and the in-depth mining of second-order implicit features, this embodiment can achieve the accurate identification of malicious domain names, ensuring that even malicious activities carried out through concealed means such as deformation and redirection can be discovered in a timely manner, providing strong technical support for network security protection.
[0057] Furthermore, step P33 of the embodiment of the present application further includes:
[0058] P33-1: Extract the first domain name request resolution record from the first-order matching result; P33-2: Perform mapping matching on the second-order mapping keyword group set according to the first-order keyword corresponding to the first domain name request resolution record to obtain the matching second-order mapping keyword group; P33-3: Perform associated matching on the matching second-order mapping keyword group and the first field of the first domain name request resolution record. If the matching is successful, obtain the first matching mapping keyword; P33-4: Add the semantic annotation corresponding to the first matching mapping keyword, the first matching mapping keyword, and the matching second-order mapping keyword group into an initially empty vector to obtain the first memory vector; P33-5: Obtain the second field of the first domain name request resolution record, perform implicit iterative matching on the second field using the first memory vector, and update the first memory vector according to the matching result to obtain the second memory vector, and so on. The memory vector obtained after the field matching is completed is used as the first second-order matching information of the first domain name request resolution record; P33-6: Perform implicit iterative matching recognition on the domain name request resolution record corresponding to the first-order matching result in combination with the second-order mapping keyword group set to obtain the second-order matching information set.
[0059] Specifically, the processing process of the domain name request resolution record in the first-order matching result can be further refined. This process goes through a series of sub-steps, and through implicit iterative matching recognition, it deeply explores potential threat information.
[0060] First, extract the first domain name request resolution record from the previously obtained first-order matching result. The purpose of this operation is to select the records in the first-order matching result that need to be further processed. These records have been preliminarily matched with malicious domain names through first-order keywords and are the basic data for further analysis.
[0061] Next, according to the first-order keyword corresponding to the extracted first domain name request resolution record, the system will perform mapping matching on the second-order mapping keyword group set. The key to this process is that the system finds the relevant second-order mapping keywords in the threat intelligence library according to the first-order keyword, so as to discover more hidden and complex malicious features. The keywords in the second-order mapping keyword group set usually contain some features that are not easily detected directly, and these features are potential signs of malicious activities.
[0062] Next, perform associated matching on the matching second-order mapping keyword group and the first field of the first domain name request resolution record. The first field here may be the domain name field, the request time, or other key data fields. In this way, the system can verify the relationship between the domain name request and the second-order keyword, and further confirm whether the domain name request is related to malicious behavior. If the matching is successful, the first matching mapping keyword will be obtained, which is the key connection between the domain name request and malicious activities.
[0063] Further, the obtained first matching mapped keyword and its corresponding semantic annotation are added together with the second-order mapped keyword group into an initially empty vector, thereby generating a first memory vector. The purpose of this step is to encode the matched keyword and its semantic information to form a vector representation for subsequent iterative matching and analysis. This memory vector not only contains the matching information but also helps the system understand the semantic meaning behind each keyword, providing richer context for subsequent matching.
[0064] Subsequently, the second field of the first domain name request parsing record is obtained, and implicit iterative matching is performed on this field using the first memory vector. Implicit iterative matching means that without directly showing all the information, based on the memory vector obtained in the previous step, further matching is performed on the second field. This process continuously updates the information in the vector and finally obtains a second memory vector. In this way, the system can continuously optimize the matching result through step-by-step iteration and obtain a more accurate malicious domain name recognition result.
[0065] Finally, the system continues to perform implicit iterative matching recognition on the domain name request parsing record corresponding to the first-order matching result in combination with the second-order mapped keyword group set, and finally obtains a second-order matching information set. This set contains all relevant information obtained through second-order keywords and implicit matching, marking the deep matching result of the domain name request in multiple dimensions. Through these second-order matching information, the system can more accurately identify potential malicious activities and provide stronger threat recognition capabilities.
[0066] Through the above steps, the matching accuracy of the domain name request parsing record can be effectively improved, not only limited to the matching of explicit features, but also deeply mining the implicit features behind the domain name request. This method provides the system with more powerful malicious domain name detection capabilities through iterative matching and vectorization techniques, can discover more complex and hidden attack behaviors, and provides more powerful data support for network security protection.
[0067] Further, step P33-5 of the embodiment of the present application further includes:
[0068] P33-51: Perform associated matching on the second field of the first domain name request parsing record using the first memory vector. If the matching is successful, obtain a second matching mapped keyword; P33-52: Add the semantic annotation corresponding to the first matching mapped keyword and the second matching mapped keyword into the first memory vector to update it, and obtain a second memory vector.
[0069] In a possible embodiment of the present application, the implicit iterative matching process between the first memory vector and the second field can be further refined. This process mainly performs associated matching on the second field of the first domain name request resolution record, continuously optimizing and updating the memory vector in order to obtain a more accurate malicious domain name recognition result.
[0070] First, use the first memory vector to perform associated matching on the second field of the first domain name request resolution record. The second field may be other information related to the domain name request, such as the source port of the request, the response time, the resolved IP address, etc. Using the information in the first memory vector, the system matches the second field to find potential features related to malicious activities therein. When the second field successfully matches the keywords in the memory vector, the system obtains the second matching mapped keywords. These mapped keywords represent the malicious features in the second field, further deepening the association between the domain name request and malicious activities.
[0071] Next, add the first matching mapped keywords and the second matching mapped keywords obtained through the matching, together with their corresponding semantic annotations, into the first memory vector and update it. The core of this process lies in continuously integrating the matched keywords and their semantic information into the memory vector, enabling the memory vector to gradually enrich and optimize during multiple iterative matching processes. This updated memory vector, namely the second memory vector, contains more context information and malicious activity features, which helps to improve the accuracy of subsequent matching.
[0072] Through the operations of these two sub-steps, the system can gradually improve the matching accuracy of the domain name request resolution record through iterative matching. Each matching and vector update will make the system's recognition of malicious domain names more in-depth and accurate, thereby enhancing the detection ability of potential threats. Finally, as the updated vector, the second memory vector will provide more accurate data support for the subsequent second-order matching information recognition, helping the system accurately identify domain name request records with threats.
[0073] P40: Based on the risk domain name request resolution record set and the unit IP library respectively, perform retrieval to determine the set of compromised units, and deploy security DNS devices at the Internet egress of each compromised unit to collect internal domain name request records, obtaining an internal domain name request record cluster.
[0074] Specifically, perform retrieval based on the risk domain name request resolution record set obtained in the previous step and the unit IP library. The purpose is to determine the set of compromised units by comparing the risk domain name requests with the public IP addresses of each unit. These units refer to those targets that may have been controlled by malware or are under cyber attacks.
[0075] First, by retrieving the unit IP library, the system can identify which units' public IP addresses match the IP addresses in the risky domain name requests. The unit IP library contains public IP address records of each unit. Through these records, the system can compare the received risky domain name requests to further locate the specific victimized units. Whenever the system finds a matching public IP address, it can determine that the unit has become the target of a cyber attack and thus include it in the set of compromised units. The network egress devices of these units may have been controlled by attackers to transmit malicious instructions or execute attack behaviors.
[0076] Exemplarily, through the unit IP table (see Table 3), the corresponding unit name can be queried by the IP address (the IP address infected with malicious programs analyzed in the analysis step), so as to locate which unit is infected with what malicious program.
[0077] Table 3: Unit IP Table (IP_Table)
[0078] Field Name Field Name Meaning Organization_Name Organization Name Address Organization Address Contacts Contacts Telephone Contact Phone Region Region Belonged To IP_address IP Address Operator Operator Belonged To
[0079] The reference SQL statement for the database operation to find the unit name from the unit IP table by the specific IP address is as follows: SELECT Organization_Name FROM IP_Table WHERE IP_address = "IP address".
[0080] Next, to further confirm and prevent these affected units, the system will deploy security DNS devices at the Internet egress of each compromised unit. The function of these security DNS devices is to monitor and record all domain name request data within the unit, so as to obtain the cluster of domain name request records within the unit. The cluster of internal domain name request records refers to the set of domain name request data initiated by all terminals in the internal network of the unit. By deploying security DNS devices at the Internet egress, the system can capture all DNS requests sent from the unit and further determine whether there are controlled terminal devices within the unit through the analysis of these requests.
[0081] The purpose of deploying a secure DNS device is as follows: on the one hand, it can collect domain name request data within the organization in real time, providing basic data for subsequent malicious domain name detection; on the other hand, the collected internal domain name request records can also be used to analyze the security status of the internal network of the organization, especially to determine whether there are compromised endpoints. During the collection process, the DNS device not only records the specific content of domain name requests but can also further determine whether there are malicious activities based on the response information of domain name resolution. Exemplarily, at the organization's exit, a secure DNS device is installed, and all domain name request records of the organization are sent to the secure DNS device for identification, storage, and disposal. The stored organization domain name resolution record table is shown in Table 4:
[0082] Table 4: Organization Domain Name Resolution Record Table (Resolution)
[0083] Field Name Field Name Meaning Organization_Name Organization Name Request_Time Domain Name Request Time Source_IP Domain Name Request Source IP Address Source_Port Source Port DNS_IP Domain Name Server IP Destination_Port Destination Port DomainName Domain Name to be Resolved Response_Time Domain Name Resolution Response Time Resolved_IP Resolved IP Address Threat_Type Threat Type Threat_Level Threat Level
[0084] Through this process, the system can achieve security monitoring from the external public network to the internal network, ensuring that compromised organizations can be accurately identified and located, and providing sufficient data support for subsequent endpoint location, risk prevention, and repair.
[0085] P50: Match the set of risk domain name request resolution records with the corresponding set of internal domain name requests in the internal domain name request record cluster to obtain the set of compromised endpoints.
[0086] Optionally, match the set of risk domain name request resolution records obtained in the previous step with the corresponding records in the internal domain name request record cluster to identify and locate the specific set of compromised endpoints.
[0087] Specifically, first, extract those domain name requests marked as having a serious threat from the set of risk domain name request resolution records. These domain name requests are usually related to known malicious activities and may include domain names from malicious control servers or domain names related to attack behaviors. Then, the system compares these risk domain name requests with the domain name request record cluster in the organization's internal network. The internal domain name request record cluster refers to all domain name request records initiated within the compromised organization, and these requests come from all terminal devices of the organization.
[0088] By comparing the set of risk domain name request resolution records and the corresponding records in the internal domain name request record cluster, the system can find out which internal terminal devices sent requests related to known malicious domain names. The key to this step is to associate external risk domain names with internal domain name request data to achieve accurate location of compromised endpoints. The successfully matched records will represent those terminals that may have been controlled by attackers, and the system will include these terminals in the set of compromised endpoints.
[0089] Once the compromised terminals are identified, the system can clarify the association between these terminals and malicious domains, and then conduct further security analysis and response. In this way, the system not only identifies the controlled terminals, but also can trace back to the attack sources involved by these terminals, helping the security team take measures promptly to prevent further damage.
[0090] In summary, by matching the risk domains with the domain name request data within the unit, the specific terminals under attack and control are accurately identified, thus providing an important basis for subsequent security repair and threat disposal.
[0091] In summary, the embodiments of the present application at least have the following technical effects:
[0092] In the present application, the upstream Internet traffic of the operator's core network equipment is split at the MAN interface, the domain name request resolution packet stream set is filtered and obtained, and then formatted into a domain name request resolution record set. Through two-stage matching with the malicious domains in the threat intelligence library, serious threats are identified and summarized into a risk domain name request resolution record set. Based on this set, a retrieval is performed with the unit IP library to determine the compromised units, and a secure DNS device is deployed to collect internal domain name requests. Finally, the risk domains are matched with the internal records to obtain the compromised terminal set and locate the controlled devices.
[0093] It achieves the technical effects of identifying and locating the compromised terminals in real time and accurately through two-stage iterative matching and real-time domain name request analysis, and improving the protection efficiency and accuracy against complex attack patterns.
[0094] Embodiment 2, based on the same inventive concept as the compromised terminal identification method for the MAN in the foregoing embodiment, as Figure 2 shown, the present application provides a compromised terminal identification system for the MAN. The system in the embodiments of the present application and the method embodiments are based on the same inventive concept. Among them, the system includes:
[0095] A domain name request resolution module 11, which is used to perform splitting processing on the upstream Internet access traffic of the operator's core network equipment at the interface of the MAN, import the optical signal into a splitter, and filter and forward the Internet access traffic in the splitter to obtain a domain name request resolution packet stream set.
[0096] A formatting processing module 12, which is used to send the domain name request resolution packet stream set to a collector for formatting processing to obtain a domain name request resolution record set.
[0097] A risk domain name identification module 13, which is used to perform two-stage iterative matching on the domain name fields in the domain name request resolution record set and the malicious domain names in the threat intelligence library, identify the threat levels of the domain name request resolution records according to the matching results, and summarize the domain name request resolution records marked as serious threats into a risk domain name request resolution record set.
[0098] An internal domain name request record collection module 14, which is used to retrieve based on the risk domain name request resolution record set and the respective unit IP library to determine the set of compromised units, and deploy security DNS devices at the Internet exits of each compromised unit to collect internal domain name request records, obtaining an internal domain name request record cluster.
[0099] A compromised terminal matching module 15, which is used to match the risk domain name request resolution record set with the corresponding internal domain name request record set in the internal domain name request record cluster to obtain a set of compromised terminals.
[0100] Furthermore, the domain name request resolution module 11 is also used to perform the following steps:
[0101] Filter the Internet access traffic that does not meet the preset requirements in the shunt using the five-tuple to obtain the domain name request resolution message stream set, where the preset requirements are that the domain name request traffic should conform to the UDP or TCP protocol, and the destination port is the corresponding UDP port or TCP port 53.
[0102] Furthermore, the formatting processing module 12 is also used to perform the following steps:
[0103] The domain name request resolution record set at least includes the domain name request time, the source IP address of the domain name request, the source port, the IP address of the domain name server, the destination port, the domain name to be resolved, the domain name resolution response time, and the resolved IP address.
[0104] Furthermore, the risk domain name identification module 13 is also used to perform the following steps:
[0105] Perform two - stage keyword extraction on the malicious domain names in the threat intelligence library to obtain a set of first - order keywords and a set of second - order mapped keyword groups; use an encoder to perform first - order matching on the domain name fields in the domain name request parsing record set according to the set of first - order keywords to obtain a first - order matching result; use the set of second - order mapped keyword groups to perform implicit iterative matching recognition on the domain name request parsing records corresponding to the first - order matching result to obtain a set of second - order matching information, where each second - order matching information corresponds to a domain name request parsing record in the first - order matching result; call a threat level identifier to perform threat level recognition on the set of second - order matching information, and perform mapping identification on the domain name request parsing records in the first - order matching result according to the recognition result to obtain a completed first - order matching result with identification; summarize the domain name request parsing records marked as serious threats in the first - order matching result with identification to obtain a set of risk domain name request parsing records.
[0106] Further, the risk domain name recognition module 13 is further configured to perform the following steps:
[0107] Extract the first domain name request parsing record from the first - order matching result; perform mapping matching on the set of second - order mapped keyword groups according to the first - order keywords corresponding to the first domain name request parsing record to obtain a matching second - order mapped keyword group; perform associated matching between the matching second - order mapped keyword group and the first field of the first domain name request parsing record. If the matching is successful, obtain a first matching mapped keyword; add the semantic annotation corresponding to the first matching mapped keyword, the first matching mapped keyword, and the matching second - order mapped keyword group into an initially empty vector to obtain a first memory vector; obtain the second field of the first domain name request parsing record, perform implicit iterative matching on the second field using the first memory vector, and update the first memory vector according to the matching result to obtain a second memory vector, and so on. The memory vector obtained after the field matching is completed is used as the first second - order matching information of the first domain name request parsing record; perform implicit iterative matching recognition on the domain name request parsing records corresponding to the first - order matching result in combination with the set of second - order mapped keyword groups to obtain a set of second - order matching information.
[0108] Further, the risk domain name recognition module 13 is further configured to perform the following steps:
[0109] Perform associated matching on the second field of the first domain name request parsing record using the first memory vector. If the matching is successful, obtain a second matching mapped keyword; add the semantic annotation corresponding to the first matching mapped keyword and the second matching mapped keyword into the first memory vector to update it and obtain a second memory vector.
[0110] Further, the risk domain name recognition module 13 is further configured to perform the following steps:
[0111] Extract malicious domain names from the threat intelligence library to obtain a set of first-order keywords; retrieve the records of the set of first-order keywords in the threat intelligence according to the preset set of implicit malicious domain name keyword types respectively to determine a set of second-order mapped keyword groups.
[0112] Further, the risk domain name recognition module 13 is further configured to perform the following steps:
[0113] The preset set of implicit malicious domain name keyword types at least includes DNS resolution, historical deformation, aliases, and jump pointers.
[0114] Embodiment 3, based on the same inventive concept as the method for identifying compromised terminals for a metropolitan area network in the foregoing embodiments, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method in Embodiment 1 is implemented.
[0115] Through the foregoing detailed description of the method for identifying compromised terminals for a metropolitan area network in this specification, those skilled in the art can clearly know the method, system, and medium for identifying compromised terminals for a metropolitan area network in this embodiment. Therefore, for the sake of simplicity of the specification, it will not be elaborated here. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0116] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0117] It should be noted that the above sequence of the embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. In addition, the specific embodiments of this specification have been described. Further, the processes depicted in the drawings do not necessarily require the particular order shown or sequential order to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0118] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
[0119] This specification and the drawings are merely exemplary descriptions of the present application and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalent technologies, the present application is intended to include these changes and modifications.
Claims
1. A method for identifying compromised terminals for a metropolitan area network, characterized in that, The method includes: At the interface of the metropolitan area network, perform optical splitting on the upstream Internet access traffic of the operator's core network devices, import the optical signal into a splitter, filter and forward the Internet access traffic in the splitter to obtain a set of domain name request resolution message streams; Send the set of domain name request resolution message streams to a collector for formatting processing to obtain a set of domain name request resolution records; Perform two-stage iterative matching on the domain name fields in the set of domain name request resolution records with malicious domain names in the threat intelligence library, identify the threat level of the domain name request resolution records according to the matching results, and summarize the domain name request resolution records marked as serious threats into a set of risk domain name request resolution records; Retrieve the unit IP library based on the set of risk domain name request resolution records to determine the set of compromised units, and deploy security DNS devices at the Internet exits of each compromised unit to collect internal domain name request records to obtain a cluster of internal domain name request records; Match the set of risk domain name request resolution records with the corresponding set of internal domain name request records in the cluster of internal domain name request records to obtain a set of compromised terminals; Among them, performing two-stage iterative matching on the domain name fields in the set of domain name request resolution records with malicious domain names in the threat intelligence library, identifying the threat level of the domain name request resolution records according to the matching results, and summarizing the domain name request resolution records marked as serious threats into a set of risk domain name request resolution records includes: Perform two-stage keyword extraction on the malicious domain names in the threat intelligence library to obtain a set of first-order keywords and a set of second-order mapped keyword groups; Use an encoder to perform first-order matching on the domain name fields in the set of domain name request resolution records according to the set of first-order keywords to obtain a first-order matching result; Use the set of second-order mapped keyword groups to perform implicit iterative matching recognition on the domain name request resolution records corresponding to the first-order matching result to obtain a set of second-order matching information, where each piece of second-order matching information corresponds to a domain name request resolution record in the first-order matching result; Call a threat level identifier to identify the threat level of the set of second-order matching information, and perform mapping identification on the domain name request resolution records in the first-order matching result according to the recognition results to obtain a completed identified first-order matching result; Summarize the domain name request resolution records marked as serious threats in the identified first-order matching result to obtain a set of risk domain name request resolution records; Among them, using the set of second-order mapped keyword groups to perform implicit iterative matching recognition on the domain name request resolution records corresponding to the first-order matching result to obtain a set of second-order matching information includes: Extract the first domain name request resolution record from the first-order matching result; Perform mapping matching on the set of second-order mapped keyword groups according to the first-order keyword corresponding to the first domain name request resolution record to obtain a matching second-order mapped keyword group; Perform associated matching on the matching second-order mapped keyword group and the first field of the first domain name request resolution record. If the matching is successful, obtain the first matching mapped keyword; Add the semantic annotation corresponding to the first matching mapped keyword, the first matching mapped keyword, and the matching second-order mapped keyword group into an initially empty vector to obtain the first memory vector; Obtain the second field of the first domain name request resolution record, perform implicit iterative matching on the second field using the first memory vector, and update the first memory vector according to the matching result to obtain the second memory vector. By analogy, use the memory vector obtained after the field matching is completed as the first second-order matching information of the first domain name request resolution record; Perform implicit iterative matching and recognition on the domain name request resolution record corresponding to the first-order matching result in combination with the set of second-order mapped keyword groups to obtain a set of second-order matching information.
2. The method for identifying compromised terminals for a metropolitan area network according to claim 1, characterized in that, Obtain the second field of the first domain name request resolution record, perform implicit iterative matching on the second field using the first memory vector, and update the first memory vector according to the matching result to obtain the second memory vector, including: Perform associated matching on the second field of the first domain name request resolution record using the first memory vector. If the matching is successful, obtain the second matching mapped keyword; Add the semantic annotation corresponding to the first matching mapped keyword and the second matching mapped keyword into the first memory vector to update it and obtain the second memory vector.
3. The method for identifying compromised terminals for a metropolitan area network according to claim 2, wherein Including: Extract malicious domain names from the threat intelligence library to obtain a set of first-order keywords; Retrieve the records of the set of first-order keywords in the threat intelligence respectively according to the preset set of implicit malicious domain name keyword types to determine the set of second-order mapped keyword groups.
4. The compromised terminal identification method for a metropolitan area network according to claim 3, wherein The preset set of implicit malicious domain name keyword types includes at least DNS resolution, historical deformation, alias, and jump destination.
5. The method for identifying compromised terminals for a metropolitan area network according to claim 1, characterized in that Use the five-tuple in the shunt to filter the Internet access traffic that does not meet the preset requirements in the shunt to obtain the set of domain name request resolution message streams, where the preset requirement is that the domain name request traffic should conform to the UDP or TCP protocol, and the destination port is the corresponding UDP port or TCP port 53.
6. The method for identifying compromised terminals for a metropolitan area network according to claim 1, wherein The set of domain name request resolution records includes at least the domain name request time, the source IP address of the domain name request, the source port, the IP address of the domain name server, the destination port, the requested domain name for resolution, the domain name resolution response time, and the resolved IP address.
7. Compromised terminal identification system for metropolitan area network, characterized in that, The system is used to implement the method for identifying compromised terminals for a metropolitan area network according to any one of claims 1-6. The system includes: A domain name request resolution module, which is used to perform optical splitting processing on the upstream Internet access traffic of the operator's core network device at the interface of the metropolitan area network, import the optical signal into the shunt, and perform filtering and forwarding on the Internet access traffic in the shunt to obtain a set of domain name request resolution message streams; A formatting processing module, which is used to send the set of domain name request resolution message streams to the collector for formatting processing to obtain a set of domain name request resolution records; A risk domain name recognition module, which is used to perform two-stage iterative matching on the domain name fields in the domain name request resolution record set and the malicious domain names in the threat intelligence library, identify the threat levels of the domain name request resolution records according to the matching results, and summarize the domain name request resolution records marked as serious threats into a risk domain name request resolution record set; An internal domain name request record collection module, which is used to retrieve the unit IP library based on the risk domain name request resolution record set, determine the set of compromised units, and deploy security DNS devices at the Internet exits of each compromised unit to collect internal domain name request records, obtaining an internal domain name request record cluster; A compromised terminal matching module, which is used to match the risk domain name request resolution record set with the corresponding internal domain name request record set in the internal domain name request record cluster to obtain a set of compromised terminals.
8. A computer-readable storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the method for identifying compromised terminals for a metropolitan area network according to any one of claims 1 to 6.
Citation Information
Patent Citations
APT detection method based on matching of flow fingerprint and communication features
CN108833437A
Detection method and device of lost terminal, electronic equipment and storage medium
CN117978472A